Target detection method and device based on rotating frame and terminal equipment
By introducing rotation box and target angle loss functions in the YOLO model, the problem of poor detection performance in complex scenarios is solved, and better target box feature extraction and detection performance improvement is achieved.
Patent Information
- Application Number
- CN202311628457.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-06
AI Technical Summary
The traditional YOLO perception algorithm based on horizontal target box has poor detection performance in complex scenarios, especially when backgrounds are mixed or there are many occlusions.
Using a rotating box-based object detection method, the prediction angle information is added by improving the network detection head of the YOLO model, and the label dimension of the target angle is added in the data set, and the cross entropy loss function is used to train the target angle loss function.
Effectively extract the features of the target box to improve the overall detection performance of the YOLO perception algorithm, especially in complex scenarios, the accuracy and robustness of the detection are improved.
Smart Images

Figure CN120107538A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of target detection, and in particular to a target detection method, device and terminal device based on a rotating frame. Background Art
[0002] Due to its high precision and efficiency, the YOLO perception algorithm is suitable for a variety of intelligent scenarios such as autonomous driving, intelligent transportation, industrial detection and security monitoring. In general, due to the high precision, high adaptability and efficiency of the YOLO perception algorithm, it is suitable for various application scenarios that require real-time target detection and scene analysis. Therefore, the YOLO perception algorithm has broad application prospects in the fields of industry, transportation, medical care, security, etc.
[0003] In the actual target detection process, due to the complexity of real scenes, the scenes to be detected may often be relatively mixed. When the background is mixed or there are many occlusions and the target object may have any direction, the target detection results of the traditional YOLO perception algorithm based on horizontal target boxes are not ideal. Therefore, it is crucial to better extract the features of the target box and improve the overall detection performance of the YOLO perception algorithm. Summary of the invention
[0004] The purpose of this application is to provide a method, apparatus, terminal device and storage medium for target detection based on a rotating frame, aiming to solve the problem of how to better extract the features of the target frame and improve the overall detection performance of the YOLOv5 perception algorithm.
[0005] A first aspect of an embodiment of the present application provides a method for object detection based on a rotating frame, the method comprising:
[0006] Acquire an image to be detected including a target pattern;
[0007] The image to be detected is input into an improved YOLO model, and the improved YOLO model outputs a target recognition result; wherein the network detection head of the improved YOLO model includes predicted angle information, and the target recognition result is target box attribute information including the target angle.
[0008] In an embodiment that can be implemented in the present application, the method further includes:
[0009] Reconstruct a standard dataset of the initial YOLO model and increase the dimension of the standard dataset by the label dimension of the target angle;
[0010] Expanding the dimension of the target angle on the data iterator to support reading the target angle;
[0011] Adding predicted angle information to the network detection head and assigning angle information to each image grid coordinate;
[0012] A target angle loss function is added to generate the improved YOLO model, wherein the target angle loss function and the classification loss function use the same cross entropy loss function.
[0013] In the embodiments that can be implemented in this application, it also includes:
[0014] Two independent network structures are set up to perform regression task prediction and classification task prediction respectively; wherein the convolutional layers of the two network structures are the same or different.
[0015] In an achievable embodiment of the present application, the reconstructing a standard data set of an initial YOLO model, increasing the dimension of the standard data set by the label dimension of the target angle, includes:
[0016] The 5-dimensional data of the standard dataset label [cls, x, y, w, h] is converted into 6-dimensional data [cls, x, y, w, h, θ] with angle information, where cls is the labeled category, w refers to the long side of the box, h refers to the other side of the box, and the angle θ refers to the angle formed by the first long side encountered by the horizontal x-axis through counterclockwise rotation, and the range is [0, 180).
[0017] In an embodiment that can be implemented in the present application, the step of adding a target angle loss function to generate the improved YOLO model includes:
[0018] Define a first rotation frame and a second rotation frame;
[0019] Take a side line segment from the first rotating frame, calculate whether there is an intersection with the second rotating frame, if not and the side line segment is inside the second rotating frame, keep the side line segment, if there is an intersection and the other end point is inside the second rotating frame, keep the line segment of the intersection with the second rotating frame, if there are two intersections, keep the line segment composed of the two intersections, and repeat traversing the line segments of the first rotating frame;
[0020] Repeat the operation from the second rotation frame to determine the corner point of the first rotation frame;
[0021] Sort the retained edges to obtain the intersection area polygon;
[0022] Use the areas of polygons to find intersection and union.
[0023] In an embodiment that can be implemented in the present application, the method further includes:
[0024] The localization loss is replaced by the error between the predicted rotated box and the calibrated rotated box.
[0025] In an embodiment that can be implemented in the present application, the convolution layer of the network structure corresponding to the regression task prediction is a 1*1 convolution layer, and the convolution layer of the network structure corresponding to the classification task prediction is a 3*3 convolution layer.
[0026] A second aspect of an embodiment of the present application provides an object detection device based on a rotating frame, comprising:
[0027] An acquisition module is used to acquire an image to be detected including a target pattern;
[0028] The recognition module inputs the image to be detected into the improved YOLO model, and the improved YOLO model outputs a target recognition result; wherein the network detection head of the improved YOLO model includes predicted angle information, and the target recognition result is target box attribute information including the target angle.
[0029] A third aspect of an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the computer program.
[0030] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0031] Beneficial effects of this application
[0032] This application inputs the acquired image to be detected into the YOLO model improved based on the rotation frame. This method facilitates better extraction of the features of the target frame and greatly improves the overall detection performance of the YOLO perception algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0034] Figure 1 A conventional target detection scene diagram based on a horizontal target frame provided in an embodiment of the present application;
[0035] Figure 2 A flowchart of a rotating frame-based target detection method provided in an embodiment of the present application;
[0036] Figure 3A scene diagram of a target detection method based on a rotating frame provided in an embodiment of the present application;
[0037] Figure 4 A flowchart of generating the improved YOLO model in a rotating frame-based target detection method provided in an embodiment of the present application;
[0038] Figure 5 A flow chart of adding a target angle loss function to generate the improved YOLO model provided in an embodiment of the present application;
[0039] Figure 6 A flowchart of the optimization loss function provided in an embodiment of the present application;
[0040] Figure 7 A schematic diagram of the structure of a rotating frame-based target detection device provided in an embodiment of the present application;
[0041] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0043] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0044] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0045] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0046] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0047] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0048] It should be understood that the size of the serial numbers of each step in this embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.
[0049] In the existing target detection technology, the horizontal target frame is basically used to locate the target object. The scene graph of the traditional method is as follows: Figure 1 As shown in the figure, each vehicle has a corresponding horizontal target box. However, since the vehicles on the road have directions, even the houses in the scene will not have only one direction. Figure 1 It can be seen that when the vehicle overtakes or tilts in a certain direction, the required horizontal target frame will become larger accordingly, and sometimes overlap with the surrounding horizontal target frames. Under such conditions, the target background of the target detection will increase, and the accuracy of the target detection result will be reduced. Therefore, in order to solve this problem, the present application proposes a method of target detection based on a rotating frame. The present application detects the target object by changing the horizontal target frame into a rotating frame based on the direction of the target object. Through this method, the rotating frame can tightly contain the target object and reduce the background information of the target object. At the same time, the present application uses the target facing direction as the angle information for annotation, and then in actual application, the vehicle heading angle information can be obtained through detection. Based on the above operation, the present application can better extract the features of the target frame on the one hand, and on the other hand, it also improves the overall detection performance of the YOLO perception algorithm.
[0050] In order to illustrate the technical solution of the present application, specific embodiments are provided below.
[0051] like Figure 2As shown, a method for object detection based on a rotating frame includes:
[0052] S201, acquiring an image to be detected including a target pattern;
[0053] S202, inputting the image to be detected into an improved YOLO model, and the improved YOLO model outputs a target recognition result; wherein the network detection head of the improved YOLO model includes predicted angle information, and the target recognition result is target box attribute information including the target angle.
[0054] For example, suppose there is a road scene Figure 3 It contains the target result, which is car No. 1. Figure 3 The image to be detected is input into the improved YOLO model, whose network detection head includes the predicted angle information. After processing, the YOLO model outputs the target vehicle, including the target frame attribute information of the target vehicle angle. The target frame attributes are: (x min ,y min , x max ,y max ), target angle information: angle value, x min is the minimum value of the horizontal coordinate of the target box, y min is the minimum value of the vertical coordinate of the target box, x max is the maximum value of the horizontal coordinate of the target box, y max is the maximum value of the vertical coordinate of the target frame. Specifically, the model will output the following target vehicle results, namely, target: Car No. 1, target frame attributes: (5.2, 2, 7.1, 3.8), target angle information: 40°. In this embodiment, the target frame attribute information in the target recognition result represents the position and size of Car No. 1 in the image, and the target angle information represents the angle of the car relative to the horizontal direction. This output result can more comprehensively understand the detected target and provide more information for subsequent processing and decision-making.
[0055] This method can not only better extract the features of the target frame, but also improve the overall detection performance of the YOLO perception algorithm.
[0056] In an embodiment that can be implemented in the present application, the method further includes:
[0057] S401, reconstructing a standard data set of an initial YOLO model, and increasing the dimension of the standard data set by the label dimension of the target angle;
[0058] S402, expanding the dimension of the target angle on the data iterator to support reading the target angle;
[0059] S403, adding predicted angle information to the network detection head, and assigning angle information to each image grid coordinate;
[0060] S404, adding a target angle loss function to generate the improved YOLO model, wherein the target angle loss function and the classification loss function use the same cross entropy loss function.
[0061] For example, suppose there is a standard dataset containing images and their corresponding target box information. Now we need to reconstruct this dataset and add the label dimension of the target angle. In this process, the data iterator will be expanded to support reading the target angle information. Next, the predicted angle information is added to the network detection head, and the angle information is assigned to each image grid coordinate. Finally, the target angle loss function is added to generate an improved YOLO model, in which the target angle loss function and the classification loss function use the same cross entropy loss function.
[0062] In deep learning, the cross entropy loss function is often used for classification tasks, which measures the difference between the probability distribution of the model output and the true label. In classification tasks, the cross entropy loss function is often used to measure the difference between the probability distribution of the model output and the true label.
[0063] For the binary classification problem, the cross entropy loss function can be expressed as:
[0064] L=-(y*log(p)+(1-y)*log(1-p))
[0065] Among them, y is the true label (0 or 1), p is the probability of the model output (a value between 0 and 1), and log represents the natural logarithm. The goal of this loss function is to make the probability of the model output as close to the true label as possible, thereby minimizing the value of the loss function.
[0066] For multi-classification problems, the form of the cross entropy loss function will be different, but the core idea is similar. It still measures the difference between the probability distribution of the model output and the true label.
[0067] The target angle loss function and the classification loss function in this application use the same cross entropy loss function. Different models or different network structures use the same cross entropy loss function to measure their performance and train them. Since the cross entropy loss function is a standard loss function for classification tasks, it is applicable to various classification models.
[0068] In addition, in some specific scenarios, the cross entropy loss function can be used to handle regression tasks, such as when dealing with angle regression problems. In this case, the angle values can be converted to category labels, and then the cross entropy loss function is used for training. The advantage of this method is that it can utilize the loss functions and optimization algorithms commonly used in classification tasks.
[0069] Another approach is to use a specially designed target angle loss function to handle angle prediction in regression tasks. This loss function usually takes into account the periodicity and continuity of the angle to better measure the difference between the predicted value and the true value.
[0070] Specifically, assuming that an image in the initial dataset contains a target (such as a car), the location information of the target box and the target angle information need to be annotated. After the dataset is reconstructed, the label data will include the location coordinates of the target box and the target angle information. Next, the YOLO model will be improved by adding a network detection head that predicts the angle information and assigning angle information to each image grid coordinate. At the same time, the target angle loss function needs to be defined, and the cross entropy loss function can be used to measure the prediction error of the target angle. In the end, an improved YOLO model will be obtained that can simultaneously predict the location of the target box and the target angle, thereby more comprehensively understanding the detected target.
[0071] Suppose the goal is to detect cars and we are also interested in the car's driving angle. The initial dataset contains the car's location information (cls, x, y, w, h) and the car's category label. Now we want to expand the dataset to add the car's angle information.
[0072] First, the format of the dataset can be modified to expand the label of each sample to (cls, x, y, w, h, θ), where θ represents the angle information of the car.
[0073] Where cls is the labeled category, w refers to the long side of the target box, h refers to the other side of the target box, and the target angle θ refers to the angle formed by the first long side encountered by the horizontal x-axis through counterclockwise rotation, and the range is [0, 180).
[0074] Next, you need to modify the data iterator to support reading the target angle information. You can modify the loading and preprocessing steps in the data iterator accordingly to ensure that the target angle information is read and processed correctly.
[0075] Then, the detection head of the YOLO model needs to be modified to add the prediction of the object's angle information. An additional output dimension can be added to the detection head to predict the angle information of the car, and the network structure can be adjusted accordingly to support this change.
[0076] Finally, a comprehensive loss function can be defined to combine the classification loss and the target angle loss, and use the cross entropy loss function to measure the classification loss part. For example, the comprehensive loss function can be defined as:
[0077] L = λ1*classification loss + λ2*target angle loss
[0078] Among them, the classification loss can use the cross entropy loss function, the target angle loss can use the loss function suitable for the angle regression task, and λ1 and λ2 are the weight parameters of the two losses.
[0079] Through such steps, an improved YOLO model can be generated, in which the target angle loss function and the classification loss function use the same cross entropy loss function. Such an improved model can better handle angle information, thereby improving the accuracy and robustness of car detection.
[0080] The present application supports reading the target angle by increasing the dimension expansion of the target angle. This step enables a more comprehensive understanding of the target information, and the target detection result is more accurate by improving the model.
[0081] In an embodiment that can be implemented in the present application, the method further includes:
[0082] Two independent network structures are set up to perform regression task prediction and classification task prediction respectively; wherein the convolutional layers of the two network structures are the same or different.
[0083] For example, suppose there is an image dataset and it is desired to simultaneously perform regression (such as target location) and classification (such as target category) tasks on the targets in the image.
[0084] Two independent network structures can be designed for prediction of regression and classification tasks respectively. These two network structures can share the same convolutional layer to extract image features, and then connect to different fully connected layers or classifiers for prediction of regression and classification tasks respectively.
[0085] Specifically, suppose a convolutional neural network (CNN) is used for image object detection and classification tasks. A network structure with a shared convolutional layer can be designed, and the subsequent branches of the network include a fully connected layer for regression tasks and a fully connected layer for classification tasks. In addition, two independent CNN network structures can be designed, which share the same convolutional layer but are connected to different fully connected layers or classifiers for prediction of regression and classification tasks respectively.
[0086] This step of the present application performs regression and classification tasks on the targets in the image respectively, so as to more comprehensively understand the content of the image to be detected.
[0087] In an achievable embodiment of the present application, the reconstructing a standard data set of an initial YOLO model, increasing the dimension of the standard data set by the label dimension of the target angle, includes:
[0088] The original 5-dimensional data of the standard dataset label [cls, x, y, w, h] is converted into 6-dimensional data [cls, x, y, w, h, θ] with angle information, where cls is the labeled category, w refers to the long side of the target box, h refers to the other side of the target box, and the target angle θ refers to the angle formed by the first long side encountered by the horizontal x-axis through counterclockwise rotation, and the range is [0, 180).
[0089] For example, suppose there is an initial YOLO model standard data set, in which each labeled target includes category, center coordinates, width and height. Now we want to increase the dimension of the label data to include target angle information, converting from 5-dimensional data to 6-dimensional data. The new label data will include category, center coordinates, width, height, and target angle.
[0090] For example, suppose an image in the initial dataset contains multiple targets (such as cars, pedestrians, etc.), and the original label data is [cls, x, y, w, h], where cls is the category, (x, y) is the center coordinate of the target box, w is the width of the target box, and h is the height of the target box. Now we need to convert this 5-dimensional data into 6-dimensional data with angle information, that is, [cls, x, y, w, h, θ], where θ is the target angle.
[0091] Assume that the category of the car is 1, the category of the pedestrian is 2, the angle information of the car target box is 30°, and the angle information of the pedestrian target box is 10°. At this time, the original standard dataset label of the car is [1, 32, 5, 3, 3], and the converted car label data will become [1, 32, 5, 3, 3, 30]. At this time, the original standard dataset label of the pedestrian is [2, 20, 2, 1, 1], and the converted pedestrian label data will become [2, 20, 2, 1, 1, 10].
[0092] In the reconstructed dataset, the label of each target includes the category, center coordinates, width, height, and target angle information, which will provide additional target angle information for the improved YOLO model, enabling the model to more comprehensively understand the detected targets.
[0093] The present application obtains the angle information of the detection target through this step. Through this information, the moving direction and path of the detection target can be predicted, and the features of the target frame can be accurately extracted.
[0094] In an embodiment that can be implemented in the present application, the step of adding a target angle loss function to generate the improved YOLO model includes:
[0095] S501, defining a first rotation frame and a second rotation frame;
[0096] S502, taking a side line segment from the first rotation frame, calculating whether there is an intersection with the second rotation frame, if there is no intersection and the side line segment is inside the second rotation frame, retaining the side line segment, if there is an intersection and the other end point is inside the second rotation frame, retaining the line segment of the intersection with the second rotation frame, if there are two intersections, retaining the line segment composed of the two intersections, and repeatedly traversing the first rotation frame line segment;
[0097] S503, repeating the operation from the second rotation frame to determine the corner point of the first rotation frame;
[0098] S504, sorting the retained edges to obtain the intersection area polygon;
[0099] S505, using the areas of the polygons to obtain the intersection and union.
[0100] For example, Figure 6 The specific process of adding the target angle loss function and generating the improved YOLO model. (1) First, define two rotating frames a and b. (2) Take a side m from a and calculate whether it has an intersection with the rotating frame b. If not and the line segment m is inside b, then keep the line segment m. If there is an intersection and the other end point is inside b, then keep the line segment that intersects with b. If there are two intersections, then keep the line segment composed of the two intersections and repeat the traversal of the rotating frame a line segment. (3) Repeat the operation in the rotating frame b to determine the corner point situation with the rotating frame a. (4) Sort the retained edges to obtain the intersection area polygon. (5) Use the area of the polygon to obtain the intersection and union.
[0101] Specifically, assume that there are rotating frames A and B, and the specific parameters are as follows: the category of rotating frame A is 1, the center coordinates are (3, 3), the width is 4, the height is 2, and the angle is 30°. The category of rotating frame B is 1, the center coordinates are (5, 5), the width is 3, the height is 3, and the angle is 45°. First, take an edge segment from rotating frame A and calculate whether it has an intersection with rotating frame B. Then perform similar operations on rotating frame B. Suppose that the intersection of them is calculated to be a triangle with an area of 2 square units. Calculate the intersection polygon and sort the retained edges to obtain the intersection area polygon. Calculate the intersection and union area. Use the area of the polygon to calculate the intersection and union area. Suppose that the union of rotating frames A and B is calculated to be a pentagon with an area of 12 square units. Finally, calculate its intersection-union ratio. In the target detection task, it is necessary to determine whether the two rotating frames A and B represent the same target. By calculating their intersection-over-union ratio, the degree of overlap between the A and B rotation frames can be determined, which facilitates the next step of target matching.
[0102] In this way, the rotation angle information of the target can be considered more comprehensively, and the areas of the intersection and union can be calculated, which helps improve the detection accuracy of the YOLO model for rotated targets.
[0103] In an embodiment that can be implemented in the present application, the method further includes:
[0104] The localization loss is replaced by the error between the predicted rotated box and the calibrated rotated box.
[0105] For example, suppose there is a target detection task that needs to detect a rotating car target. The dimension of the standard dataset has been increased by the label dimension of the target angle. Now it is hoped that the loss function of the YOLO model will be modified to take into account the error between the predicted rotation box and the calibrated rotation box.
[0106] In this case, a loss function suitable for rotated targets, such as the angle regression loss function, can be used. This loss function can measure the model's prediction accuracy of the target rotation angle, thereby helping the model better understand the position and angle information of the rotated target.
[0107] Specifically, the mean absolute error (MAE) or mean square error (MSE) can be used as the angle regression loss function to calculate the error between the rotation angle predicted by the model and the calibrated rotation angle. This angle regression loss function is then combined with other loss functions of the YOLO model (such as the classification loss function and the target box coordinate loss function) to form a new comprehensive loss function.
[0108] In this way, the present application can better consider the rotation angle information of the target, thereby improving the detection accuracy of the model for rotating targets.
[0109] In an embodiment that can be implemented in the present application, the convolution layer of the network structure corresponding to the regression task prediction is a 1*1 convolution layer, and the convolution layer of the network structure corresponding to the classification task prediction is a 3*3 convolution layer.
[0110] For example, suppose there is a target detection task that needs to detect car targets. Use the YOLO model for target detection, which includes classification tasks and regression tasks. The classification task is used to predict which category the target belongs to, and the regression task is used to predict the location and size of the target. In the YOLO model, classification tasks and regression tasks use different convolutional layers for prediction. The classification task uses a 3*3 convolutional layer for prediction, while the regression task uses a 1*1 convolutional layer for prediction. This is because the classification task needs to consider the different categories of the target and requires a larger convolution kernel to extract the features of the target. The regression task only needs to predict the location and size of the target, and does not need to consider the category of the target, so a smaller convolution kernel can be used to extract features. For example, in the YOLOv5 model, the classification task uses the convolutional layer in the Darknet-53 network, which contains multiple 3*3 convolutional layers. The regression task uses a 1*1 convolutional layer for prediction.
[0111] Specifically, for classification tasks, we need to predict the category of the car target, such as "sedan", "SUV", "truck", etc. In order to extract the features of the target and classify it, we can use 3*3 convolutional layers to capture the local features and contextual information of the target. These convolutional layers can help the network learn the features of different categories of cars.
[0112] For regression tasks, it is necessary to predict the position and size of the car target. In this case, a 1*1 convolution layer can be used for regression prediction. 1*1 convolution layers can help the network learn the position and size information of the target because they do not need to consider the local features of the target, but only need to perform a linear transformation on each pixel to achieve regression prediction of position and size. Since the prediction of classification tasks and regression tasks focuses on different tasks and information extraction requirements, the convolution layer sizes of the corresponding network structures are different in this process.
[0113] The method of the present application reduces the amount of calculation and model size in the detection process while ensuring the accuracy of target detection.
[0114] Figure 7 A structural schematic diagram of a rotating frame-based target detection device provided in an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0115] A rotating frame-based object detection device 700 may specifically include the following modules:
[0116] An acquisition module 701 is used to acquire an image to be detected including a target pattern;
[0117] The recognition module 702 is used to input the image to be detected into the improved YOLO model, and the improved YOLO model outputs a target recognition result; wherein the network detection head of the improved YOLO model includes predicted angle information, and the target recognition result is target box attribute information including the target angle.
[0118] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0119] In the above embodiments, the description of each embodiment has its own emphasis. If a rated part is not described or recorded in detail in a certain embodiment, reference may be made to the relevant description of other embodiments.
[0120] Figure 8 Schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 8 As shown, the electronic device 3000 of this embodiment includes: at least one processor 3001 ( Figure 8 Only one is shown), a memory 3002, and a computer program 3003 stored in the memory 3002 and executable on at least one processor 3001, where the processor 3001 implements the steps in the above embodiments when executing the computer program 3003.
[0121] The processor 3001 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0122] In some embodiments, the memory 3002 may be an internal storage unit of the electronic device 3000, such as a hard disk or memory of the electronic device 3000. In other embodiments, the memory 3002 may also be an external storage device of the electronic device 3000, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 3000. Further, the memory 3002 may also include both an internal storage unit of the electronic device 3000 and an external storage device. The memory 3002 is used to store an operating system, an application program, a boot loader (Boot Loader) data, and other programs, such as program codes of a computer program, etc. The memory 3002 may also be used to temporarily store data that has been output or is to be output.
[0123] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0124] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0125] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0126] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0127] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0128] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0129] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0130] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for object detection based on rotating frame, It is characterized in that include: Acquire an image to be detected including a target pattern; The image to be detected is input into an improved YOLO model, and the improved YOLO model outputs a target recognition result; wherein the network detection head of the improved YOLO model includes predicted angle information, and the target recognition result is target box attribute information including the target angle.
2. The method according to claim 1, It is characterized in that The method further comprises: Reconstruct a standard dataset of the initial YOLO model and increase the dimension of the standard dataset by the label dimension of the target angle; Expanding the dimension of the target angle on the data iterator to support reading the target angle; Adding predicted angle information to the network detection head and assigning angle information to each image grid coordinate; A target angle loss function is added to generate the improved YOLO model, wherein the target angle loss function and the classification loss function use the same cross entropy loss function.
3. The method according to claim 2, It is characterized in that Also includes: Two independent network structures are set up to perform regression task prediction and classification task prediction respectively; wherein the convolutional layers of the two network structures are the same or different.
4. The method according to claim 2, It is characterized in that The reconstructing of a standard data set of an initial YOLO model, increasing the dimension of the standard data set by the label dimension of the target angle, includes: The 5-dimensional data of the standard dataset label [cls, x, y, w, h] is converted into 6-dimensional data [cls, x, y, w, h, θ] with angle information, where cls is the labeled category, w refers to the long side of the box, h refers to the other side of the box, and the angle θ refers to the angle formed by the first long side encountered by the horizontal x-axis through counterclockwise rotation, and the range is [0, 180).
5. The method according to claim 2, It is characterized in that The adding of the target angle loss function to generate the improved YOLO model includes: Define a first rotation frame and a second rotation frame; Take a side line segment from the first rotating frame, calculate whether there is an intersection with the second rotating frame, if not and the side line segment is inside the second rotating frame, keep the side line segment, if there is an intersection and the other end point is inside the second rotating frame, keep the line segment of the intersection with the second rotating frame, if there are two intersections, keep the line segment composed of the two intersections, and repeat traversing the line segments of the first rotating frame; Repeat the operation from the second rotation frame to determine the corner point of the first rotation frame; Sort the retained edges to obtain the intersection area polygon; Use the areas of polygons to find intersection and union.
6. The method according to claim 2, It is characterized in that The method further comprises: The localization loss is replaced by the error between the predicted rotated box and the calibrated rotated box.
7. The method according to claim 3, It is characterized in that The convolution layer of the network structure corresponding to the regression task prediction is a 1*1 convolution layer, and the convolution layer of the network structure corresponding to the classification task prediction is a 3*3 convolution layer.
8. A target detection device based on a rotating frame, It is characterized in that include: An acquisition module is used to acquire an image to be detected including a target pattern; The recognition module inputs the image to be detected into the improved YOLO model, and the improved YOLO model outputs a target recognition result; wherein the network detection head of the improved YOLO model includes predicted angle information, and the target recognition result is target box attribute information including the target angle.
9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.