Neural network model training method and device and image target detection method
By designing the first loss function, the prediction bounding box allows the prediction bounding box to approximate the real bounding box in an external form, solving the problem that the existing IoU loss function cannot fully include the target boundary, and improving the accuracy of image object detection of neural network models, especially in applications such as face key point detection and optical character recognition.
Patent Information
- Application Number
- CN202510339637.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
When training neural network models, the existing IoU loss function fails to fully consider whether the prediction bounding box can fully contain the real boundary of the target, resulting in a decrease in the accuracy of image object detection, especially in downstream applications such as face key point detection and optical character recognition.
A new first loss function is designed to calculate the probability that the real bounding box falls outside the predicted bounding box, allowing the predicted bounding box to approximate the real bounding box in an external form, reducing the loss of the target detection pixel point and improving detection accuracy.
By optimizing the position of the prediction bounding box, the target contour bounding can be included more accurately, improving the detection accuracy of the neural network model, and meeting the requirements of downstream applications such as face key point detection and optical character recognition.
Smart Images

Figure CN120279375A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a method for training a neural network model, an apparatus, and an image target detection method. Background Art
[0002] The image target detection task is to detect a specific target as accurately as possible from an image or video, and obtain the classification (or type) of the target and the position of the target object in the video or image. The classification of the target covers various objects, such as various animals, vehicles, airplanes, ships, human bodies, human faces, etc. The position of the target is described in the industry by a bounding box (i.e., bounding box) (as Figure 1 shown), and the bounding box is a region of a specific shape of the object in the video or image.
[0003] Image target detection based on a neural network model is widely applied to various detection and recognition tasks such as human face key point detection, face recognition, and optical character recognition. For example, in the human face key point detection task using a neural network model, first, it is necessary to identify a human face in an original video or image containing multiple objects, obtain the predicted bounding box of the human face, and then extract the local image containing the human face according to the position information given by the predicted bounding box, so as to further perform human face key point detection (as Figure 2 shown). The performance of the neural network model affects the detection accuracy of the image target detection task, and during the training process of the neural network model, the loss function used to optimize the parameters of the neural network model affects the performance of the neural network model.
[0004] Currently, the Intersection over Union (IoU) loss function is often used to optimize the parameters of a neural network model for image target detection. IoU is defined as the ratio of the intersection area of the predicted bounding box and the ground truth bounding box to their union area. The calculation formula corresponding to the IoU loss function IoU Loss is IoU Loss = 1 - IoU, with the goal of training the neural network to make the predicted bounding box overlap with the ground truth bounding box as much as possible. However, the IoU loss function has defects, which affect the performance of the trained neural network model and reduce the detection accuracy of the image target detection task. Summary of the Invention
[0005] This application provides a method for training a neural network model, an apparatus, and an image target detection method, which can reduce the loss of target detection pixel points during the process of using the neural network model for image target detection and improve the detection accuracy of target detection.
[0006] In a first aspect, a method for training a neural network model is provided, which is applied to image target detection and includes:
[0007] Obtain a labeled training data set, where the training data set includes sample images and annotation data, and the annotation data is used to label the position information of the true bounding box corresponding to the detection target in the sample image;
[0008] Input the sample image into the neural network model to obtain the position information of the predicted bounding box corresponding to the detection target output by the neural network model;
[0009] Determine a first loss function according to the position information of the true bounding box and the position information of the predicted bounding box, and the first loss function is used to train the neural network model with the goal of reducing the probability that the true bounding box falls outside the predicted bounding box;
[0010] Train the neural network model based on the first loss function.
[0011] In a feasible design, determine the first loss function according to the position information of the true bounding box and the position information of the predicted bounding box:
[0012] Determine the first area of the true bounding box according to the position information of the true bounding box;
[0013] Determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the position information of the true bounding box and the position information of the predicted bounding box;
[0014] Determine the difference between the first area and the second area as the third area where the true bounding box falls outside the predicted bounding box;
[0015] Determine the first loss function as the ratio of the third area to the first area.
[0016] In a feasible design, when both the true bounding box and the predicted bounding box are rectangles, determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the position information of the true bounding box and the position information of the predicted bounding box, including:
[0017] Determine the first corner coordinates and the second corner coordinates of the true bounding box according to the position information of the true bounding box, where the first corner coordinates are the upper left vertex coordinates of the true bounding box, and the second corner coordinates are the lower right vertex coordinates of the true bounding box;
[0018] Determine the third corner coordinates and the fourth corner coordinates of the predicted bounding box according to the position information of the predicted bounding box, where the third corner coordinates are the upper left vertex coordinates of the predicted bounding box, and the fourth corner coordinates are the lower right vertex coordinates of the predicted bounding box;
[0019] Determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the first corner coordinates, the second corner coordinates, the third corner coordinates, and the fourth corner coordinates.
[0020] In a feasible design, when both the true bounding box and the predicted bounding box are circular, the position information includes the center point coordinates and radius of the bounding box. Determining the second area of the overlapping part between the true bounding box and the predicted bounding box according to the position information of the true bounding box and the position information of the predicted bounding box includes:
[0021] Determine the center point coordinates of the true bounding box as the first center point coordinates, and determine the radius of the true bounding box as the first radius;
[0022] Determine the center point coordinates of the predicted bounding box as the second center point coordinates, and determine the radius of the predicted bounding box as the second radius;
[0023] Determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the first center point coordinates, the first radius, the second center point coordinates, and the second radius.
[0024] In a feasible design, determining the second area of the overlapping part between the true bounding box and the predicted bounding box according to the first corner coordinates, the second corner coordinates, the third corner coordinates, and the fourth corner coordinates, using the following formula:
[0025] B is =(min(x pred2 ,x gt2 )-max(x pred1 ,x gt1 ))×(min(y pred2 ,y gt2 )-max(y pred1 ,y gt1 ));
[0026] Wherein, B is represents the second area, x gt1 represents the abscissa of the first corner coordinate, y gt1 represents the ordinate of the first corner coordinate, x gt2 represents the abscissa of the second corner coordinate, y gt2 represents the ordinate of the second corner coordinate, x pred1 represents the abscissa of the third corner coordinate, y pred1 represents the ordinate of the third corner coordinate, x pred2 represents the abscissa of the fourth corner coordinate, y pred2 represents the ordinate of the fourth corner coordinate.
[0027] In a feasible design, training the neural network model based on the first loss function includes:
[0028] Determine a combined loss function according to a first loss function and a second loss function, where the second loss function is used to train a neural network model with the goal of increasing the degree of overlap between the predicted bounding box and the ground truth bounding box;
[0029] Train the neural network model using the combined loss function.
[0030] In a feasible design, determining the combined loss function according to the first loss function and the second loss function includes:
[0031] Weight the first loss function using a first parameter;
[0032] Weight the second loss function using a second parameter;
[0033] Determine the sum of the weighted first loss function and the weighted second loss function as the combined loss function.
[0034] In a feasible design, after training the neural network model using the combined loss function, the method further includes:
[0035] Adjust the value of the first parameter and / or the value of the second parameter to obtain a new combined loss function;
[0036] Train the neural network model again using the new combined loss function so that the performance of the neural network model meets the requirements of downstream applications.
[0037] In a second aspect, a training apparatus for a neural network model is provided, which is applied to image object detection and includes:
[0038] A training dataset acquisition module, configured to acquire a labeled training dataset, where the training dataset includes sample images and annotation data, and the annotation data is used to annotate the position information of the ground truth bounding box corresponding to the detection target in the sample images;
[0039] A loss function determination module, configured to input a sample image into the neural network model to obtain the position information of the predicted bounding box corresponding to the detection target output by the neural network model;
[0040] The loss function determination module is further configured to determine a first loss function according to the position information of the ground truth bounding box and the position information of the predicted bounding box, where the first loss function is used to train the neural network model with the goal of reducing the probability that the ground truth bounding box falls outside the predicted bounding box;
[0041] A model training module, configured to train the neural network model based on the first loss function.
[0042] In a third aspect, an image object detection method is provided, including:
[0043] Obtain the image to be detected;
[0044] Input the image to be detected into the neural network model trained by the method described in the first aspect, and obtain the detection result of the target object output by the neural network model.
[0045] Currently, the target localization loss function iteratively optimizes the model with the goal of reducing the area of the predicted bounding box in the decreasing direction, suppressing the increase of the predicted bounding box area and the approximation of the predicted bounding box to the true bounding box in an external circumscribing form, resulting in the problem of decreased accuracy due to the loss of target detection pixel points during the image target detection process using the neural network model. Based on this, the present application creatively proposes the design concept of the first loss function, that is, training the neural network model with the goal of reducing the probability that the true bounding box falls outside the predicted bounding box. This first loss function allows the predicted bounding box to approximate the true bounding box in an external circumscribing form, enabling the predicted bounding box to more accurately contain the target contour boundary, improving the accuracy of the predicted bounding box position, thereby improving the detection accuracy of the neural network model for image targets, and better meeting the requirements of downstream applications such as face key point detection, face recognition, and optical character recognition. Brief Description of the Drawings
[0046] To more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a schematic diagram of an image target detection provided by an exemplary embodiment of the present application;
[0048] Figure 2 It is a schematic diagram of an example of face key points provided by an exemplary embodiment of the present application;
[0049] Figure 3 It is a schematic diagram of the positional relationship between a predicted bounding box and a true bounding box provided by an exemplary embodiment of the present application;
[0050] Figure 4 It is a schematic diagram of a predicted bounding box and a true bounding box provided by an exemplary embodiment of the present application;
[0051] Figure 5 It is a schematic diagram of a training method of a neural network model provided by an exemplary embodiment of the present application;
[0052] Figure 6 It is a schematic diagram of the loss visualization of the first loss function provided by an exemplary embodiment of the present application;
[0053] Figure 7It is a schematic structural diagram of a training device for a neural network model provided by an exemplary embodiment of the present application. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0055] In the image target detection task based on a neural network model, the accuracy of predicting the position of the bounding box affects the accuracy of the target detection result. For example, in the face key point detection task, first, a neural network model is used to identify the face in an image with complex content, and the predicted bounding box of the face is obtained, which can remove redundant and irrelevant information in the original image and improve the accuracy of subsequent key point detection. Then, a local image with the face as the main content is extracted according to the position information given by the predicted bounding box. Finally, face key points are detected and extracted on the local image. Since the key points of the face are distributed in parts such as the contour, eyebrows, eyes, nose, and mouth of the face, inaccurate prediction of the position of the bounding box will cause some key points of the face to be missed on the local image, resulting in a decrease in the accuracy of face key point detection. Therefore, the accuracy of predicting the position of the bounding box affects the accuracy of the target detection result. Another example is that in the optical character recognition task, usually, it is first detected whether there is a text area in a large-sized image. If there is a text area, the predicted bounding box containing the text is extracted by the cropping method, and then the image content within the cropped predicted bounding box is recognized. The accuracy of predicting the position of the bounding box also affects the accuracy of text content recognition in the optical character recognition task. It can be seen that determining the predicted bounding box is an important part of the image target detection task.
[0056] The design of the loss function is also crucial in the design of the neural network model. The loss function sets the optimization goal of the model, reflects the error between the current inference result of the model and the correct value, and calculates the best path for updating and optimizing the neural network model parameters by solving the gradient of the loss function, so as to update the network parameters through backpropagation and continuously optimize the model in iteration. The loss function of the target detection algorithm based on the neural network model usually includes a classification loss function and a target localization loss function.
[0057] After in-depth research on the impact of the position of the predicted bounding box on the image target detection task, the present application puts forward the following viewpoints on the positional relationship between the predicted bounding box and the ground truth bounding box. Herein, the present application does not limit the shape of the bounding box, which can be a polygon such as a rectangle, or a circle, etc.:
[0058] For downstream practical applications such as face key point detection, face recognition, optical character recognition, etc., in order to improve the accuracy of determining the position of the predicted bounding box and thus improve the accuracy of object detection, it is acceptable to moderately allow the predicted bounding box to approximate the true bounding box in an enclosing form, while the approximation of the predicted bounding box to the true bounding box in an inscribed form should be avoided as much as possible.
[0059] The following takes the rectangular bounding box as an example to further prove and explain the above viewpoints:
[0060] As Figure 3 shown in Region 1 of gt1 , the larger bounding box is the true bounding box, whose upper left corner coordinates are (x gt1 , y gt2 ), and its lower right corner coordinates are (x gt2 ). The smaller bounding box is the predicted bounding box, whose upper left corner coordinates are (x pred1 , y pred1 ), and its lower right corner coordinates are (x pred2 , y pred2 ). It can be seen that the predicted bounding box being smaller than the true bounding box will cause the true boundary of the face contour to fall outside the predicted bounding box and some key points to be missed, which will affect the accuracy of face key point detection and face recognition.
[0061] As Figure 3 shown in Region 2 of gt1 , the smaller bounding box is the true bounding box, whose upper left corner coordinates are (x′ gt1 ), and its lower right corner coordinates are (x′ gt2 , y′ gt2 ). The larger bounding box is the predicted bounding box, whose upper left corner coordinates are (x′ pred1 , y′ pred1 ), and its lower right corner coordinates are (x′ pred2 , y′ pred2 ). It can be seen that the predicted bounding box being slightly larger than the true bounding box can contain the face contour and will not miss the key points on the face contour, which is acceptable for downstream practical applications.
[0062] However, the traditional object localization loss function iteratively optimizes the model with the goal of reducing the area of the predicted bounding box, which suppresses the increase in the area of the predicted bounding box. It does not impose more constraints and precautions on the approximation of the predicted bounding box to the true bounding box in an inscribed form, and treats the approximation in an inscribed form and the approximation in an enclosing form equally, lacking the necessary emphasis and trade-off.
[0063] For example, the IoU loss function, which is widely used in object detection algorithms, is used to evaluate the overlap between the predicted bounding box and the ground truth bounding box. Among them, IoU is defined as the ratio of the intersection area of the predicted bounding box and the ground truth bounding box to their union area, and the calculation formula is shown in the following formula (1):
[0064]
[0065] Among them, S pred represents the predicted bounding box, and S gt represents the ground truth bounding box.
[0066] The IoU loss function is shown in the following formula (2):
[0067]
[0068] Among them, IoU Loss represents the IoU loss function.
[0069] Take Figure 4 as an example to illustrate the IoU loss function. Among them, the upper left coordinate of the predicted bounding box is (x pred1 , y pred1 ), and the lower right coordinate is (x pred2 , y pred2 ); the upper left coordinate of the ground truth bounding box is (x gt1 , y gt1 ), and the lower right coordinate is (x gt2 , y gt2 ).
[0070] The intersection area of the predicted bounding box and the ground truth bounding box is shown in the following formula (3):
[0071] S pred ∩S gt =(min(x pred2 , x gt2 ) - max(x pred1 , x gt1 )) × (min(y pred2 , y gt2 ) - max(y pred1 , y gt1 ), formula (3);
[0072] Among them, min represents taking the minimum value, and max represents taking the maximum value.
[0073] The union area of the predicted bounding box and the ground truth bounding box is shown in the following formula (4):
[0074] S pred ∪S gt =(xgt2 -x gt1 )×(y gt2 -y gt1 )+(x pred2 -x pred1 )×(y pred2 -y pred1 )-S pred ∩S gt , formula (4).
[0075] As can be seen from the above formula, the network optimization goal essentially set by the IoU loss function is to optimize the model to make the area of the predicted bounding box as close as possible to the area of the true bounding box. However, the predicted bounding box does not always guarantee complete overlap with the true bounding box. When the predicted bounding box and the true bounding box cannot overlap 100%, a part of the true bounding box, such as the boundary of the target object, may fall outside the predicted bounding box; or the predicted bounding box contains more regions that do not belong to the target object. Still taking Figure 4 as an example, a part of the true target boundary, such as the face contour, including the face key points on this part of the contour, will fall outside the predicted bounding box, thus affecting the accuracy of face key point detection and face recognition.
[0076] In summary, the traditional IoU-based loss function aims to train a neural network to make the predicted bounding box overlap as much as possible with the true bounding box, without particularly fully considering whether the predicted bounding box can completely contain the true boundary of the target. Therefore, the IoU-based loss function has defects.
[0077] To solve the problem of the decrease in accuracy caused by the target detection pixel point loss during the process of using a neural network model for image target detection in downstream applications, based on the above view of the positional relationship between the predicted bounding box and the true bounding box, this application proposes a training method for a neural network model, which is applied to image target detection, such as Figure 5 shown, the method includes:
[0078] S110, obtaining a labeled training data set.
[0079] Among them, the training data set includes sample images and annotation data, and the annotation data is used to annotate the position information of the true bounding box corresponding to the detection target in the sample images.
[0080] Exemplarily, in the case where the bounding box is a rectangle, the position information of the true bounding box includes the first corner coordinate and the second corner coordinate, the first corner coordinate is the upper left vertex coordinate of the true bounding box, and the second corner coordinate is the lower right vertex coordinate of the true bounding box.
[0081] Exemplarily, in the case where the bounding box is rectangular, the position information of the true bounding box includes the coordinates of the center point of the true bounding box in the image, the length and width of the true bounding box.
[0082] Exemplarily, in the case where the bounding box is circular, the position information of the true bounding box includes the coordinates of the center point (i.e., the center of the circle) and the radius.
[0083] S120. Input a sample image into the neural network model to obtain the position information of the predicted bounding box corresponding to the detection target output by the neural network model.
[0084] Among them, the structure of the neural network model is not limited in this application, as long as it can implement the function of image target detection.
[0085] Exemplarily, in the case where the bounding box is rectangular, the position information of the predicted bounding box includes the coordinates of the center point of the predicted bounding box in the image, the length and width of the predicted bounding box.
[0086] Exemplarily, in the case where the bounding box is rectangular, the position information of the predicted bounding box includes the third corner coordinates and the fourth corner coordinates. The third corner coordinates are the coordinates of the upper left vertex of the predicted bounding box, and the fourth corner coordinates are the coordinates of the lower right vertex of the predicted bounding box.
[0087] Exemplarily, in the case where the bounding box is circular, the position information of the predicted bounding box includes the coordinates of the center point (i.e., the center of the circle) and the radius.
[0088] S130. Determine the first loss function according to the position information of the true bounding box and the position information of the predicted bounding box.
[0089] Among them, the first loss function is used to train the neural network model with the goal of reducing the probability that the true bounding box falls outside the predicted bounding box.
[0090] This application also creatively puts forward the following viewpoints through in-depth research:
[0091] (1) Existing object localization loss functions do not particularly consider the situation where the true boundary of the object falls outside the predicted bounding box;
[0092] (2) For many downstream applications, in the case where the predicted bounding box and the true bounding box cannot completely overlap, it is acceptable that the predicted bounding box contains the true bounding box, but it is unacceptable that the predicted bounding box does not completely contain the true bounding box;
[0093] (3) Whether the true bounding box of the object falls outside the predicted bounding box can be mathematically characterized;
[0094] (4) The degree to which the true bounding box of the target falls within the predicted bounding box can be quantitatively characterized mathematically (which means that the degree to which the true bounding box of the target falls outside the predicted bounding box can be quantitatively characterized mathematically);
[0095] Based on this, the present application is implemented in the following manner. According to the position information of the true bounding box and the position information of the predicted bounding box, a first loss function is determined:
[0096] According to the position information of the true bounding box, the first area of the true bounding box is determined;
[0097] According to the position information of the true bounding box and the position information of the predicted bounding box, the second area of the overlapping part between the true bounding box and the predicted bounding box is determined;
[0098] The difference between the first area and the second area is determined as the third area where the true bounding box falls outside the predicted bounding box;
[0099] The first loss function is determined as the ratio of the third area to the first area.
[0100] Among them, the third area can be understood as the area of the part where the true bounding box does not overlap with the predicted bounding box.
[0101] In the above example, the ratio of the third area where the true bounding box falls outside the predicted bounding box to the first area of the true bounding box is determined as the first loss function, realizing the quantitative mathematical characterization of the degree to which the true bounding box of the target falls outside the predicted bounding box, that is, regarding the part where the true bounding box falls outside the predicted bounding box as a loss, such as Figure 6 the area surrounded by the bold and black lines (therefore, the present application also refers to the first loss function as the out-of-box loss function). In model training, the present application calculates the gradient of the first loss function and updates parameters such as the weights of the neural network model through backpropagation, making the loss function gradually decrease, that is, making the degree to which the true bounding box falls outside the predicted bounding box gradually decrease, thereby minimizing the probability that the true bounding box of the target falls outside the predicted bounding box to the greatest extent.
[0102] It can be seen that the first loss function in the above example can allow the predicted bounding box to moderately approximate the true bounding box in an external circumscribed form when the predicted bounding box cannot completely overlap the true bounding box, while focusing more on suppressing the approximation of the predicted bounding box to the true bounding box in an internal inscribed form. Therefore, the above example fully considers whether the predicted bounding box can completely contain the true boundary of the target, achieving the improvement of the accuracy of downstream practical applications such as face key point detection, face recognition, and optical character recognition.
[0103] In a feasible design, when both the true bounding box and the predicted bounding box are rectangles, the determination of the first loss function is achieved through the following method:
[0104] Based on the position information of the true bounding box, determine the first corner coordinates and the second corner coordinates of the true bounding box. The first corner coordinates are the upper left vertex coordinates of the true bounding box, and the second corner coordinates are the lower right vertex coordinates of the true bounding box;
[0105] Based on the position information of the predicted bounding box, determine the third corner coordinates and the fourth corner coordinates of the predicted bounding box. The third corner coordinates are the upper left vertex coordinates of the predicted bounding box, and the fourth corner coordinates are the lower right vertex coordinates of the predicted bounding box;
[0106] Based on the first corner coordinates and the second corner coordinates, determine the first area of the true bounding box;
[0107] Based on the first corner coordinates, the second corner coordinates, the third corner coordinates, and the fourth corner coordinates, determine the second area of the overlapping part between the true bounding box and the predicted bounding box;
[0108] Determine the difference between the first area and the second area as the third area where the true bounding box falls outside the predicted bounding box;
[0109] Determine the first loss function as the ratio of the third area to the first area.
[0110] Exemplarily, as Figure 4 shown, the first corner coordinates are the upper left coordinates of the true bounding box, and the second corner coordinates are the lower right coordinates of the true bounding box. When the position information of the true bounding box includes the coordinates of the center point of the true bounding box in the image, the length and width of the true bounding box, calculate the first corner coordinates and the second corner coordinates of the true bounding box through the following formula:
[0111] The first corner coordinates of the true bounding box are calculated through the following formula (5) and formula (6):
[0112] x gt1 = x c1 - 0.5 × width1, formula (5);
[0113] y gt1 = y c1 + 0.5 × length1, formula (6);
[0114] Among them, x gt1 represents the abscissa of the first corner coordinates, y gt1 represents the ordinate of the first corner coordinates, x c1 represents the abscissa of the center point of the true bounding box, y c1 represents the ordinate of the center point of the true bounding box, width1 represents the width of the true bounding box, and length1 represents the length of the true bounding box.
[0115] The second corner coordinates of the true bounding box are calculated by the following formulas (7) and (8):
[0116] x gt2 = x c1 + 0.5 × width1, formula (7);
[0117] y gt2 = y c1 - 0.5 × length1, formula (8);
[0118] where, x gt2 represents the abscissa of the second corner coordinate, and y gt2 represents the ordinate of the second corner coordinate.
[0119] The above example uses formulas (5) - (8) to convert the position information of the true bounding box into the coordinates required for the subsequent steps.
[0120] Exemplarily, as Figure 4 shown, the third corner coordinate is the upper left corner coordinate of the predicted bounding box, and the fourth corner coordinate is the lower right corner coordinate of the true bounding box. In the case where the position information includes the coordinates of the center point of the predicted bounding box in the image, the length and width of the predicted bounding box, the third corner coordinate and the fourth corner coordinate of the predicted bounding box are calculated according to the position information of the predicted bounding box by the following formulas:
[0121] The third corner coordinate of the predicted bounding box is calculated by the following formulas (9) and (10):
[0122] x pred1 = x c2 - 0.5 × width2, formula (9);
[0123] y pred1 = y c2 + 0.5 × length2, formula (10);
[0124] where, x pred1 represents the abscissa of the third corner coordinate, y pred1 represents the ordinate of the third corner coordinate, x c2 represents the abscissa of the center point of the predicted bounding box, y c2 represents the ordinate of the center point of the predicted bounding box, width2 represents the width of the predicted bounding box, and length2 represents the length of the predicted bounding box.
[0125] The fourth corner coordinate of the predicted bounding box is calculated by the following formulas (11) and (12):
[0126] x pred2 = x c2 + 0.5 × width2, formula (11);
[0127] y pred2 = y c2 -0.5 × length2, Equation (12);
[0128] where x pred2 represents the abscissa of the fourth corner coordinate, and y pred2 represents the ordinate of the fourth corner coordinate.
[0129] The above example uses Equations (9) - (12) to convert the position information of the predicted bounding box into the coordinates required for the subsequent steps.
[0130] Exemplarily, when the bounding box is rectangular and the position information of the ground truth bounding box includes the coordinates of the center point of the ground truth bounding box in the image, the length and width of the ground truth bounding box, the product of the length and width of the ground truth bounding box is determined as the first area of the ground truth bounding box.
[0131] Exemplarily, when the bounding box is rectangular and the position information of the ground truth bounding box includes the first corner coordinate and the second corner coordinate, the first area of the ground truth bounding box is determined according to the first corner coordinate and the second corner coordinate, and is implemented using Equation (13):
[0132] B gt = (x gt2 - x gt1 ) × (y gt2 - y gt1 ), Equation (13);
[0133] where B gt represents the first area, x gt1 represents the abscissa of the first corner coordinate, y gt1 represents the ordinate of the first corner coordinate, x gt2 represents the abscissa of the second corner coordinate, and y gt2 represents the ordinate of the second corner coordinate.
[0134] In a feasible design, when the bounding box is rectangular and the position information of the ground truth bounding box and the predicted bounding box includes corner coordinates, the second area of the overlapping part between the ground truth bounding box and the predicted bounding box is determined according to the first corner coordinate, the second corner coordinate, the third corner coordinate, and the fourth corner coordinate, and is implemented using Equation (14):
[0135] B is = B pred ∩ B gt = (min(x pred2 , x gt2 ) - max(x pred1 , x gt1 )) × (min(y pred2, y gt2 ) - max(y pred1 , y gt1 ), formula (14);
[0136] Among them, B is represents the second area, and B pred represents the area of the predicted bounding box (for the calculation method, see B gt ), x pred1 represents the abscissa of the third corner coordinate, and y pred1 represents the ordinate of the third corner coordinate, x pred2 represents the abscissa of the fourth corner coordinate, and y pred2 represents the ordinate of the fourth corner coordinate.
[0137] Exemplarily, it is determined that the first loss function is the ratio of the third area to the first area, and it is implemented using formula (15):
[0138]
[0139] Among them, Loss represents the first loss function.
[0140] The above example can calculate the loss outside the box through formulas (13) - (15). The calculation method is simple and has low requirements for hardware computing power. Moreover, compared with the calculation formula of the IoU loss function, the calculation complexity of the first loss function is lower and does not increase the algorithm complexity.
[0141] In the case where both the true bounding box and the predicted bounding box are rectangles, the above example can efficiently and accurately determine the first area and the second area according to the corner coordinates of the two bounding boxes, so as to accurately determine the first loss function. It can be seen that when the predicted bounding box cannot completely overlap the true bounding box, the first loss function of the above example allows the predicted bounding box to moderately approximate the true bounding box in the form of an outer rectangle, while more focusing on suppressing the approximation of the predicted bounding box to the true bounding box in the form of an inner rectangle. Therefore, the above example fully considers whether the predicted rectangular bounding box can completely contain the true boundary of the target, and realizes the improvement of the accuracy of downstream practical applications such as face keypoint detection, face recognition, and optical character recognition.
[0142] In a feasible design, in the case where both the true bounding box and the predicted bounding box are circles, the position information includes the center point coordinates and radius of the bounding box, and the first loss function is determined in the following way:
[0143] Determine the center point coordinates of the true bounding box as the first center point coordinates, and determine the radius of the true bounding box as the first radius;
[0144] Determine the center point coordinates of the predicted bounding box as the second center point coordinates, and determine the radius of the predicted bounding box as the second radius;
[0145] According to the first center point coordinates and the first radius, determine the first area of the true bounding box;
[0146] According to the first center point coordinates, the first radius, the second center point coordinates and the second radius, determine the second area of the overlapping part between the true bounding box and the predicted bounding box;
[0147] Determine the difference between the first area and the second area as the third area where the true bounding box falls outside the predicted bounding box;
[0148] Determine the first loss function as the ratio of the third area to the first area.
[0149] In the case where the bounding box is circular, there are currently various formulas for implementing the area of a circular bounding box and the overlapping area of two circular bounding boxes, and this application does not limit this.
[0150] Taking the distance between the first center point coordinates and the second center point coordinates as d, the first radius as r1, and the second radius as r2 as an example, the second area B of the overlapping part is calculated is for exemplary illustration.
[0151] Exemplarily, if d > r1 + r2, the two circles do not intersect and the overlapping area is 0. If d < |r1 - r2|, that is, one circle is completely contained in the other circle, the overlapping area is the area of the smaller circle.
[0152] If |r1 - r2| ≤ d ≤ r1 + r2, that is, the two circles overlap, this application provides an implementation formula for reference, as shown in the following formula (16):
[0153]
[0154] In the case where both the true bounding box and the predicted bounding box are circular, the above example can efficiently and accurately determine the first area and the second area according to the center point coordinates and radii of the two bounding boxes, so as to accurately determine the first loss function. It can be seen that the first loss function in the above example can allow the predicted bounding box to approximate the true bounding box in the form of a circumscribed circle moderately when the predicted bounding box cannot completely overlap the true bounding box, and more focuses on suppressing the approximation of the predicted bounding box to the true bounding box in the form of an inscribed circle. Therefore, the above example fully considers whether the circular predicted bounding box can completely contain the true boundary of the target, and realizes the improvement of the accuracy of downstream practical applications such as face key point detection, face recognition, and optical character recognition.
[0155] S140, train the neural network model based on the first loss function.
[0156] Specifically, by taking the gradient of the first loss function and through backpropagation, parameters such as the weights of the neural network model are updated to train the neural network model.
[0157] In a feasible design, training the neural network model based on the first loss function is achieved through the following means:
[0158] Determine a combined loss function according to the first loss function and the second loss function, where the second loss function is used to train the neural network model with the goal of increasing the degree of overlap between the predicted bounding box and the ground truth bounding box;
[0159] Use the combined loss function to train the neural network model.
[0160] Exemplarily, the second loss function is the Complete IoU (CIoU) loss function.
[0161] The above example trains the neural network model by jointly using the first loss function and the second loss function, enabling the training of the neural network model to gradually approach the area of the predicted bounding box to that of the ground truth bounding box while reducing the probability that the ground truth bounding box falls outside the predicted bounding box, comprehensively improving the performance and accuracy of the neural network model.
[0162] In a feasible design, determining the combined loss function according to the first loss function and the second loss function includes:
[0163] Weight the first loss function using the first parameter;
[0164] Weight the second loss function using the second parameter;
[0165] Determine the sum of the weighted first loss function and the weighted second loss function as the combined loss function.
[0166] For example, the combined loss function is determined by the following formula (17):
[0167] Loss loc = αCIoU + βLoss, formula (17);
[0168] where Loss loc represents the combined loss function, CIoU is the second loss function, β represents the first parameter, and α represents the second parameter.
[0169] The above example can make the training objective of the neural network model have different focuses by weighting the first loss function and the second loss function respectively to meet different requirements.
[0170] In a feasible design, after training a neural network model using a combined loss function, the method further includes:
[0171] Adjusting the value of the first parameter and / or the value of the second parameter to obtain a new combined loss function;
[0172] Using the new combined loss function to train the neural network model again so that the performance of the neural network model meets the requirements of downstream applications.
[0173] The above example iterates by adjusting the first parameter and the second parameter in the combined loss function to obtain the best effect. For example, when the first parameter tends to 0 and the second parameter tends to 1, the performance of the loss function is close to the currently widely used CioU loss function; while reducing the setting of the second parameter and increasing the setting of the first parameter can relax the suppression of the increase in the predicted bounding box area and avoid the loss of the target boundary as much as possible. Therefore, by adjusting the first parameter and the second parameter, the best performance balance point can be obtained for specific downstream applications.
[0174] The current object localization loss function continuously iteratively optimizes the model with the goal of the predicted bounding box area decreasing, suppressing the increase in the predicted bounding box area and the predicted bounding box approaching the true bounding box in an external circumscribed form, resulting in a problem of decreased accuracy due to the loss of object detection pixels during the process of using the neural network model for image object detection. Based on this, the present application creatively proposes the design concept of the first loss function, that is, training the neural network model with the goal of reducing the probability that the true bounding box falls outside the predicted bounding box. This first loss function allows the predicted bounding box to approach the true bounding box in an external circumscribed form, enabling the predicted bounding box to more accurately contain the target contour boundary, improving the accuracy of the predicted bounding box position, thereby improving the detection accuracy of the neural network model for image objects, and better meeting the requirements of downstream applications such as face key point detection, face recognition, and optical character recognition.
[0175] As Figure 7 shown, the present application also provides a training device for a neural network model, applied to image object detection, including:
[0176] A training data set acquisition module for acquiring a labeled training data set, where the training data set includes sample images and annotation data, and the annotation data is used to annotate the position information of the true bounding box corresponding to the detection target in the sample images;
[0177] A loss function determination module for inputting the sample images into the neural network model to obtain the position information of the predicted bounding box corresponding to the detection target output by the neural network model;
[0178] The loss function determination module is further configured to determine a first loss function according to the position information of the true bounding box and the position information of the predicted bounding box, where the first loss function is used to train the neural network model with the goal of reducing the probability that the true bounding box falls outside the predicted bounding box;
[0179] The model training module is configured to train the neural network model based on the first loss function.
[0180] In a feasible design, the loss function determination module is implemented by the following method: determining the first loss function according to the position information of the true bounding box and the position information of the predicted bounding box:
[0181] Determine the first area of the true bounding box according to the position information of the true bounding box;
[0182] Determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the position information of the true bounding box and the position information of the predicted bounding box;
[0183] Determine the difference between the first area and the second area as the third area where the true bounding box falls outside the predicted bounding box;
[0184] Determine the first loss function as the ratio of the third area to the first area.
[0185] In a feasible design, when both the true bounding box and the predicted bounding box are rectangles, the loss function determination module is implemented by the following method: determining the second area of the overlapping part between the true bounding box and the predicted bounding box according to the position information of the true bounding box and the position information of the predicted bounding box:
[0186] Determine the first corner coordinate and the second corner coordinate of the true bounding box according to the position information of the true bounding box, where the first corner coordinate is the upper left vertex coordinate of the true bounding box, and the second corner coordinate is the lower right vertex coordinate of the true bounding box;
[0187] Determine the third corner coordinate and the fourth corner coordinate of the predicted bounding box according to the position information of the predicted bounding box, where the third corner coordinate is the upper left vertex coordinate of the predicted bounding box, and the fourth corner coordinate is the lower right vertex coordinate of the predicted bounding box;
[0188] Determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the first corner coordinate, the second corner coordinate, the third corner coordinate, and the fourth corner coordinate.
[0189] In a feasible design, when both the true bounding box and the predicted bounding box are circles, the position information includes the center point coordinate and the radius of the bounding box. The loss function determination module is implemented by the following method: determining the second area of the overlapping part between the true bounding box and the predicted bounding box according to the position information of the true bounding box and the position information of the predicted bounding box:
[0190] Determine the center point coordinates of the true bounding box as the first center point coordinates, and determine the radius of the true bounding box as the first radius;
[0191] Determine the center point coordinates of the predicted bounding box as the second center point coordinates, and determine the radius of the predicted bounding box as the second radius;
[0192] According to the first center point coordinates, the first radius, the second center point coordinates and the second radius, determine the second area of the overlapping part between the true bounding box and the predicted bounding box.
[0193] In a feasible design, the loss function determination module is implemented by the following formula. According to the first corner coordinates, the second corner coordinates, the third corner coordinates and the fourth corner coordinates, determine the second area of the overlapping part between the true bounding box and the predicted bounding box:
[0194] B is =(min(x pred2 ,x gt2 )-max(x pred1 ,x gt1 ))×(min(y pred2 ,y gt2 )-max(y pred1 ,y gt1 ));
[0195] Wherein, B is represents the second area, x gt1 represents the abscissa of the first corner coordinates, y gt1 represents the ordinate of the first corner coordinates, x gt2 represents the abscissa of the second corner coordinates, y gt2 represents the ordinate of the second corner coordinates, x pred1 represents the abscissa of the third corner coordinates, y pred1 represents the ordinate of the third corner coordinates, x pred2 represents the abscissa of the fourth corner coordinates, y pred2 represents the ordinate of the fourth corner coordinates.
[0196] In a feasible design, the model training module is implemented in the following manner. Based on the first loss function, train the neural network model:
[0197] Determine the joint loss function according to the first loss function and the second loss function. The second loss function is used to train the neural network model with the goal of increasing the degree of overlap between the predicted bounding box and the true bounding box;
[0198] Use the joint loss function to train the neural network model.
[0199] In a feasible design, the model training module is implemented as follows. The combined loss function is determined based on the first loss function and the second loss function:
[0200] The first loss function is weighted using a first parameter;
[0201] The second loss function is weighted using a second parameter;
[0202] The sum of the weighted first loss function and the weighted second loss function is determined as the combined loss function.
[0203] In a feasible design, after training the neural network model using the combined loss function, the model training module is further configured to adjust the value of the first parameter and / or the value of the second parameter to obtain a new combined loss function;
[0204] The neural network model is trained again using the new combined loss function so that the performance of the neural network model meets the requirements of the downstream application.
[0205] For other embodiments and effects of the above device, refer to the description in the embodiments of the method for training a neural network model, which will not be elaborated here.
[0206] This application also provides an image target detection method, including:
[0207] Obtain the image to be detected;
[0208] Input the image to be detected into the neural network model trained by the method such as S110 - S140 to obtain the detection result of the target object output by the neural network model.
[0209] The basic principles of this application are described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in this application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of this application. Additionally, the above - disclosed specific details are only for illustrative and easy - to - understand purposes, rather than limitations. These details do not limit this application to necessarily implement using the above - specific details.
[0210] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least some of the sub-steps or stages of other steps or other steps.
[0211] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0212] It should also be noted that in the devices, equipment, and methods of this application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this application.
[0213] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0214] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A training method for a neural network model, applied to image target detection, characterized in that, Including: Obtain a labeled training data set, where the training data set includes sample images and annotation data, and the annotation data is used to annotate the position information of the true bounding box corresponding to the detection target in the sample image; Input the sample image into the neural network model to obtain the position information of the predicted bounding box corresponding to the detection target output by the neural network model; Determine a first loss function according to the position information of the true bounding box and the position information of the predicted bounding box, where the first loss function is used to train the neural network model with the goal of reducing the probability that the true bounding box falls outside the predicted bounding box; Train the neural network model based on the first loss function.
2. The method according to claim 1, wherein, The step of determining the first loss function according to the position information of the true bounding box and the position information of the predicted bounding box includes: Determine the first area of the true bounding box according to the position information of the true bounding box; Determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the position information of the true bounding box and the position information of the predicted bounding box; Determine the difference between the first area and the second area as the third area where the true bounding box falls outside the predicted bounding box; Determine the first loss function as the ratio of the third area to the first area.
3. The method according to claim 2, wherein When both the true bounding box and the predicted bounding box are rectangles, the step of determining the second area of the overlapping part between the true bounding box and the predicted bounding box according to the position information of the true bounding box and the position information of the predicted bounding box includes: Determine the first corner coordinate and the second corner coordinate of the true bounding box according to the position information of the true bounding box, where the first corner coordinate is the upper left vertex coordinate of the true bounding box, and the second corner coordinate is the lower right vertex coordinate of the true bounding box; Determine the third corner coordinate and the fourth corner coordinate of the predicted bounding box according to the position information of the predicted bounding box, where the third corner coordinate is the upper left vertex coordinate of the predicted bounding box, and the fourth corner coordinate is the lower right vertex coordinate of the predicted bounding box; Determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the first corner coordinate, the second corner coordinate, the third corner coordinate, and the fourth corner coordinate.
4. The method according to claim 2, wherein When both the true bounding box and the predicted bounding box are circles, the position information includes the center point coordinate and the radius of the bounding box. The step of determining the second area of the overlapping part between the true bounding box and the predicted bounding box according to the position information of the true bounding box and the position information of the predicted bounding box includes: Determine the center point coordinate of the true bounding box as the first center point coordinate, and determine the radius of the true bounding box as the first radius; Determine the center point coordinate of the predicted bounding box as the second center point coordinate, and determine the radius of the predicted bounding box as the second radius; Determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the first center point coordinate, the first radius, the second center point coordinate, and the second radius.
5. The method according to claim 3, characterized in that, Determine the second area of the overlapping part between the true bounding box and the predicted bounding box according to the first angular coordinate, the second angular coordinate, the third angular coordinate, and the fourth angular coordinate, using the following formula: B is = (min(x pred2 , x gt2 )) - max(x pred1 , x gt1 )) × (min(y pred2 , y gt2 )) - max(y pred1 , y gt1 )); Among them, B is represents the second area, x gt1 represents the abscissa of the first angular coordinate, y gt1 represents the ordinate of the first angular coordinate, x gt2 represents the abscissa of the second angular coordinate, y gt2 represents the ordinate of the second angular coordinate, x pred1 represents the abscissa of the third angular coordinate, y pred1 represents the ordinate of the third angular coordinate, x pred2 represents the abscissa of the fourth angular coordinate, y pred2 represents the ordinate of the fourth angular coordinate.
6. The method according to any one of claims 1-5, characterized in that Training the neural network model based on the first loss function includes: Determine a combined loss function according to the first loss function and the second loss function, where the second loss function is used to train the neural network model with the goal of increasing the degree of overlap between the predicted bounding box and the true bounding box; Use the combined loss function to train the neural network model.
7. The method according to claim 6, wherein Determining the combined loss function according to the first loss function and the second loss function includes: Weight the first loss function using a first parameter; Weight the second loss function using a second parameter; Determine the sum of the weighted first loss function and the weighted second loss function as the combined loss function.
8. The method according to claim 7, wherein After training the neural network model using the combined loss function, the method further includes: Adjust the value of the first parameter and / or the value of the second parameter to obtain a new combined loss function; Use the new combined loss function to train the neural network model again to make the performance of the neural network model meet the requirements of downstream applications.
9. A training device for a neural network model, which is applied to image object detection, is characterized in that, including: A training dataset acquisition module for acquiring a labeled training dataset, where the training dataset includes sample images and annotation data, and the annotation data is used to annotate the position information of the true bounding box corresponding to the detection target in the sample image; A loss function determination module for inputting the sample image into the neural network model to obtain the position information of the predicted bounding box corresponding to the detection target output by the neural network model; The loss function determination module is further configured to determine a first loss function according to the position information of the true bounding box and the position information of the predicted bounding box, where the first loss function is used to train the neural network model with the goal of reducing the probability that the true bounding box falls outside the predicted bounding box; A model training module for training the neural network model based on the first loss function.
10. An image target detection method, characterized in that, including: Obtain the image to be detected; Input the image to be detected into the neural network model trained by the method according to any one of claims 1-8 to obtain the detection result of the target object output by the neural network model.
Citation Information
Patent Citations
Road target detection method and device, electronic equipment and storage medium
CN111062413A
Target detection model training method, and target detection method and device
CN112329873A
Cervical cytology image abnormal region positioning method and device based on fusion attention
CN114897779A
Single-stage target detection regression position loss algorithm based on intersection-parallel ratio
CN116597206A
Target detection method based on angle, diagonal and target foreground information
CN117541891A