Directional rotation target detection method, system and device, storage medium and product

By introducing coordinates of direction points and other information of the rotation frame in the rotation object detection method, the head and tail orientation of the target object can be accurately inferred, solving the problem that the prior art cannot judge the direction of the target object, improving the detection accuracy and promoting the intelligence of industrial automation.

CN120147599APending Publication Date: 2025-06-13SHENZHEN HANS GREEN POWER LIGHTING TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510206811.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing rotation object detection method cannot effectively determine the orientation of at least one of the head and tail of the target object, limiting its application in a scenario where the direction of the target object needs to be determined.

Method used

A directed rotation object detection method is proposed. By acquiring image data and inputting it into the rotation object detection model, features are extracted, fused features and predicting the object detection result, including the center point of the rotation box, the rotation angle, width, height and coordinates of the direction point of the rotation box. The target detection result is then decoded to infer the orientation of at least one of the head and tail.

Benefits of technology

The coordinates of the direction point indicate the position of at least one of the head and tail of the target object, and combined with other information of the rotating box, the orientation of at least one of the head and tail can be accurately inferred, improving the accuracy of target detection, and bringing deeper intelligence to industrial automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147599A_ABST
    Figure CN120147599A_ABST
Patent Text Reader

Abstract

The invention discloses a directed rotation target detection method, system and device, a storage medium and a product, and the method comprises the following steps: obtaining image data which comprises a target object; inputting the image data into a rotating target detection model, and performing feature extraction, feature fusion and prediction on the image data by the rotating target detection model so as to obtain a target detection result; wherein the target detection result comprises the coordinate (x0, y0) of the center point C of the rotating frame, the rotating angle theta of the rotating frame, the width w of the rotating frame, the height h of the rotating frame and the coordinate of the direction point, and the coordinate of the direction point is used for marking the position of at least one of the head and the tail of the target object; and decoding the target detection result so as to deduce the orientation of at least one of the head part and the tail part. According to the directed rotation target detection method, the orientation of at least one of the head and the tail of the target object in the image can be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular, to an oriented rotation target detection method, system, device, storage medium and product. Background Art

[0002] In many applications in today's industrial field, the oriented rotation target detection technology is gradually becoming a key technology. The core advantage of this technology is that it can accurately identify and locate the target object, and can obtain the direction information of the target object, which is crucial for realizing the automation and intelligence of industrial systems.

[0003] Traditional target detection methods mainly rely on horizontal bounding boxes to locate target objects in images. However, when dealing with target objects with uncertain directions, dense distributions, large aspect ratios or complex backgrounds, the accuracy of this method is often not high and the generalization ability is limited. To overcome these limitations, an oriented bounding box (OBB) is introduced. By adding a parameter for the rotation angle, it can more accurately describe the orientation and shape of the target object, and shows significant advantages especially in the field of target detection such as remote sensing images.

[0004] In related technologies, although the rotation target detection method based on Probabilistic IoU Loss (ProbIoULoss) has achieved remarkable results, it still has limitations and cannot determine the orientation of at least one of the head and tail of the target object, which limits its application in scenarios where the orientation of the head and tail of the target object needs to be determined. Summary of the Invention

[0005] Therefore, the present application proposes an oriented rotation target detection method that can detect the orientation of at least one of the head and tail of the target object in an image.

[0006] The present application also proposes an oriented rotation target detection system.

[0007] The present application also proposes an oriented rotation target detection device.

[0008] The present application also proposes a computer-readable storage medium.

[0009] The present application also proposes a computer program product.

[0010] The oriented rotation target detection method according to the first aspect embodiment of the present application includes the following steps:

[0011] Obtain image data, where the image data includes a target object;

[0012] Input the image data into a rotated object detection model, which can perform feature extraction, feature fusion, and prediction on the image data to obtain an object detection result. Among them, the object detection result includes the coordinates (x 0 , y 0 ) of the center point C of the rotated bounding box, the rotation angle θ of the rotated bounding box, the width w of the rotated bounding box, the height h of the rotated bounding box, and the coordinates of the direction point, where the coordinates of the direction point are used to indicate the position of at least one of the head and the tail of the target object;

[0013] Decode the object detection result, and then infer the orientation of at least one of the head and the tail.

[0014] According to the directed rotated object detection method of the embodiments of the present application, it has at least the following beneficial effects: By indicating the position of at least one of the head and the tail of the target object through the coordinates of the direction point, and then combining the coordinates (x 0 , y 0 ) of the center point C of the rotated bounding box, the rotation angle θ of the rotated bounding box, the width w of the rotated bounding box, and the height h of the rotated bounding box and other information, the orientation of at least one of the head and the tail can be inferred, which can not only improve the accuracy of object detection, but also bring deeper intelligence to industrial automation.

[0015] According to some embodiments of the present application, the "decoding the object detection result and then inferring the orientation of at least one of the head and the tail" includes the following steps:

[0016] The direction point includes a head direction point, and obtain the coordinates (x 3 , y 3 ) of the head direction point;

[0017] Assume that the four vertices of the rotated bounding box after rotating by the rotation angle θ are LT′, RT′, RB′, and LB′ respectively, and calculate the coordinates of the four vertices LT′, RT′, RB′, and LB′;

[0018] Calculate the distances from the four vertices LT′, RT′, RB′, and LB′ of the rotated bounding box to the head direction point, and set them as d 10 , d 20 , d 30 , and d 40 respectively;

[0019] Compare the magnitudes of d 10 , d 20 , d 30 , and d 40 , and obtain the two vertices corresponding to the two smallest distances;

[0020] From the two vertices corresponding to the two smallest distances in terms of value, the orientation of the corrected head is obtained.

[0021] According to some embodiments of the present application, the "calculating the coordinates of the four vertices LT′, RT′, RB′, and LB′" includes the steps of:

[0022] Assume that when the rotation box is in a horizontal state, the upper left vertex, upper right vertex, lower right vertex, and lower left vertex are LT, RT, RB, and LB respectively, and obtain the coordinates of LT, RT, RB, and LB relative to the center point C. Among them, the relative coordinate of LT is (-w / 2, -h / 2), the relative coordinate of RT is (w / 2, -h / 2), the relative coordinate of RB is (w / 2, h / 2), and the coordinate of LB is (-w / 2, h / 2);

[0023] Assume that after the rotation box rotates by the rotation angle θ, the upper left vertex, upper right vertex, lower right vertex, and lower left vertex are LT′, RT′, RB′, and LB′ respectively;

[0024] According to the rotation angle θ, the obtained rotation matrix R is

[0025] According to the coordinates of the four vertices LT, RT, RB, and LB relative to the center point C, and the rotation matrix R, calculate the coordinates of LT′, RT′, RB′, and LB′. Among them, the coordinate of LT′ is The coordinate of RT′ is The coordinate of RB′ is The coordinate of LB′ is

[0026] According to some embodiments of the present application, the "decoding the target detection result and then inferring the orientation of at least one of the head and the tail" includes the following steps:

[0027] The direction points include the tail direction point, and obtain the coordinates (x 4 , y 4 ) of the tail direction point;

[0028] Assume that after the rotation box rotates by the rotation angle θ, the four vertices are LT′, RT′, RB′, and LB′ respectively, and calculate the coordinates of the four vertices LT′, RT′, RB′, and LB′;

[0029] Calculate the distances from the four vertices LT′, RT′, RB′, and LB′ of the rotation box to the tail direction point, and set them as d 10 , d 20, d 30 and d 40 ;

[0030] Compare d 10 , d 20 , d 30 and d 40 in terms of magnitude, and obtain the two vertices corresponding to the two distances with the largest values;

[0031] Based on the two vertices corresponding to the two distances with the largest values, determine the orientation of the corrected tail.

[0032] According to some embodiments of the present application, the loss function of the rotation target detection model is Loss = αLoss Box + βLoss Cls + γLoss DFL + δLoss Kpts + εLoss Kpts_obj , where Loss Box represents the rotation box regression loss, Loss Cls represents the class loss, Loss DFL represents the distribution focal loss, Loss Kpts represents the direction point regression loss, Loss Kpts_obj represents the direction point target loss, and α, β, γ, δ, and ε respectively represent the weights of each loss.

[0033] According to some embodiments of the present application, the probability intersection over union loss function is used to calculate the rotation box regression loss, which includes the following steps:

[0034] Model the rotation box as a Gaussian rectangle box;

[0035] Calculate the distance between the two Gaussian rectangle boxes;

[0036] Based on the distance between the two Gaussian rectangle boxes, calculate the probability intersection over union loss ProbIoU.

[0037] According to some embodiments of the present application, the "calculating the distance between the two Gaussian rectangle boxes" includes the following steps:

[0038] Calculate the covariance matrix of the X coordinates and Y coordinates of all points within the rotation box where the range of the rotation angle θ is [0°, 90°];

[0039] To measure the distance between two Gaussian rectangle boxes, the Bhattacharyya distance Bd(p,q) is used to represent the distance between two two-dimensional probability density functions p(x) and q(x), where, (x1 , y 1 ) are the coordinates of the rotation center point of p(x), (x 2 , y 2 ) are the coordinates of the rotation center point of q(x), (a 1 , b 1 ) is the variance in the horizontal direction of the rotation box, (a 2 , b 2 ) is the variance in the vertical direction of the rotation box, c 1 is the covariance in the horizontal direction of the rotation box, c 2 is the covariance in the vertical direction of the rotation box;

[0040] Use the Hellinger distance Hd(p, q) that satisfies the triangle inequality to represent the distance between two two-dimensional probability density functions p(x) and q(x), where,

[0041] According to some embodiments of the present application, calculate the direction point regression loss Loss OKS = 1 - OKS, where, Kpts , where, d is the Euclidean distance between the predicted position and the true position of the direction point, S is the area of the target object; σ is the standard deviation parameter related to the scale of the direction point; δ is an indicator function, δ indicates whether the direction point is visible, when v > 0, δ is 1, indicating that the direction point is visible, when v ≤ 0, δ is 0, indicating that the direction point is invisible.

[0042] According to the second aspect embodiment of the present application, a directed rotation target detection system includes:

[0043] An acquisition module, configured to acquire image data, where the image data includes a target object;

[0044] A detection module, configured to input the image data into a rotation target detection model, and the rotation target detection model can perform feature extraction, feature fusion and prediction on the image data to obtain a target detection result; where the target detection result includes the coordinates (x 0 , y 0 ) of the center point C of the rotation box, the rotation angle θ of the rotation box, the width w of the rotation box, the height h of the rotation box, and the coordinates of the direction point, and the coordinates of the direction point are used to mark the position of at least one of the head and the tail of the target object;

[0045] A decoding module, configured to decode the target detection result, and further infer the orientation of at least one of the head and the tail.

[0046] The directed rotation target detection system according to the embodiments of the present application has at least the following beneficial effects: By indicating the position of at least one of the head and tail of the target object through the coordinates of the direction points, and then combining the information such as the coordinates (x 0 , y 0 ) of the center point C of the rotation box, the rotation angle θ of the rotation box, the width w of the rotation box, and the height h of the rotation box, the orientation of at least one of the head and tail can be inferred, which can not only improve the accuracy of target detection, but also bring deeper intelligence to industrial automation.

[0047] A directed rotation target detection device according to the third aspect embodiment of the present application includes:

[0048] A memory storing a computer program;

[0049] A processor, when the processor executes the computer program, can implement the steps of the directed rotation target detection method as described above.

[0050] The directed rotation target detection device according to the embodiments of the present application has at least the following beneficial effects: By implementing the directed rotation target detection method as described above, the orientation of at least one of the head and tail can be inferred, which can not only improve the accuracy of target detection, but also bring deeper intelligence to industrial automation.

[0051] A computer-readable storage medium according to the fourth aspect embodiment of the present application, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the directed rotation target detection method as described above.

[0052] The computer-readable storage medium according to the embodiments of the present application has at least the following beneficial effects: By implementing the directed rotation target detection method as described above, the orientation of at least one of the head and tail can be inferred, which can not only improve the accuracy of target detection, but also bring deeper intelligence to industrial automation.

[0053] A computer program product according to the fifth aspect embodiment of the present application includes a computer program, and when the computer program is executed by a processor, it implements the steps of the directed rotation target detection method as described above.

[0054] The computer program product according to the embodiments of the present application has at least the following beneficial effects: By implementing the directed rotation target detection method as described above, the orientation of at least one of the head and tail can be inferred, which can not only improve the accuracy of target detection, but also bring deeper intelligence to industrial automation.

[0055] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings

[0056] The following further describes the present application in conjunction with the drawings and embodiments, where:

[0057] Figure 1 is a flowchart of the directed rotation target detection method for the first embodiment of the present application;

[0058] Figure 2 is a flowchart for determining the accurate orientation of the head of a target object through the head direction point;

[0059] Figure 3 is a schematic diagram for correcting the position of the head direction point;

[0060] Figure 4 is a schematic diagram when the rotation box is in a horizontal state;

[0061] Figure 5 is a flowchart for determining the accurate orientation of the head of a target object through the tail direction point;

[0062] Figure 6 is a schematic diagram for modeling the rotation box into a Gaussian rectangular box;

[0063] Figure 7 is a flowchart of the directed rotation target detection method for the second embodiment of the present application;

[0064] Figure 8 is a network structure diagram of a deep learning algorithm;

[0065] Figure 9 is a schematic diagram for marking direction points for a target object;

[0066] Figure 10 is a comparison diagram of whether a target object has direction points;

[0067] Figure 11 is a schematic diagram where one target object has one direction point and another target object has two direction points;

[0068] Figure 12 is the prediction result of the rotation target detection model based on ProbIoU;

[0069] Figure 13 is the effect diagram of the directed rotation target detection method for the second embodiment of the present application;

[0070] Figure 14 is a schematic diagram of the directed rotation target detection system for the embodiment of the present application;

[0071] Figure 15 Schematic diagram of the directed rotation target detection device according to an embodiment of the present application.

[0072] Reference numerals: directed rotation target detection system 100, acquisition module 110, detection module 120, decoding module 130;

[0073] Directed rotation target detection device 200, memory 210, processor 220. Detailed implementation manners

[0074] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as a limitation to the present application.

[0075] In the description of the present application, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.

[0076] In the description of the present application, the meaning of "a number of" is one or more, the meaning of "a plurality of" is two or more, "greater than", "less than", "exceeding", etc. are understood as not including the number itself, and "above", "below", "within", etc. are understood as not including the number itself. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0077] In the description of the present application, unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.

[0078] In the description of the present application, the description referring to terms such as "an embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0079] Reference Figure 1 , according to the directed rotation target detection method of the first aspect embodiment of the present application, it includes the following steps:

[0080] S100. Obtain image data, where the image data contains target objects;

[0081] S200. Input the image data into a rotation target detection model, and the rotation target detection model can perform feature extraction, feature fusion and prediction on the image data to obtain a target detection result; where the target detection result includes the coordinates (x 0 , y 0 ) of the center point C of the rotation box, the rotation angle θ of the rotation box, the width w of the rotation box, the height h of the rotation box, and the coordinates of the direction point, and the coordinates of the direction point are used to indicate the position of at least one of the head and tail of the target object;

[0082] S300. Decode the target detection result, and then infer the orientation of at least one of the head and tail.

[0083] According to the directed rotation target detection method of the embodiment of the present application, it has at least the following beneficial effects: By indicating the position of at least one of the head and tail of the target object through the coordinates of the direction point, and then combining the coordinates (x 0 , y 0 ) of the center point C of the rotation box, the rotation angle θ of the rotation box, the width w of the rotation box, and the height h of the rotation box and other information, the orientation of at least one of the head and tail can be inferred, which can not only improve the accuracy of target detection, but also bring deeper intelligence to industrial automation.

[0084] Reference Figure 2 and Figure 3 , in some embodiments of the present application, "decoding the target detection result and then inferring the orientation of at least one of the head and tail" includes the following steps:

[0085] The direction point includes a head direction point, and obtain the coordinates (x 3 , y 3 ) of the head direction point;

[0086] Assume that the four vertices of the rotation box after rotating the rotation angle θ are LT′, RT′, RB′ and LB′ respectively, and calculate the coordinates of the four vertices LT′, RT′, RB′ and LB′;

[0087] Calculate the distances from the four vertices LT′, RT′, RB′ and LB′ of the rotation box to the head direction point, and set them as d 10 , d 20 , d 30 and d 40 ;

[0088] Compare d 10 , d 20 , d 30 and d 40 to obtain the two vertices corresponding to the two smallest distances;

[0089] Based on the two vertices corresponding to the two smallest distances, determine the orientation of the corrected head.

[0090] Since the head direction point is just a certain point in the head of the target object (such as Figure 3 the point Head0), the head direction point only indicates the approximate position of the head of the target object, and cannot accurately express the position and orientation of the head of the target object. By correcting the position of the head through the two vertices closest to the head direction point, the accurate position of the head of the target object can be obtained.

[0091] Specifically, the distances from the four vertices LT′, RT′, RB′, and LB′ to the head direction point can be the Euclidean distance. Referring to Figure 2 , let the coordinates of LT′ be (LT′ x , LT′ y ), the coordinates of RT′ be (RT′ x , RT′ y ), the coordinates of RB′ be (RB′ x , RB′ y ), and the coordinates of LB′ be (LB′ x , LB′ y ), from which it can be known that

[0092] Specifically, referring to Figure 3 , by obtaining the coordinates of the point Head0′, the accurate position of the head of the target object can be obtained. Let RT′ and RB′ be the two vertices closest to the head direction point Head0, and let the coordinates of RT′ be (RT′ x , RT′ y ), and the coordinates of RB′ be (RB′ x , RB′ y ), then the coordinates of the corrected head direction point Head0′ are

[0093] After obtaining the coordinates of the head direction point Head0′, combined with the coordinates (x 0 , y 0 ) of the center point C of the rotation box, the accurate orientation of the head of the target object can be known.

[0094] Referring to Figure 3 andFigure 4 , in the improvement scheme of the above embodiment, "calculating the coordinates of the four vertices LT′, RT′, RB′, and LB′" includes the steps of:

[0095] Assume that the upper left vertex, upper right vertex, lower right vertex, and lower left vertex of the rotation box when it is in a horizontal state are LT, RT, RB, and LB respectively (refer to Figure 4 ), obtain the coordinates of LT, RT, RB, and LB relative to the center point C, where the relative coordinates of LT are (-w / 2, -h / 2), the relative coordinates of RT are (w / 2, -h / 2), the relative coordinates of RB are (w / 2, h / 2), and the coordinates of LB are (-w / 2, h / 2);

[0096] Assume that the upper left vertex, upper right vertex, lower right vertex, and lower left vertex of the rotation box after rotating by the rotation angle θ are LT′, RT′, RB′, and LB′ respectively (refer to Figure 3 );

[0097] According to the rotation angle θ, the obtained rotation matrix R is

[0098] According to the coordinates of the four vertices LT, RT, RB, and LB relative to the center point C, and the rotation matrix R, calculate the coordinates of LT′, RT′, RB′, and LB′, where the coordinates of LT′ are The coordinates of RT′ are The coordinates of RB′ are The coordinates of LB′ are

[0099] By multiplying the coordinates of the four vertices LT, RT, RB, and LB relative to the center point C by the rotation matrix R, and then adding the coordinates of the center point C, the coordinates of LT′, RT′, RB′, and LB′ can be obtained. The calculation method is accurate and convenient, which is beneficial to obtaining the accurate coordinates of LT′, RT′, RB′, and LB′.

[0100] Refer to Figure 5 , in some embodiments of the present application, "decoding the target detection result and then inferring the orientation of at least one of the head and the tail" includes the following steps:

[0101] The direction points include the tail direction points, and obtain the coordinates of the tail direction points (x 4 , y 4 );

[0102] Assume that the four vertices of the rotation box after rotating by the rotation angle θ are LT′, RT′, RB′, and LB′ respectively, and calculate the coordinates of the four vertices LT′, RT′, RB′, and LB′;

[0103] Calculate the distances from the four vertices LT′, RT′, RB′, and LB′ of the rotated bounding box to the tail direction point, and denote them as d 10 , d 20 , d 30 , and d 40 ;

[0104] Compare the magnitudes of d 10 , d 20 , d 30 , and d 40 , and obtain the two vertices corresponding to the two largest distances;

[0105] Based on the two vertices corresponding to the two largest distances, determine the orientation of the corrected tail.

[0106] Similarly, by knowing the two vertices farthest from the tail direction point through the tail direction point, the accurate orientation of the tail of the target object can be obtained.

[0107] Specifically, the distances from the four vertices LT′, RT′, RB′, and LB′ to the tail direction point can use the Euclidean distance, and the specific calculation method can refer to the calculation method of the distances from the four vertices LT′, RT′, RB′, and LB′ to the head direction point.

[0108] In some embodiments of the present application, the loss function of the rotation target detection model is Loss = αLoss Box + βLoss Cls + γLoss DFL + δLoss Kpts + εLoss Kpts_obj , where Loss Box represents the rotation bounding box regression loss, Loss Cls represents the class loss, Loss DFL represents the distribution focal loss, Loss Kpts represents the direction point regression loss, Loss Kpts_obj represents the direction point target loss, and α, β, γ, δ, ε respectively represent the weights of each loss.

[0109] By designing the loss function, the rotated bounding box can be made closer to the labeled bounding box, thereby obtaining the accurate position and orientation of the target object.

[0110] Specifically, the initial values of α, β, γ, δ, and ε are 7.5, 0.5, 1.5, 12.0, and 1.0 respectively. The values of α, β, γ, δ, and ε can also be other values. For example, the value of α can be 6.5, 7, 8, 8.5, or other values, the value of β can also be 0.3, 0.4, 0.6, or other values, the value of γ can also be 1.3, 1.4, 1.6, or other values, the value of δ can also be 10, 10.5, 11, 11.5, 12.5, or other values, and the value of ε can also be 0.8, 0.9, 1.1, or other values.

[0111] Refer to Figure 6 , it should be noted that the elliptical frame in the figure represents a Gaussian rectangle frame, and the rectangular frame represents a rotated frame. In the improved solution of the above embodiment, a probability intersection over union loss function is used to calculate the rotated frame regression loss, which includes the following steps:

[0112] Model the rotated frame as a Gaussian rectangle frame;

[0113] Calculate the distance between two Gaussian rectangle frames;

[0114] Calculate the probability intersection over union loss ProbIoU based on the distance between two Gaussian rectangle frames.

[0115] The Gaussian rectangle frame can better fit the position of the target object than the rotated frame, and the rotated frame regression loss calculated using the probability intersection over union loss function is more accurate.

[0116] In the improved solution of the above embodiment, "calculating the distance between two Gaussian rectangle frames" includes the following steps:

[0117] Calculate the covariance matrix of the X coordinates and Y coordinates of all points within the rotated frame where the range of the rotation angle θ is [0°, 90°];

[0118] To measure the distance between two Gaussian rectangle frames, the Bhattacharyya distance Bd(p,q) is used to represent the distance between two two-dimensional probability density functions p(x) and q(x), where (x 1 , y 1 ) is the coordinate of the rotation center point of p(x), (x 2 , y 2 ) is the coordinate of the rotation center point of q(x), (a 1 , b 1 ) is the variance in the horizontal direction of the rotated frame, (a 2 , b 2 ) is the variance in the vertical direction of the rotated frame, c 1 is the covariance in the horizontal direction of the rotated frame, c2 is the covariance in the vertical direction of the rotation box;

[0119] The Hellinger distance Hd(p,q) that satisfies the triangle inequality is used to represent the distance between two two-dimensional probability density functions p(x) and q(x), where,

[0120] Through the above steps, the distance between two Gaussian rectangular boxes can be accurately calculated.

[0121] Specifically, "calculating the Probability Intersection over Union Loss ProbIoU according to the distance between two Gaussian rectangular boxes" includes the steps of calculating ProbIoU through the formula ProbIoU = 1 - Hd(p,q).

[0122] In some embodiments of the present application, through the formula Loss OKS = 1 - OKS calculates the orientation point regression loss Loss Kpts , where, d is the Euclidean distance between the predicted position and the true position of the orientation point, S is the area of the target object; σ is the standard deviation parameter related to the scale of the orientation point; δ is an indicator function, δ indicates whether the orientation point is visible, when v > 0, δ is 1, indicating that the orientation point is visible, when v ≤ 0, δ is 0, indicating that the orientation point is invisible.

[0123] Through the above formula, the orientation point regression loss Loss can be calculated. Kpts .

[0124] Specifically, Loss Cls is calculated through the Binary Cross Entropy Loss (BCELoss) function to determine each predicted class and its confidence. Loss DFL is calculated through the Cross Entropy Loss (CE Loss) function to enable the model to better predict the width and height of the target object. Loss Kpts_obj is also calculated using the Binary Cross Entropy Loss function to enable the model to learn which are visible points and which are invisible points.

[0125] The following combines Figures 2 to 13 to illustrate the oriented rotation target detection method of a specific embodiment of the present application.

[0126] The technical solution adopted by the oriented rotation target detection method of this embodiment is as Figure 7 shown and is generally divided into three steps.

[0127] Step 1: Data acquisition and image preprocessing stage.

[0128] The image is captured in real time by a high-resolution industrial camera, then formatted into a unified size, and the image is normalized. This step helps to reduce the processing time of the model and improve the processing consistency.

[0129] Step 2: Design a deep learning algorithm for rotating object detection that includes a positive direction information prediction branch (i.e., the orientation prediction branch of the head or tail of the target object). Inputting the processed picture will obtain the coordinates (x 0 , y 0 ) of the center point C of the rotated bounding box, the rotation angle θ of the rotated bounding box, the width w of the rotated bounding box, the height h of the rotated bounding box, and the coordinates of the direction point. The following mainly introduces the network structure design and loss function design of the deep learning algorithm.

[0130] The design of the deep learning algorithm includes three parts, namely the Backbone, Neck, and Head layers, as Figure 8 shown.

[0131] The Backbone layer is the main network and is used to extract features from the input picture. The main network is mainly composed of CBS modules, C2f modules, and SPPF modules. The CBS module represents Conv, BN, and SILU, and is generally used to extract features from the input image or feature map. The result of convolution is a decrease in resolution and an increase in the number of channels. The Backbone contains a total of 5 CBS modules, that is, 5 times of downsampling. The input image resolution changes from 640*640 to 20*20 after 5 times of downsampling, and the input number of channels changes from 3 to 512.

[0132] The Neck layer is the neck network and is composed of an FPN (Feature Pyramid Network) module and a PANet (Path-Aggregation Network) structure. It is used to perform feature fusion on the feature maps output by the main network, that is, to connect (Concat) the features extracted from three different scales output in the Backbone layer through upsampling. After two upsamplings, the first feature map (80×80) is output to the Head layer, and then the feature map is downsampled through the CBS module to output the remaining two feature maps (40×40 and 20×20) to the Head layer.

[0133] The Head layer is the head network, which mainly uses multi-scale feature maps to predict the extracted features. Feature maps of different scales have different receptive fields. In this embodiment, detections at each scale are used to predict the position information and positive direction information of the target. The position information includes the rotation center and category (Rotation Center), rotation angle (RotationAngel), and height / width of the rotation bounding box. The positive direction information mainly includes the orientation point (OrientationPoint). Generally, only one point is needed to determine the positive direction. However, to ensure that the object is truncated and the positive direction cannot be correctly detected, the Head orientation point and Tail orientation point need to be set.

[0134] The following introduces the design of the loss function.

[0135] The loss function is designed as Loss = αLoss Box + βLoss Cls + γLoss DFL + δLoss Kpts + εLoss Kpts_obj 。As can be seen from the formula, the total loss consists of the rotation box regression loss Loss Box , the category loss Loss Cls and the distribution focal loss (DFL loss) Loss DFL , the orientation point regression loss Loss Kpts and the orientation point target loss Loss Kpts_obj , where α, β, γ, δ, and ε respectively represent the weights of each loss, and the initial values are default set to 7.5, 0.5, 1.5, 12.0, and 1.0. It can be seen that the weights of Loss Box and Loss Kpts are very large, and these two losses are also important losses for implementing the directed rotation target detection method.

[0136] To calculate the rotation box regression loss, ProbIoU Loss is used to calculate and quantify the difference between the predicted box and the label box. First, the rotated rectangle box is modeled as a Gaussian bounding box (GBB), and then the distance and other metrics between two Gaussian bounding boxes are used as the rotation box regression loss. The following introduces these two processes.

[0137] To model the rotated bounding box as a Gaussian rectangle box, the covariance matrix of the X and Y coordinates of all points within the rotation box needs to be calculated. After expanding the covariance matrix, it is where. θ represents the rotation angle, and the range of the rotation angle θ is [0°, 90°], and h and w respectively represent the height and width of the rotation box.

[0138] FromFigure 6 It can be seen that modeling with a Gaussian rectangle can better fit the position of the target object than a rotated rectangle. Among them, Figure 6 the ellipse in [] represents the Gaussian rectangle, and the rectangle represents the rotated box.

[0139] To measure the distance between two Gaussian rectangles, the Bhattacharyya distance Bd(p,q) is used to represent the distance between two two-dimensional probability density functions p(x) and q(x). The calculation formula of the Bhattacharyya distance Bd(p,q) is where the rotation center point coordinates of p(x) and q(x) are (x 1 , y 1 ) and (x 2 , y 2 ), respectively. The variances of the rotated box in the horizontal and vertical directions are (a 1 , b 1 ) and (a 2 , b 2 ), respectively. The covariances of the rotated box in the horizontal and vertical directions are c 1 and c 2 .

[0140] Since the Bhattacharyya distance is not an actual distance, the Hellinger distance Hd(p,q) that satisfies the triangle inequality is defined to represent the distance between two probability distributions. The calculation formula of the Hellinger distance is Therefore, the final ProbIoU loss is obtained: ProbIoU = 1 - Hd(p,q).

[0141] The orientation point regression loss is the loss that measures the deviation degree of the model's predicted orientation points and is specifically used for the orientation point detection task. It is mainly calculated by the formula Loss OKS = 1 - OKS, where d represents the Euclidean distance between the predicted and true positions of the orientation point, S is the area of the target object, σ is a constant for each orientation point, and when there are two orientation points, the value of each σ is 0.5. δ is an indicator function that is 1 when v > 0 and 0 otherwise, indicating whether the orientation point is visible.

[0142] To determine the positive direction, it is necessary to mark the orientation points (Orientation Point) on the original data. To handle the situation where the target object is truncated in the picture, two orientation points are designed, including Head and Tail, as Figure 9 shown. Among them, Head represents the head position of the target, and Tail represents the tail position of the target. The information of each orientation point includes (x,y) and Conf, representing the coordinate information and confidence information of the orientation point, respectively.

[0143] Loss Cls Calculate using the Binary Cross Entropy Loss (BCE Loss). Loss Cls is used to judge each predicted class and its confidence. Loss DFL Calculate using the CrossEntropy Loss (CE Loss) so that the model can better predict the width and height of the target. Loss Kpts_obj Also calculate using the Binary Cross Entropy Loss so that the model can learn which are visible points and which are invisible points.

[0144] The ProbIoU Loss may not fully cover the range of 0 - 360° in some cases because it is mainly designed to handle errors within the range of 0 - 180°. To solve this problem and improve the ability to identify the positive direction of the target, it can be achieved by combining the results of direction points and rotated object detection. The core of this method lies in that the direction points provide specific direction information of the target object, such as the bow or the stern of a ship. By combining these direction points with the rotated object detection algorithm, the positive direction of the target object can be determined more accurately, thus achieving a 360° all-round angle judgment.

[0145] Figure 10 is a comparison diagram of whether there are direction points. Without direction points, the positive direction of the target cannot be determined, and at this time, all four directions may be the positive direction (refer to the upper part of Figure 10 ). When using the Head or Tail direction points, the positive direction position can be discriminated. Refer to the lower part of Figure 10 , that is, using the Head direction point, the positive direction can be roughly discriminated.

[0146] Step 3: Use the neural network output result obtained in the previous step to decode the result to obtain the rotation box angle, rotation center, width and height values of the rotation box, and the position of the direction points of each target, so as to infer the positive direction of the target.

[0147] Among them, according to the different images, the inference of the positive direction of the target includes three cases. First, both the Head and Tail direction points are included, which means the image of the target is complete. Second, only the Head or Tail direction point is included, which means the target in the image is truncated. Third, neither the Head nor the Tail direction point is included, which means the target is completely truncated, and at this time, there is no need to judge the positive direction anymore.

[0148] Such as Figure 11As shown, assume there are two prediction results, Object0 and Object1. Among them, Object0 contains both Head and Tail direction points. In this case, directly select the Head direction point to determine the positive direction. Object1 is truncated and only contains the Tail direction point. The only point used to judge the positive direction is Tail1. As long as there are still direction points, the positive direction of the target can be determined.

[0149] Regarding the positioning of the positive direction, a more detailed explanation is given here.

[0150] Assume the coordinates of the center point C of the rotated bounding box are (x 0 , y 0 ), the rotation angle of the rotated bounding box is θ, and the height and width of the rotated bounding box are h and w respectively. From the rotation angle θ, the rotation matrix R can be known as LT, RT, RB, and LB represent the upper left vertex, upper right vertex, lower right vertex, and lower left vertex of the horizontal prediction bounding box before rotation respectively. Obtain the coordinates of LT, RT, RB, and LB relative to the center point C. Among them, the relative coordinates of LT are (-w / 2, -h / 2), the relative coordinates of RT are (w / 2, -h / 2), the relative coordinates of RB are (w / 2, h / 2), and the coordinates of LB are (-w / 2, h / 2).

[0151] It can be seen from this that the coordinates of LT′ are The coordinates of RT′ are The coordinates of RB′ are LB v The coordinates of are

[0152] For the first case, which contains both Head and Tail direction points, the process is as Figure 2 shown.

[0153] To determine the positive direction, first filter out the Tail direction points and select the Head direction points to determine the positive direction. The coordinates of the Head direction point are (x 3 , y 3 ). Let the coordinates of LT′ be (LT′ x , LT′ y ), the coordinates of RT′ be (RT′ x , RT′ y ), the coordinates of RB′ be (RB′ x , RB′ y ), and the coordinates of LB′ be (LB′ x , LB′ y ). It can be seen from this that the Euclidean distances from the four vertices LT′, RT′, RB′, and LB′ to the head direction point are respectively

[0154] Then compare d 10 , d 20 , d 30 and d 40 to select the two sides with the shortest distance to confirm the position of the positive direction. As shown in Figure 3 , RT and RB are the two vertices closest to the direction point Head 0 . Thus, the corrected direction point Head 0 ′ can be determined, and the coordinates of Head 0 ′ are

[0155] For the second case, there is only one direction point, which may be the Head direction point or the Tail direction point. If it only contains the Head direction point, the positive direction is determined according to the above method. If it only contains the Tail direction point, the two sides with the longest distance need to be selected to confirm the position of the positive direction. The corrected direction point Head 1 ′ is as shown in Figure 5 .

[0156] For the third case, that is, there is neither a Head direction point nor a Tail direction point, which means that both the head and tail of the target are truncated, and there is no meaning in discriminating the positive direction at this time.

[0157] For the technical effect, the embodiments of the present application aim to solve the problem that the positive direction cannot be determined in the rotation target detection. The direction point information is added to the model output layer, which can be used to predict the head or tail direction point of the target. Together with the information such as the center of the rotation box and the rotation angle, the target can be easily located in the post-processing, and the positive direction of the target can be identified.

[0158] Although the rotation target detection can accurately locate the target, for the targets that need to distinguish the head and tail directions, the existing rotation target detection network based on ProbIoU cannot distinguish the positive direction. As shown in Figure 12 , the rotation angles of the three ships are approximately θ, but the head and tail positions of the three ships are different, which requires further positioning of the target and adding a positive direction to the position of each target.

[0159] Embodiments of the present application are mainly applied to the scenario of detecting rotating targets that require distinguishing the positive direction. Common rotating target detection algorithms can only detect the position of an object and determine the angle by which the object rotates counterclockwise around the center point, with the range being [-45°, 135°]. It can be seen from this that common rotating target detection algorithms can only determine the angle within 180°, and cannot determine the angle within the range of 180° to 360°. Therefore, the embodiments of the present application propose a directed rotating target detection algorithm, which mainly introduces direction points to determine the positive direction of the target. The direction points can only roughly determine the positive direction, and then through calibration, a rotating target detection that can finally determine the direction is obtained. The algorithm effect is as Figure 13 shown.

[0160] Referring Figure 14 , according to the directed rotating target detection system 100 of the second aspect embodiment of the present application, it includes an acquisition module 110, a detection module 120, and a decoding module 130. The acquisition module 110 is used to acquire image data, where the image data contains a target object. The detection module 120 is used to input the image data into a rotating target detection model, and the rotating target detection model can perform feature extraction, feature fusion, and prediction on the image data to obtain a target detection result; where the target detection result includes the coordinates (x 0 , y 0 ) of the center point C of the rotating bounding box, the rotation angle θ of the rotating bounding box, the width w of the rotating bounding box, the height h of the rotating bounding box, and the coordinates of the direction point, and the coordinates of the direction point are used to mark the position of at least one of the head and tail of the target object. The decoding module 130 is used to decode the target detection result, and then infer the orientation of at least one of the head and tail.

[0161] According to the directed rotating target detection system 100 of the embodiment of the present application, it has at least the following beneficial effects: By marking the position of at least one of the head and tail of the target object through the coordinates of the direction point, and then combining the coordinates (x 0 , y 0 ) of the center point C of the rotating bounding box, the rotation angle θ of the rotating bounding box, the width w of the rotating bounding box, and the height h of the rotating bounding box and other information, the orientation of at least one of the head and tail can be inferred, which can not only improve the accuracy of target detection, but also bring deeper intelligence to industrial automation.

[0162] Referring Figure 15 , according to a directed rotating target detection device 200 of the third aspect embodiment of the present application, it includes a memory 210 and a processor 220. The memory 210 stores a computer program. When the processor 220 executes the computer program, it can implement the steps of the directed rotating target detection method as described above.

[0163] The directed rotation target detection device 200 according to the embodiments of the present application has at least the following beneficial effects: By implementing the above-mentioned directed rotation target detection method, the orientation of at least one of the head and the tail can be inferred, which can not only improve the accuracy of target detection, but also bring deeper intelligence to industrial automation.

[0164] The memory 210 includes at least one type of readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 210 may be an internal storage unit of the directed rotation target detection device 200, such as the hard disk or memory of the directed rotation target detection device 200. In other embodiments, the memory 210 may also be an external storage device of the directed rotation target detection device 200, such as a plug-in hard disk equipped on the directed rotation target detection device 200, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the memory 210 may also include both the internal storage unit and the external storage device of the directed rotation target detection device 200. In this embodiment, the memory 210 is generally used to store the operating system and various application software installed in the directed rotation target detection device 200, such as the program code of the directed rotation target detection method. In addition, the memory 210 may also be used to temporarily store various types of data that have been output or will be output.

[0165] In some embodiments, the processor 220 may be a central processing unit 220 (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 220 is generally used to control the overall operation of the directed rotation target detection device 200. In this embodiment, the processor 220 is used to run the program code stored in the memory 210 or process data, such as running the program code of the directed rotation target detection method.

[0166] In addition, the directed rotation target detection device 200 generally further includes a display, a communication interface, and a bus. The memory 210, the processor 220, the display, and the communication interface can complete mutual communication through the bus. The display screen is set to display the preset user operation interface in the initial setting mode, and at the same time, the display screen can also display the process control window. The communication interface may include a wireless network interface or a wired network interface.

[0167] A computer-readable storage medium according to an embodiment of the fourth aspect of the present application has a computer program stored thereon, and when the computer program is executed by a processor 220, the steps of the above-described directed rotation target detection method are implemented.

[0168] The computer-readable storage medium according to an embodiment of the present application has at least the following beneficial effects: By implementing the above-described directed rotation target detection method, the orientation of at least one of the head and the tail can be inferred, which can not only improve the accuracy of target detection, but also bring deeper intelligence to industrial automation.

[0169] The computer-readable storage medium provided by the embodiment of the present application may be a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific embodiments of the computer-readable storage medium may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, or magnetic storage devices, or any suitable combination of the above.

[0170] In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0171] The above computer-readable storage medium may be included in an electronic device or may exist separately without being loaded into the electronic device.

[0172] The computer program for executing the present application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0173] A computer program product according to an embodiment of the fifth aspect of the present application includes a computer program, and when the computer program is executed by a processor 220, it implements the steps of the directed rotation target detection method as described above.

[0174] The computer program product according to the embodiment of the present application has at least the following beneficial effects: By implementing the above-mentioned directed rotation target detection method, the orientation of at least one of the head and the tail can be inferred, which can not only improve the accuracy of target detection, but also bring deeper intelligence to industrial automation.

[0175] The embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present application within the scope of knowledge possessed by those of ordinary skill in the art. In addition, the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

Claims

1. A directed rotating target detection method, characterized in that: The following steps are involved: Acquire image data, wherein the image data includes a target object; The image data is input into a rotating target detection model, and the rotating target detection model can perform feature extraction, feature fusion and prediction on the image data, so as to obtain a target detection result; wherein the target detection result includes the coordinates (x0, y0) of the center point C of the rotating frame, the rotation angle θ of the rotating frame, the width w of the rotating frame, the height h of the rotating frame and the coordinates of the direction point, and the coordinates of the direction point are used to indicate the position of at least one of the head and the tail of the target object; The target detection result is decoded to infer the orientation of at least one of the head and the tail.

2. The method for detecting a directed rotating target according to claim 1, characterized in that: The "decoding the target detection result to infer the orientation of at least one of the head and the tail" comprises the following steps: The direction point includes a head direction point, and the coordinates (x3, y3) of the head direction point are obtained; Assume that the four vertices of the rotating frame after rotating by the rotation angle θ are LT′, RT′, RB′ and LB′ respectively, and calculate the coordinates of the four vertices LT′, RT′, RB′ and LB′; The distances from the four vertices LT′, RT′, RB′ and LB′ of the rotation frame to the head direction point are calculated and set as d 10 d 20 d 30 and d 40 ; Compare 10 ,d 20 ,d 30 and d 40 , obtaining the two vertices corresponding to the two smallest distances; The corrected orientation of the head is obtained through the two vertices corresponding to the two distances with the smallest values.

3. The method for detecting a directed rotating target according to claim 2, characterized in that: The "calculating the coordinates of the four vertices LT', RT', RB' and LB'" comprises the steps of: Assume that the upper left vertex, upper right vertex, lower right vertex and lower left vertex of the rotating frame when it is in a horizontal state are LT, RT, RB and LB respectively, and obtain the coordinates of LT, RT, RB and LB relative to the center point C, wherein the relative coordinate of LT is (-w / 2, -h / 2), the relative coordinate of RT is (w / 2, -h / 2), the relative coordinate of RB is (w / 2, h / 2), and the coordinate of LB is (-w / 2, h / 2); Assume that the upper left vertex, the upper right vertex, the lower right vertex and the lower left vertex of the rotating frame after rotating by the rotation angle θ are LT′, RT′, RB′ and LB′ respectively; According to the rotation angle θ, the obtained rotation matrix R is According to the coordinates of the four vertices LT, RT, RB and LB relative to the center point C and the rotation matrix R, the coordinates of LT′, RT′, RB′ and LB′ are calculated, where the coordinate of LT′ is The coordinates of RT′ are The coordinates of RB′ are The coordinates of LB′ are 4. The method for detecting a directed rotating target according to claim 1, characterized in that: The "decoding the target detection result to infer the orientation of at least one of the head and the tail" comprises the following steps: The direction point includes a tail direction point, and the coordinates (x4, y4) of the tail direction point are obtained; Assume that the four vertices of the rotating frame after rotating by the rotation angle θ are LT′, RT′, RB′ and LB′ respectively, and calculate the coordinates of the four vertices LT′, RT′, RB′ and LB′; The distances from the four vertices LT′, RT′, RB′ and LB′ of the rotating frame to the tail direction point are calculated and set as d 10 ,d 20 ,d 30 and d 40 ; Compare 10 d 20 d 30 and d 40 , obtaining the two vertices corresponding to the two distances with the largest values; The corrected direction of the tail is obtained through the two vertices corresponding to the two distances with the largest values.

5. The method for detecting a directed rotating target according to claim 1, characterized in that: The loss function of the rotating target detection model is Loss = αLoss Box +βLoss Cls +λLoss DFL +δLoss Kpts +εLoss Kpts_obj , Among them, Loss Box Represents the rotation box regression loss, Loss Cls Represents category loss, Loss DFL Represents distribution focus loss, Loss Kpts Represents the direction point regression loss, Loss Kpts_obj represents the direction point target loss, and α, β, γ, δ, and ε represent the weights of each loss respectively.

6. The method for detecting a directed rotating target according to claim 5, characterized in that: The rotation frame regression loss is calculated using a probability intersection-over-union loss function, which includes the following steps: Model the rotating box as a Gaussian rectangular box; Calculate the distance between the two Gaussian rectangular frames; According to the distance between the two Gaussian rectangular boxes, the probability intersection-over-union loss ProbIoU is calculated.

7. The method for detecting a directed rotating target according to claim 6, characterized in that: The "calculating the distance between the two Gaussian rectangular frames" comprises the following steps: Calculate the covariance matrix of the X and Y coordinates of all points in the rotation frame Among them, the range of the rotation angle θ is [0°, 90°]; To measure the distance between two Gaussian rectangles, the Bhattacharyya distance Bd(p,q) is used to represent the distance between two two-dimensional probability density functions p(x) and q(x). Wherein, (x1, y1) are the coordinates of the rotation center point of p(x), (x2, y2) are the coordinates of the rotation center point of q(x), (a1, b1) are the variance of the rotating frame in the horizontal direction, (a2, b2) are the variance of the rotating frame in the vertical direction, c1 is the covariance of the rotating frame in the horizontal direction, and c2 is the covariance of the rotating frame in the vertical direction; The Hellinger distance Hd(p,q) that satisfies the triangle inequality is used to represent the distance between two two-dimensional probability density functions p(x) and q(x), where 8. The method for detecting a directed rotating target according to claim 5, characterized in that: Through the formula Loss OKS =1-OKS Calculate the direction point regression loss Loss Kpts ,in, d is the Euclidean distance between the predicted position and the actual position of the direction point, S is the area of ​​the target object; σ is the standard deviation parameter related to the scale of the direction point; δ is an indicator function, δ indicates whether the direction point is visible, when v>0, δ is 1, indicating that the direction point is visible, when v≤0, δ is 0, indicating that the direction point is invisible.

9. A directed rotating target detection system, characterized in that: include: An acquisition module, used for acquiring image data, wherein the image data includes a target object; a detection module, configured to input the image data into a rotating target detection model, wherein the rotating target detection model can perform feature extraction, feature fusion and prediction on the image data, thereby obtaining a target detection result; wherein the target detection result includes the coordinates (x0, y0) of the center point C of the rotating frame, the rotation angle θ of the rotating frame, the width w of the rotating frame, the height h of the rotating frame and the coordinates of the direction point, wherein the coordinates of the direction point are used to indicate the position of at least one of the head and the tail of the target object; A decoding module is used to decode the target detection result to infer the direction of at least one of the head and the tail.

10. A directed rotating target detection device, characterized in that: include: a memory storing a computer program; A processor, wherein when executing the computer program, the processor can implement the steps of the directed rotating target detection method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the directed rotating target detection method according to any one of claims 1 to 8 are implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the directional rotating target detection method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Disaster reduction rescue early warning method and system based on artificial intelligence

    CN120894902A

  • Target positive direction detection method and system based on rotation frame and key point fusion

    CN121170238A

  • Target front direction detection method and system based on fusion of rotating box and key points

    CN121170238B