Labeling method and labeling apparatus

By acquiring multiple video frames from multiple perspectives and calculating the proportional relationship between the images and the actual space, and combining Euclidean distance and UAV trajectory displacement information, the problem of large measurement error under single perspective was solved, and high-precision three-dimensional dimension measurement of objects to be marked in power plant survey was realized.

CN116596988BActive Publication Date: 2026-02-17HEFEI SUNGROW RENEWABLE ENERGY SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310524174.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-09
Publication Date
2026-02-17
Estimated Expiration
2043-05-09

AI Technical Summary

Technical Problem

In existing technologies, when measuring the size of an object to be labeled based on a single survey image, the single perspective leads to larger measurement errors at distances far from the object or close to the horizon, resulting in lower labeling accuracy and precision.

Method used

By acquiring multiple video frames from multiple perspectives, and based on the proportional relationship between the images of these frames and the actual space, the first three-dimensional dimensions of the object to be labeled are determined. The three-dimensional dimensions of the target are then calculated by averaging the multiple video frames. Combined with the Euclidean distance calculation method and UAV trajectory displacement information, a more accurate three-dimensional dimension is obtained.

Benefits of technology

It improves the accuracy and precision of annotation, solves the problem of large measurement errors when the object distance is far or close to the horizon under single-view conditions, and is suitable for a variety of scenarios, with high flexibility and universality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116596988B_ABST
    Figure CN116596988B_ABST
Patent Text Reader

Abstract

The application discloses a labeling method and device, and belongs to the field of power station surveying. The labeling method comprises the following steps: acquiring a plurality of video frames corresponding to a to-be-labeled object under a plurality of perspectives, wherein the plurality of perspectives correspond to the plurality of video frames one by one; determining a first three-dimensional size corresponding to the to-be-labeled object based on a proportional relationship of an image corresponding to a target video frame in the plurality of video frames being mapped to an actual space; the first three-dimensional size corresponds to the target video frame, and the proportional relationship of the image being mapped to the actual space is a proportional relationship of the image of the to-be-labeled object in a target dimension being mapped to the actual space; and determining a target three-dimensional size of the to-be-labeled object based on a mean value of a plurality of first three-dimensional sizes corresponding to the plurality of video frames. The labeling method can acquire the three-dimensional size of the to-be-labeled object under multiple perspectives, improves the labeling precision and accuracy, and solves the problem of large measurement error of an object far away or close to a horizon under single perspective.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of power station surveying, and particularly relates to a labeling method and a labeling device. BACKGROUND

[0002] In the early construction of a power station, the actual three-dimensional size of a labeled object needs to be surveyed. In related technologies, there is a method for measuring the size of a labeled object based on a single surveying picture. The corresponding angle of view of the surveying picture is single. The commonly used labeling method has a large measurement error for a labeled object far away or close to the horizon when measuring the distance, and has low labeling precision and accuracy. SUMMARY

[0003] The present application aims to at least solve one of the technical problems in the prior art. To this end, the present application provides a labeling method and a labeling device, which can obtain the three-dimensional size of a labeled object under multiple angles of view, improve the labeling precision and accuracy, and solve the problem of a large measurement error for a labeled object far away or close to the horizon under a single angle of view.

[0004] In a first aspect, the present application provides a labeling method, which comprises:

[0005] obtaining a plurality of video frames corresponding to a labeled object under a plurality of angles of view, the plurality of angles of view corresponding one-to-one to the plurality of video frames;

[0006] determining a first three-dimensional size corresponding to the labeled object based on a proportional relationship of an image corresponding to a target video frame in the plurality of video frames being mapped to an actual space, the first three-dimensional size corresponding to the target video frame, and the proportional relationship of the image being mapped to the actual space being a proportional relationship of the image of the labeled object in a target dimension being mapped to the actual space;

[0007] determining a target three-dimensional size of the labeled object based on a mean value of a plurality of first three-dimensional sizes corresponding to the plurality of video frames.

[0008] According to the labeling method provided by the embodiments of the present application, the surveying video collected is decomposed into a plurality of video frames under different angles of view, then the labeled object in each video frame is measured respectively, the first three-dimensional size of the labeled object corresponding to a target video frame in the plurality of video frames is obtained based on the proportional relationship of the labeled object in a target dimension, and then the target three-dimensional size of the labeled object is determined based on the mean value of a plurality of first three-dimensional sizes corresponding to the plurality of video frames. The three-dimensional size of the labeled object under multiple angles of view can be obtained, the labeling precision and accuracy are improved, and the problem of a large measurement error for a labeled object far away or close to the horizon under a single angle of view is solved.

[0009] The labeling method of one embodiment of the application determines the first three-dimensional size corresponding to the to-be-labeled object based on the proportional relationship of the image corresponding to the target video frame in the multiple video frames being mapped to the actual space, and includes:

[0010] The proportional relationship corresponding to the target dimension is determined based on the imaging information corresponding to the target reference object and the three-dimensional information corresponding to the target reference object; the imaging information is determined based on the video frame;

[0011] The first three-dimensional size is obtained by processing the multiple proportional relationships corresponding to the multiple dimensions using the Euclidean distance calculation method.

[0012] The labeling method of one embodiment of the application determines the proportional relationship corresponding to the target dimension based on the imaging information corresponding to the target reference object and the three-dimensional information corresponding to the target reference object, and includes:

[0013] In the case where the target reference object is a standard reference object, the imaging information is the imaging size of the standard reference object on the target video frame, and the three-dimensional information is the second three-dimensional size of the standard reference object in the actual space;

[0014] Alternatively,

[0015] In the case where the target reference object is a UAV, the imaging information is the pixel trajectory displacement of the UAV in the acquisition time period corresponding to the multiple video frames, and the three-dimensional information is the actual trajectory displacement of the UAV in the acquisition time period corresponding to the multiple video frames.

[0016] The labeling method of one embodiment of the application determines the second three-dimensional size by the following method:

[0017] The actual measurement size of the standard reference object is determined as the second three-dimensional size;

[0018] Alternatively,

[0019] The second three-dimensional size is obtained based on the observation length of the distance between the two target points in the vertical direction between the standard reference object and the to-be-labeled object at different time instants; the vertical direction is perpendicular to the standard reference object and the to-be-labeled object, and the vertical direction, the standard reference object and the to-be-labeled object are in the same plane.

[0020] The labeling method of one embodiment of the application obtains the second three-dimensional size based on the observation length of the distance between the two target points in the vertical direction between the standard reference object and the to-be-labeled object at different time instants, and includes:

[0021] acquire the two target points in the vertical direction between the standard reference and the to-be-labeled object;

[0022] based on the flight direction of the unmanned aerial vehicle, acquire a first observation length corresponding to a distance between the two target points within a target time length during flight of the unmanned aerial vehicle in the flight direction at a first time, and a second observation length corresponding to the distance between the two target points at a second time;

[0023] based on the first observation length and the second observation length, determine a horizontal distance between the standard reference and the to-be-labeled object;

[0024] based on the horizontal distance, determine the second three-dimensional size.

[0025] The labeling method of one embodiment of the present application, the pixel trajectory displacement is determined based on the following steps:

[0026] based on position information of a first pixel point of a same target monitoring point extracted in each of the video frames in the video frames, reorganize a plurality of target monitoring points in the plurality of video frames to acquire a first curve image;

[0027] take a target pixel point in a plurality of first pixel points in the first curve image as a reference, filter at least part of the pixel points in a target filter window in the first curve image to acquire a second pixel point;

[0028] fit a plurality of the second pixel points to determine a flight trajectory curve of the unmanned aerial vehicle;

[0029] based on a pixel distance of the flight trajectory curve, determine the pixel trajectory displacement of the unmanned aerial vehicle.

[0030] The labeling method of one embodiment of the present application, the plurality of proportion relationships corresponding to the plurality of dimensions are processed by using the Euclidean distance calculation method to acquire the first three-dimensional size, comprising:

[0031] based on the formula:

[0032]

[0033] determine the first three-dimensional size, wherein L is the first three-dimensional size, k1 is a proportion relationship corresponding to a first dimension, k2 is a proportion relationship corresponding to a second dimension, k3 is a proportion relationship corresponding to a third dimension, Δx is a unit distance in the first dimension, Δy is a unit distance in the second dimension, and Δz is a unit distance in the third dimension.

[0034] The labeling method of one embodiment of the present application, the plurality of video frames corresponding to the to-be-labeled object under a plurality of perspectives are acquired, comprising:

[0035] obtain a plurality of initial video frames collected by the unmanned aerial vehicle;

[0036] perform inverse direction translation transformation on each of the initial video frames based on a shaking direction of the plurality of initial video frames, to obtain the plurality of video frames.

[0037] In a second aspect, the present application provides a labeling device, which comprises:

[0038] a first processing module, configured to obtain a plurality of video frames corresponding to a to-be-labeled object under a plurality of perspectives, the plurality of perspectives corresponding to the plurality of video frames one by one;

[0039] a second processing module, configured to determine a first three-dimensional size corresponding to the to-be-labeled object based on a proportional relationship of an image corresponding to a target video frame in the plurality of video frames being mapped to an actual space, the first three-dimensional size corresponding to the target video frame, and the proportional relationship of the image being mapped to the actual space being a proportional relationship of an image of the to-be-labeled object in a target dimension being mapped to the actual space;

[0040] a third processing module, configured to determine a target three-dimensional size of the to-be-labeled object based on a mean value of a plurality of first three-dimensional sizes corresponding to the plurality of video frames.

[0041] According to the labeling device provided in the embodiments of the present application, the survey video collected is decomposed into a plurality of video frames under different perspectives, then the to-be-labeled object in each video frame is measured respectively, and based on the proportional relationship of the to-be-labeled object in the target dimension, the first three-dimensional size of the to-be-labeled object corresponding to the target video frame in the plurality of video frames is obtained, and then based on the mean value of the plurality of first three-dimensional sizes corresponding to the plurality of video frames, the target three-dimensional size of the to-be-labeled object is determined, so that the three-dimensional size of the to-be-labeled object under multiple perspectives can be obtained, the labeling precision and accuracy are improved, and the problem of large measurement error of the object at a distance or close to the horizon under single perspective is solved.

[0042] In a third aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the labeling method of the first aspect described above when executing the computer program.

[0043] In a fourth aspect, the present application provides a non-transitory computer readable storage medium, having a computer program stored thereon, and the computer program is executed by a processor to implement the labeling method of the first aspect described above.

[0044] In a fifth aspect, the present application provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the labeling method of the first aspect described above.

[0045] The one or more technical solutions in the embodiments of the present application have at least one of the following technical effects.

[0046] By decomposing the collected survey video into multiple video frames under different perspectives, then measuring the to-be-labeled object in each video frame, and based on the proportional relationship of the to-be-labeled object in the target dimension, the first three-dimensional size of the to-be-labeled object corresponding to the target video frame in the multiple video frames is obtained, and then based on the mean value of the multiple first three-dimensional sizes corresponding to the multiple video frames, the target three-dimensional size of the to-be-labeled object is determined, the three-dimensional size of the to-be-labeled object under multiple perspectives can be obtained, the labeling accuracy and precision are improved, and the problem of large measurement error under single perspective when the object is far away or close to the horizon is solved.

[0047] Further, based on the imaging information corresponding to the target reference object and the three-dimensional information corresponding to the target reference object, the proportional relationship corresponding to the target dimension is determined, multiple proportional relationships corresponding to multiple dimensions are obtained, and then the proportional relationships of multiple dimensions are fused to obtain the first three-dimensional size corresponding to the to-be-labeled object, which can obtain more real three-dimensional size of the to-be-labeled object, improve the labeling accuracy and precision, and solve the problem of large measurement error under single perspective when the object is far away or close to the horizon.

[0048] Still further, by dividing the target reference object into a standard reference object and a UAV, and in the case that the target reference object is the standard reference object, the imaging size of the standard reference object on the target video frame is determined as the imaging information, and the second three-dimensional size of the standard reference object in the actual space is determined as the three-dimensional information; in the case that the target reference object is the UAV, the pixel trajectory displacement of the UAV is determined as the imaging information, and the actual trajectory displacement of the UAV is determined as the three-dimensional information, which provides multiple ways to obtain the proportional relationship, facilitates users to select the best way based on the actual use scene, has high flexibility, is suitable for multiple different situations, has high universality, and has a wide range of application scenarios.

[0049] Still further, by extracting the first pixel point of the same target monitoring point in each video frame to obtain a first curve image, then performing mean filtering processing on at least part of the pixel points in the first curve image to obtain a second pixel point, and then fitting multiple second pixel points to determine the flight trajectory curve of the UAV, and then determining the pixel trajectory displacement of the UAV, noise data in the collected data can be filtered out to reconstruct the flight trajectory of the UAV under multiple perspectives, the noise interference and the error of data measurement are reduced, the regression error of the flight trajectory of the UAV is reduced, and the precision and accuracy of the final labeling result are improved.

[0050] Additional aspects and advantages of the application will be made apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0051] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings in which:

[0052] Figure 1 is one of the flow schematic diagrams of the labeling method provided by the embodiments of the present application;

[0053] Figure 2 is the second flow schematic diagram of the labeling method provided by the embodiments of the present application;

[0054] Figure 3 is one of the principle schematic diagrams of the labeling method provided by the embodiments of the present application;

[0055] Figure 4 is the second principle schematic diagram of the labeling method provided by the embodiments of the present application;

[0056] Figure 5 is the third principle schematic diagram of the labeling method provided by the embodiments of the present application;

[0057] Figure 6 is the third flow schematic diagram of the labeling method provided by the embodiments of the present application;

[0058] Figure 7 is the fourth principle schematic diagram of the labeling method provided by the embodiments of the present application;

[0059] Figure 8 is the fifth principle schematic diagram of the labeling method provided by the embodiments of the present application;

[0060] Figure 9 is the structural schematic diagram of the labeling device provided by the embodiments of the present application;

[0061] Figure 10 is the structural schematic diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0062] The technical solutions in the embodiments of the present application will be described clearly below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0063] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / " generally indicates that the front and rear associated objects are in an "or" relationship.

[0064] The following will be described in conjunction with Figures 1 to 8 The labeling method of the embodiments of the present application is described.

[0065] It should be noted that the execution subject of the labeling method can be a server, or can be a labeling device, or can also be a user's terminal, including but not limited to mobile terminals and non-mobile terminals.

[0066] For example, mobile terminals include but are not limited to mobile phones, PDA smart terminals, tablet computers, and vehicle-mounted smart terminals, etc.; non-mobile terminals include but are not limited to PC terminals, etc.

[0067] As Figure 1 The labeling method comprises steps 110, 120 and 130.

[0068] Step 110, obtaining a plurality of video frames corresponding to a to-be-labeled object under a plurality of perspectives, the plurality of perspectives and the plurality of video frames corresponding one by one.

[0069] In this step, the to-be-labeled object is an object in the video frame whose actual size is to be measured.

[0070] For example, in the environment of power station surveying, the plurality of video frames can be video frames collected when the power station construction environment is surveyed, and the to-be-labeled object can be an object in the video frame for evaluating the installation capacity of the power station.

[0071] The plurality of video frames can be extracted from the surveying video, for example, the collected surveying video can be decomposed to obtain all video frames arranged in time sequence, and in actual execution process, part of the video frames can be extracted from all the video frames as the plurality of video frames; wherein, part of the video frames can be randomly extracted from all the video frames as the plurality of video frames, or part of the video frames can be extracted from all the video frames based on a target interval as the plurality of video frames.

[0072] The plurality of video frames can also be collected by a drone based on a certain time interval; the acquisition method of the plurality of video frames can be selected based on actual user demand, which is not limited by the present application.

[0073] In the process of collecting the survey video, each frame of the survey video corresponds to a different view angle, and multiple view angles correspond to multiple frames of video one by one.

[0074] In actual application, the collected multiple frames of video can be stored in a local database or a cloud database, and can be retrieved when needed later.

[0075] As shown in FIG. 1, in some embodiments, step 110 can include: Figure 2

[0076] Obtaining multiple frames of initial video frames collected by the unmanned aerial vehicle;

[0077] Based on the shaking direction of the multiple frames of initial video frames, performing inverse direction translation transformation on each frame of initial video frame to obtain multiple frames of video.

[0078] In this embodiment, the unmanned aerial vehicle is used to collect the survey video.

[0079] The initial video frame is an unprocessed video frame.

[0080] In the survey process, the collected survey video can have shaking, and multiple frames of initial video frames in the survey video can have relatively severe translation shaking.

[0081] The shaking direction of the multiple frames of video is the shaking direction of the survey video, and the multiple frames of initial video frames can be subjected to translation transformation relative to the shaking direction.

[0082] In actual execution, the collection interval between adjacent video frames is short, and in this time period, the flight trajectory of the unmanned aerial vehicle is a straight line, and the shaking between the video frames is translation shaking.

[0083] Obtaining the pixel position (a i ,b i ) of the target point on the i-th frame of video and the pixel position (a j ,b j ) of the target point on the j-th frame of video;

[0084] The shaking distance of the j-th frame of video relative to the i-th frame of video in the horizontal direction is a j -a i , and the shaking distance in the vertical direction is b j -b i ;

[0085] Performing translation transformation on the multiple frames of initial video relative to the shaking direction, that is, performing (x, y) <- (x-a j +a i , y-b j +b i ​Alternatively, you can perform the operation (x,y)<-(xa) on each pixel (x,y) of the i-th video frame. i +a j ,yb i +b j operate;

[0086] According to the size determination method provided in the embodiments of this application, based on the jitter direction of multiple initial video frames, the multiple initial video frames are translated in the opposite direction to obtain multiple video frames. This can eliminate the jitter of the video frames, obtain smoother and more realistic video frames, and thus improve the accuracy and precision of the final determined size.

[0087] Step 120: Based on the proportional relationship between the image mapping of the target video frame in multiple video frames and the actual space, determine the first three-dimensional size of the object to be labeled; the first three-dimensional size corresponds to the target video frame, and the proportional relationship between the image mapping of the object to be labeled and the actual space is the proportional relationship between the image mapping of the object to be labeled in the target dimension and the actual space.

[0088] In this step, the target video frame is any one of the multiple video frames, which can be user-defined and is not limited in this application.

[0089] The proportional relationship between the image and the actual space is the proportional relationship between the image of the object to be labeled in the target dimension and the actual space.

[0090] The actual space can be three-dimensional, and the actual space includes multiple dimensions. For example, in a three-dimensional space, multiple dimensions can include the three dimensions of X direction, Y direction and Z direction.

[0091] The target dimension can be any dimension in the actual space, and can be user-defined; this application does not impose any restrictions.

[0092] The object to be labeled is the object in the survey video whose actual size needs to be measured.

[0093] The first three-dimensional dimension is obtained based on the pixel size of the object to be labeled in the target video frame.

[0094] A multi-frame video can correspond to multiple first three-dimensional dimensions.

[0095] The principle behind the proportional relationship between an image and its actual space is explained below.

[0096] like Figure 6 As shown, the correspondence between the actual size and the image size can be obtained based on the Gaussian imaging principle.

[0097] like Figure 7 As shown, the relationship between the actual projected size and the image size can be obtained:

[0098]

[0099] Where d is the distance between the object projection and the lens, and v is the distance between the object image and the lens.

[0100] There is a linear mapping relationship between the physical object and its projection, which can be expressed as:

[0101] Actual size * cos(ω) = Actual projected size

[0102] Where ω is the projection angle, which is the angle between the tilt direction of the object and the normal plane pointing in the direction of the camera lens.

[0103] Based on the above formula, we can further obtain:

[0104]

[0105] Where ω is the projection angle, d is the distance between the object projection and the lens, and v is the distance between the object image and the lens; this formula is used to characterize the one-to-one mapping relationship from a two-dimensional image to three-dimensional space.

[0106] The projection angle ω can be obtained as follows:

[0107] like Figure 8 As shown, based on spatial analytic geometry, the lengths of line segments AB, BC, CD, and DE in the figure are respectively:

[0108]

[0109] Where v is the distance between the object image and the lens, d is the distance between the object projection and the lens, D1 is the width of the rectangle shown in the figure, D2 is the length of the rectangle shown in the figure, and d>>max(D1,D2), |D1-D2| / min(D1,D2)<ε, ε is the threshold constant, θ is the pitch angle shown in the figure, and β is the yaw angle shown in the figure.

[0110] It can also be known that:

[0111]

[0112] Where α1 and α2 are the included angles shown in the figure, which can be measured, and α1, α2 > 0.

[0113] Provided the difference between the lengths of the two sides of the rectangle is no greater than a preset threshold, the yaw angle β and pitch angle θ can be obtained based on the following formulas:

[0114]

[0115] Where α1 and α2 are the included angles shown in the figure, β is the yaw angle, and θ is the pitch angle.

[0116] Based on the formula: actual size cos(ω) = actual projected size, the length and width of the rectangle can be obtained as follows:

[0117]

[0118] Where D1 is the width of the rectangle shown in the figure, D2 is the length of the rectangle shown in the figure, d is the distance between the object projection and the lens, v is the distance between the object image and the lens, ω1 is the projection angle corresponding to D1, and ω2 is the projection angle corresponding to D2.

[0119] based on Figure 8 The spatial analytical geometric relationships shown can be used to obtain the relationship between the projection angle ω1 and the pitch angle θ and yaw angle β:

[0120]

[0121] And the relationship between the projection angle ω2 and the pitch angle θ and yaw angle β:

[0122]

[0123] Where ω1 is the projection angle corresponding to D1, ω2 is the projection angle corresponding to D2, β is the yaw angle, and θ is the pitch angle.

[0124] In some embodiments, step 120 may include:

[0125] Based on the imaging information and three-dimensional information corresponding to the target reference object, the proportional relationship corresponding to the target dimension is determined; the imaging information is determined based on video frames.

[0126] The Euclidean distance calculation method is used to process multiple proportional relationships corresponding to multiple dimensions to obtain the first three-dimensional dimension.

[0127] In this embodiment, the target reference is an object used as a labeling reference.

[0128] The imaging information corresponding to the target reference object is the imaging information of the target reference object on the two-dimensional plane. For example, the imaging information corresponding to the target reference object can be the size information of the target reference object on the two-dimensional plane.

[0129] The imaging information is determined based on video frames.

[0130] The three-dimensional information corresponding to the target reference object is the three-dimensional information of the target reference object in three-dimensional space. For example, the three-dimensional information corresponding to the target reference object can be the actual size information of the target reference object in three-dimensional space.

[0131] The proportional relationship corresponding to the target dimension is the proportional relationship between the image of the object to be labeled in the target dimension and the actual space.

[0132] It is understandable that the object to be labeled corresponds to multiple proportional relationships between images mapped to the actual space in multiple dimensions.

[0133] Euclidean distance is the true distance between two points in a multidimensional space.

[0134] The first three-dimensional dimension can be obtained based on multiple proportional relationships corresponding to multiple dimensions and Euclidean distance.

[0135] According to the annotation method provided in the embodiments of this application, based on the imaging information and three-dimensional information corresponding to the target reference object, the proportional relationship corresponding to the target dimension is determined, multiple proportional relationships corresponding to multiple dimensions are obtained, and then the proportional relationships of multiple dimensions are fused to obtain the first three-dimensional dimension corresponding to the object to be annotated. This method can obtain a more realistic three-dimensional dimension of the object to be annotated, improves the annotation accuracy and precision, and solves the problem of large measurement errors when the object distance is far or close to the horizon line under a single viewpoint.

[0136] like Figure 6 As shown, in some embodiments, the Euclidean distance calculation method is used to process multiple proportional relationships corresponding to multiple dimensions to obtain the first three-dimensional dimension, which may include:

[0137] Based on the formula:

[0138]

[0139] Determine the first three-dimensional dimension, where L is the first three-dimensional dimension, k1 is the proportional relationship corresponding to the first dimension, k2 is the proportional relationship corresponding to the second dimension, k3 is the proportional relationship corresponding to the third dimension, Δx is the unit distance in the first dimension, Δy is the unit distance in the second dimension, and Δz is the unit distance in the third dimension.

[0140] In this embodiment, the multidimensional space corresponds to multiple dimensions, and each dimension corresponds to the proportional relationship between the image and the actual space.

[0141] Based on Euclidean distance and multiple proportional relationships corresponding to multiple dimensions, the first three-dimensional dimension can be obtained.

[0142] In actual implementation, such as Figure 5 As shown, the dimensions of the object to be labeled on the two-dimensional plane can be obtained based on the following formula:

[0143]

[0144] Where L is the size of the object to be labeled on the two-dimensional plane, k1 is the proportional relationship corresponding to the first dimension of the two-dimensional plane, k2 is the proportional relationship corresponding to the second dimension of the two-dimensional plane, Δx is the unit distance in the first dimension, and Δy is the unit distance in the second dimension.

[0145] Based on the Euclidean distance calculation method, the first three-dimensional dimension of the object to be labeled in the actual three-dimensional space can be determined using the following formula:

[0146]

[0147] Where L is the first three-dimensional dimension, k1 is the proportional relationship corresponding to the first dimension, k2 is the proportional relationship corresponding to the second dimension, k3 is the proportional relationship corresponding to the third dimension, Δx is the unit distance in the first dimension, Δy is the unit distance in the second dimension, and Δz is the unit distance in the third dimension.

[0148] According to the annotation method provided in the embodiments of this application, multiple proportional relationships corresponding to multiple dimensions are processed by the Euclidean distance calculation method to determine the first three-dimensional size of the object to be annotated in the three-dimensional actual space. The calculation method is convenient and easy to implement, thereby improving annotation efficiency and saving annotation time.

[0149] Step 130: Determine the target three-dimensional size of the object to be labeled based on the average of multiple first three-dimensional sizes corresponding to multiple video frames.

[0150] In this step, multiple video frames can correspond to multiple first three-dimensional dimensions.

[0151] The target's three-dimensional dimensions are the actual dimensions of the object to be labeled in real space.

[0152] The average of multiple first three-dimensional dimensions is obtained and determined as the target three-dimensional dimension of the object to be labeled.

[0153] In actual implementation, such as Figure 2 As shown, the survey video can be decomposed into multiple frames to obtain multiple video frames, where the multiple video frames correspond to different viewpoints.

[0154] Based on the proportional relationship between the image mapping of the target video frame in multiple video frames and the actual space, the first three-dimensional size of the object to be labeled on the target video frame is determined.

[0155] Obtain multiple first three-dimensional dimensions corresponding to multiple video frames, and then average the multiple first three-dimensional dimensions to obtain the mean value corresponding to the multiple first three-dimensional dimensions;

[0156] The average value corresponding to multiple first three-dimensional dimensions is determined as the target three-dimensional dimension of the object to be labeled.

[0157] During the research and development process, the inventors discovered that in related technologies, the size measurement of the object to be labeled is based on a single survey image. Since the survey image corresponds to a single perspective, this commonly used method has a large measurement error when measuring distances to objects that are far away or close to the horizon, resulting in low labeling accuracy and precision.

[0158] In this application, the acquired survey video is decomposed into multiple video frames from different perspectives. Then, the object to be labeled in each video frame is measured. Based on the proportional relationship of the object to be labeled in the target dimension, the first three-dimensional dimension of the object to be labeled corresponding to the target video frame in the multiple video frames is obtained. Then, based on the average of the multiple first three-dimensional dimensions corresponding to the multiple video frames, the target three-dimensional dimension of the object to be labeled is determined. This method can obtain the three-dimensional dimension of the object to be labeled from multiple perspectives, improves the labeling accuracy and precision, and solves the problem of large measurement errors when the object distance is far or close to the horizon line under a single perspective.

[0159] According to the annotation method provided in the embodiments of this application, the collected survey video is decomposed into multiple video frames from different perspectives. Then, the object to be annotated in each video frame is measured. Based on the proportional relationship of the object to be annotated in the target dimension, the first three-dimensional dimension of the object to be annotated corresponding to the target video frame in the multiple video frames is obtained. Then, based on the average of the multiple first three-dimensional dimensions corresponding to the multiple video frames, the target three-dimensional dimension of the object to be annotated is determined. This method can obtain the three-dimensional dimension of the object to be annotated from multiple perspectives, improves the annotation accuracy and precision, and solves the problem of large measurement errors when the object distance is far or close to the horizon line under a single perspective.

[0160] The following specific examples illustrate how to determine the proportional relationship between an image and its actual spatial mapping.

[0161] I. The target reference object is the standard reference object.

[0162] In some embodiments, determining the proportional relationship corresponding to the target dimension based on the imaging information and the three-dimensional information corresponding to the target reference object may include:

[0163] When the target reference object is a standard reference object, the imaging information is the imaging size of the standard reference object on the target video frame, and the three-dimensional information is the second three-dimensional size of the standard reference object in actual space.

[0164] In this embodiment, when the target reference object is a standard reference object, the imaging information is the imaging size of the standard reference object on the target video frame, and the three-dimensional information is the second three-dimensional size of the standard reference object in actual space.

[0165] The imaging size is the size of the standard reference object on the target video frame, which can be directly measured.

[0166] The second three-dimensional dimension is the size information of the standard reference object in actual space.

[0167] According to the annotation method provided in the embodiments of this application, when the target reference object is a standard reference object, the imaging size of the standard reference object on the target video frame can be directly determined as imaging information, and the second three-dimensional size of the standard reference object in actual space can be determined as three-dimensional information. In subsequent applications, the proportional relationship corresponding to the target dimension can be determined based on the imaging information and the three-dimensional information, which is convenient and fast, and the final annotation accuracy and precision are high.

[0168] Understandably, in actual implementation, the second and third dimensions of the standard reference object may be directly obtained or calculated based on analytical geometry.

[0169] The following sections will explain the methods for obtaining the second three-dimensional dimension from two different perspectives.

[0170] Firstly, obtain directly

[0171] In some embodiments, the second three-dimensional dimension can be determined in the following manner:

[0172] The actual measured dimensions of the standard reference object can be determined as the second three-dimensional dimension.

[0173] In this embodiment, the second three-dimensional dimension is the size information of the standard reference object in actual space.

[0174] The actual measured dimensions of the standard reference object can be obtained by the user.

[0175] According to the annotation method provided in the embodiments of this application, by determining the actual measured size of the standard reference object as the second three-dimensional size, the second three-dimensional size can be quickly obtained when the actual measured size of the standard reference object is known, thereby determining the proportional relationship and improving the calculation efficiency.

[0176] Secondly, it is obtained based on analytical geometry calculations.

[0177] This scenario is mainly used when it is impossible to directly obtain a standard reference object for actual measurement dimensions.

[0178] In some embodiments, the second three-dimensional dimension can be determined in the following manner:

[0179] The second three-dimensional dimension is obtained based on the observed length corresponding to the distance between two target points in the vertical direction between the standard reference object and the object to be labeled at different times; the vertical direction is perpendicular to the standard reference object and the object to be labeled, and the vertical direction, the standard reference object, and the object to be labeled are in the same plane.

[0180] In this embodiment, the vertical direction is the direction perpendicular to both the standard reference and the object to be labeled, such as... Figure 3 As shown, the vertical direction can be the direction perpendicular to the flight trajectory.

[0181] The vertical direction, the standard reference object, and the object to be labeled are all on the same plane.

[0182] The two target points can be user-defined. The two target points should be on the same vertical line. For example, two points close to the object to be labeled can be selected as the two target points.

[0183] At different times, the observation points for the drone-based flight are different, and at different times, the observation lengths corresponding to the two target points are different.

[0184] The observation length is the observation distance between two target points.

[0185] In actual implementation, such as Figure 3 As shown, a time period in which the standard reference object and the object to be labeled are relatively horizontal can be selected. Within a short period of time, it can be assumed that the drone flies in a straight line during this period.

[0186] Two target points, M and N, can be selected on the side closest to the object to be labeled;

[0187] When the UAV flies to point O, the observation length corresponding to the two target points at that moment can be L1; when the UAV flies to point B, the observation length corresponding to the two target points at that moment can be L2.

[0188] Based on L1 and L2, obtain the second three-dimensional dimension.

[0189] According to the annotation method provided in the embodiments of this application, a second three-dimensional dimension is obtained by measuring the distance between two target points in the vertical direction between the standard reference object and the object to be annotated at different times when the actual measured size of the standard reference object is unknown. This method can also obtain the actual size of the object to be annotated when the actual measured size of the standard reference object is unknown, thereby determining the proportional relationship. It has a wide range of applicable scenarios.

[0190] In some embodiments, obtaining the second three-dimensional dimension based on the observed length corresponding to the distance between two target points along the vertical direction between the standard reference object and the object to be labeled at different times may include:

[0191] Obtain two target points along the vertical direction between the standard reference object and the object to be labeled;

[0192] Based on the flight direction of the UAV, the first observation length corresponding to the distance between two target points at the first moment and the second observation length corresponding to the distance between the two target points at the second moment are obtained during the flight of the UAV along the flight direction.

[0193] Based on the first and second observation lengths, determine the horizontal distance between the standard reference object and the object to be labeled;

[0194] The second three-dimensional dimension is determined based on the horizontal distance.

[0195] In this embodiment, the flight direction of the drone can be as follows: Figure 3 Example flight trajectory direction.

[0196] The target duration is any time period during the drone's flight along the flight direction.

[0197] In a short period of time, it can be assumed that the drone is flying in a straight line.

[0198] like Figure 3 As shown, the first moment can be the moment when the UAV flies to point O, and the first observation length is the observation distance between the two target points corresponding to point O, which can be L1;

[0199] The second moment can be the moment when the UAV flies to point B, and the second observation length is the observation distance between the two target points corresponding to point B, which can be L2.

[0200] Horizontal distance is the distance between a standard reference object and the object to be labeled, such as... Figure 3 In this context, X represents the horizontal distance between the standard reference object and the object to be labeled.

[0201] In actual implementation, it can be based on geometric relationships:

[0202]

[0203]

[0204] Determine the second three-dimensional dimension, where α1 is the angle between ON and AN, α2 is the angle between BN and AN, a is the length of the measurement segment, b is the vertical distance between the two target points, x is the horizontal distance between the standard reference and the object to be labeled, l1 is the first observation distance, l2 is the second observation distance, d1 is the length of ON, and d2 is the length of BN.

[0205] Let tanα1 = , tanα2 = 2t, assume Substituting tanα1=, tanα2=2t and β into the above formula, we can calculate t and the horizontal distance X, and thus obtain the second three-dimensional dimension.

[0206] According to the annotation method provided in the embodiments of this application, by establishing an analytical geometric model, the horizontal distance between the standard reference object and the object to be annotated is determined based on the first observation length and the second observation length corresponding to two target points, and then the second three-dimensional dimension is determined. Even when the actual measured dimension of the reference object is unknown, the second three-dimensional dimension of the standard reference object can be obtained, and then the actual three-dimensional dimension of the object to be annotated can be obtained, thus broadening the application scenarios and improving the accuracy and precision of annotation.

[0207] In this application, the actual measured size of the standard reference object is determined as the second three-dimensional size when the actual measured size of the standard reference object is known; and when the actual measured size of the standard reference object is unknown, the second three-dimensional size is obtained based on the observation length corresponding to the distance between two target points in the vertical direction between the standard reference object and the object to be labeled at different times. Different methods can be selected to obtain the second three-dimensional size according to different situations, which is applicable to a variety of different situations and has high universality. It can also obtain the actual size of the object to be labeled when the actual measured size of the standard reference object is unknown, and has a wide range of applicable scenarios.

[0208] II. The target reference object is a drone.

[0209] This embodiment is mainly used in situations where a standard reference object cannot be obtained, and the UAV can be approximately determined as the target reference object for calculation.

[0210] In some embodiments, determining the proportional relationship corresponding to the target dimension based on the imaging information and the three-dimensional information corresponding to the target reference object may include:

[0211] When the target reference is a drone, the imaging information is the pixel trajectory displacement of the drone within the acquisition time period corresponding to multiple video frames, and the three-dimensional information is the actual trajectory displacement of the drone within the acquisition time period corresponding to multiple video frames.

[0212] In this embodiment, when the target reference is a drone, the imaging information is the pixel trajectory displacement of the drone within the acquisition time period corresponding to multiple video frames, and the three-dimensional information is the actual trajectory displacement of the drone within the acquisition time period corresponding to multiple video frames.

[0213] Drones are used to collect survey videos.

[0214] The acquisition time period refers to the time taken by the drone to acquire multiple video frames.

[0215] The pixel trajectory is the movement trajectory curve on the two-dimensional plane corresponding to the drone's movement when acquiring multiple video frames within the acquisition time period.

[0216] The pixel trajectory displacement is the perimeter corresponding to the pixel trajectory.

[0217] The actual trajectory is the curve of the drone's real movement in three-dimensional space corresponding to the multiple video frames collected during the acquisition period.

[0218] The actual trajectory displacement is the perimeter corresponding to the actual trajectory.

[0219] According to the annotation method provided in the embodiments of this application, when the target reference object is a drone, the pixel trajectory displacement of the drone is determined as imaging information, and the actual trajectory displacement of the drone is determined as three-dimensional information. Even when there is no standard reference object in the video frame, the proportional relationship corresponding to the target dimension can be obtained. It has high flexibility and universality and is applicable to a wide range of scenarios.

[0220] In some embodiments, pixel trajectory displacement can be determined based on the following steps:

[0221] Based on the position information of the first pixel of the same target monitoring point extracted from each video frame, multiple target monitoring points in multiple video frames are recombined to obtain the first curve image.

[0222] Using the target pixel among multiple first pixels in the first curve image as a reference, at least some pixels in the target filtering window in the first curve image are filtered to obtain the second pixel.

[0223] By fitting multiple second pixel points, the flight trajectory curve of the drone is determined;

[0224] The pixel trajectory displacement of the UAV is determined based on the pixel distance of the flight trajectory curve.

[0225] In this embodiment, the target monitoring point is any point on the video frame, which can be user-defined and is not limited in this application.

[0226] The first pixel is the pixel of the target monitoring point in the video frame.

[0227] The position information of the first pixel can be the coordinate information corresponding to the first pixel.

[0228] The first curve image is obtained by reconstructing multiple target monitoring points from multiple video frames.

[0229] The target pixel is a pixel among multiple first pixels. For example, it can be the pixel corresponding to the median of a sequence obtained by sorting the position information of multiple first pixels based on the acquisition order of video frames.

[0230] The target filtering window can be determined based on the target pixels, and the length of the target filtering window can be user-defined, which is not limited in this application.

[0231] The second pixel is obtained by filtering at least some of the pixels within the target window.

[0232] By filtering multiple first pixels separately, multiple second pixels can be obtained.

[0233] The flight trajectory curve of the drone is obtained by fitting multiple second pixel points.

[0234] The flight trajectory curve of a drone is used to characterize the flight trajectory of the drone.

[0235] The pixel distance of the flight trajectory curve is the distance of the flight trajectory curve on the two-dimensional plane.

[0236] In actual execution, the position information of the first pixel can be represented as (x, y). Based on the position information of the first pixel of the same target monitoring point extracted from each video frame, multiple target monitoring points in multiple video frames are reconstructed, such as... Figure 4 As shown in (a), the first curve image is then obtained;

[0237] The target pixel (x) among multiple first pixel points in the first curve image i ,y i Based on the above, mean filtering is applied to at least a portion of the pixels within the target filtering window in the first curve image according to the following formula:

[0238]

[0239]

[0240] Where i>k>0, 2k+1 is the length of the target filtering window, and then the second pixel is obtained. Multiple second pixels are as follows: Figure 4 As shown in (b), multiple second pixel points are fitted to determine the flight trajectory curve of the UAV, as follows. Figure 4 As shown in the curve in (c), the pixel trajectory displacement of the UAV is determined based on the pixel distance of the flight trajectory curve, where the pixel distance of the flight trajectory curve can be calculated by computer.

[0241] According to the annotation method provided in the embodiments of this application, a first curve image is obtained by extracting the first pixel point of the same target monitoring point in each video frame. Then, at least some pixels in the first curve image are subjected to mean filtering to obtain second pixels. Multiple second pixels are then fitted to determine the flight trajectory curve of the UAV, thereby determining the pixel trajectory displacement of the UAV. This method can filter out noise data in the collected data to reconstruct the UAV flight trajectory from multiple perspectives, reduce noise interference and data measurement errors, reduce the regression error of the UAV flight trajectory, and improve the accuracy and precision of the final annotation result.

[0242] This application provides multiple ways to obtain the ratio, making it easy for users to choose the best method based on their actual usage scenarios. It is highly flexible, applicable to a variety of different situations, and has high universality and wide applicability.

[0243] Continue to refer to Figure 6 In some embodiments, the annotation method may further include:

[0244] Edge features are extracted from the objects to be labeled in the video frame to obtain the contour information corresponding to the objects to be labeled.

[0245] Based on the contour information, obtain the pixel size corresponding to the object to be labeled.

[0246] In this embodiment, parameters such as the location information of the object to be labeled in the video frame can be extracted based on object detection technology. The object detection technology can be traditional manual labeling, or it can be an intelligent classifier based on machine learning combined with neural networks. It can be user-defined and is not limited in this application.

[0247] The contour information of the object to be labeled can be obtained based on edge extraction technology.

[0248] According to the annotation method provided in the embodiments of this application, edge features of the object to be annotated in the video frame are extracted to obtain the contour information corresponding to the object to be annotated. Then, based on the contour information, the pixel size corresponding to the object to be annotated is obtained. This method can efficiently and accurately track the position of the object to be annotated, thereby improving the accuracy and precision of the annotation.

[0249] The annotation apparatus provided in this application is described below. The annotation apparatus described below can be referred to in correspondence with the annotation method described above.

[0250] The annotation method provided in this application can be executed by an annotation device. This application uses an annotation device executing the annotation method as an example to illustrate the annotation device provided in this application.

[0251] This application also provides a labeling device.

[0252] like Figure 9 As shown, the labeling device includes: a first processing module 910, a second processing module 920 and a third processing module 930.

[0253] The first processing module 910 is used to acquire multiple video frames corresponding to the object to be labeled from multiple perspectives, with each perspective corresponding to a video frame.

[0254] The second processing module 920 is used to determine the first three-dimensional size of the object to be labeled based on the proportional relationship between the image mapping of the target video frame in multiple video frames and the actual space; the first three-dimensional size corresponds to the target video frame, and the proportional relationship between the image mapping of the object to be labeled and the actual space is the proportional relationship between the image mapping of the object to be labeled in the target dimension and the actual space.

[0255] The third processing module 930 is used to determine the target three-dimensional size of the object to be labeled based on the average of multiple first three-dimensional sizes corresponding to multiple video frames.

[0256] According to the annotation device provided in the embodiments of this application, the collected survey video is decomposed into multiple video frames from different perspectives. Then, the object to be annotated in each video frame is measured. Based on the proportional relationship of the object to be annotated in the target dimension, the first three-dimensional dimension of the object to be annotated corresponding to the target video frame in the multiple video frames is obtained. Then, based on the average of the multiple first three-dimensional dimensions corresponding to the multiple video frames, the target three-dimensional dimension of the object to be annotated is determined. This device can obtain the three-dimensional dimension of the object to be annotated from multiple perspectives, improves the annotation accuracy and precision, and solves the problem of large measurement errors when the object distance is far or close to the horizon line under a single perspective.

[0257] In some embodiments, the labeling device may further include:

[0258] The fourth processing module is used to determine the proportional relationship of the target dimension based on the imaging information and the three-dimensional information of the target reference object; the imaging information is determined based on video frames.

[0259] The fifth processing module is used to process multiple proportional relationships corresponding to multiple dimensions using the Euclidean distance calculation method to obtain the first three-dimensional dimension.

[0260] In some embodiments, the fourth processing module can also be used for:

[0261] When the target reference object is a standard reference object, the imaging information is the imaging size of the standard reference object on the target video frame, and the three-dimensional information is the second three-dimensional size of the standard reference object in actual space.

[0262] or,

[0263] When the target reference is a drone, the imaging information is the pixel trajectory displacement of the drone within the acquisition time period corresponding to multiple video frames, and the three-dimensional information is the actual trajectory displacement of the drone within the acquisition time period corresponding to multiple video frames.

[0264] In some embodiments, the annotation device may further include a sixth processing module for determining a second three-dimensional dimension:

[0265] The actual measured dimensions of the standard reference object are determined as the second three-dimensional dimension;

[0266] or,

[0267] The second three-dimensional dimension is obtained based on the observed length corresponding to the distance between two target points in the vertical direction between the standard reference object and the object to be labeled at different times; the vertical direction is perpendicular to the standard reference object and the object to be labeled, and the vertical direction, the standard reference object, and the object to be labeled are in the same plane.

[0268] In some embodiments, the labeling device may further include a seventh processing module for:

[0269] Obtain two target points along the vertical direction between the standard reference object and the object to be labeled;

[0270] Based on the flight direction of the UAV, the first observation length corresponding to the distance between two target points at the first moment and the second observation length corresponding to the distance between the two target points at the second moment are obtained during the flight of the UAV along the flight direction.

[0271] Based on the first and second observation lengths, determine the horizontal distance between the standard reference object and the object to be labeled;

[0272] The second three-dimensional dimension is determined based on the horizontal distance.

[0273] In some embodiments, the annotation device may further include an eighth processing module for determining pixel trajectory displacement:

[0274] Based on the position information of the first pixel of the same target monitoring point extracted from each video frame, multiple target monitoring points in multiple video frames are recombined to obtain the first curve image.

[0275] Using the target pixel among multiple first pixels in the first curve image as a reference, at least some pixels in the target filtering window in the first curve image are filtered to obtain the second pixel.

[0276] By fitting multiple second pixel points, the flight trajectory curve of the drone is determined;

[0277] The pixel trajectory displacement of the UAV is determined based on the pixel distance of the flight trajectory curve.

[0278] In some embodiments, the fifth processing module may also be used for:

[0279] Based on the formula:

[0280]

[0281] Determine the first three-dimensional dimension, where L is the first three-dimensional dimension, k1 is the proportional relationship corresponding to the first dimension, k2 is the proportional relationship corresponding to the second dimension, k3 is the proportional relationship corresponding to the third dimension, Δx is the unit distance in the first dimension, Δy is the unit distance in the second dimension, and Δz is the unit distance in the third dimension.

[0282] In some embodiments, the first processing module 910 may also be used for:

[0283] Acquire multiple initial video frames captured by the drone;

[0284] Based on the jitter direction of the initial video frames, the initial video frames are translated in the opposite direction to obtain the multiple video frames.

[0285] The labeling device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0286] The labeling device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0287] The labeling device provided in this application embodiment can achieve... Figures 1 to 8 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0288] In some embodiments, such as Figure 10As shown, this application embodiment also provides an electronic device 1000, including a processor 1001, a memory 1002, and a computer program stored on the memory 1002 and executable on the processor 1001. When the program is executed by the processor 1001, it implements the various processes of the above-described annotated method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0289] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0290] On the other hand, this application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the various processes of the above-described annotation method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0291] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it is implemented to perform the various processes of the above-described annotated method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0292] In another aspect, this application embodiment provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned annotation method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0293] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0294] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0295] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0296] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A labeling method, characterized in that, include: Acquire multiple video frames corresponding to the object to be labeled from multiple perspectives, wherein the multiple perspectives correspond one-to-one with the multiple video frames; Based on the proportional relationship between the image mapping of the target video frame in the multi-frame video and the actual space, the first three-dimensional dimension of the object to be labeled is determined. The first three-dimensional dimension corresponds to the target video frame, and the proportional relationship of the image mapping to the actual space is the proportional relationship of the image of the object to be labeled in the target dimension mapping to the actual space. The target three-dimensional dimensions of the object to be labeled are determined based on the average of multiple first three-dimensional dimensions corresponding to the multiple video frames. The step of determining the first three-dimensional dimension of the object to be labeled based on the proportional relationship between the image mapping of the target video frame in the multi-frame video and the actual space includes: Based on the imaging information corresponding to the target reference object and the three-dimensional information corresponding to the target reference object, the proportional relationship corresponding to the target dimension is determined; the imaging information is determined based on the video frame; The first three-dimensional dimension is obtained by using the Euclidean distance calculation method to process multiple proportional relationships corresponding to multiple dimensions.

2. The annotation method according to claim 1, characterized in that, The step of determining the proportional relationship corresponding to the target dimension based on the imaging information corresponding to the target reference object and the three-dimensional information corresponding to the target reference object includes: When the target reference object is a standard reference object, the imaging information is the imaging size of the standard reference object on the target video frame, and the three-dimensional information is the second three-dimensional size of the standard reference object in the actual space; or, When the target reference object is a drone, the imaging information is the pixel trajectory displacement of the drone within the acquisition time period corresponding to the multi-frame video frame, and the three-dimensional information is the actual trajectory displacement of the drone within the acquisition time period corresponding to the multi-frame video frame.

3. The annotation method according to claim 2, characterized in that, The second three-dimensional dimension is determined as follows: The actual measured dimensions of the standard reference object are determined as the second three-dimensional dimension; or, The second three-dimensional dimension is obtained based on the observed lengths corresponding to the distance between two target points along the vertical direction between the standard reference object and the object to be labeled at different times; the vertical direction is perpendicular to the standard reference object and the object to be labeled, and the vertical direction, the standard reference object, and the object to be labeled are in the same plane.

4. The annotation method according to claim 3, characterized in that, The method of obtaining the second three-dimensional dimension based on the observed length corresponding to the distance between two target points in the vertical direction between the standard reference object and the object to be labeled at different times includes: Obtain the two target points along the vertical direction between the standard reference object and the object to be labeled; Based on the flight direction of the UAV, the first observation length corresponding to the distance between the two target points at the first moment and the second observation length corresponding to the distance between the two target points at the second moment are obtained during the flight of the UAV along the flight direction within the target time. Based on the first observation length and the second observation length, determine the horizontal distance between the standard reference object and the object to be labeled; The second three-dimensional dimension is determined based on the horizontal distance.

5. The annotation method according to claim 2, characterized in that, The pixel trajectory displacement is determined based on the following steps: Based on the position information of the first pixel of the same target monitoring point extracted from each video frame, multiple target monitoring points in the multi-frame video are reconstructed to obtain a first curve image. Using the target pixel among the multiple first pixel points in the first curve image as a reference, at least some of the pixel points in the target filtering window in the first curve image are filtered to obtain the second pixel point. By fitting multiple second pixel points, the flight trajectory curve of the drone is determined; The pixel trajectory displacement of the UAV is determined based on the pixel distance of the flight trajectory curve.

6. The annotation method according to claim 1, characterized in that, The step of using the Euclidean distance calculation method to process multiple proportional relationships corresponding to multiple dimensions to obtain the first three-dimensional dimension includes: Based on the formula: Determine the first three-dimensional dimensions, wherein, The first three-dimensional dimension, This represents the proportional relationship corresponding to the first dimension. This represents the proportional relationship corresponding to the second dimension. This represents the proportional relationship corresponding to the third dimension. The unit distance in the first dimension. The unit distance in the second dimension. The unit distance in the third dimension.

7. The annotation method according to any one of claims 1-6, characterized in that, The process of acquiring multiple video frames corresponding to the object to be labeled from multiple viewpoints includes: Acquire multiple initial video frames captured by the drone; Based on the jitter direction of the initial video frames, each of the initial video frames is translated in the opposite direction to obtain the multiple video frames.

8. A labeling device, characterized in that, include: The first processing module is used to acquire multiple video frames corresponding to the object to be labeled from multiple perspectives, wherein the multiple perspectives correspond one-to-one with the multiple video frames. The second processing module is used to determine the first three-dimensional dimension of the object to be labeled based on the proportional relationship between the image corresponding to the target video frame in the multi-frame video and the actual space; The first three-dimensional dimension corresponds to the target video frame, and the proportional relationship of the image mapping to the actual space is the proportional relationship of the image of the object to be labeled in the target dimension mapping to the actual space. The third processing module is used to determine the target three-dimensional size of the object to be labeled based on the average of multiple first three-dimensional sizes corresponding to the multiple video frames. The second processing module is specifically used to: determine the proportional relationship corresponding to the target dimension based on the imaging information corresponding to the target reference object and the three-dimensional information corresponding to the target reference object; the imaging information is determined based on the video frame; and use the Euclidean distance calculation method to process multiple proportional relationships corresponding to multiple dimensions to obtain the first three-dimensional size.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the annotation method as described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the annotation method as described in any one of claims 1-7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the annotation method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Three-dimensional measurement technology-based system and method for measuring surface area of object

    WO2018049818A1