Video labeling method and device, electronic equipment and medium
By establishing vector coordinate transformation relationships in rotatable and scalable camera videos and calculating the current pixel coordinates, the problem of automatic marker following is solved, achieving efficient video annotation and reducing costs.
Patent Information
- Application Number
- CN202310706332.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-14
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-06-14
AI Technical Summary
Existing technologies struggle to achieve automatic marker following in videos captured by rotatable and scalable cameras, especially when marking the location of objects in 3D scenes. Existing technologies cannot complete marker following without re-identifying the objects to be marked.
By acquiring the initial vector and pixel coordinates of the labeled image, establishing the vector coordinate transformation relationship, calculating the difference between the current vector and the initial vector of the image to be labeled, and then obtaining the current pixel coordinates of the image to be labeled, the automatic following of the marker is achieved.
It eliminates the need for real-time detection and labeling of every frame, improving video labeling efficiency and reducing usage costs. It is especially suitable for efficient labeling of buildings, roads, etc. in high-altitude eagle-eye cameras.
Smart Images

Figure CN116740716B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data labeling, and specifically provides a video labeling method and device, an electronic device and a medium. BACKGROUND
[0002] With the popularity of cameras in life, many scenes need video content shot by cameras to assist work, and thus the demand for labeling and recognizing video content arises at the historic moment. After a 360-degree video of the surrounding scene is obtained by using a high-altitude hawk-eye monitoring camera, buildings, roads, regions and other markers need to be labeled.
[0003] In the prior art, it is relatively easy to label real objects in a fixed picture video. However, with the upgrading of photography and monitoring devices, a rotatable and zoomable camera can rotate to shoot the direction required by a user, and the application scenarios have become more and more widespread.
[0004] However, since labeling real objects on a fixed picture is equivalent to labeling on a 2D fixed plane, labeling real objects on a video shot by a rotatable and zoomable camera is equivalent to determining the position of a real object in a 3D scene. Compared with labeling real objects on a fixed picture, it is more complex to label real objects on a video shot by a rotatable and zoomable camera, and to realize that the labeling can follow the position of a real object in the video in real time.
[0005] When labeling a video shot by a rotatable and zoomable camera, the position of a real object to be labeled in the screen will also change due to the movement of the picture, and the prior art is difficult to complete the follow-up of labeling without re-identifying the real object to be labeled.
[0006] Correspondingly, there is a need in the art for a new video labeling scheme to solve the above problems. SUMMARY
[0007] In order to overcome the above defects, the present application is proposed, and a video labeling method, device, electronic device and medium are provided to solve or at least partially solve the technical problem of how to realize automatic follow-up of labeling in video data labeling.
[0008] In a first aspect, the present application provides a video labeling method, comprising:
[0009] Based on the labeled initial labeling picture, an initial vector and an initial pixel coordinate of each marker are obtained, and a vector coordinate conversion relationship is obtained;
[0010] Based on the to-be-labeled picture, a vector difference between a current vector and an initial vector of each marker is obtained;
[0011] Based on the initial vector, the initial pixel coordinate and the vector difference, the current pixel coordinate of each mark in the to-be-labeled picture is obtained through the vector coordinate conversion relationship.
[0012] In one of the above technical solutions of the video labeling method, based on the initial labeled picture, the initial vector and the initial pixel coordinate of each mark are obtained, and the vector coordinate conversion relationship comprises:
[0013] A frame of picture in a video is obtained as an initial labeled picture, and a camera rotation angle corresponding to the initial labeled picture is obtained;
[0014] A plurality of marks are made along the contour of the to-be-labeled real object in the initial labeled picture, and the initial pixel coordinate of each mark in the player is obtained;
[0015] The initial vector of each mark is obtained;
[0016] Based on the camera rotation angle corresponding to the initial labeled picture, and the initial vector and the initial pixel coordinate of each mark, a vector coordinate conversion relationship is obtained.
[0017] In one of the above technical solutions of the video labeling method, the method further comprises:
[0018] A camera scaling factor corresponding to the initial labeled picture is obtained;
[0019] And,
[0020] Based on the camera rotation angle and the scaling factor corresponding to the initial labeled picture, and the initial vector and the initial pixel coordinate of each mark, a vector coordinate conversion relationship is obtained.
[0021] In one of the above technical solutions of the video labeling method, the initial vector of each mark comprises:
[0022] A three-dimensional camera coordinate system is established with the viewpoint of the camera as the origin;
[0023] The coordinates of each mark in the camera coordinate system are obtained;
[0024] Based on the coordinates in the camera coordinate system, the initial vector of each mark is obtained.
[0025] In one of the above technical solutions of the video labeling method, the vector difference between the current vector and the initial vector of each mark based on the to-be-labeled picture comprises:
[0026] A camera rotation angle corresponding to the to-be-labeled picture is obtained;
[0027] based on a change of the camera rotation angle corresponding to the to-be-labeled picture and the camera rotation angle corresponding to the initial labeled picture, a vector difference between the current vector and the initial vector of each marker is obtained.
[0028] In one of the technical solutions of the video labeling method, the method further includes:
[0029] a screen size corresponding to the initial labeled picture is obtained;
[0030] a two-dimensional pixel coordinate system is established based on the screen size corresponding to the initial labeled picture;
[0031] coordinates of each marker in the initial labeled picture in the pixel coordinate system are obtained as initial pixel coordinates;
[0032] and / or,
[0033] a screen size corresponding to the to-be-labeled picture is obtained;
[0034] based on the initial vector, the initial pixel coordinates, the vector difference, and the screen size corresponding to the to-be-labeled picture, the current pixel coordinates of each marker in the to-be-labeled picture are obtained through the vector coordinate conversion relationship.
[0035] In one of the technical solutions of the video labeling method, the method further includes:
[0036] one or more pictures in the to-be-labeled video are obtained as to-be-labeled pictures;
[0037] pixel coordinates of each marker in all to-be-labeled pictures are obtained respectively;
[0038] based on the pixel coordinates of each marker in all to-be-labeled pictures, markers are formed respectively to complete labeling of the to-be-labeled video.
[0039] In a second aspect, the present application provides a video labeling device, including a vector coordinate conversion module, a vector difference obtaining module, and a coordinate generating module;
[0040] the vector coordinate conversion module is configured to obtain an initial vector and initial pixel coordinates of each marker and a vector coordinate conversion relationship based on an initial labeled picture that has been labeled;
[0041] the vector difference obtaining module is configured to obtain a vector difference between a current vector and an initial vector of each marker based on a to-be-labeled picture;
[0042] the coordinate generating module is configured to obtain current pixel coordinates of each marker in the to-be-labeled picture through the vector coordinate conversion relationship based on the initial vector, the initial pixel coordinates, and the vector difference.
[0043] In a third aspect, an electronic device is provided, comprising a processor and a memory, the memory being adapted to store a plurality of program codes, the program codes being adapted to be loaded and run by the processor to execute the video labeling method according to any one of the technical solutions of the video labeling method.
[0044] In a fourth aspect, a computer readable storage medium is provided, the computer readable storage medium having stored therein a plurality of program codes, the program codes being adapted to be loaded and run by a processor to execute the video labeling method according to any one of the technical solutions of the video labeling method.
[0045] The above one or more technical solutions of the present application have at least one or more of the following advantages
[0046] Advantages:
[0047] In the implementation of the technical solutions of the present application, when a fixed object is labeled by video labeling on the picture taken by the rotating zoom camera, real-time detection and labeling of each frame of picture is not required, and the automatic following in the video data labeling can be realized by calculation. The video labeling efficiency is effectively improved, and the use cost is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0048] The disclosure of the present application will become more readily apparent from the following description of the drawings. As those skilled in the art will readily appreciate, the drawings are not intended to limit the present application in any way. Furthermore, like reference numerals are intended to represent similar components throughout the figures. In the drawings:
[0049] Figure 1 is a main step flow chart of the video labeling method of an embodiment of the present application;
[0050] Figure 2 is a three-dimensional schematic diagram of coordinate system conversion of an embodiment of the present application;
[0051] Figure 3 is a schematic diagram of the global coordinate system and the camera coordinate system in the xz plane of an embodiment of the present application;
[0052] Figure 4 is a schematic diagram of the yz plane of the camera coordinate system of an embodiment of the present application;
[0053] Figure 5 is a main structure block diagram of a video labeling device according to an embodiment of the present application;
[0054] Figure 6 is a main structure block diagram of an electronic device for executing the video labeling method of the present application.
[0055] List of reference signs :
[0056] 51: vector coordinate conversion module; 52: vector difference acquisition module; 53: coordinate generation module. DETAILED DESCRIPTION
[0057] Some embodiments of the present application will be described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present application, and are not intended to limit the protection scope of the present application.
[0058] In the description of the present application, "module" and "processor" can include hardware, software or a combination of both. A module can include hardware circuit, various suitable sensors, communication port, memory, and can also include software part such as program code, and can be a combination of software and hardware. The processor can be a central processor, microprocessor, image processor, digital signal processor or any other suitable processor. The processor has data and / or signal processing functions. The processor can be implemented in software, hardware or a combination of both. The non-transitory computer readable storage medium includes any suitable medium that can store program code, such as magnetic disk, hard disk, optical disk, flash memory, read-only memory, random access memory, etc. The term "A and / or B" means all possible combinations of A and B, such as only A, only B or both A and B. The term "at least one of A or B" or "at least one of A and B" has similar meaning as "A and / or B", and can include only A, only B or both A and B. The singular form of the term "one", "this" can also include plural forms.
[0059] Referring to the accompanying drawings Figure 1 , Figure 1 is a schematic diagram of the main steps of a video labeling method according to an embodiment of the present application. The method is suitable for video data labeling on web, pc, mobile terminal and other terminals, and can adapt to various camera video sizes and client player screen sizes.
[0060] As shown in Figure 1 , the video labeling method in the embodiment of the present application mainly includes the following steps S11-S13.
[0061] Step S11, based on the initial labeled picture, the initial vector and initial pixel coordinates of each label are obtained, and the vector coordinate conversion relationship is obtained.
[0062] The initial vector is a vector of each mark in the initial labeled picture, the initial pixel coordinate is a two-dimensional coordinate of each mark in the initial labeled picture on the screen, and the vector coordinate conversion relationship is a conversion relationship corresponding to the vector and the pixel coordinate of each mark. In the same video, the vector coordinate conversion relationship is constant, and thus the vector coordinate conversion relationship can be obtained based on the initial vector and the initial pixel coordinate.
[0063] In one embodiment, based on the labeled initial labeled picture, the initial vector and the initial pixel coordinate of each mark, and the vector coordinate conversion relationship are obtained through the following steps S21-S24:
[0064] In step S21, a frame of picture in the video is obtained as the initial labeled picture, and a camera rotation angle corresponding to the initial labeled picture is obtained.
[0065] In this embodiment, the first frame of picture of the video is taken as the initial labeled picture. In actual operation, any frame of picture can be selected as the initial labeled picture according to actual needs, and the first frame of picture in the example is not limited.
[0066] In step S22, a plurality of marks are made along the contour of the real object to be labeled in the initial labeled picture, and the initial pixel coordinate of each mark in the player is obtained.
[0067] For example, the number of marks can be set according to actual needs by those skilled in the art. In the player screen, the marks can be directly displayed in the form of points, or the plurality of marks can be connected by lines to form a labeled frame for the user to observe and perform subsequent operations.
[0068] Further, the screen size corresponding to the initial labeled picture is obtained, a two-dimensional pixel coordinate system is established based on the screen size corresponding to the initial labeled picture, and the coordinate of each mark in the initial labeled picture in the pixel coordinate system is obtained as the initial pixel coordinate.
[0069] For example, the origin of the pixel coordinate system is set at the lower left corner of the rectangular screen, the x-axis of the pixel coordinate system extends along the lower frame of the screen, and the y-axis extends along the left frame of the screen. In actual operation, the pixel coordinate system can be established according to actual needs by those skilled in the art, and the example is not limited.
[0070] Please refer to Figure 2 , Figure 2is a three-dimensional schematic diagram of coordinate system conversion of an embodiment of the present application. Wherein, the pixel coordinate system is arranged on the plane of the projection screen, and the coordinates in the pixel coordinate system are (sx, sy). When the screen size corresponding to the initial labeled picture is acquired, that is, the screen height (screenHeight) and the screen width (screenWidth) are acquired.
[0071] Wherein, the screen size refers to the screen size of the actually played video. For example, if only part of the large screen is used to play the video, the screen size refers to the size of the part of the screen playing the video.
[0072] For the initial pixel coordinates, they can be acquired by measuring the real object to be labeled on the screen. For the pixel coordinates in the subsequent pictures to be labeled, they can be obtained according to the calculation.
[0073] Step S23, acquiring the initial vector of each marker.
[0074] Specifically, in an embodiment, the acquiring of the initial vector of each marker comprises: establishing a three-dimensional camera coordinate system with the viewpoint of the camera as the origin; acquiring the coordinates of each marker in the camera coordinate system; and acquiring the initial vector of each marker based on the coordinates in the camera coordinate system.
[0075] Please refer to Figure 2 , the camera coordinate system, that is, the coordinate system of the camera visual area, is a three-dimensional coordinate system, the right side of which is the positive direction of the cx axis, the upper side is the positive direction of the cy axis, and the front side is the positive direction of the cz axis. The coordinates in the camera coordinate system are (cx, cy, cz). In the subsequent steps, two dimensions can be used for calculation only for the convenience of calculation. In this embodiment, the vector of the marker is n = (x, y, z). When (cx, cy, cz) is the initial camera coordinate, (x, y, z) is the initial vector.
[0076] The present application does not have restrictions on the vector coordinate system and the origin of the initial vector, and it is only necessary to ensure that the vector coordinate system does not change, but in order to facilitate the subsequent acquisition of the vector coordinate conversion relationship, the camera coordinate system can be selected as the vector coordinate system, the origin of the camera coordinate system can be selected as the origin of the initial vector, and the bases of the vector can be arranged along the positive directions of the x, y, and z axes of the camera coordinate system, respectively.
[0077] Further, in an embodiment, before acquiring the coordinates of each marker in the camera coordinate system, the global coordinates can be acquired first, comprising the following steps:
[0078] establish a three-dimensional global coordinate system and obtain the coordinates of each marker in the global coordinate system; and transform the coordinates of each marker in the global coordinate system into coordinates in a camera coordinate system based on a view matrix transformation according to a rotation angle of the camera corresponding to the initial marker image.
[0079] Please refer to Figure 2 , Figure 2 is a three-dimensional diagram of coordinate system conversion of an embodiment of the application.
[0080] wherein the global coordinate system is a reference coordinate system for subsequent rotation angles of the camera, and is a three-dimensional coordinate system, and the coordinates in the global coordinate system are (gx, gy, gz).
[0081] In one embodiment, the global coordinate system coincides with an initial camera coordinate system, the positive direction of the gx axis is to the right of the initial position of the camera, the positive direction of the gy axis is upward, the positive direction of the gz axis is forward, and the origin is the rotation axis base point, i.e., coincides with the viewpoint of the camera.
[0082] In another embodiment, the global coordinate system can also adopt a geographic coordinate system, and in this case, the global coordinate system does not coincide with the initial camera coordinate system.
[0083] In this embodiment, the global coordinate system coincides with the initial camera coordinate system, and when the camera rotates, the rotation of the camera causes a deviation between the camera coordinate system and the global coordinate system, and the deviation is the rotation angle of the camera. In other embodiments, if the global coordinate system does not coincide with the initial camera coordinate system, the deviations between the camera coordinate system in the current and initial states and the global coordinate system need to be obtained respectively, and the difference between the two is the rotation angle of the camera.
[0084] In one embodiment, the coordinates in the global coordinate system are converted into coordinates in the camera coordinate system, and the view matrix transformation is as follows:
[0085] Please refer to Figure 2 and Figure 3 , Figure 3 is a diagram of the global coordinate system and the camera coordinate system in the xz plane of an embodiment of the application.
[0086] Suppose that a ray is emitted from the viewpoint to the marker position, the ray is a three-dimensional vector, and any point on the ray that hits the screen is at the same position. If the vector is decomposed into xz and yz components, Figure 3 then the xz component is in the plane.
[0087] As Figure 3As shown, taking the view matrix transformation of a marked view as an example, the coordinates of a point on the ray in the global coordinate system are (gx, gz), and the coordinates in the camera coordinate system are (cx, cz). The angle offset between the camera coordinate system and the global coordinate system in the xz plane is θ angle (that is, the angle of the coordinate axis counterclockwise rotation about the cy axis, and the angle in this method follows the right-hand screw rule). In this embodiment, the camera view point and the rotation axis base point are the same point, and then the view matrix transformation relationship formula (1) (2) for converting the coordinates in the global coordinate system into the coordinates in the camera coordinate system in the xz plane can be obtained:
[0088] cx = gx*cosθ + gz*sinθ (1)
[0089] cz = -gx*sinθ + gz*cosθ (2)
[0090] After step S23 is performed, step S24 is continued, and the vector coordinate conversion relationship is obtained based on the camera rotation angle corresponding to the initial marked picture, and the initial vector and initial pixel coordinates of each marker.
[0091] The derivation process of the vector coordinate conversion relationship is as follows:
[0092] If the three-dimensional coordinate point in the camera coordinate system is projected on the screen to obtain a two-dimensional coordinate, the distance from the origin in the xz plane is sx'. In combination with the distance near of the view point from the screen, the relationship formula (3) can be obtained:
[0093] sx' = near*cx / cz (3)
[0094] Further, the distance from the origin is converted into pixel coordinates:
[0095] Suppose kx is the slope of the ray in the xz plane, that is, kx = gx / gz (4); then the following relationship formula (5) can be obtained:
[0096] sx' = near*(kx+tanθ) / (1–kx*tanθ) (5)
[0097] Since sx' is the position information in the camera screen, a proportionality coefficient needs to be added between sx' and the pixel coordinates sx in the client player. Taking the example that the camera video is horizontally stretched and vertically empty in the player, then: the proportionality coefficient coefficient = playerWidth / screenWidth (6)
[0098] That is, the following relationship formula (7) is obtained:
[0099] sx = near * (playerWidth / screenWidth) * (kx+tanθ) / (1-kx*tanθ) (7)
[0100] Similarly, the following relationship (8) can be obtained:
[0101] sy = near * (playerWidth / screenWidth) * (ky+tanβ) / (1-ky*tanβ) (8)
[0102] where ky is the slope of the ray in the yz plane, and β is the angle offset of the camera coordinate system and the global coordinate system in the yz plane.
[0103] If the camera only rotates without scaling during the entire labeling process, the scaling coefficient does not need to be introduced, and the vector coordinate conversion relationship can be obtained through the above relationship (7) (8).
[0104] In another embodiment of the application, when the camera rotates and scales during shooting, the scaling coefficient needs to be introduced, and the method further comprises: obtaining the camera scaling coefficient corresponding to the initial labeling picture; and based on the camera rotation angle and the scaling coefficient corresponding to the initial labeling picture, and the initial vector and the initial pixel coordinates of each marker, obtaining the vector coordinate conversion relationship.
[0105] In one embodiment, when the camera adjusts the focal length, the video content presents the effect of zooming in and out, please refer to the attached Figure 4 , Figure 4 is a schematic view of the yz plane of the camera coordinate system of one embodiment of the application. When the fov (Field of View, field of view) decreases, the video content is stretched to achieve the effect of zooming in, and when the fov increases, the video content is compressed to achieve the effect of zooming out. Thus, it is concluded that when the fov decreases, the video content is zoomed in, and when the fov increases, the video content is zoomed out, and the relationship is obtained as follows:
[0106] screenHeight / 2 = near * tan(fov / 2) (9)
[0107] In this embodiment, fov refers to the vertical field of view.
[0108] The relationship (9) is brought into the relationship (7) to obtain:
[0109] sx = (screenHeight / screenWidth) * (playerWidth / (2*tan(fov / 2))) * (kx+tanθ) / (1-kx*tanθ) (10)
[0110] Let the scaling factor rate = (screenHeight / screenWidth)*playerWidth / 2,
[0111] The scaling factor rate is consistent in the xz plane and the yz plane, and then:
[0112] sx = (rate / tan(0.5*fov))*(kx+tanθ) / (1-kx*tanθ) (11)
[0113] kx = (sx*tan(0.5*fov)-rate*tanθ) / (rate+sx*tan(0.5*fov)*tanθ) (12)
[0114] The corresponding ky ray in the yz plane, the angle offset β of the camera coordinate system and the global coordinate system in the yz plane, and then:
[0115] sy = (rate / tan(0.5*fov))*(ky+tanβ) / (1-ky*tanβ) (13)
[0116] ky = (sy*tan(0.5*fov)-rate*tanβ) / (rate+sy*tan(0.5*fov)*tanβ) (14)
[0117] Through the above relationship (11)-(14), the vector coordinate conversion relationship of the embodiment can be obtained. Specifically, the initial vector of each marker, the initial pixel coordinate can be brought into the relationship (11)-(14), and the slope kx, ky can be calculated to obtain the vector coordinate conversion relationship in the video to be labeled
[0118] After the above step is completed, step S12 is continued, and the vector difference between the current vector and the initial vector of each marker is obtained based on the to-be-labeled picture.
[0119] In one embodiment, obtaining the vector difference between the current vector and the initial vector of each marker based on the to-be-labeled picture comprises:
[0120] Obtaining the camera rotation angle corresponding to the to-be-labeled picture;
[0121] Based on the change of the camera rotation angle corresponding to the to-be-labeled picture and the camera rotation angle corresponding to the initial labeled picture, the vector difference between the current vector and the initial vector of each marker is obtained.
[0122] Since the real objects to be labeled are fixed real objects such as buildings, trees, and streets, and the camera only rotates around the viewpoint, the vector difference can be directly obtained based on the change of the rotation angle of the camera. In the embodiment, the mark corresponding to the rotation angle of the camera of the labeling picture and the rotation angle of the camera corresponding to the initial labeling picture can be directly obtained by the camera.
[0123] After the vector difference is obtained, the current vector of each mark in the picture to be labeled can be directly obtained through the vector difference and the initial vector.
[0124] In step S13, the current pixel coordinates of each mark in the picture to be labeled are obtained through the vector coordinate conversion relationship based on the initial vector, the initial pixel coordinates, and the vector difference.
[0125] Further, the method further comprises: obtaining the screen size corresponding to the picture to be labeled; and obtaining the current pixel coordinates of each mark in the picture to be labeled through the vector coordinate conversion relationship based on the initial vector, the initial pixel coordinates, the vector difference, and the screen size corresponding to the picture to be labeled.
[0126] Specifically, the obtained initial vector and initial pixel coordinates of each mark are brought into the relationship (11)-(14), and then the slope kx and ky can be calculated to obtain the vector coordinate conversion relationship in the video to be labeled. Then, the current vector is obtained through the vector difference, and the current vector is brought into the above vector coordinate conversion relationship, and then the current pixel coordinates of each mark in the picture to be labeled can be obtained.
[0127] Further, in an embodiment, the method further comprises:
[0128] obtaining one or more pictures in the video to be labeled as the picture to be labeled;
[0129] obtaining the pixel coordinates of each mark in all pictures to be labeled;
[0130] forming marks based on the pixel coordinates of each mark in all pictures to be labeled to complete the labeling of the video to be labeled.
[0131] The video labeling method of the application is simple and easy to use, and can be used not only for labeling of complete videos, but also for real-time labeling of videos taken by a single camera monitoring device, and any frame can be conveniently viewed for subsequent operation by a user.
[0132] Based on the above steps S11-S13, the video annotation is performed on the picture taken by the rotating zoom camera, so that when the fixed real object is marked, it is not necessary to detect and mark each frame of picture in real time, and the automatic marking following can be realized in the video data annotation through calculation. The video marking efficiency is effectively improved, and the use cost is reduced.
[0133] The above technical solution only needs to mark the initial marking picture once through manual marking or AI recognition. As long as the to-be-marked real object does not change, the subsequent camera rotation and zoom can quickly calculate and obtain the marking position through the above method. Compared with the real-time AI recognition, the speed is faster, and the accuracy is higher. It can be applied to marking buildings, roads, areas and the like in high-altitude hawk-eye cameras, and the efficiency is higher.
[0134] It should be noted that although the steps are described in a specific order in the above embodiments, those skilled in the art can understand that, in order to achieve the effect of the present application, the different steps do not have to be executed in such an order, they can be executed at the same time (in parallel) or in other orders, and these changes are within the protection scope of the present application.
[0135] Further, the present application also provides a video annotation device.
[0136] Reference is made to the accompanying drawings Figure 5 , Figure 5 is the main structure block diagram of the video annotation device according to an embodiment of the present application. As shown in Figure 5 , the video annotation device in the embodiment of the present application mainly includes a vector coordinate conversion module 51, a vector difference obtaining module 52 and a coordinate generating module 53. In some embodiments, one or more of the vector coordinate conversion module 51, the vector difference obtaining module 52 and the coordinate generating module 53 can be combined together to become one module.
[0137] In some embodiments, the vector coordinate conversion module 51 is configured to obtain an initial vector and an initial pixel coordinate of each marker in the marked initial marking picture, and a vector coordinate conversion relationship based on the initial marking picture; the vector difference obtaining module 52 is configured to obtain a vector difference between a current vector and the initial vector of each marker based on the to-be-marked picture; and the coordinate generating module 53 is configured to obtain a current pixel coordinate of each marker in the to-be-marked picture based on the initial vector, the initial pixel coordinate and the vector difference through the vector coordinate conversion relationship.
[0138] In one embodiment, the description of the specific implementation function can be referred to the steps S11-S13.
[0139] The above video annotation device is used to execute Figure 1The video labeling method embodiment shown, the technical principles, the technical problems solved and the technical effects generated are similar, and the person skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process and related description of the video labeling device can refer to the description of the video labeling method embodiment, which will not be repeated here.
[0140] Further, it should be understood that, since the setting of each module is only for illustrating the functional units of the system of the present application, the physical device corresponding to the module can be the processor itself, or a part of software, hardware or a combination of software and hardware in the processor. Therefore, the number of each module in the figure is only illustrative.
[0141] The person skilled in the art can understand that each module in the device can be adaptively split or combined. Such splitting or combining of specific modules does not cause the technical solution to deviate from the principles of the present application, therefore, the technical solution after splitting or combining will fall within the protection scope of the present application.
[0142] The person skilled in the art can understand that all or part of the processes in the method of the above embodiment of the present application can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer readable storage medium can include any entity or device, medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code. It should be noted that the contents included in the computer readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable storage medium does not include electrical carrier signals and telecommunication signals.
[0143] Further, the present application also provides an electronic device. Please refer to the attached Figure 6 , Figure 6 is the main structure block diagram of the electronic device for executing the video labeling method of the present application.
[0144] As Figure 6As shown, in an electronic device embodiment according to the present application, the electronic device includes a processor 601 and a memory 602, the memory 602 can be configured to store program code 603 of the video annotation method of the above-mentioned method embodiments, and the processor 601 can be configured to execute the program code 603 in the memory 602, which includes but is not limited to the program code 603 of the video annotation method of the above-mentioned method embodiments. For the convenience of illustration, only the parts related to the embodiments of the present application are shown, and for the technical details not disclosed, please refer to the method part of the embodiments of the present application.
[0145] Further, the present application also provides a computer readable storage medium. In a computer readable storage medium embodiment according to the present application, the computer readable storage medium can be configured to store a program of the video annotation method of the above-mentioned method embodiments, which can be loaded and run by a processor to implement the above-mentioned video annotation method. For the convenience of illustration, only the parts related to the embodiments of the present application are shown, and for the technical details not disclosed, please refer to the method part of the embodiments of the present application. The computer readable storage medium can be a storage device formed by various electronic devices, and optionally, the computer readable storage medium in the embodiments of the present application is a non-transitory computer readable storage medium.
[0146] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical solutions after these changes or replacements will all fall within the protection scope of the present application.
Claims
1. A video annotation method, characterized by, The method comprises: Based on the annotated initial annotation picture, the initial vector and the initial pixel coordinates of each mark are obtained, and the vector coordinate conversion relationship is obtained, which comprises: A frame of picture in the video is obtained as an initial annotation picture, and the camera rotation angle corresponding to the initial annotation picture is obtained; The camera scaling coefficient corresponding to the initial annotation picture is obtained; and Based on the camera rotation angle and the scaling coefficient corresponding to the initial annotation picture, and the initial vector and the initial pixel coordinates of each mark, the vector coordinate conversion relationship is obtained; Based on the to-be-annotated picture, the vector difference between the current vector and the initial vector of each mark is obtained; Based on the initial vector, the initial pixel coordinates and the vector difference, the current pixel coordinates of each mark in the to-be-annotated picture are obtained through the vector coordinate conversion relationship.
2. The method of claim 1, wherein, Based on the annotated initial annotation picture, the initial vector and the initial pixel coordinates of each mark are obtained, which comprises: A plurality of marks are made along the contour of the to-be-annotated real object in the initial annotation picture, and the initial pixel coordinates of each mark in the player are obtained; The initial vector of each mark is obtained.
3. The method of claim 2, wherein, The initial vector of each mark is obtained, which comprises: A three-dimensional camera coordinate system is established with the viewpoint of the camera as the origin; The coordinates of each mark in the camera coordinate system are obtained; Based on the coordinates in the camera coordinate system, the initial vector of each mark is obtained.
4. The method of claim 2, wherein, Based on the to-be-annotated picture, the vector difference between the current vector and the initial vector of each mark is obtained, which comprises: The camera rotation angle corresponding to the to-be-annotated picture is obtained; Based on the change of the camera rotation angle corresponding to the to-be-annotated picture and the camera rotation angle corresponding to the initial annotation picture, the vector difference between the current vector and the initial vector of each mark is obtained.
5. The method of claim 1, wherein, The method further comprises: The screen size corresponding to the initial annotation picture is obtained; A two-dimensional pixel coordinate system is established based on the screen size corresponding to the initial annotation picture; The coordinates of each mark in the initial annotation picture in the pixel coordinate system are obtained as initial pixel coordinates; And / or, The screen size corresponding to the to-be-annotated picture is obtained; Based on the initial vector, the initial pixel coordinates, the vector difference and the screen size corresponding to the to-be-annotated picture, the current pixel coordinates of each mark in the to-be-annotated picture are obtained through the vector coordinate conversion relationship.
6. The method of claim 1, wherein, The method further comprises: One or more frames of pictures in the to-be-annotated video are obtained as to-be-annotated pictures; The pixel coordinates of each mark in all to-be-annotated pictures are obtained respectively; Based on the pixel coordinates of each mark in all to-be-annotated pictures, marks are formed respectively to complete the annotation of the to-be-annotated video.
7. A video annotation apparatus characterized by comprising: The method comprises a vector coordinate conversion module, a vector difference acquisition module and a coordinate generation module; The vector coordinate conversion module is configured to obtain an initial vector and an initial pixel coordinate of each marker based on the labeled initial labeled picture, and obtain a vector coordinate conversion relationship, including: obtaining a picture in a frame of video as the initial labeled picture, and obtaining a camera rotation angle corresponding to the initial labeled picture; obtaining a camera scaling coefficient corresponding to the initial labeled picture; and obtaining the vector coordinate conversion relationship based on the camera rotation angle and the scaling coefficient corresponding to the initial labeled picture, and the initial vector and the initial pixel coordinate of each marker; The vector difference obtaining module is configured to obtain a vector difference between a current vector and an initial vector of each marker based on the to-be-labeled picture; The coordinate generating module is configured to obtain a current pixel coordinate of each marker in the to-be-labeled picture based on the initial vector, the initial pixel coordinate, and the vector difference through the vector coordinate conversion relationship.
8. An electronic device comprising a processor and a memory, the memory being adapted to store a plurality of program codes, characterized in that, The program code is adapted to be loaded and run by the processor to execute the video labeling method of any one of claims 1 to 7.
9. A computer readable storage medium having stored therein a plurality of program codes, characterized in that, The program code is adapted to be loaded and run by the processor to execute the video labeling method of any one of claims 1 to 7.
Citation Information
Patent Citations
Video tag adjusting method, system and equipment
CN110298889A