A dynamic multi-target labeling method, system and device

Through dynamic multi-target labeling method and hardware assistance, video data labeling is automatically processed, solving the problems of low manual labeling efficiency and low accuracy of target tracking algorithms, and achieving efficient and accurate video data labeling.

CN114820694BActive Publication Date: 2025-09-02CHANGZHOU XINGYU AUTOMOTIVE LIGHTING SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110037869.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-12
Publication Date
2025-09-02
Estimated Expiration
2041-01-12

AI Technical Summary

Technical Problem

Existing image and video data annotations rely on manual operation efficiency, and simple target tracking algorithms have low accuracy and poor data quality in target locations where motion changes are complex.

Method used

The dynamic multi-objective annotation method is adopted to read video records, extract the logo frame images, mark the target location and types, calculate the target location and domain of interest of successive frames, and generate record files. The target motion trajectory is predicted using the Lucas-Kanade optical flow algorithm, and automatically annotated with hardware devices such as external hard disk seats, video editing keyboards and dual display screens.

Benefits of technology

It improves the accuracy and labeling efficiency of the domain of interest at the target location, realizes automated video data labeling, and improves data quality and labeling speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114820694B_ABST
    Figure CN114820694B_ABST
Patent Text Reader

Abstract

The present invention provides a dynamic multi-target labeling method, system and device. The labeling method includes the following steps: reading video records; extracting marker frame images; marking the target position interest domain and type of the marker frame; calculating the target position interest domain of continuous frames; after labeling, generating images of the marker frame and the continuous frames, and generating a record file. The method solves the problems of manual labor and low efficiency in existing video data labeling, and low accuracy and poor data quality of existing simple target tracking algorithms for target position interest domains with complex motion changes. The method can automatically realize video data labeling, and improve the accuracy and labeling efficiency of the target position interest domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a dynamic multi-target labeling method, system and device supporting customized motion trajectories. Background Art

[0002] Products utilizing artificial intelligence (AI) image recognition technology are increasingly being used across various industries. Currently, image recognition technology is primarily achieved through training on massive amounts of labeled image and video data. However, existing image and video data annotation involves manually selecting the target's location and region of interest (ROI) for each frame and adding relevant classification information. This is labor-intensive and extremely inefficient.

[0003] Although auxiliary annotation tools using simple target tracking algorithms (such as interpolation algorithms) have been released, for targets with complex motion changes, there are problems such as low accuracy of target location and poor data quality. Summary of the Invention

[0004] The present invention provides a dynamic multi-target labeling method, system and device, which solves the problems that existing video data labeling is time-consuming and labor-intensive and inefficient, and existing simple target tracking algorithms have low accuracy and poor data quality for target locations with complex motion changes. The method can automatically implement video data labeling, and improve the accuracy and labeling efficiency of target location domains of interest.

[0005] To achieve the above object, the technical solution of the present invention is specifically implemented as follows:

[0006] In one aspect, the present invention discloses a dynamic multi-target labeling method, comprising the following steps:

[0007] Read video records;

[0008] Extracting the marker frame image;

[0009] Mark the target position, domain of interest and type of the marker frame;

[0010] Calculate the target position region of interest in consecutive frames;

[0011] After the marking is completed, the marked frame and continuous frames are generated into images and a record file is generated.

[0012] Furthermore, the calculation of the target position region of interest of the continuous frames includes the following steps:

[0013] Convert the RGB images of the marker frame and the continuous frames into grayscale images;

[0014] In the region of interest of the grayscale image of the marker frame, it is necessary to find the corner points with high brightness;

[0015] In the continuous frame grayscale image, find new corner points with the positions of each corner point of the marker frame as the center of the circle;

[0016] Calculate the brightness deviation and gradient change direction between the corner points of the continuous frames and the corner points of the landmark frame, and predict the position of the corner points of the next frame;

[0017] Calculate the target coordinates based on the positions of multiple new corner points and find their center point as the motion trajectory point;

[0018] The motion trajectory points are brought into the next continuous frame grayscale image to calculate the corner point position, target coordinates and trajectory points.

[0019] On the other hand, the present invention discloses a dynamic multi-target labeling system, which includes a video reading module, an extraction module, a marking module, a calculation module and an image generation module, wherein the video reading module is used to read video records; the extraction module is used to extract marker frame images; the marking module is used to mark the target position interest domain and type of the marker frame; the calculation module is used to calculate the target position interest domain of continuous frames; the image generation module is used to generate images from the marker frame and continuous frames, and generate record files.

[0020] On the other hand, the present invention discloses a dynamic multi-target labeling device, including an external hard disk seat, a video editing keyboard, a mouse, a keyboard, a computer host, a left display screen and a right display screen, wherein the external hard disk seat is used to read video records; the video editing keyboard is used to extract marker frame images; the mouse is used to manually modify the target's motion trajectory to improve the accuracy of the target position domain of interest; the keyboard and the mouse cooperate to mark the target position domain of interest and type of the marker frame; the computer host is used to calculate and process the target position domain of interest of continuous frames; the left display screen is used to display video records; and the right display screen is used to display the target's motion trajectory and target position domain of interest.

[0021] Furthermore, the marker frame is a video frame in which the motion trajectory changes significantly.

[0022] Furthermore, the video frames between the two marker frames are continuous frames.

[0023] Beneficial technical effects:

[0024] 1. The present invention discloses a dynamic multi-target labeling method, comprising the following steps:

[0025] Read video records;

[0026] Extracting the marker frame image;

[0027] Mark the target position, domain of interest and type of the marker frame;

[0028] Calculate the target position region of interest in consecutive frames;

[0029] After the annotation is completed, the marked frame and continuous frames are generated into images and a record file is generated. This solves the problems of existing video data annotation, which is time-consuming and labor-intensive and inefficient, and the low accuracy and poor data quality of the existing simple target tracking algorithm for target locations with complex motion changes. It can automatically implement video data annotation, improving the accuracy and efficiency of target location areas of interest.

[0030] 2. The present invention discloses a dynamic multi-target labeling device, comprising an external hard disk seat, a video editing keyboard, a mouse, a keyboard, a computer host, a left display screen and a right display screen, wherein the external hard disk seat is used to read video records; the video editing keyboard is used to extract marker frame images; the mouse is used to manually modify the target's motion trajectory to improve the accuracy of the target position domain of interest; the keyboard and the mouse cooperate to mark the target position domain of interest and type of the marker frame; the computer host is used to calculate and process the target position domain of interest of continuous frames; the left display screen is used to display video records; the right display screen is used to display the target's motion trajectory and target position domain of interest; using the video editing keyboard, mouse and dual display screens, both hands can be operated simultaneously, which effectively improves the labeling efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0032] Figure 1 A schematic diagram of a dynamic multi-target labeling device provided by an embodiment of the present invention;

[0033] Figure 2 A flowchart of a dynamic multi-target labeling method provided by an embodiment of the present invention.

[0034] Among them, 1-external hard drive holder, 2-computer host, 3-video editing keyboard, 4-mouse, 5-keyboard, 6-left display screen, 7-right display screen. DETAILED DESCRIPTION

[0035] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0036] In addition, it should be noted that the use of terms such as "first" and "second" to limit components is only for the convenience of distinguishing the corresponding components. Unless otherwise stated, the above terms have no special meaning and therefore cannot be understood as limiting the scope of protection of the present invention.

[0037] Unless otherwise specifically stated, the relative arrangement of the parts and steps, the numerical expressions and the numerical values ​​set forth in these embodiments do not limit the scope of the present invention. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The techniques, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific values ​​should be interpreted as being merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar numbers and letters represent similar items in the following figures, and therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.

[0038] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0039] The present invention discloses a method for dynamic multi-target marking, which specifically includes the following steps: Figure 2 :

[0040] S1: read video records;

[0041] Read video records via external hard drive;

[0042] S2: logo frame image;

[0043] Use the knob button on the video editing keyboard to extract the marker frame image on the left display screen. The marker frame is a video frame in which the target motion in the video changes significantly.

[0044] S3: Mark the target position, domain of interest and type of the marker frame;

[0045] Use the keyboard and mouse to manually mark the target location of the landmark frame and the area of ​​interest and category.

[0046] S4: Calculate the target position region of interest in consecutive frames;

[0047] The video frames between two marker frames are continuous frames.

[0048] S5: Once labeling is complete, images are generated for the marked frame and consecutive frames, and a record file is created. After labeling is complete, images are generated for each video frame, and record files of various target locations and types are generated. This can be used as a training or testing image database and substituted into an AI image recognition model to complete model training.

[0049] As an embodiment of the present invention, the step of calculating the target position region of interest of consecutive frames includes:

[0050] Convert the RGB images of the marker frame and the continuous frames into grayscale images;

[0051] Specifically, the grayscale value is calculated based on the RGB value of each pixel:

[0052] Gray=(R*30+G*587+B*11+50) / 100

[0053] Find multiple high-brightness corner points within the region of interest (the rectangular area enclosed by the upper left corner coordinates and the lower right corner coordinates) of the marker frame grayscale image and save the positions of these corner points.

[0054] In the continuous frame grayscale image, the position of each corner point of the marker frame is used as the center of the circle, and new corner points are found within a certain range;

[0055] According to the LK (Lucas-Kanade) optical flow algorithm, the brightness deviation and gradient change direction of the corner points of the continuous frames and the landmark frame are calculated, and the position of the corner points in the next frame is predicted;

[0056] Calculate the target coordinates based on the positions of multiple new corner points and find their center point as the motion trajectory point. The calculation method is as follows:

[0057]

[0058]

[0059] Substitute the next continuous frame grayscale image, calculate the corner point position, target coordinates and trajectory points, and terminate at the next marker frame.

[0060] It should be noted that if the calculated target position is deviated, the trajectory point or target position can be manually modified, and the targets of subsequent consecutive frames will be automatically corrected according to the optical flow algorithm.

[0061] Another aspect of the present invention discloses a system for dynamic multi-target labeling, including a video reading module, an extraction module, a marking module, a calculation module and an image generation module, wherein the video reading module is used to read video records; the extraction module is used to extract marker frame images; the marking module is used to mark the target position interest domain and type of the marker frame; the calculation module is used to calculate the target position interest domain of continuous frames; the image generation module is used to generate images from the marker frame and continuous frames, and generate record files.

[0062] Another aspect of the present invention discloses a device for dynamic multi-target labeling, see Figure 1 , specifically including an external hard disk seat, a video editing keyboard, a mouse, a keyboard, a computer host, a left display screen and a right display screen, wherein the external hard disk seat is used to read the video record; the video editing keyboard is used to extract the marker frame image; the mouse is used to manually modify the target's motion trajectory to improve the accuracy of the target position interest domain; the keyboard and the mouse are used to mark the target position interest domain and type of the marker frame; the computer host is used to calculate and process the target position interest domain of consecutive frames; the left display screen is used to display the video record; the right display screen is used to display the target's motion trajectory and target position interest domain. Specifically, the hard disk with video data recorded is connected to the computer host through the external hard disk seat. After starting the video annotation program, the left display screen displays the video track, marker Mark frame image and annotation frame information. Use the knob on the video editing keyboard to quickly browse the video track and select the marker frame image where the target motion shows a significant change. Select the target's location interest region on the image and enter the target type and related attributes. The annotation program uses the Lucas-Kanade optical flow algorithm by default to calculate the motion trajectory and target position of consecutive frames (between the starting frame and this marker frame). The right display screen will use continuous bubble points to display the motion trajectory and the corresponding target location interest region box. If the target selection is found to be inaccurate, you can use the mouse to click on the trajectory point to modify it. After the modification, the new target location interest region will also be highlighted. You can also use the knob on the left display screen to select the frame with inaccurate location interest region annotation in this frame and manually modify the target position. Other consecutive frames are recalculated according to the LK optical flow algorithm to generate new trajectories and target positions. After the entire segment annotation is completed and confirmed, lock the annotation of one target, return to the starting position of the video, and repeat the above steps to annotate other targets.

[0063] It should be noted that for the target of the marker frame, the information that needs to be marked includes: the coordinates of the upper left corner of the target, the coordinates of the lower right corner and the type of the target.

[0064] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0065] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0066] The above embodiments are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineers and technicians in this field should fall within the scope of protection determined by the claims of the present invention.

Claims

1. A dynamic multi-target labeling method, characterized in that: The following steps are involved: Read video records; Extracting the marker frame image; Mark the target position, domain of interest and type of the marker frame; Calculate the target position region of interest in consecutive frames; After the marking is completed, the marked frame and the continuous frames are generated into images and a record file is generated; The marker frame is a video frame in which the motion trajectory changes significantly; The video frames between the two marker frames are continuous frames; Calculating the target position region of interest of consecutive frames comprises the following steps: Convert the RGB images of the marker frame and the continuous frames into grayscale images; In the region of interest of the grayscale image of the marker frame, it is necessary to find the corner points with high brightness; In the continuous frame grayscale image, find new corner points with the positions of each corner point of the marker frame as the center of the circle; Calculate the brightness deviation and gradient change direction between the corner points of the continuous frames and the corner points of the landmark frame, and predict the position of the corner points of the next frame; Calculate the target coordinates based on the positions of multiple new corner points and find their center point as the motion trajectory point; The motion trajectory points are brought into the next continuous frame grayscale image to calculate the corner point position, target coordinates and trajectory points.

2. A dynamic multi-target labeling system, characterized in that: include: Video reading module, used to read video records; An extraction module, used for extracting a marker frame image; a marking module for marking the target position, domain of interest and type of the marker frame; A calculation module for calculating the target position region of interest of consecutive frames; An image generation module is used to generate images from the marker frame and the continuous frames and generate a record file; The marker frame is a video frame in which the motion trajectory changes significantly; The video frames between the two marker frames are continuous frames; Calculating the target position region of interest of consecutive frames comprises the following steps: Convert the RGB images of the marker frame and the continuous frames into grayscale images; In the region of interest of the grayscale image of the marker frame, it is necessary to find the corner points with high brightness; In the continuous frame grayscale image, find new corner points with the positions of each corner point of the marker frame as the center of the circle; Calculate the brightness deviation and gradient change direction between the corner points of the continuous frames and the corner points of the landmark frame, and predict the position of the corner points of the next frame; Calculate the target coordinates based on the positions of multiple new corner points and find their center point as the motion trajectory point; The motion trajectory points are brought into the next continuous frame grayscale image to calculate the corner point position, target coordinates and trajectory points.

3. A dynamic multi-target labeling device for the dynamic multi-target labeling method according to claim 1, characterized in that: include: External hard drive dock for reading video records; Video editing keyboard, used to extract the marker frame image; Mouse, used to manually modify the target's trajectory to improve the accuracy of the target's location in the region of interest; A keyboard, which cooperates with the mouse to mark the target position, field of interest and type of the marker frame; A computer host for calculating and processing target position regions of interest in successive frames; The left display screen is used to display video records; The display screen on the right is used to display the target's motion trajectory and the target position area of ​​interest.

Citation Information

Patent Citations

  • Identity labeling method of face images and face identity recognition method of face images

    CN103793697A

  • Real-time continuous frame embedded information recognition system

    CN108833964A