Artificial intelligence rapid target labeling method

By manually selecting the target of the first frame image in the video and using the feature cross-correlation coupling detector neural network, the automatic labeling of the target in the video is achieved, solving the problems of high and low human labeling costs, improving data labeling efficiency and reducing costs.

CN119992415APending Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510060179.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, artificial intelligence target recognition technology relies on a large number of marked data sets, and manpower labeling is costly and inefficient. Especially when the amount of data is huge, a large amount of manpower is required to perform high-intensity data labeling.

Method used

An artificial intelligence fast target labeling method is adopted to manually select the target to be labeled through the first frame image of the video data, and a symmetric feature cross-correlation coupling detector neural network is used to automatically mark the position, shape and size of the target in the video, and generate the JSON file record labeling results.

Benefits of technology

Automatic labeling of targets in videos is achieved, reducing manpower operations, improving data labeling efficiency, reducing labor costs, and can be used for video files obtained by imagers of multiple bands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992415A_ABST
    Figure CN119992415A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision and artificial intelligence, in particular to an artificial intelligence rapid target labeling method, which comprises the following steps that: in the traditional artificial intelligence recognition technology, different types of targets in a mass data set obtained by splitting frames of a video need to be manually selected and labeled in a manual mode; for different types of targets in a picture data set obtained through video frame splitting, the coordinate position of each target in pictures obtained through frame splitting is automatically calculated through artificial intelligence, the size of a labeling box is adaptively calculated according to the actual size of the target in the pictures, and the target to be labeled only needs to be selected in the video once; no matter how many massive pictures can be split into frames in the video source, the follow-up labeling of the selected target can be automatically completed, and the target labeling efficiency is greatly improved; any model pre-training does not need to be carried out on the target needing to be labeled, and the method can be used for targets of any category and any scale.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and artificial intelligence, and in particular to an artificial intelligence fast target labeling method. Background Art

[0002] Artificial intelligence is a new type of productivity technology that is highly data-driven. Neural networks achieve reliable and stable autonomous work by learning from massive data sets.

[0003] At present, artificial intelligence-based target recognition technology is the most widely used and most reliable technology in the field of industrial vision. It has been widely deployed in various industrial vision occasions such as urban monitoring, road monitoring, public security arrests, and personnel verification at stations and airports. It has greatly improved the work efficiency that could only rely on manpower to identify targets in the past. The core foundation for the widespread application of artificial intelligence-based target recognition technology is to continuously provide the neural network with a steady stream of data sets of labeled target categories and target areas. Only when the data set continues to increase can target recognition become more and more accurate. However, most of the work of adding annotations to the data set is still done by manual annotation. In actual operation, first, a long video is shot by a well-arranged camera, and then a single picture is obtained by decomposing the video frame. Then, the data annotator goes through these pictures one by one and uses a special marking tool to mark the targets in the picture one by one. Calculated at a frame rate of 30 frames per second for a single camera, 1 second of video contains 30 separate pictures, and 1 minute of video can obtain 1,800 pictures. By analogy, shooting a few hours of video can obtain hundreds of thousands or even tens of millions of pictures. Such a large amount of data is a huge workload for manual annotation. Therefore, many artificial intelligence companies have recruited a large number of full-time or part-time data annotators to carry out high-intensity data annotation work, and the labor cost is extremely high.

[0004] At the same time, after obtaining a certain amount of labeled data sets, some companies will first use them to train neural networks, and then use the trained networks to identify unlabeled image sets. After labeling the untrained data sets with the training results, they will send them to the network for incremental training. This also improves the efficiency of data labeling and saves labor costs. However, due to the small amount of initial training data sets, the recognition and labeling method will lead to misidentification and missed recognition in the early stages of training. In the early stages of training, human intervention is still required to pay attention to the correctness of data labeling. Therefore, the demand for manpower consumption by this labeling method is inversely proportional to the amount of labeled image data. Only when the amount of data increases to a certain amount can manpower be freed up. Summary of the invention

[0005] The purpose of the present invention is to provide a rapid target labeling method based on artificial intelligence to solve the problems raised in the above background technology.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] An artificial intelligence fast target labeling method, the method comprising:

[0008] S100, manually selecting a target to be labeled based on the first frame image of the video data and segmenting it separately;

[0009] S200, sending the segmented target image and the whole image to the upper and lower convolutional neural networks respectively for feature calculation to confirm the position of the target image in the original image;

[0010] S300: In the first frame image, based on the position of the target image in the original image, the absolute coordinates and size relationship of the target in the subsequent frames are confirmed.

[0011] Preferably, a symmetrical feature cross-correlation coupled detector neural network is proposed, which does not pay attention to the selected target category, but only calculates the own features of the target in the selected picture and the coordinate position of the feature with the highest similarity in the subsequent frames of the video. At the same time, according to the actual shape and size of the target in the current frame, the center coordinates of the target and the coordinates of the vertices of the rectangular annotation box closest to the size and shape of the target are jointly given. Figure 1 The corresponding form of the annotation is recorded in the JSON file.

[0012] Preferably, S100 includes:

[0013] S101, acquiring surveillance video data, and manually selecting a target to be marked in the first frame of the video; the target to be marked is selected based on the criterion that it cannot be recognized by common recognition technology;

[0014] S102, the target to be marked is segmented separately from the original image, and based on the segmented image, the starting center coordinates (X0, Y0), the left vertex coordinates (X1, Y1) of the selection box, and the current scale Δx and Δy of the target in the first frame are obtained.

[0015] Preferably, the convolutional neural network in S200 includes:

[0016] The convolutional neural network consists of 4 layers of convolution kernels. The first layer uses a 7×7 convolution kernel with a total of 64 dimensions, the second layer uses a 3×3 convolution kernel with a total of 256 dimensions, the third layer uses a 3×3 convolution kernel with a total of 512 dimensions, and the fourth layer uses a 1×1 convolution kernel with a total of 256 dimensions.

[0017] Preferably, S200 includes:

[0018] S201, sending the segmented target image and the whole image to upper and lower convolutional neural networks respectively, wherein the upper and lower convolutional neural networks are completely identical networks, and coupling of the two images is achieved in the calculation by parameter sharing;

[0019] S202, after the image passes through the complete 4-layer neural network, the segmented image is obtained with a size of 15×15 and a feature map array with 256 channels, and the original image is obtained with a size of 31×31 and a feature map array with 256 channels; and the feature map arrays of the segmented image and the original image are subjected to channel cross-correlation calculation to confirm the cross-correlation of the feature map of each channel;

[0020] S203. A 17×17 feature vector set with a dimension of 256 is obtained through channel cross-correlation calculation. Each 1×1×256 feature vector in the feature vector set is a cross-correlation calculation value of the feature of the segmented image appearing at the current position in the set, and the position of an arbitrarily segmented target in the original image is obtained.

[0021] Preferably, S300 includes:

[0022] S301, according to the starting center coordinates (X0, Y0) of the target in the first frame, the coordinates of the left vertex of the selection box (X1, Y1), and the current scale Δx, Δy of the target, as well as the relative coordinates (ΔX0, ΔY0) in the feature vector set, the relative coordinates of the left vertex of the selection box (ΔX1, ΔY1), and the relative scale Δdx, Δdy of the target, the absolute coordinates and size relationship of the target in subsequent frames can be quickly calculated according to the following formula:

[0023]

[0024] Where Xc and Yc are the center coordinates of the target in the subsequent frame, Xd and Yd are the left vertex coordinates of the target's annotation box, and Xs and Ys are the actual scales of the target in the subsequent frame.

[0025] S302, after the target to be marked is selected in the first frame, the network will record the target in real time frame by frame when it appears at any position in the video.

[0026] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the above-mentioned artificial intelligence rapid target labeling method.

[0027] A computer device includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the program, the steps in the above-mentioned artificial intelligence rapid target labeling method are implemented.

[0028] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0029] The core idea of ​​the present invention is that any target in any video image can be automatically labeled in any picture in the current video, completely eliminating the need to manually label a single image after splitting the frame. It only needs to read in the complete video to be labeled and select the target to be labeled on the first frame of the video. The subsequent target labeling process is automatically completed by the method of the present invention. When the complete video ends, a JSON file of the labeled box coordinates of the selected target in each frame of the video and the corresponding video split frame picture are automatically generated to form a data set. This method can be used to label a data set from scratch, and can also be used to supplement the labeling of targets that need to be added, which is convenient and fast. The method of the present invention can be applied to video files acquired by imagers of multiple bands such as visible light and infrared light for rapid labeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0031] Figure 1 It is a structural diagram of a symmetrical characteristic coupling cross-correlation neural network of the present invention;

[0032] Figure 2 It is a schematic diagram of different shapes and sizes of targets in the cross-correlation value calculation feature of the present invention. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] See also Figure 1-Figure 2 , the present invention provides a technical solution:

[0035] Embodiment 1:

[0036] The invention discloses an artificial intelligence fast target labeling method, which is a non-deterministic target tracking technology based on artificial intelligence. Its core framework is a symmetrical feature cross-correlation coupled detector neural network.

[0037] The network does not pay attention to the selected target category, but only calculates the features of the target in the selected image and the coordinate position of the feature with the highest similarity in the subsequent frames of the video. At the same time, according to the actual shape and size of the target in the current frame, the center coordinates of the target and the coordinates of the vertices of the rectangular annotation box that is closest to the size and shape of the target are jointly given. Figure 1 The corresponding form of the annotation is recorded in the JSON file.

[0038] The network structure of the feature cross-correlation coupling detector neural network is as follows Figure 1 As shown:

[0039] There are three targets in the current image, among which the vehicle target and the person riding the electric bike are common targets. The commonly used recognition technologies can complete the recognition of these two targets, but the third target driving an agricultural tricycle is not common in daily life, so the commonly used recognition technology cannot recognize this target. This technology can be used to quickly mark this target.

[0040] Figure 1 In the figure, 1 is the target segmentation image to be marked in the first frame image, 2 is the complete original image, 3 is the backbone neural network structure of the present invention, 4 is the channel cross-correlation calculation, and 5 is the result feature vector set.

[0041] Specific methods include:

[0042] S100, manually selecting a target to be labeled based on the first frame image of the video data and segmenting it separately;

[0043] There is no need to manually annotate the images in the video frame by frame, but the position, shape and size of the target to be marked in each frame of the video can be calculated by artificial intelligence to automatically complete the annotation.

[0044] Preferably, S100 includes:

[0045] S101, acquiring surveillance video data, and manually selecting a target to be marked in the first frame of the video; the target to be marked is selected based on the criterion that it cannot be recognized by common recognition technology;

[0046] S102, the target to be marked is segmented separately from the original image, and based on the segmented image, the starting center coordinates (X0, Y0), the left vertex coordinates (X1, Y1) of the selection box, and the current scale Δx and Δy of the target in the first frame are obtained.

[0047] S200, sending the segmented target image and the whole image to the upper and lower convolutional neural networks respectively for feature calculation to confirm the position of the target image in the original image;

[0048] Among them, the feature vector sets of the segmented image and the original image are calculated respectively, and the position and size of the target in the current frame are determined by cross-correlation value calculation, so as to obtain the position and size information of the annotation box;

[0049] Preferably, the convolutional neural network in S200 includes:

[0050] The convolutional neural network consists of 4 layers of convolution kernels. The first layer uses a 7×7 convolution kernel with a total of 64 dimensions, the second layer uses a 3×3 convolution kernel with a total of 256 dimensions, the third layer uses a 3×3 convolution kernel with a total of 512 dimensions, and the fourth layer uses a 1×1 convolution kernel with a total of 256 dimensions.

[0051] The network contains a total of 4 convolutional layers, which calculate the feature vector sets of the segmented image and the original image respectively, and determine the position and size of the target in the current frame through cross-correlation value calculation, thereby obtaining the position and size information of the annotation box.

[0052] Preferably, S200 includes:

[0053] S201, sending the segmented target image and the whole image to upper and lower convolutional neural networks respectively, wherein the upper and lower convolutional neural networks are completely identical networks, and coupling of the two images is achieved in the calculation by parameter sharing;

[0054] As can be seen in the figure, the two convolutional neural networks are composed of two identical convolutional architectures. The segmented target image and the entire image are simultaneously sent to the upper and lower convolutional neural networks for feature calculation. Since the upper and lower convolutional neural networks are exactly the same, the two images are coupled in the calculation by sharing parameters.

[0055] S202, after the image passes through the complete 4-layer neural network, the segmented image is obtained with a size of 15×15 and a feature map array of 256 channels, and the original image is obtained with a size of 31×31 and a feature map array of 256 channels; and the feature map arrays of the segmented image and the original image are subjected to channel cross-correlation calculation, because the number of channels of the feature maps of the segmented image 4 and the original image 2 are both 256, so the channel cross-correlation calculation is to calculate the feature map of each channel. Figure 1 1. performing cross-correlation calculation;

[0056] S203, obtaining a 17×17 feature vector set with a dimension of 256 through channel cross-correlation calculation, wherein each 1×1×256 feature vector in the feature vector set is a cross-correlation calculation value of a feature of the segmented image at a current position in the set, thereby obtaining a position of an arbitrarily segmented target in the original image;

[0057] The larger the cross-correlation calculation value is, the closer the vector at that position is to the coordinate position of the segmented target in the original image. Since the segmented image will not be a single-pixel target and must contain several pixel areas, the cross-correlation calculation value of a certain area in the set must be very close and in a larger range. In this way, the position of an arbitrarily segmented target in the original image is obtained.

[0058] S300, in the first frame image, based on the position of the target image in the original image, confirm the absolute coordinates and size relationship of the target in the subsequent frames;

[0059] Preferably, S300 includes:

[0060] Because the convolutional neural network is calculating by segmenting the target selected in the first frame and sending it together with the original image into the network for calculation, the coordinates in the obtained feature vector set are relative values ​​rather than absolute values;

[0061] S301, according to the starting center coordinates (X0, Y0) of the target in the first frame, the coordinates of the left vertex of the selection box (X1, Y1), and the current scale Δx, Δy of the target, as well as the relative coordinates (ΔX0, ΔY0) in the feature vector set, the relative coordinates of the left vertex of the selection box (ΔX1, ΔY1), and the relative scale Δdx, Δdy of the target, the absolute coordinates and size relationship of the target in subsequent frames can be quickly calculated according to the following formula:

[0062]

[0063] Where Xc and Yc are the center coordinates of the target in the subsequent frame, Xd and Yd are the left vertex coordinates of the target's annotation box, and Xs and Ys are the actual scales of the target in the subsequent frame.

[0064] S302, after the target to be marked is selected in the first frame, the network will record the target in real time frame by frame when it appears at any position in the video.

[0065] You only need to manually select the target to be annotated in the first frame of the video. No human operation is required. After all video frames are calculated, the video will be automatically deframed and the JSON of the annotation file will be generated. It also has strong resistance to deformation and scale changes, avoiding the problems of missed detection and false detection in other processes of identifying targets and assisting in annotation through pre-trained models.

[0066] Since the target's movement speed will not exceed the frame rate during actual motion, otherwise the target will be blurred and invisible, the target's movement is considered to be a slow movement with little change between frames. Even if the shape and size of the target change during the movement, when the cross-correlation value is calculated in the network, the cross-correlation value areas of different sizes will be calculated according to the actual shape and size of the target, and the final selection box size and other information results will be sent out, without the phenomenon of missed detection or false detection due to the failure to identify the target in the traditional artificial intelligence labeling method. Figure 2 This is an illustration of the different shapes and sizes of targets in the cross-correlation value calculation features.

[0067] The present invention is applicable to various images that can be imaged today, including but not limited to visible light images, infrared thermal images, ultraviolet images, millimeter wave images, terahertz images, etc., which can be presented in the form of images for target recognition and rapid marking. The method of the present invention can be used to quickly and effectively mark the target to be marked in the collected video, without the need for heavy manual frame-by-frame marking after frame splitting, and also avoids the phenomenon of missed detection and false detection in the use of pre-trained model recognition and marking, providing a new and fast data marking method for industrial target recognition technology.

[0068] Embodiment 2:

[0069] The computer-readable storage medium of this embodiment stores a computer program, which, when executed by a processor, implements the steps of an artificial intelligence rapid target labeling method in embodiment 1.

[0070] The computer-readable storage medium of this embodiment may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal; the computer-readable storage medium of this embodiment may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card, a secure digital card, a flash memory card, etc. equipped on the terminal; further, the computer-readable storage medium may also include both an internal storage unit of the terminal and an external storage device.

[0071] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.

[0072] Embodiment 3:

[0073] The computer device of this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of an artificial intelligence rapid target labeling method in Embodiment 1 are implemented.

[0074] In this embodiment, the processor may be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, readily available programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0075] Those skilled in the art will appreciate that the disclosed content of the embodiments may be provided as methods, systems, or computer program products. Therefore, the present solution may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Moreover, the present solution may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.

[0076] The present solution is described with reference to the method according to the embodiment of the present solution and the flowchart and / or block diagram of the computer program product. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions; these computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 one or more processes and / or methods Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0077] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 one or more processes and / or methods Figure 1 A function specified in one or more boxes.

[0078] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 one or more processes and / or methods Figure 1The steps for the functions specified in one or more boxes.

[0079] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0080] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An artificial intelligence rapid target labeling method, characterized by: The method comprises: S100, manually selecting a target to be labeled based on the first frame image of the video data and segmenting it separately; S200, sending the segmented target image and the whole image to the upper and lower convolutional neural networks respectively for feature calculation to confirm the position of the target image in the original image; S300: In the first frame image, based on the position of the target image in the original image, the absolute coordinates and size relationship of the target in the subsequent frames are confirmed.

2. The artificial intelligence rapid target labeling method according to claim 1, characterized in that: A symmetric feature cross-correlation coupled detector neural network is proposed. The network does not pay attention to the selected target category, but only calculates the features of the target in the selected image and the coordinate position of the feature with the highest similarity in the subsequent frames of the video. At the same time, the center coordinates of the target and the vertex coordinates of the rectangular annotation box that is closest to the size of the target are given according to the actual size of the target in the current frame, and recorded in a JSON file in the form of one image and one annotation.

3. The artificial intelligence rapid target labeling method according to claim 1, characterized in that: The S100 includes: S101, acquiring surveillance video data, and manually selecting a target to be marked in the first frame of the video; the target to be marked is selected based on the criterion that it cannot be recognized by common recognition technology; S102, the target to be marked is segmented separately from the original image, and based on the segmented image, the starting center coordinates (X0, Y0), the left vertex coordinates (X1, Y1) of the selection box, and the current scale Δx and Δy of the target in the first frame are obtained.

4. The artificial intelligence rapid target labeling method according to claim 1, characterized in that: The convolutional neural network in S200 includes: The convolutional neural network consists of 4 layers of convolution kernels. The first layer uses a 7×7 convolution kernel with a total of 64 dimensions, the second layer uses a 3×3 convolution kernel with a total of 256 dimensions, the third layer uses a 3×3 convolution kernel with a total of 512 dimensions, and the fourth layer uses a 1×1 convolution kernel with a total of 256 dimensions.

5. The artificial intelligence rapid target labeling method according to claim 1, characterized in that: The S200 includes: S201, sending the segmented target image and the whole image to upper and lower convolutional neural networks respectively, wherein the upper and lower convolutional neural networks are completely identical networks, and coupling of the two images is achieved in the calculation by parameter sharing; S202, after the image passes through the complete 4-layer neural network, the segmented image is obtained with a size of 15×15 and a feature map array with 256 channels, and the original image is obtained with a size of 31×31 and a feature map array with 256 channels; and the feature map arrays of the segmented image and the original image are subjected to channel cross-correlation calculation to confirm the cross-correlation of the feature map of each channel; S203. A 17×17 feature vector set with a dimension of 256 is obtained through channel cross-correlation calculation. Each 1×1×256 feature vector in the feature vector set is a cross-correlation calculation value of the feature of the segmented image appearing at the current position in the set, and the position of an arbitrarily segmented target in the original image is obtained.

6. The artificial intelligence rapid target labeling method according to claim 1, characterized in that: The S300 includes: S301, according to the starting center coordinates (X0, Y0) of the target in the first frame, the coordinates of the left vertex of the selection box (X1, Y1), and the current scale Δx, Δy of the target, as well as the relative coordinates (ΔX0, ΔY0) in the feature vector set, the relative coordinates of the left vertex of the selection box (ΔX1, ΔY1), and the relative scale Δdx, Δdy of the target, the absolute coordinates and size relationship of the target in subsequent frames can be quickly calculated according to the following formula: Where Xc and Yc are the center coordinates of the target in the subsequent frame, Xd and Yd are the left vertex coordinates of the target's annotation box, and Xs and Ys are the actual scales of the target in the subsequent frame. S302, after the target to be marked is selected in the first frame, the network will record the target in real time frame by frame when it appears at any position in the video.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in an artificial intelligence rapid target labeling method as described in any one of claims 1 to 6 are implemented.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps of the artificial intelligence rapid target labeling method as described in any one of claims 1-6 are implemented.