Video data labeling method, device, apparatus, and storage medium
By training a multi-task model to label video frame images, generating combined feature information and integrating it into a second training model, the problems of high computational resource consumption and low labeling accuracy are solved, achieving efficient and accurate video data labeling.
Patent Information
- Application Number
- CN202111217497.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2041-10-19
AI Technical Summary
Existing technologies suffer from high computational resource consumption and low accuracy in video data labeling.
A multi-task training model is used to label video frame images. The first training model generates combined feature information, which is then input into the second training model to obtain the labeling results of the video data. The multi-task processing mechanism and multi-dimensional information fusion are used to improve the labeling accuracy.
It significantly reduces computing resource consumption and improves the efficiency and accuracy of video data labeling.
Smart Images

Figure CN114022807B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of computer, and particularly relate to a video data labeling method, device, equipment and storage medium. BACKGROUND
[0002] With the rise of short video social entertainment, a large amount of unlabeled video data will be generated in the video network platform every day. For the video platform, the value of the unlabeled video data is relatively limited, and the way of manually labeling the video data needs to consume a lot of manpower and material resources.
[0003] In the prior art, the labeling processing of the massive video data is mainly realized by using a deep learning model. According to the different input contents of the deep learning model, the labeling method is basically divided into a multi-modal labeling method and a labeling method of directly classifying video data. The multi-modal labeling method needs to obtain information data in multiple dimensions, such as audio information, text information and video pictures in the video data, and needs a high computing power, large running power consumption, and a large number of machine resources in the training and actual deployment process. The labeling method of directly classifying video data has a small amount of calculation, can meet the real-time requirement of real-time video data labeling, but the labeling result accuracy is low, and there are a large number of error labeling samples. SUMMARY
[0004] Embodiments of the present application provide a video data labeling method, device, equipment and storage medium, which solve the problems of large computing resource consumption and low labeling result accuracy in the prior art, and improve the labeling efficiency and accuracy of the video data.
[0005] In a first aspect, embodiments of the present application provide a video data labeling method, which comprises:
[0006] inputting a plurality of video frame images in the video data into a first training model to obtain a labeling result of different types corresponding to each video frame image, the first training model being a multi-task training model trained based on different types of labeling tasks;
[0007] generating corresponding combined feature information according to the labeling result of each video frame image;
[0008] inputting the combined feature information into a second training model to obtain a labeling result of the video data.
[0009] In a second aspect, embodiments of the present application further provide a video data labeling device, which comprises:
[0010] The first model processing module is configured to input a plurality of video frame images in the video data into a first training model to obtain different types of labeling results corresponding to each video frame image, wherein the first training model is a multi-task training model trained based on different types of labeling tasks.
[0011] The feature information generation module is configured to generate corresponding combined feature information according to the labeling result of each video frame image.
[0012] The second model processing module is configured to input the combined feature information into a second training model to obtain the labeling result of the video data.
[0013] In a third aspect, an embodiment of the present application further provides a video data labeling device, which comprises:
[0014] One or more processors;
[0015] A storage device configured to store one or more programs,
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the video data labeling method provided by the embodiment of the present application.
[0017] In a fourth aspect, an embodiment of the present application further provides a storage medium storing computer executable instructions, which, when executed by a computer processor, are used to perform the video data labeling method provided by the embodiment of the present application.
[0018] In the embodiment of the present application, after a plurality of video frame images in the video data are input into a first training model to obtain different types of labeling results corresponding to each video frame image, combined feature information corresponding to each video frame image is generated according to the labeling result of each video frame image, and the combined feature information is input into a second training model to obtain the labeling result of the video data, wherein the first training model is a multi-task training model trained based on different types of labeling tasks, thereby solving the problems of large consumption of computing resources and low accuracy of labeling results in the prior art when video data is labeled, and improving the labeling efficiency and accuracy of the video data. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 A flowchart of a video data labeling method provided by an embodiment of the present application is shown in FIG. 1;
[0020] Figure 2 A flowchart of another video data labeling method provided by an embodiment of the present application is shown in FIG. 2;
[0021] Figure 3 A flowchart of another video data labeling method provided by an embodiment of the present application is shown in FIG. 3;
[0022] Figure 4 A flowchart of another video data marking method provided by an embodiment of the present application is shown in FIG. 6.
[0023] Figure 5 A flowchart of another video data marking method provided by an embodiment of the present application is shown in FIG. 6.
[0024] Figure 6 A module schematic diagram of a video data marking device provided by an embodiment of the present application is shown in FIG. 7.
[0025] Figure 7 A structural schematic diagram of a video data marking device provided by an embodiment of the present application is shown in FIG. 8. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be described clearly below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art belong to the scope of protection of the present application.
[0027] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the objects before and after it.
[0028] The video data marking method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings, specific embodiments and application scenarios.
[0029] Figure 1 A flowchart of a video data marking method provided by an embodiment of the present application is shown in FIG. 5. The embodiment can realize efficient, automatic and accurate marking of video data. The method can be executed by a device with computing function, such as a server, a notebook computer, a desktop computer, etc. The specific steps include the following steps:
[0030] In step S101, a plurality of video frame images in the video data are input to a first training model to obtain a marking result of different types corresponding to each video frame image.
[0031] The video data is data that needs to be marked to obtain a marking result. Taking a video live broadcast platform as an example, the video data can be video live broadcast data in a video live broadcast process. In an embodiment, after obtaining live video data of different live broadcast users, the video data is marked by the video data marking method of the embodiment to obtain a corresponding marking result, and it is determined that the video data belongs to which type, whether it is compliant, or whether it contains a preset type of person object, article object, etc.
[0032] In an embodiment, for the obtained video data, frame extraction processing is first performed, such as automatic frame extraction processing of the video data by a video decoding frame extraction module to obtain a plurality of video frame images. After obtaining the plurality of video frame images, they are input into a first training model to obtain different types of marking results corresponding to each video frame image. The first training model is a multi-task training model trained based on different types of marking tasks. That is, the first training model can output different types of marking results corresponding to different marking tasks for each video frame image.
[0033] For example, during the training of the first training model, two marking tasks are set, such as a first type of marking task one and a second type of marking task two. Through the training of the first training model, it outputs the first type of marking result and the second type of marking result for the input video frame image, respectively.
[0034] Step S102, generating corresponding combined feature information according to the marking result of each video frame image.
[0035] In an embodiment, based on the marking result of each video frame image output by the first training model, the marking results of each video frame image are fused to obtain combined feature information. Optionally, taking the first marking result and the second marking result including the first marking task and the second marking task as an example, the first marking result and the second marking result are fused to obtain two-dimensional combined feature information.
[0036] In an embodiment, the first marking result can be a score marking result, and the second marking result can be a type marking result. The score marking result and the type marking result are combined to obtain two-dimensional combined feature information containing scores and categories.
[0037] Step S103, inputting the combined feature information into a second training model to obtain a marking result of the video data.
[0038] After determining the combined feature information containing different marking results, it is input into a second training model to obtain the final marking result of the video data. The second training model is a pre-trained model that takes multi-dimensional marking results as input and outputs corresponding marking results of the video data.
[0039] For example, different labeling tasks are set to represent the characteristics of the preset type of person object in the video data, such as part labeling task, appearance labeling task, clothing labeling task, bad tendency labeling task, and temperament labeling task. After obtaining the labeling results of the part labeling task, the appearance labeling task, the clothing labeling task, the bad tendency labeling task, and the temperament labeling task, the five-dimensional combined feature information is obtained by combining them, and then input into the second training model to obtain the final result of whether the preset type of person object is contained in the video data.
[0040] In another embodiment, the video data can also be labeled to determine whether it contains a preset type of object. Corresponding multiple labeling tasks corresponding to the object are set. After obtaining the labeling results of the multiple labeling tasks based on the first training model, the combined information is obtained by combining them, and then input into the second training network to output the labeling result of whether the video data contains the preset type of object.
[0041] As can be seen from the above, by inputting multiple video frame images in the video data into the first training model, obtaining the labeling results of different types corresponding to each video frame image, and generating the corresponding combined feature information according to the labeling results of each video frame image, the labeling result of the video data is obtained by inputting the combined feature information into the second training model. The first training model is a multi-task training model trained based on different types of labeling tasks. In one training model, a multi-task processing mechanism is used to realize the multi-task labeling result output of the video frame image, and after integration, another machine learning model is used to output the labeling result of the video data. The problem of large consumption of computing resources and low accuracy of labeling result in the prior art is solved, and the labeling efficiency and accuracy of the video data are improved.
[0042] Figure 2 Another flowchart of a video data labeling method provided by the embodiment of the present application is provided. A specific method for training the first training network is given, as shown in Figure 2 The specific steps include:
[0043] Step S201, determine the training samples of different types of labeling tasks, and train the first training model through the training samples, wherein the first training model includes a backbone network, convolutional layers of different layers, and a fully connected layer.
[0044] In one embodiment, the model training is first performed to obtain a first training model. When the training sample selection is performed, the corresponding training samples are selected according to different types of labeling tasks. Taking whether a preset type of human object is contained in the video data as an example, multi-task decomposition is performed to obtain five different types of labeling tasks, i.e., a part labeling task, a look labeling task, a clothing labeling task, a bad tendency labeling task, and a temperament labeling task. The corresponding training samples are selected for each labeling task, i.e., for the part labeling task, the video images containing the corresponding part are selected in the video live streaming platform to obtain the training samples corresponding to the part labeling task through artificial labeling; similarly, for the look labeling task, the clothing labeling task, the bad tendency labeling task, and the temperament labeling task, the video images containing the corresponding labeling task are selected in the video live streaming platform to obtain the training samples corresponding to the look labeling task, the clothing labeling task, the bad tendency labeling task, and the temperament labeling task, respectively, after labeling.
[0045] After the training samples corresponding to different labeling tasks are determined, the first training model is trained. In one embodiment, the first training model includes a backbone network and convolution layers and full connection layers of different numbers of layers, and the backbone network can be an efficientnet_b1 network, for example, and the convolution layers of different numbers of layers correspond to different labeling tasks. Optionally, the convolution layers of the numbers of layers with good final recognition effect are selected for different labeling task types. The numbers of layers of the convolution layers corresponding to different labeling tasks can be set according to the actual training process, and the optimal selection is performed according to the recognition accuracy, convergence speed, and the like of the training result. In one embodiment, when the first training model is trained, the training samples of the corresponding labeling tasks are input into the convolution layers of different numbers of layers according to the different convolution layer inputs set, so as to train the same.
[0046] Step S202, inputting a plurality of video frame images in the video data into the first training model to obtain different types of labeling results corresponding to each video frame image.
[0047] Step S203, generating corresponding combined feature information according to the labeling result of each video frame image.
[0048] Step S204, inputting the combined feature information into the second training model to obtain the labeling result of the video data.
[0049] From the above, the first training model in the scheme adopts a main network combined with convolution layers of different layers to realize the output of multiple different types of marking results of the video frame image, so as to obtain multi-dimensional information, and then obtain the marking result of the frequency data through the second training model. The first training model used in the scheme can significantly improve the learning and training speed, reduce the data operation amount, and complete the learning and training work with a small amount of computing power. Meanwhile, the multi-task learning mode is adopted to improve the accuracy of the video data marking result.
[0050] Figure 3 The flowchart of another video data marking method provided by the embodiment of the application is shown in FIG. 6. Different types of marking tasks include first marking type, second marking type and third marking type marking tasks. The embodiment gives a specific method for training the first training network using different loss functions, as shown in FIG. 6. Figure 3 As described above, the specific steps include:
[0051] In step S301, the training samples of different types of marking tasks are determined, and the first training model is trained based on the training samples, the mean square error loss function of the first marking type, the multi-classification cross-entropy loss function of the second marking type and the binary classification cross-entropy loss function of the third marking type.
[0052] In one embodiment, when marking the video data, the multi-task decomposition method is used to mark and identify different types of marking tasks respectively, and then the second training model is used to output the marking result of the integrated combined feature information.
[0053] Optionally, taking whether the video data contains a preset type of character object as an example, multi-task decomposition is performed to obtain the first marking type, the second marking type and the third marking type marking tasks. The first marking type marking task is exemplarily a part marking task, a feature marking task and a temperament marking task. The second marking type marking task is exemplarily a clothing marking task, which includes a clothing upper garment marking task and a clothing lower garment marking task. The third marking type marking task is exemplarily an undesirable tendency marking task.
[0054] The mean square error loss function is selected for the first marking type to train the corresponding convolutional layer network, the multi-classification cross-entropy loss function is selected for the marking task of the second marking type to train the corresponding convolutional layer network, and the cross-entropy loss function is selected for the marking task of the third marking type to train the corresponding convolutional layer network. The loss function is a mathematical parameter used to evaluate the degree of inconsistency between the predicted value and the true value of the model. The more reasonable the loss function is set, the more accurate the output result of the model obtained by the final training is. In the present scheme, different loss functions are set for different marking tasks to obtain marking results of multiple different task branches under the same backbone network.
[0055] Specifically, the mean square error loss function is used for the part marking task, the appearance marking task and the temperament marking task of the first marking type, which is represented as Lr1, Lr2 and Lr3 respectively, the binary classification cross-entropy loss function is used for the clothing marking task of the second marking type, which is represented as Ls, and the binary classification cross-entropy loss function is used for the adverse tendency marking task of the third marking type, which is represented as Lb.
[0056] In one embodiment, when the first training network is trained, the same backbone network is used to train the convolutional layers of different layers using the loss function corresponding to the marking task type. In another embodiment, a joint loss function is obtained according to the loss functions of different marking tasks, and the first training network is trained as a whole using the joint loss function. Taking the loss functions Lr1, Lr2, Lr3, Ls and Lb as examples, the joint loss function is represented as Loss = k1*Lr1 + k2*Lr2 + k3*Lr3 + k4*Ls + k5*Lb, wherein k1, k2, k3, k4 and k5 are floating-point hyperparameters between 0 and 1, and the sum is 1, which is used to alleviate the problem of data imbalance between training samples. Specifically, k1, k2, k3, k4 and k5 can be adaptively adjusted according to the actual training process. For example, k1, k2, k3, k4 and k5 are 0.2, 0.3, 0.1, 0.25 and 0.35 respectively.
[0057] Step S302, inputting a plurality of video frame images in the video data into the first training model to obtain marking results of different types corresponding to each video frame image.
[0058] Step S303, generating corresponding combined feature information according to the marking results of each video frame image.
[0059] Step S304, inputting the combined feature information into the second training model to obtain the marking result of the video data.
[0060] From the above, in the process of labeling video data, a plurality of different types of labeling tasks are obtained by adopting multi-task decomposition, and the first training model is obtained by training based on different loss functions for different types of labeling tasks. While ensuring less computing power, the labeling accuracy of video data is improved, and the video data labeling mechanism is significantly optimized.
[0061] Figure 4 Another flowchart of a video data labeling method is provided for the embodiments of the application. A specific method for generating training samples of the first training model is given, as shown in Figure 4 As described above, specifically includes:
[0062] Step S401, according to different types of labeling tasks, the corresponding original image is obtained, the original image is annotated with scores and classified, and the annotation results are packaged with the original video to obtain training samples.
[0063] In one embodiment, when labeling video data, the first training model is a multi-task training model. Optionally, according to different types of labeling tasks, the corresponding original image is obtained, and the identification of whether the video data contains a preset type of human object is taken as an example, multi-task decomposition is performed to obtain human part labeling task, appearance labeling task, clothing labeling task, bad tendency labeling task and temperament labeling task. For part labeling task, appearance labeling task and temperament labeling task, manual score annotation method is adopted, and the score value range is 0-100 points. For clothing labeling task and bad tendency labeling task, classification annotation method is adopted, and the corresponding classification result is given, such as clothing of a preset type of human object, the classification result is marked as yes, otherwise as no. Similarly, when the bad tendency of a preset type of human object meets the preset type, the classification result is marked as yes, otherwise as no.
[0064] Step S402, training the first training model through the training sample, the first training model includes a backbone network, convolution layers with different layers and a full connection layer.
[0065] Step S403, inputting a plurality of video frame images in the video data into the first training model to obtain different types of labeling results corresponding to each video frame image.
[0066] Step S404, generating corresponding combined feature information according to the labeling result of each video frame image.
[0067] In one embodiment, the different types of marking tasks are respectively marked with score results and category results, wherein the score results can be normalized to a value between 0 and 1, and the score results and the category results are combined, for example, the score results have three scores of part marking tasks, appearance marking tasks and temperament marking tasks, and the category results have two classification results of clothing marking tasks and bad tendency marking tasks, and then a 5-dimensional combined feature information is formed, which contains the marking results of the part marking tasks, the appearance marking tasks, the clothing marking tasks, the bad tendency marking tasks and the temperament marking tasks.
[0068] In step S405, the combined feature information is input into a second training model to obtain the marking result of the video data.
[0069] The second training model can be an XGBoost model, a network model based on a random forest or a support vector machine, and the embodiment is not limited.
[0070] As described above, when marking the video data, different types of different marking methods are obtained by multi-task decomposition to train the model, and the marking results of different task types are obtained when the video data is recognized, and then the marking results are input into the machine learning model to output the marking results, which reduces the overall operation data amount of the video data recognition and ensures the recognition accuracy.
[0071] Figure 5 Another flowchart of a video data marking method provided by the embodiment of the application is provided. A specific method for generating training samples of the first training model is given, as shown in the following table. Figure 5 As described above, the specific method includes the following steps.
[0072] In step S501, a plurality of video frame images in the video data are input into the first training model to obtain different types of marking results corresponding to each video frame image.
[0073] In step S502, the corresponding combined feature information is generated according to the marking result of each video frame image.
[0074] In step S503, a preset number of sample frame images in the training sample video are randomly extracted, the sample frame images are scored and classified based on different dimensions, the scoring results and the classification results are combined as training samples of the second training model, and the second training model is trained.
[0075] In one embodiment, 20 frames of images are extracted for each training sample video, and each frame of image is scored and classified from five different dimensions respectively. For example, in the case of identification of whether a preset type of human object is contained in the video data, five dimensions are respectively part labeling dimension, appearance labeling dimension, clothing labeling dimension, bad tendency labeling dimension and temperament labeling dimension, and a total of 100-dimensional features are obtained as the training data of the second training model. Optionally, the training data is packaged into libSVM data format to accelerate the training process.
[0076] Step S504, inputting the combined feature information into the second training model to obtain the labeling result of the video data.
[0077] As can be seen from the above, by training the second training model, the labeling result under multiple dimensions can be identified to output the labeling result of the video data, and the comprehensive labeling result of the video data is obtained by identifying the combined information of the task results under multiple dimensions, thereby solving the problems of large consumption of computing resources and low accuracy of labeling result in the prior art when labeling the video data, and improving the labeling efficiency and accuracy of the video data.
[0078] Figure 6 A module schematic diagram of a video data labeling device provided by the embodiment of the present application is shown in the figure. The device is used to execute the video data labeling method described above, and has the corresponding function modules and beneficial effects of the execution method. As shown in the figure, the device specifically comprises: a first model processing module 101, a feature information generation module 102 and a second model processing module 103, wherein, Figure 6
[0079] The first model processing module 101 is used to input a plurality of video frame images in the video data into a first training model to obtain different types of labeling results corresponding to each video frame image, and the first training model is a multi-task training model trained based on different types of labeling tasks;
[0080] The feature information generation module 102 is used to generate corresponding combined feature information according to the labeling result of each video frame image;
[0081] The second model processing module 103 is used to input the combined feature information into a second training model to obtain the labeling result of the video data.
[0082] It can be known from the above scheme that after inputting a plurality of video frame images in video data into the first training model to obtain different types of labeling results corresponding to each video frame image, generating corresponding combined feature information according to the labeling result of each video frame image, and inputting the combined feature information into the second training model to obtain the labeling result of the video data, the first training model is a multi-task training model trained based on different types of labeling tasks, which solves the problems of large consumption of computing resources and low accuracy of labeling results in the prior art when video data is labeled, and improves the labeling efficiency and accuracy of video data.
[0083] In one possible embodiment, the first model processing module 101 is further configured to:
[0084] Before inputting the plurality of video frame images in the video data into the first training model, training samples of different types of labeling tasks are determined.
[0085] The first training model is trained through the training samples, and the first training model includes a backbone network, convolutional layers of different layers, and a fully connected layer, wherein the convolutional layers of different layers and the fully connected layer correspond to the training learning of one labeling task respectively.
[0086] In one possible embodiment, the different types of labeling tasks include labeling tasks of a first labeling type, a second labeling type, and a third labeling type, and the first model processing module 101 is specifically configured to:
[0087] The first training model is trained based on the training samples, a mean square error loss function of the first labeling type, a multi-classification cross-entropy loss function of the second labeling type, and a binary classification cross-entropy loss function of the third labeling type.
[0088] In one possible embodiment, the first model processing module 101 is specifically configured to:
[0089] According to different types of labeling tasks, corresponding original images are obtained, score labeling and category labeling are performed on the original images, and the labeling result is packaged with the original video to obtain training samples.
[0090] In one possible embodiment, the feature information generation module 102 is specifically configured to:
[0091] The score labeling result and the category labeling result are combined to obtain multi-dimensional combined feature information.
[0092] In one possible embodiment, the second model processing module 103 is further configured to:
[0093] Before inputting the combined feature information into the second training model to obtain the labeling result of the video data, a preset number of sample frame images in the training sample video are randomly extracted;
[0094] The sample frame images are scored and classified based on different dimensions, and the scoring result and the classification result are combined as a training sample of the second training model, and the second training model is trained, wherein the loss function of the second training model is a mean square error loss function.
[0095] In one possible embodiment, the different types of labeling tasks include at least two of a part labeling task of a human body, a feature labeling task, a clothing labeling task, a bad tendency labeling task, and a temperament labeling task.
[0096] Figure 7 A structural schematic diagram of a video data labeling device provided by an embodiment of the present application is shown in FIG. 1. Figure 7 As shown in the figure, the device includes a processor 201 and a memory 202; the number of processors 201 in the device can be one or more, Figure 7 and one processor 201 is taken as an example in the figure; the processor 201 and the memory 202 in the device can be connected through a bus or other means, Figure 7 and the connection through the bus is taken as an example in the figure. The memory 202 is a kind of computer readable storage medium, which can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the video data labeling method in the embodiment of the present application. The processor 201 executes the software programs, instructions and modules stored in the memory 202, thereby performing various functional applications and data processing of the device, that is, realizing the video data labeling method described above.
[0097] The embodiment of the present application also provides a storage medium containing computer executable instructions, which can be stored in the form of a server application, and the computer executable instructions are used to execute a video data labeling method when executed by a computer processor, and the method includes:
[0098] Input a plurality of video frame images in the video data into a first training model to obtain a labeling result corresponding to each video frame image of different types, and the first training model is a multi-task training model trained based on different types of labeling tasks;
[0099] According to the labeling result of each video frame image, corresponding combined feature information is generated;
[0100] The combined feature information is input into a second training model to obtain the labeling result of the video data.
[0101] It should be noted that, in the present document, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element. Furthermore, it is to be understood that the method and apparatus of the present application can be carried out by more than one process, method, article, or apparatus either simultaneously, concurrently or with intermediate steps missing or in reverse order, without departing from the scope of the application described herein. In addition, features described in relation to one example can be combined with features described in relation to other examples.
[0102] From the above description of the embodiments, it is apparent that the above-described method of the embodiments can be implemented by means of software and the necessary universal hardware platform, of course, can also be implemented by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product in essence or in the form of a part of the prior art that makes a contribution. The computer software product is stored in a storage medium (such as a ROM / RAM, a magnetic disc, an optical disc), and includes a plurality of instructions for causing a terminal (which can be an unmanned device, a mobile phone, a computer, a server, or a network device) to execute the method described in the various embodiments of the present application.
[0103] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative rather than limiting, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.
Claims
1. A method of marking video data, characterized by, The method comprises the following steps: determining training samples of different types of marking tasks, wherein the different types of marking tasks comprise marking tasks of a first marking type, a second marking type and a third marking type; training a first training model through the training samples, wherein the first training model comprises a backbone network, convolutional layers of different layers and a full connection layer, the convolutional layers of different layers and the full connection layer correspond to the training learning of one type of marking task respectively; the training of the first training model through the training samples comprises training the first training model based on the training samples, a mean square error loss function of the first marking type, a multi-classification cross-entropy loss function of the second marking type and a binary classification cross-entropy loss function of the third marking type; inputting a plurality of video frame images in video data into the first training model to obtain different types of marking results corresponding to each video frame image, wherein the marking result of each video frame image comprises a score marking result and a category marking result, the score marking result comprises marking results of a part marking task, a feature marking task and a temperament marking task, and the category marking result comprises marking results of a clothing marking task and an undesirable tendency marking task, and the first training model is a multi-task training model trained based on different types of marking tasks; generating corresponding combined feature information according to the marking result of each video frame image, wherein the combined feature information comprises a multi-dimensional combined feature information obtained by combining the score marking result and the category marking result; randomly extracting a preset number of sample frame images in a training sample video; scoring and classifying the sample frame images based on different dimensions, combining the scoring result and the classification result as training samples of a second training model, training the second training model, wherein a loss function of the second training model is a mean square error loss function; inputting the combined feature information into the second training model to obtain the marking result of the video data.
2. The video data labeling method of claim 1, wherein, The method comprises the following steps: obtaining corresponding original images according to different types of marking tasks, performing score annotation and category annotation on the original images, packing the annotation results and original videos to obtain training samples.
3. The method of claim 1 or 2, wherein, The different types of marking tasks comprise at least two of a part marking task, a feature marking task, a clothing marking task, an undesirable tendency marking task and a temperament marking task.
4. A video data marking apparatus characterized by comprising: The method comprises the following steps: The first model processing module is configured to input a plurality of video frame images in the video data into a first training model to obtain different types of labeling results corresponding to each video frame image. The labeling result of each video frame image includes a score labeling result and a category labeling result. The score labeling result includes labeling results of a part labeling task, a feature labeling task, and a temperament labeling task. The category labeling result includes labeling results of a clothing labeling task and an undesirable tendency labeling task. The first training model is a multi-task training model trained based on different types of labeling tasks. The first model processing module is further configured to determine training samples of different types of labeling tasks, and train the first training model based on the training samples. The first training model includes a backbone network, convolution layers with different numbers of layers, and a full connection layer. The convolution layers with different numbers of layers and the full connection layer correspond to training learning of one labeling task respectively. The different types of labeling tasks include labeling tasks of a first labeling type, a second labeling type, and a third labeling type. The first model processing module is specifically configured to train the first training model based on the training samples, a mean square error loss function of the first labeling type, a multi-classification cross-entropy loss function of the second labeling type, and a binary-classification cross-entropy loss function of the third labeling type. The feature information generation module is configured to generate corresponding combined feature information based on the labeling result of each video frame image. The feature information generation module is specifically configured to combine the score labeling result and the category labeling result to obtain multi-dimensional combined feature information. The second model processing module is configured to input the combined feature information into a second training model to obtain a labeling result of the video data. The second model processing module is further configured to randomly extract a preset number of sample frame images in a training sample video, score and classify the sample frame images based on different dimensions, combine a scoring result and a classification result as a training sample of the second training model, and train the second training model. A loss function of the second training model is a mean square error loss function.
5. Video data marking apparatus, the apparatus comprising: One or more processors; A storage device configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the video data labeling method in any one of claims 1-4.
6. A storage medium storing computer-executable instructions for performing the video data labeling method in any one of claims 1-4 when executed by a computer processor.
Citation Information
Patent Citations
Fundus image lesion labeling method and device based on feature visualization and medium
CN110264443A
Image recognition method and device, electronic equipment and computer readable medium
CN112381074A
Video classification method and device, terminal and storage medium
CN113158710A
Spread spectrum signal recognition method based on multi-dimensional parameter extraction and support vector machine
CN113408420A