A video recognition method, device, and electronic device
Through multiple feature extraction models and feature fusion models, the problem of low degree of automation of video advertising segment recognition in the prior art is solved, and efficient and accurate video tag recognition is achieved.
Patent Information
- Application Number
- CN202011133415.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-10-21
AI Technical Summary
The prior art is difficult to realize the device automatically and intelligently screens video advertising clips, resulting in a low degree of automation in the recognition of advertising clips.
The video tag of the target video is identified through multiple feature extraction models, including the first image feature extraction model, the second image feature extraction model, the first text feature extraction model and the second text feature extraction model. Combining the image feature fusion model and the text feature fusion model, the automation and intelligence of video tag recognition are improved.
The video tag identification process of target videos is achieved efficiently, the accuracy and efficiency of recognition are improved, and the recognition methods of video tags are enriched through diversified feature expression.
Smart Images

Figure CN112149632B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a video recognition method, apparatus, and electronic device. Background Art
[0002] With the development of computer technology, electronic devices have become increasingly popular. There are a large number of video playback platforms on electronic devices, and the vast amount of videos provided have enriched people's daily lives. However, the embedded advertisements in the videos have seriously affected people's viewing experience.
[0003] Currently, when recognizing videos, taking the embedded advertisements in the videos as an example. For videos without labels for embedded advertisements, the embedded advertisements in the videos can be filtered through crowdsourcing. That is, the advertisement filtering task is posted to the video platform, and users are asked to label it, and certain material rewards are given to the users. However, the automatic intelligent screening of the advertisement segments of the video cannot be achieved through manual means, resulting in a low degree of automation in the recognition of advertisement segments. Summary of the Invention
[0004] Embodiments of this application provide a video recognition method, apparatus, and electronic device. It can effectively improve the automation and intelligence level of the video label recognition process of the target video.
[0005] On the one hand, embodiments of this application provide a video recognition method, and the method includes:
[0006] Obtain a target video, where the target video includes video frame images and target text;
[0007] Call a first image feature extraction model to extract first image features of the video frame images; the first image feature extraction model is an image feature extraction model trained based on a first classification task;
[0008] Call a second image feature extraction model to extract second image features of the video frame images; the second image feature extraction model is an image feature extraction model trained based on the first classification task and a second classification task;
[0009] Call a first text feature extraction model to extract first text features of the target text; the first text feature extraction model is a text feature extraction model trained based on the first classification task;
[0010] Call a second text feature extraction model to extract second text features of the target text; the second text feature extraction model is a text feature extraction model trained based on the first classification task and a third classification task;
[0011] Determine the video label of the target video according to the first image feature, the second image feature, the first text feature, and the second text feature, and determining that the video label of the target video belongs to the first classification task.
[0012] On the one hand, an embodiment of the present application provides a video recognition device, the device includes:
[0013] An acquisition unit, configured to acquire a target video, where the target video includes video frame images and target text;
[0014] A processing unit, configured to call a first image feature extraction model to extract a first image feature of the video frame image; the first image feature extraction model is an image feature extraction model trained based on a first classification task;
[0015] The processing unit is further configured to call a second image feature extraction model to extract a second image feature of the video frame image; the second image feature extraction model is an image feature extraction model trained based on the first classification task and a second classification task;
[0016] The processing unit is further configured to call a first text feature extraction model to extract a first text feature of the target text; the first text feature extraction model is a text feature extraction model trained based on the first classification task;
[0017] The processing unit is further configured to call a second text feature extraction model to extract a second text feature of the target text; the second text feature extraction model is a text feature extraction model trained based on the first classification task and a third classification task;
[0018] A determination unit, configured to determine the video label of the target video according to the first image feature, the second image feature, the first text feature, and the second text feature, and determining that the video label of the target video belongs to the first classification task.
[0019] On the one hand, an embodiment of the present application provides an electronic device, including a processor, a memory, a communication interface, and one or more programs, wherein the above one or more programs are stored in the above memory and are configured to be executed by the above processor, and the above programs include instructions for performing the steps in the above method.
[0020] Correspondingly, an embodiment of the present application provides a computer-readable storage medium, configured to store computer program instructions for a terminal device, which includes programs involved in performing the steps in the above method.
[0021] Correspondingly, the embodiments of the present application provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. When the computer instructions are executed by a processor of a computer device, the methods in the above embodiments are executed.
[0022] It can be seen that in the embodiments of the present application, video tags of a target video are recognized through multiple feature extraction models, realizing an automatic recognition process with high efficiency and high automation. Moreover, the models for extracting the second image features and the second text features are jointly trained models based on the first classification task and the second classification task. That is, the extracted image features and text features can not only represent the features in the field of the first classification task, but also represent the features in the field of the second classification task. This can not only enrich the expression ways of the image features and text features and the recognition ways of video tags, but also improve the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0024] Figure 1 It is a schematic diagram of the target video recognition process provided by the embodiments of the present application;
[0025] Figure 2 It is a schematic flowchart of a video recognition method provided by the embodiments of the present application;
[0026] Figure 3A It is a schematic flowchart of another video recognition method provided by the embodiments of the present application;
[0027] Figure 3B It is a schematic diagram of obtaining a first predicted label by training an intermediate model provided by the embodiments of the present application;
[0028] Figure 3C It is a schematic diagram of obtaining a video label by using a total model provided by the embodiments of the present application;
[0029] Figure 3D It is a schematic structural diagram of a second image to-be-trained model provided by the embodiments of the present application;
[0030] Figure 3E It is a schematic structural diagram of a module one provided by the embodiments of the present application;
[0031] Figure 3FIt is a schematic structural diagram of a first image to-be-trained model provided by an embodiment of the present application;
[0032] Figure 3G It is a schematic structural diagram of a second text to-be-trained model provided by an embodiment of the present application;
[0033] Figure 3H It is a schematic structural diagram of a module two provided by an embodiment of the present application;
[0034] Figure 3I It is a schematic structural diagram of a first text to-be-trained model provided by an embodiment of the present application;
[0035] Figure 3J It is a schematic structural diagram of a total model provided by an embodiment of the present application;
[0036] Figure 4 It is a schematic diagram of the functional units of an image recognition device provided by an embodiment of the present application;
[0037] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0038] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.
[0039] Terms such as "first" and "second" in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0040] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in conjunction with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0041] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, detection, and measurement, and further performs image processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0042] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0043] This application is combined with machine learning, and artificial neural network technology is used to construct and train multiple models used in this application, resulting in multiple models with strong image recognition capabilities. And this application is combined with CV technology, and ORC technology is used to extract files from video frame images to obtain target texts. The video content is recognized to determine the video tags of the target video. The intelligence and automation degree of the video tag recognition process are improved.
[0044] The embodiments of this application provide a method for video recognition, which is applied to a video recognition device. The video recognition device can be an internal device of an electronic device or an external device of the electronic device. The following is a detailed introduction with reference to the accompanying drawings.
[0045] First, please refer to Figure 1 the schematic diagram of the target video recognition process shown. The recognition process of the target video includes a first image feature extraction model, a second image feature extraction model, a first text feature extraction model, and a second text feature extraction model.
[0046] For a target video, the target video includes video frame images and target text. Invoke a first image feature extraction model to extract first image features of the video frame images; the first image feature extraction model is an image feature extraction model trained based on a first classification task; invoke a second image feature extraction model to extract second image features of the video frame images; the second image feature extraction model is an image feature extraction model trained based on the first classification task and a second classification task; invoke a first text feature extraction model to extract first text features of the target text; the first text feature extraction model is a text feature extraction model trained based on the first classification task; invoke a second text feature extraction model to extract second text features of the target text; the second text feature extraction model is a text feature extraction model trained based on the first classification task and a third classification task; finally, a video label of the target video can be determined according to the first image features, the second image features, the first text features, and the second text features, and it is determined that the video label of the target video belongs to the first classification task.
[0047] The above-mentioned first image feature extraction model, second image feature extraction model, first text feature extraction model, and second text feature extraction model can be any one or more of recurrent neural networks (RNNs), convolutional neural networks (CNNs), deep belief neural networks, generative adversarial networks, autoencoders (AEs), and recurrent neural networks.
[0048] The above-mentioned electronic device can include, for example, a distributed storage server, a traditional server, a large storage system, a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart phone, a portable digital player, a smart watch, and a smart bracelet, etc.
[0049] The technical solution of the embodiment of the present application can be based on Figure 1 the schematic diagram or deformation schematic diagram of the video process shown in the example for specific implementation.
[0050] See Figure 2 , Figure 2 is a schematic flow chart of a video recognition method provided by an embodiment of the present application, which can be applied to a video recognition device. This method may include but is not limited to the following steps:
[0051] 201. Obtain a target video, where the target video includes video frame images and target text.
[0052] Specifically, for a complete video, the complete video can be divided into individual video segments according to a preset method. For example, every 5 seconds, or 2 seconds, or 3 seconds, etc. can be used as a video segment. The video recognition device can use any one of the video segments as the target video. It can be understood that the target video includes video frame images and target text, and the target text is the caption included in each video frame image.
[0053] 202. Invoke the first image feature extraction model to extract the first image feature of the video frame image. The first image feature extraction model is an image feature extraction model trained based on the first classification task.
[0054] Specifically, the video recognition device can invoke the first image feature extraction model to extract the first image feature from the video frame image. The first image feature extraction model is an image feature extraction model trained based on the first classification task. The first classification task can be understood as distinguishing whether the target video is an advertisement video or not, which can be used for video advertisement classification. The corresponding first image feature is the image feature that can determine whether the target video corresponding to the video frame image is an advertisement video or not.
[0055] 203. Invoke the second image feature extraction model to extract the second image feature of the video frame image. The second image feature extraction model is an image feature extraction model trained based on the first classification task and the second classification task.
[0056] Specifically, the video recognition device can invoke the second image feature extraction model to extract the second image feature from the video frame image. The second image feature extraction model is an image feature extraction model trained based on the first classification task and the second classification task. The first classification task is as described above and can be understood as distinguishing whether the target video is an advertisement video or not. The corresponding first image feature is the image feature that can determine whether the target video corresponding to the video frame image is an advertisement video or not.
[0057] In addition, the second classification task can be understood as distinguishing different objects in the video or different behaviors of the objects based on the video frame image, which can be used as a general video classification. The corresponding second image feature can not only be the image feature that can determine whether the target video corresponding to the video frame image is an advertisement video or not, but also be the image feature that can determine the objects in the video corresponding to the video frame image and the different behaviors of the objects.
[0058] 204. Invoke the first text feature extraction model to extract the first text feature of the target text. The first text feature extraction model is a text feature extraction model trained based on the first classification task.
[0059] Specifically, since the target video includes target text, the video recognition device calls the first text feature extraction model to extract first text features from the target text. The first text feature extraction model is a text feature extraction model trained based on a first classification task. The first classification task can be understood as distinguishing whether the target video is an advertisement video or not, which can be used as video advertisement classification. The corresponding first text features are the text features that can determine whether the target video corresponding to the target text is an advertisement video or not.
[0060] 205. Call a second text feature extraction model to extract second text features of the target text. The second text feature extraction model is a text feature extraction model trained based on the first classification task and a third classification task.
[0061] Specifically, the video recognition device calls the second text feature extraction model to extract second text features from the target text. The second text feature extraction model is a text feature extraction model trained based on the first classification task and a second classification task. The first classification task is as described above and can be understood as distinguishing whether the target video is an advertisement video or not. The corresponding first text features are the text features that can determine whether the target video corresponding to the target text is an advertisement video or not.
[0062] In addition, the third classification task can be understood as distinguishing different objects or different behaviors of objects in the video based on the target text, which can be used as a general video text classification. The corresponding second text features are the features that can determine the objects and different behaviors of the objects in the video corresponding to the target text. The third classification task can be the same as or different from the second classification task.
[0063] 206. Determine the video label of the target video according to the first image feature, the second image feature, the first text feature, and the second text feature, and determine that the video label of the target video belongs to the first classification task.
[0064] Specifically, the video device can determine the video label of the target video according to the first image feature, the second image feature, the first text feature, and the second text feature, and determine that the video label of the target video belongs to the above-mentioned first classification task, that is, it can be a video advertisement classification task. It can also be understood that determining the video label of the target video means that it can be determined whether the target video is an advertisement video through the video label. Of course, if the first classification task is other classification tasks, such as determining whether the video is an entertainment video, a funny video, a news and information video, a war history video, etc., it can also be determined whether the target video is a video of the corresponding category through the video label.
[0065] Optionally, if it is determined that the target video is an advertisement video segment according to the video tag of the target video, the target video can be deleted from the complete video, so that the complete video is a normal video without the advertisement video segment, realizing the filtering process of the complete video. Reducing the interference of the advertisement video segment to the complete video and improving the video playback effect.
[0066] It can be seen that in the embodiments of the present application, in order to complete the video tag recognition belonging to the first classification task, it is not only necessary to extract the image features and text feature problems in the field of the first classification task, but also necessary to refer to the image features and text feature problems in the fields of other classification tasks (i.e., the second classification task and the third classification task). Based on the video tags obtained by the diversified feature recognition in multiple classification task fields, the recognition accuracy can be guaranteed.
[0067] Consistent with the above Figure 2 shown embodiment, please refer to Figure 3A , Figure 3A is a schematic flowchart of another video recognition method provided by the embodiments of the present application. This method is applied to a mini-program generation device, and this method may include but is not limited to the following steps:
[0068] 301. Obtain a target video, where the target video includes video frame images and target text.
[0069] 302. Invoke a first image feature extraction model to extract the first image feature of the video frame image. The first image feature extraction model is an image feature extraction model trained based on the first classification task.
[0070] 303. Invoke a second image feature extraction model to extract the second image feature of the video frame image. The second image feature extraction model is an image feature extraction model trained based on the first classification task and the second classification task.
[0071] 304. Invoke a first text feature extraction model to extract the first text feature of the target text. The first text feature extraction model is a text feature extraction model trained based on the first classification task.
[0072] 305. Invoke a second text feature extraction model to extract the second text feature of the target text. The second text feature extraction model is a text feature extraction model trained based on the first classification task and the third classification task.
[0073] Steps 301-305 refer to the above steps 201-205 and will not be elaborated here.
[0074] 306. Invoke an image feature fusion model to fuse the first image feature and the second image feature into a first feature.
[0075] Specifically, the image feature fusion model may include at least two non-linear transformation layers, with a fully connected layer connected between the two non-linear transformation layers. One non-linear transformation layer is connected to the first image feature extraction model and the second image feature extraction model, and the other non-linear transformation layer is connected to the label recognition model. The image feature fusion model can be used to fuse the first image feature and the second image feature into a first feature. For example, the first image feature is the advertisement video feature, and the second image feature is the human motion image feature. The fused first feature may include the human motion image feature and the advertisement video image feature.
[0076] In addition, the image feature fusion model can be an Inside-Outside Net (ION), pixel-level fusion, feature-level fusion, decision-level fusion, etc. It can also be an Early fusion model. Early fusion is to fuse the features of multiple layers first, and then train a predictor on the fused features (detection is only unified after complete fusion). This type of method is also called skip connection, that is, using concat (concatenation) and add (addition) fusion methods. Concat is a series feature fusion that directly concatenates two features. If the dimensions of two input features x and y are p and q respectively, the dimension of the output feature z is p + q. Add is a parallel strategy that combines these two feature vectors into a complex vector. For input features x and y, z = x + iy, where i is the imaginary unit.
[0077] 307. Invoke the text feature fusion model to fuse the first text feature and the second text feature into a second feature.
[0078] Specifically, the text feature fusion model includes at least two non-linear transformation layers, with a fully connected layer connected between the two non-linear transformation layers. One non-linear transformation layer is connected to the first text feature extraction model and the second text feature extraction model, and the other non-linear transformation layer is connected to the label recognition model. It can also be an Early fusion model.
[0079] The video recognition device invokes the text feature fusion model to fuse the first text feature and the second text feature into a second feature. For example, the first text feature is the advertisement video feature, and the second text feature is the human motion text feature. The fused second feature may include the human motion text feature and the advertisement video text feature.
[0080] 308. Invoke the label recognition model to recognize the first feature and the second feature to obtain the video label of the target video. Among them, the image feature fusion model, the text feature fusion model, and the label recognition model are models trained based on the first classification task.
[0081] Specifically, the image recognition device can call the label recognition model to recognize the first feature and the second feature, and obtain the video label of the target video. Taking the first feature as the advertisement video feature including the human motion image feature and the second feature as the advertisement video feature including the human motion text feature as an example, after the image recognition device recognizes the first feature and the second feature, the video label of the target video obtained is the advertisement video feature including the human motion feature. Further, the target video can be determined as an advertisement video through this video label.
[0082] Among them, the image feature fusion model, the text feature fusion model, and the label recognition model are models trained based on the first classification task. The first classification task is as described above and will not be elaborated here.
[0083] It can be seen that after the video recognition device obtains the target video, it can respectively call the first image feature extraction model to extract the first image feature; call the second image feature extraction model to extract the second image feature of the video frame image. Call the first text feature extraction model to extract the first text feature of the target text; call the second text feature extraction model to extract the second text feature of the target text. That is, based on diverse feature extraction models, diverse features of the target video are obtained. Further, call the image feature fusion model to fuse the first image feature and the second image feature into the first feature; call the text feature fusion model to fuse the first text feature and the second text feature into the second feature; call the label recognition model to recognize the first feature and the second feature, and obtain the video label of the target video. Fusing diverse features based on diverse feature fusion models can effectively improve the accuracy of video label recognition.
[0084] Moreover, since the second image feature extraction model and the second text feature extraction model are trained not only based on the first classification task but also using general videos based on the second classification task, the training samples used are more easily obtained, which can effectively make up for the problem of insufficient sample size in the first classification task and significantly improve the training effect. As a result, when they finally perform target video recognition, the recognition effect of the overall model can be improved.
[0085] In a possible embodiment, the obtaining of the target video includes: obtaining the video frame image, recognizing the text in the video frame image, and using the recognized text as the target text; combining the video frame image and the target text into the target video.
[0086] Specifically, since the target video is composed of video frame images frame by frame, to obtain the target video, it is necessary to obtain the video frame images that make up the target video. After obtaining the video frame images, it is necessary to recognize the text in the video frame images and use the recognized text as the target text, that is, the text contained in the target video. Recognizing the text in the video frame images means extracting the subtitles in the video frame images. For the subtitles embedded in the video frame images, the Optical Character Recognition (OCR) technology needs to be used to extract the subtitles contained in each video segment. If some of the subtitles are in separate subtitle files, the text can be directly extracted from the files. By analyzing the video frame image information of the target video itself and the linguistic information of the target video subtitles, the image features and text features of the extracted target video are used to recognize the video tags. Further, the video frame images and the target text can be combined into the target video.
[0087] It can be seen that according to the obtained video frame images, the text in the video frame images is recognized, and the recognized text is used as the target text; further, the video frame images and the target text are combined into the target video. Subsequently, not only the video frame images of the target video but also the target text of the target video are used, so that when multiple feature extraction models perform feature extraction, diverse features can be extracted, thereby improving the accuracy of video tag determination.
[0088] In a possible embodiment, the method further includes: obtaining first sample data for a first classification task, where the first sample data includes first sample video frame images and first sample texts; invoking a first intermediate model to be trained for images to extract first sample image features of the first sample video frame images, and invoking a second intermediate model to be trained for images to extract second sample image features of the first sample video frame images; invoking a first intermediate model to be trained for texts to extract first sample text features of the first sample texts, and invoking a second intermediate model to be trained for texts to extract second sample text features of the first sample texts; invoking a to-be-trained image feature fusion model to fuse the first sample image features and the second sample image features into first sample features; invoking a to-be-trained text feature fusion model to fuse the first sample text features and the second sample text features into second sample features; invoking a to-be-trained label recognition model to recognize the first sample features and the second sample features, and obtain a first predicted label of the first sample data; obtaining a first sample label of the first sample data, and training the first intermediate model to be trained for images, the second intermediate model to be trained for images, the first intermediate model to be trained for texts, the second intermediate model to be trained for texts, the to-be-trained image feature fusion model, the to-be-trained text feature fusion model, and the to-be-trained label recognition model according to the first predicted label and the first sample label, so as to obtain a first image feature extraction model, a second image feature extraction model, a first text feature extraction model, a second text feature extraction model, an image feature fusion model, a text feature fusion model, and a label recognition model.
[0089] Specifically, it can be understood that the feature extraction model used in the video recognition stage is obtained by training an intermediate model with the first sample data. The first sample data can be a video with advertisement annotation or a video without advertisement annotation, that is, the true label of the first sample data is used to identify whether the first sample data is an advertisement video or not. The first sample data is used for the first classification task, and the first classification task can be an advertisement classification task, that is, to determine whether the video is an advertisement video or a non-advertisement video. As Figure 3B shown Figure 3BThe process of training an intermediate model based on first sample data to obtain first prediction labels. The first sample data includes first sample video frame images and first sample texts. The video recognition device calls a first image intermediate model to be trained to extract first sample image features of the first sample video frame images, and calls a second image intermediate model to be trained to extract second sample image features of the first sample video frame images. It calls a first text intermediate model to be trained to extract first sample text features of the first sample texts, and calls a second text intermediate model to be trained to extract second sample text features of the first sample texts. It calls an image feature fusion model to be trained to fuse the first sample image features and the second sample image features into first sample features. For example, the first sample image features are advertising video features, and the second sample image features are moving person image features. The first sample features obtained after the image feature fusion model to be trained fuses the two can be advertising video features containing moving person image features.
[0090] In addition, the video recognition device calls a text feature fusion model to be trained to fuse the first sample text features and the second sample text features into second sample features. For example, the first sample text features are advertising texts, and the second sample text features can be sentiment features. Such as positive sentiment features, negative sentiment features, etc. The second sample features obtained after the text feature fusion model to be trained fuses the two can be advertising texts with inspiring features.
[0091] Furthermore, since the label recognition model to be trained is a label recognition model based on a first classification task, therefore, it calls the label recognition model to be trained to recognize the first sample features and the second sample features, and the first prediction label of the first sample data can obtain the prediction result of the first classification task. For example, the first prediction label is an advertising label or a non-advertising label. Of course, the first prediction label can also carry other labels, such as a person movement label extracted based on image features, or a sentiment label extracted based on text features, etc.
[0092] Furthermore, the video recognition device obtains the first sample label of the first sample data (i.e., the true label of the first sample data), and trains the first intermediate model to be trained for images, the second intermediate model to be trained for images, the first intermediate model to be trained for text, the second intermediate model to be trained for text, the intermediate model to be trained for image feature fusion, the intermediate model to be trained for text feature fusion, and the label recognition model to be trained based on the first prediction label and the first sample label. That is, according to the difference between the first prediction label and the first sample label (i.e., the error), that is, according to the loss function of the above intermediate model or the model to be trained, the parameters of the above model are adjusted so that the above model gradually reaches the model convergence condition. The model convergence condition can be any one or more of the following: the loss value (i.e., the error) is less than a preset error threshold; or, the weight change (parameters) between two iterations is already very small, and a threshold can be set. When the weight change value is less than the parameter threshold, the training is stopped; or, the maximum number of iterations is set, and when the iteration exceeds the maximum number of times, the training is stopped, which can be regarded as reaching the model convergence condition. After convergence, the first image feature extraction model corresponding to the first intermediate model to be trained for images, the second image feature extraction model corresponding to the second intermediate model to be trained for images, the first text feature extraction model corresponding to the first intermediate model to be trained for text, the second text feature extraction model corresponding to the second intermediate model to be trained for text, the image feature fusion model corresponding to the intermediate model to be trained for image feature fusion, the text feature fusion model corresponding to the intermediate model to be trained for text feature fusion, and the label recognition model corresponding to the label recognition model to be trained are obtained.
[0093] After the training is completed, the comprehensive model that can recognize the video label of the target video can be as Figure 3C shown. Among them, the image feature fusion model includes at least two non-linear transformation layers, and a fully connected layer is connected between the two non-linear transformation layers. One non-linear transformation layer connects the first image feature extraction model and the second image feature extraction model, and the other non-linear transformation layer connects the label recognition model. Similarly, the text feature fusion model includes at least two non-linear transformation layers, and a fully connected layer is connected between the two non-linear transformation layers. One non-linear transformation layer connects the first text feature extraction model and the second text feature extraction model, and the other non-linear transformation layer connects the label recognition model. The label recognition model includes at least one fully connected layer.
[0094] It can be seen that the video recognition device calls multiple intermediate models based on the first classification task and jointly outputs the first prediction label of the first sample data. Then, the first sample label of the first sample data is obtained, and the above intermediate models are trained according to the first prediction label and the first sample label, so that each model can converge as much as possible. The video recognition ability of each model for the first classification task is improved, and the accuracy of each model in recognizing the target video is enhanced.
[0095] The following describes the specific process of obtaining the intermediate model to be trained for the second image:
[0096] In a possible embodiment, it further includes: obtaining second sample data for the second classification task; the second sample data includes second sample video frame images; based on the intermediate model to be trained for the second image, identifying the second predicted labels of the second sample video frame images; according to the second sample labels and the second predicted labels of the second sample data, training the intermediate model to be trained for the second image to obtain the intermediate model to be trained for the second image, where the number of the second sample data is greater than the number of the first sample data.
[0097] Specifically, the image recognition device obtains second sample data for the second classification task. As described above, the second classification task can be understood as distinguishing different objects or different behaviors of objects in a video based on video frame images, and can be used as a general video classification. Different from the first sample data that needs to be marked with advertisements, therefore, the number of the second sample data is much larger than the number of the first sample data.
[0098] The second sample data includes second sample video frame images. The intermediate model to be trained for the second image can be as Figure 3D shown, including Module 1, and further including at least two fully connected layers, and the two fully connected layers are connected by a non-linear transformation layer. Any one of the fully connected layers is connected to Module 1.
[0099] In addition, Module 1 is a video vector representation model, as Figure 3E shown, including at least one three-dimensional convolutional neural network (Convolutional Neural Networks, 3D-CNN), and further including at least two fully connected layers, and the two fully connected layers are connected by a non-linear transformation layer. Any one of the fully connected layers is connected to the 3D-CNN network. Since a complete video contains many video segments, for example, every 5 seconds, 3 seconds, 4 seconds, etc. are counted as a video segment. The second sample data contains a large number of video segments. This module will perform content analysis on the segmented video segments, and the 3D-CNN network will model the video segments. Finally, Module 1 will convert each video segment into a video vector, and this video vector is used as the representation of the video content.
[0100] Further, based on the second image to-be-trained model, identify the second predicted label of the second sample video frame image. That is, Module 1 inputs the video vector of the second sample video frame image into the second image to-be-trained model, and the second image to-be-trained model identifies the second predicted label of the second sample video frame image. Then, according to the second sample label and the second predicted label of the second sample data, determine the loss value of the second image to-be-trained model, and adjust the parameter value of the second image to-be-trained model according to this loss value, so that the second image to-be-trained model converges completely, and obtain the second intermediate model of the image to-be-trained. Since the second sample data is sample data for general image classification, the second sample data can be a large amount of data, such as 100,000 video segments, or 1,000,000 video segments, etc. Therefore, the second image to-be-trained model can be trained until it converges completely, so that it has good general video recognition ability.
[0101] It can be seen that based on a large amount of second sample data, the second image to-be-trained model is pre-trained until it converges completely to obtain the second intermediate model of the image to-be-trained. Improve the general video classification ability of the second intermediate model of the image to-be-trained. Further, using the first sample data to train the second intermediate model of the image to-be-trained again, the obtained second image feature extraction model can have better second image feature extraction ability, which is beneficial to improving the accuracy of target video label recognition for the first classification task.
[0102] The following describes the specific process of how to obtain the first intermediate model of the image to-be-trained:
[0103] In a possible embodiment, it further includes: based on the first image to-be-trained model, identify the original image predicted label of the first sample video frame image; according to the sample label and the original image predicted label of the first sample data, train the first image to-be-trained model to obtain the first intermediate model of the image to-be-trained.
[0104] Specifically, similar to the structure of the second image to-be-trained model, the structure of the first image to-be-trained model can be as Figure 3F shown, including Module 1, and further including at least two fully connected layers, which are connected by a non-linear transformation layer between the two fully connected layers. Any one of the fully connected layers is connected to Module 1. The structure of Module 1 is as described above and will not be elaborated here.
[0105] Further, based on the first image to-be-trained model, the original image prediction label of the first sample video frame image is recognized. That is, Module 1 inputs the video vector of the first sample video frame image into the first image to-be-trained model, and the first image to-be-trained model recognizes the original image prediction label of the first sample video frame image. Then, according to the sample label and the original image prediction label of the first sample data, the loss function of the first image to-be-trained model is determined, and the parameter values of the first image to-be-trained model are adjusted according to this loss function, so that the first image to-be-trained model converges as much as possible (that is, it is trained until all the first sample data participate in the training of the first image to-be-trained model), and the first image to-be-trained intermediate model is obtained. However, since the first sample data comes from manual annotation, it will annotate whether each video segment is an advertisement segment. However, since the annotation data needs to be done manually, it is difficult, time-consuming and costly, so we can only collect less data in this part. For example, 50 video segments, 100, or 500, 20, etc. Therefore, it cannot be guaranteed that the first image to-be-trained model can be trained to complete convergence using the first sample data, that is, the first image to-be-trained intermediate model is not necessarily completely convergent.
[0106] It can be seen that the first image to-be-trained model is pre-trained to converge as much as possible based on a small amount of first sample data to obtain the first image to-be-trained intermediate model. Improving the ability of the first classification task of the first image to-be-trained intermediate model can be the advertisement classification ability. So that, subsequently, the first sample data is further used to train the first image to-be-trained intermediate model and the second image to-be-trained intermediate model simultaneously again to obtain the first image feature extraction model and the second image feature extraction model. Strengthening the image feature extraction ability of the overall model through the second image feature extraction model helps the video label recognition result of the first classification task for the target video to be more accurate.
[0107] The following describes the specific process of how to obtain the second text to-be-trained intermediate model:
[0108] In a possible embodiment, it further includes: obtaining third sample data for a third classification task; the third sample data includes third sample texts; based on the second text to-be-trained model, recognizing the third prediction label of the third sample texts; training the second text to-be-trained model according to the third sample label and the third prediction label of the third sample data to obtain the second text to-be-trained intermediate model, where the quantity of the third sample data is greater than the quantity of the first sample data.
[0109] Specifically, the image recognition device obtains third sample data for the third classification task. The third classification task can be the same as or different from the second classification task. It can be understood that different objects, different behaviors corresponding to different objects, or different emotions of the text are distinguished from the sample text based on the third sample data. It can be used as a general text classification. Different from the first sample data that requires advertisement annotation, the third sample data does not require any supervised data and can directly extract text captions from all videos for text input and vectorized output of the text. Therefore, the quantity of the third sample data is much larger than that of the first sample data.
[0110] The third sample data includes third sample text. The second text to-be-trained model can be as Figure 3G shown on the left, including Module 2 and at least two fully connected layers, which are connected by a non-linear transformation layer between the two fully connected layers. Any one of the fully connected layers is connected to Module 2.
[0111] In addition, Module 2 is a text vectorization representation model. And to improve the accuracy and efficiency of the text vectorization representation of Module 2, the text vectorization representation model can be pre-trained. The structure of this model is as Figure 3H shown. When training the text vectorization representation model, an unsupervised task can be performed. That is, input the sample text, which can be the third sample text or the first sample text. The sample text is transformed into a vector through a Recurrent Neural Network (RNN), and the text vector is output; the output text vector is input into another RNN to reconstruct the original sentence, that is, the sample text. Training the text vectorization representation model based on a large amount of data can enable the model to have the ability to understand text and a strong text vectorization representation ability.
[0112] Furthermore, based on the second text to-be-trained model, the third predicted label of the third sample text is recognized. That is, Module 2 inputs the text vector of the third sample text into the second text to-be-trained model, and the second text to-be-trained model recognizes the third predicted label of the third sample text. Then, according to the third sample label and the third predicted label of the third sample data, the loss function of the second text to-be-trained model is determined, and the parameter values of the second text to-be-trained model are adjusted according to this loss function so that the second image to-be-trained model converges completely, obtaining the second image to-be-trained intermediate model. Since the second sample data is a large amount of data, such as 200,000 video segments, 500,000, or 1,000,000 video segments, etc., the second text to-be-trained model can be trained until it converges completely, enabling it to have good general text recognition ability.
[0113] It can be seen that based on a large amount of third-sample data, the second text to-be-trained model is pre-trained until it converges completely to obtain the second text to-be-trained intermediate model. The general text classification ability of the second text to-be-trained intermediate model is improved. Further, using the first-sample data to train the second text to-be-trained intermediate model again, the obtained second text feature extraction model can have better second text feature extraction ability, which is beneficial to improving the accuracy of target video label recognition for the first classification task.
[0114] The following describes the specific process of how to obtain the first text to-be-trained intermediate model:
[0115] In a possible embodiment, it further includes: based on the first text to-be-trained model, identifying the original text prediction label of the first sample text; according to the sample label and the original text prediction label of the first sample data, training the first text to-be-trained model to obtain the first text to-be-trained intermediate model.
[0116] Specifically, similar to the structure of the second text to-be-trained model, the structure of the first text to-be-trained model can be as Figure 3I shown, including module two, and further including at least two fully connected layers, which are connected by a non-linear transformation layer between the two fully connected layers. Any one of the fully connected layers is connected to module two. The structure of module two is as described above and will not be elaborated here. Through module two, the first sample text can be transformed into a corresponding text vector.
[0117] Furthermore, based on the first text to-be-trained model, the original text prediction label of the first sample text is identified. That is, module two inputs the text vector of the first sample text into the first text to-be-trained model, and the first text to-be-trained model identifies the original text prediction label of the first sample text. Then, according to the sample label and the original text prediction label of the first sample data, the loss function of the first text to-be-trained model is determined, and the parameter values of the first text to-be-trained model are adjusted according to this loss function, so that the first text to-be-trained model converges as much as possible to obtain the first text to-be-trained intermediate model. However, since the first sample data comes from manual annotation, it will be marked whether each video segment is an advertisement segment. However, since the annotation data requires manual operation, it is difficult, time-consuming and costly, so we can only collect less data in this part. For example, 30 video segments, 100, or 500, 20, etc. Therefore, it cannot be guaranteed that the first text to-be-trained model can be trained until it converges completely using the first sample data, that is, the first text to-be-trained intermediate model is not necessarily completely convergent.
[0118] It can be seen that by pre-training the first text model to be trained based on a small amount of first sample data until it converges as much as possible, the first intermediate text model to be trained is obtained, which can effectively improve the ability of the first intermediate text model in the first classification task. The first classification task can be an advertisement classification task. Then, the first intermediate text model to be trained and the second intermediate text model to be trained can be further trained simultaneously using the first sample data to obtain the first text feature extraction model and the second text feature extraction model. Strengthening the text feature extraction ability of the overall model through the second text feature extraction model helps to make the video label recognition result of the first classification task for the target video more accurate.
[0119] Summarizing the above process, first is the separate training of each model to be trained, including training the first image model to be trained and the first text model to be trained based on the first classification task, training the first text model to be trained based on the second classification task, and training the second text model to be trained based on the third classification task, so that each model to be trained converges as much as possible to obtain the intermediate model corresponding to each model to be trained and improve the image recognition ability of the intermediate model.
[0120] Then, the intermediate models are jointly trained based on the first classification task. During the joint training, it can be considered that the intermediate models have undergone the above-mentioned pre-training. Therefore, even if a small amount of first samples participate in the joint training, the model convergence condition can be achieved. That is, the present application reduces the requirement for the amount of first sample data belonging to the first classification task and instead pre-trains the models to be trained with data from other classification fields (i.e., the second classification task and the third classification task) to obtain the intermediate models. Other classification fields can be understood as general classification fields.
[0121] In terms of the volume of sample data, the training model is a semi-supervised model, that is, both advertisement-annotated video data and a large amount of general classification video (without advertisement annotation) data are used to pre-train the models to be trained, improving the image and video capabilities of each intermediate model to assist in completing the first classification task.
[0122] After the model training is completed, the overall diagram obtained can be as Figure 3J shown, Figure 3J in which each feature extraction model, each feature fusion model, and the label recognition model are used for the first classification task, and the first classification task can be an advertisement video classification task.
[0123] Among them, the second image feature extraction model is obtained by training the second intermediate image model to be trained. Since the second intermediate image model to be trained is obtained by training the second image model to be trained with general video data until it converges, the second intermediate image model to be trained has the general video classification ability. Therefore,Figure 3J The second image feature extraction model marked in the bid is used for general video classification tasks, which can be understood as being completed by the second intermediate model to be trained for images. Both the first image feature extraction model and the second image feature extraction model include Module 1.
[0124] Similarly, the second text feature extraction model is obtained after being trained by the second intermediate model to be trained for texts. Since the second intermediate model to be trained for texts is obtained by training the second model to be trained for texts with general text data until convergence, the second intermediate model to be trained for texts has the ability of general text classification. Therefore, Figure 3J The second text feature extraction model marked in the bid is used for general text classification tasks, which can be understood as being completed by the second intermediate model to be trained for texts. Both the first text feature extraction model and the second text feature extraction model include Module 2.
[0125] In addition, when using the total model to identify the video label of the target video for the first classification task, the image feature fusion model can fuse the image features output by the first image feature extraction model and the second image feature extraction model to obtain the first feature, and input the fused first feature into the label recognition model. The label recognition model may include at least one fully connected layer. The text feature fusion model can fuse the text features output by the first text feature extraction model and the second text feature extraction model to obtain the second feature, and input the fused second feature into the label recognition model. The label recognition model outputs the video label based on the first feature and the second feature. Thus, it can be seen that the video label is obtained based on the first classification task. Therefore, it can be determined whether the target video is an advertising video through the video label.
[0126] Please refer to Figure 4 which is a schematic diagram of the functional units of an image recognition device 400 according to an embodiment of the present invention. The image recognition device 400 according to an embodiment of the present application may be the image recognition device in the corresponding Figures 1 - 3J embodiment. The image recognition device 400 may be a computer program (including program code) running in a computer device. For example, the image recognition device is an application software.
[0127] In one implementation manner of the device according to an embodiment of the present invention, the device includes:
[0128] An obtaining unit 410, configured to obtain a target video, where the target video includes video frame images and target texts;
[0129] A processing unit 420, configured to call the first image feature extraction model to extract the first image feature of the video frame image; the first image feature extraction model is an image feature extraction model trained based on the first classification task;
[0130] The processing unit 420 is further configured to call a second image feature extraction model to extract second image features of the video frame image; the second image feature extraction model is an image feature extraction model trained based on the first classification task and the second classification task;
[0131] The processing unit 420 is further configured to call a first text feature extraction model to extract first text features of the target text; the first text feature extraction model is a text feature extraction model trained based on the first classification task;
[0132] The processing unit 420 is further configured to call a second text feature extraction model to extract second text features of the target text; the second text feature extraction model is a text feature extraction model trained based on the first classification task and the third classification task;
[0133] The determination unit 430 is configured to determine a video label of the target video according to the first image feature, the second image feature, the first text feature, and the second text feature, and determine that the video label of the target video belongs to the first classification task.
[0134] In a possible embodiment, in terms of determining the video label of the target video according to the first image feature, the second image feature, the first text feature, and the second text feature, the determination unit 430 is specifically configured to: call an image feature fusion model to fuse the first image feature and the second image feature into a first feature; call a text feature fusion model to fuse the first text feature and the second text feature into a second feature; call a label recognition model to recognize the first feature and the second feature to obtain the video label of the target video; wherein, the image feature fusion model, the text feature fusion model, and the label recognition model are models trained based on the first classification task.
[0135] In a possible embodiment, in terms of obtaining the target video, the obtaining unit 410 is specifically configured to: obtain the video frame image, recognize the text in the video frame image, and use the recognized text as the target text; combine the video frame image and the target text into the target video.
[0136] In a possible embodiment, the processing unit 420 is further configured to: obtain first sample data for a first classification task, where the first sample data includes first sample video frame images and first sample text; call a first intermediate model to be trained for images to extract first sample image features of the first sample video frame images, and call a second intermediate model to be trained for images to extract second sample image features of the first sample video frame images; call a first intermediate model to be trained for text to extract first sample text features of the first sample text, and call a second intermediate model to be trained for text to extract second sample text features of the first sample text; call a to-be-trained image feature fusion model to fuse the first sample image features and the second sample image features into first sample features; call a to-be-trained text feature fusion model to fuse the first sample text features and the second sample text features into second sample features; call a to-be-trained label recognition model to recognize the first sample features and the second sample features, and obtain a first predicted label of the first sample data; obtain a first sample label of the first sample data, and train the first intermediate model to be trained for images, the second intermediate model to be trained for images, the first intermediate model to be trained for text, the second intermediate model to be trained for text, the to-be-trained image feature fusion model, the to-be-trained text feature fusion model, and the to-be-trained label recognition model according to the first predicted label and the first sample label, so as to obtain a first image feature extraction model, a second image feature extraction model, a first text feature extraction model, a second text feature extraction model, an image feature fusion model, a text feature fusion model, and a label recognition model.
[0137] In a possible embodiment, the processing unit 420 is further configured to: obtain second sample data for a second classification task; the second sample data includes second sample video frame images; based on a second to-be-trained model for images, recognize a second predicted label of the second sample video frame images; and train the second to-be-trained model for images according to a second sample label and the second predicted label of the second sample data to obtain the second intermediate model to be trained for images, where the quantity of the second sample data is greater than that of the first sample data.
[0138] In a possible embodiment, the processing unit 420 is further configured to: based on a first to-be-trained model for images, recognize an original image predicted label of the first sample video frame images; and train the first to-be-trained model for images according to a sample label and the original image predicted label of the first sample data to obtain the first intermediate model to be trained for images.
[0139] In a possible embodiment, the processing unit 420 is further configured to: obtain third sample data for a third classification task; the third sample data includes third sample texts; based on a second text to-be-trained model, identify third prediction labels of the third sample texts; and train the second text to-be-trained model according to the third sample labels and the third prediction labels of the third sample data to obtain a second text to-be-trained intermediate model, wherein the quantity of the third sample data is greater than that of the first sample data.
[0140] In a possible embodiment, the processing unit 420 is further configured to: based on a first text to-be-trained model, identify original text prediction labels of the first sample texts; and train the first text to-be-trained model according to the sample labels and the original text prediction labels of the first sample data to obtain a first text to-be-trained intermediate model.
[0141] In some embodiments, the video recognition device may further include an input / output interface, a communication interface, a power supply, and a communication bus.
[0142] Embodiments of the present application may perform a functional unit division on the video recognition device according to the above method examples. For example, each functional unit may be corresponding to each function, or two or more functions may be integrated into one processing unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, and is only a logical functional division. There may be other division methods in actual implementation.
[0143] Please refer again to Figure 5 , which is a schematic structural diagram of an electronic device according to an embodiment of the present invention. The foregoing Figures 1 - 3J image recognition device in the corresponding embodiment may be applied to the electronic device. The electronic device includes structures such as a power supply module, and includes a processor 501, a storage device 502, and a communication interface 503. Data can be exchanged between the processor 501, the storage device 502, and the communication interface 503.
[0144] The storage device 502 may include a volatile memory, such as a random-access memory (RAM); the storage device 502 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the storage device 502 may further include a combination of the above types of memories. The communication interface 503 is an interface for data interaction between internal devices of the electronic device, such as between the storage device 502 and the processor 501.
[0145] The processor 501 may be a central processing unit (CPU). In one embodiment, the processor 501 may also be a Graphics Processing Unit (GPU). The processor 501 may also be a combination of a CPU and a GPU. In one embodiment, the storage device 502 is used to store program instructions. The processor 501 may call the program instructions to perform the following steps:
[0146] Obtain a target video, where the target video includes video frame images and target text;
[0147] Call a first image feature extraction model to extract first image features of the video frame images; the first image feature extraction model is an image feature extraction model trained based on a first classification task;
[0148] Call a second image feature extraction model to extract second image features of the video frame images; the second image feature extraction model is an image feature extraction model trained based on the first classification task and a second classification task;
[0149] Call a first text feature extraction model to extract first text features of the target text; the first text feature extraction model is a text feature extraction model trained based on the first classification task;
[0150] Call a second text feature extraction model to extract second text features of the target text; the second text feature extraction model is a text feature extraction model trained based on the first classification task and a third classification task;
[0151] Determine a video label of the target video according to the first image features, the second image features, the first text features, and the second text features, and determine that the video label of the target video belongs to the first classification task.
[0152] In a possible embodiment, when determining the video label of the target video according to the first image feature, the second image feature, the first text feature, and the second text feature, the processor 501 is specifically configured to: call an image feature fusion model to fuse the first image feature and the second image feature into a first feature; call a text feature fusion model to fuse the first text feature and the second text feature into a second feature; call a label recognition model to recognize the first feature and the second feature to obtain the video label of the target video; wherein, the image feature fusion model, the text feature fusion model, and the label recognition model are models trained based on the first classification task.
[0153] In a possible embodiment, when obtaining the target video, the processor 501 is specifically configured to: obtain the video frame image, recognize the text in the video frame image, and use the recognized text as the target text; combine the video frame image and the target text into the target video.
[0154] In a possible embodiment, the processor 501 is further configured to: obtain first sample data for a first classification task, where the first sample data includes a first sample video frame image and a first sample text; call a first image intermediate model to be trained to extract a first sample image feature of the first sample video frame image, and call a second image intermediate model to be trained to extract a second sample image feature of the first sample video frame image; call a first text intermediate model to be trained to extract a first sample text feature of the first sample text, and call a second text intermediate model to be trained to extract a second sample text feature of the first sample text; call an image feature fusion model to be trained to fuse the first sample image feature and the second sample image feature into a first sample feature; call a text feature fusion model to be trained to fuse the first sample text feature and the second sample text feature into a second sample feature; call a label recognition model to be trained to recognize the first sample feature and the second sample feature to obtain a first predicted label of the first sample data; obtain a first sample label of the first sample data, and train the first image intermediate model to be trained, the second image intermediate model to be trained, the first text intermediate model to be trained, the second text intermediate model to be trained, the image feature fusion model to be trained, the text feature fusion model to be trained, and the label recognition model to be trained according to the first predicted label and the first sample label, so as to obtain a first image feature extraction model, a second image feature extraction model, a first text feature extraction model, a second text feature extraction model, an image feature fusion model, a text feature fusion model, and a label recognition model.
[0155] In a possible embodiment, the processor 501 is further configured to: obtain second sample data for a second classification task; the second sample data includes second sample video frame images; based on a second image to-be-trained model, identify second predicted labels of the second sample video frame images; according to the second sample labels and the second predicted labels of the second sample data, train the second image to-be-trained model to obtain the second image to-be-trained intermediate model, where the quantity of the second sample data is greater than that of the first sample data.
[0156] In a possible embodiment, the processor 501 is further configured to: based on a first image to-be-trained model, identify original image predicted labels of the first sample video frame images; according to the sample labels and the original image predicted labels of the first sample data, train the first image to-be-trained model to obtain the first image to-be-trained intermediate model.
[0157] In a possible embodiment, the processor 501 is further configured to: obtain third sample data for a third classification task; the third sample data includes third sample texts; based on a second text to-be-trained model, identify third predicted labels of the third sample texts; according to the third sample labels and the third predicted labels of the third sample data, train the second text to-be-trained model to obtain the second text to-be-trained intermediate model, where the quantity of the third sample data is greater than that of the first sample data.
[0158] In a possible embodiment, the processor 501 is further configured to: based on a first text to-be-trained model, identify original text predicted labels of the first sample texts; according to the sample labels and the original text predicted labels of the first sample data, train the first text to-be-trained model to obtain the first text to-be-trained intermediate model.
[0159] The embodiment of the present application further provides a computer storage medium, where the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute some or all of the steps of any method described in the above method embodiments.
[0160] The embodiment of the present application further provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes some or all of the steps of any method described in the above method embodiments.
[0161] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0162] The above-disclosed are only some embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A video recognition method, characterized in that, The method includes: Obtain a target video, where the target video includes video frame images and target text; Call a first image feature extraction model to extract first image features of the video frame images; the first image feature extraction model is an image feature extraction model trained based on a first classification task; the first classification task is a task of distinguishing whether the target video is an advertisement video; Call a second image feature extraction model to extract second image features of the video frame images; the second image feature extraction model is an image feature extraction model trained based on the first classification task and a second classification task; the second image features obtained by the image feature extraction model trained based on the first classification task and the second classification task are used to determine whether the target video is an advertisement video and to determine the objects in the target video and different behaviors of the objects; Call a first text feature extraction model to extract first text features of the target text; the first text feature extraction model is a text feature extraction model trained based on the first classification task; Call a second text feature extraction model to extract second text features of the target text; the second text feature extraction model is a text feature extraction model trained based on the first classification task and a third classification task; Determine a video label of the target video according to the first image features, the second image features, the first text features, and the second text features, and determine that the video label of the target video belongs to the first classification task, and the video label is used to determine whether the target video is an advertisement video.
2. The method according to claim 1, characterized in that, The determining the video label of the target video according to the first image features, the second image features, the first text features, and the second text features includes: Call an image feature fusion model to fuse the first image features and the second image features into first features; Call a text feature fusion model to fuse the first text features and the second text features into second features; Call a label recognition model to recognize the first features and the second features to obtain the video label of the target video; Wherein, the image feature fusion model, the text feature fusion model, and the label recognition model are models trained based on the first classification task.
3. The method according to claim 1, characterized in that, The obtaining the target video includes: Obtain the video frame images, recognize the text in the video frame images, and use the recognized text as the target text; Combine the video frame images and the target text into the target video.
4. The method according to claim 1, wherein The method further includes: Obtain first sample data for the first classification task, where the first sample data includes first sample video frame images and first sample text; Call a first image intermediate model to be trained to extract first sample image features of the first sample video frame images, and call a second image intermediate model to be trained to extract second sample image features of the first sample video frame images; Call a first text intermediate model to be trained to extract first sample text features of the first sample text, and call a second text intermediate model to be trained to extract second sample text features of the first sample text; Call the image feature fusion model to be trained to fuse the first sample image feature and the second sample image feature into a first sample feature; Call the text feature fusion model to be trained to fuse the first sample text feature and the second sample text feature into a second sample feature; Call the label recognition model to be trained to recognize the first sample feature and the second sample feature, and obtain a first predicted label of the first sample data; Obtain the first sample label of the first sample data, and train the first image intermediate model to be trained, the second image intermediate model to be trained, the first text intermediate model to be trained, the second text intermediate model to be trained, the image feature fusion model to be trained, the text feature fusion model to be trained, and the label recognition model to be trained according to the first predicted label and the first sample label, so as to obtain a first image feature extraction model, a second image feature extraction model, a first text feature extraction model, a second text feature extraction model, an image feature fusion model, a text feature fusion model, and a label recognition model.
5. The method according to claim 4, wherein It further includes: Obtain second sample data for a second classification task; the second sample data includes second sample video frame images; Based on the second image model to be trained, recognize the second predicted label of the second sample video frame image; Train the second image model to be trained according to the second sample label and the second predicted label of the second sample data to obtain the second image intermediate model to be trained, where the number of the second sample data is greater than the number of the first sample data.
6. The method according to claim 4, wherein It further includes: Based on the first image model to be trained, recognize the original image predicted label of the first sample video frame image; Train the first image model to be trained according to the sample label and the original image predicted label of the first sample data to obtain the first image intermediate model to be trained.
7. The method according to claim 4, characterized in that It further includes: Obtain third sample data for a third classification task; the third sample data includes third sample texts; Based on the second text model to be trained, recognize the third predicted label of the third sample text; Train the second text model to be trained according to the third sample label and the third predicted label of the third sample data to obtain the second text intermediate model to be trained, where the number of the third sample data is greater than the number of the first sample data.
8. The method according to claim 4, characterized in that It further includes: Based on the first text model to be trained, recognize the original text predicted label of the first sample text; Train the first text model to be trained according to the sample label and the original text predicted label of the first sample data to obtain the first text intermediate model to be trained.
9. A method for a video recognition device, characterized in that The device includes: An acquisition unit, configured to acquire a target video, where the target video includes video frame images and target texts; A processing unit, configured to call a first image feature extraction model to extract a first image feature of the video frame image; the first image feature extraction model is an image feature extraction model trained based on a first classification task; the first classification task is a task of distinguishing whether the target video is an advertisement video; The processing unit is further configured to call a second image feature extraction model to extract second image features of the video frame image; the second image feature extraction model is an image feature extraction model trained based on the first classification task and the second classification task; the second image features obtained by the image feature extraction model trained based on the first classification task and the second classification task are used to determine whether the target video is an advertisement video, and are used to determine the objects in the target video and different behaviors of the objects. The processing unit is further configured to call a first text feature extraction model to extract first text features of the target text; the first text feature extraction model is a text feature extraction model trained based on the first classification task. The processing unit is further configured to call a second text feature extraction model to extract second text features of the target text; the second text feature extraction model is a text feature extraction model trained based on the first classification task and the third classification task. The determining unit is configured to determine a video label of the target video according to the first image feature, the second image feature, the first text feature, and the second text feature, determine that the video label of the target video belongs to the first classification task, and the video label is used to determine whether the target video is an advertisement video.
10. An electronic device, characterized in that, It includes a processor, a storage device, a communication interface, and one or more programs, the one or more programs are stored in the storage device, and are configured to be executed by the processor to perform the method according to any one of claims 1-8.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Video classification method and device, storage medium and server
CN111209970A
Video splitting method and device, electronic equipment and storage medium
CN111541939A