Video recognition method, video recognition model training method, medium and electronic device
By extracting video frames and combining them with video violation categories as prior information, and using a video recognition model for feature fusion, the problem of insufficient video recognition accuracy in existing technologies is solved, achieving higher recognition accuracy and consistency.
Patent Information
- Application Number
- CN202210964357.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-11
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-08-11
AI Technical Summary
Existing technologies are insufficient to effectively identify illegal content in internet videos, especially violent and pornographic videos, and their accuracy and consistency in identification are inadequate.
By extracting image frames from the video to be identified, and taking the video violation category and the image frames as input, the video recognition model is used to perform feature fusion. The video violation category is combined as prior information to extract the violation location and category of the image frames.
This improves the accuracy of video recognition models in identifying image frames, ensuring that the location and category of the violation image remain consistent with the violation category in the video, thus enhancing the accuracy and consistency of the recognition.
Smart Images

Figure CN115294501B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image technology, and more specifically, to a video recognition method, a video recognition model training method, an apparatus, a medium, and an electronic device. Background Technology
[0002] With the rapid development of internet technology, online streaming media resources have exploded. At the same time, a large number of videos involving violence, pornography, and other illegal content are also spreading rapidly online. Therefore, higher requirements are being placed on the identification of video content. Summary of the Invention
[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, embodiments of this disclosure provide a video recognition method, including:
[0005] Extract image frames from the video to be identified;
[0006] The video violation category and the image frame are used as input to the video recognition model to obtain the image recognition result of the image frame. The image recognition result includes the violation image location corresponding to the image frame and / or the image violation category to which the violation image location belongs. The video violation category is used to make the features extracted by the video recognition model from the image frame related to the video violation category.
[0007] Secondly, embodiments of this disclosure provide a video recognition model training method, including:
[0008] Obtain a training image set, wherein the training image set includes at least one image sample, the image sample having a first label and a second label, wherein the first label includes a bounding box for marking the location of the violation image in the image sample and / or the image violation category to which the bounding box belongs, and the second label includes the video violation category corresponding to the image sample;
[0009] The machine learning model is trained using the training image set to obtain a video recognition model.
[0010] Thirdly, embodiments of this disclosure provide a video recognition device, including:
[0011] The determination module is configured to determine the video violation category corresponding to the video to be identified;
[0012] The extraction module is configured to extract image frames from the video to be identified;
[0013] The recognition module is configured to take the video violation category and the image frame as input to the video recognition model to obtain the image recognition result of the image frame. The image recognition result includes the violation image location corresponding to the image frame and / or the image violation category to which the violation image location belongs. The video violation category is used to make the features extracted by the video recognition model from the image frame related to the video violation category.
[0014] Fourthly, embodiments of this disclosure provide a video recognition model training apparatus, comprising:
[0015] The acquisition module is configured to acquire a training image set, wherein the training image set includes at least one image sample, the image sample has a first label and a second label, wherein the first label includes a bounding box for marking the location of the violation image in the image sample and the image violation category of the image to which the bounding box belongs, and the second label includes the video violation category corresponding to the image sample;
[0016] The training module is configured to train the machine learning model using the training image set to obtain a video recognition model.
[0017] Fifthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the video recognition method described in the first aspect, or the steps of the video recognition model training method described in the second aspect.
[0018] Sixthly, embodiments of this disclosure provide an electronic device, including:
[0019] A storage device on which computer programs are stored;
[0020] A processing device is configured to execute the computer program in the storage device to implement the steps of the video recognition method described in the first aspect, or to implement the steps of the video recognition model training method described in the second aspect.
[0021] Based on the above technical solution, by using video violation categories and image frames as input to the video recognition model, image recognition results for the image frames are obtained. Since the video violation category provides prior information for the video recognition model to recognize the image frames, the features extracted by the video recognition model from the image frames are related to the video violation category, thereby making the obtained violation image location and / or image violation category more accurate. Furthermore, it ensures that the violation image location and / or image violation category of the obtained image frames are consistent with the video violation category of the video to be recognized. For example, the violation image location and / or image violation category output by the video recognition model can be consistent with the video violation category.
[0022] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0023] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0024] Figure 1 This is a flowchart illustrating a video recognition method according to some embodiments.
[0025] Figure 2 This is a schematic diagram illustrating an application scenario of a video recognition method based on some embodiments.
[0026] Figure 3 This is a structural schematic diagram of a video recognition model according to some embodiments.
[0027] Figure 4 This is a structural schematic diagram of a video recognition model according to some other embodiments.
[0028] Figure 5 This is a structural schematic diagram of a video recognition model according to some other embodiments.
[0029] Figure 6 This is a flowchart illustrating a video recognition model training method according to some embodiments.
[0030] Figure 7 This is a schematic diagram of the module connection of a video recognition device according to some embodiments.
[0031] Figure 8 This is a schematic diagram of the module connections of a video recognition model training device according to some embodiments.
[0032] Figure 9This is a schematic diagram of the structure of an electronic device according to some embodiments. Detailed Implementation
[0033] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0034] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0035] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0036] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0037] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0038] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0039] Figure 1 This is a flowchart illustrating a video recognition method according to some embodiments. For example... Figure 1 As shown, this disclosure provides a video recognition and playback method, which can be executed by an electronic device, specifically by a video recognition and playback device. This device can be implemented in software and / or hardware and configured within the electronic device. Figure 1 As shown, the method may include the following steps.
[0040] In step 110, the video violation category corresponding to the video to be identified is determined.
[0041] Here, the video to be identified can refer to any video uploaded to the internet by a user through various video applications. The video violation category refers to the overall violation category to which the video to be identified belongs. For example, video violation categories may include categories such as "sale of prohibited items," "pornography," "violence," and "vulgarity." It should be understood that there can be one or more video violation categories.
[0042] In some embodiments, the video to be identified can be used as input to a video detection model to obtain the video violation category.
[0043] The video detection model is obtained by training a machine learning model using video samples labeled with video violation categories.
[0044] Here, the video detection model is a pre-trained neural network model capable of accurately scoring videos under different violation categories, such as deep neural networks (DNN) or convolutional neural networks (CNN). By using the video to be identified as input to the video detection model, the model outputs the corresponding video violation category.
[0045] In step 120, image frames are extracted from the video to be identified.
[0046] Here, the image frame is a video frame extracted from the video to be identified. The image frame can be a video frame at certain moments in the video to be identified that can cover most of the screen features of the video to be identified.
[0047] As some examples, each video frame in the video to be identified can be extracted as an image frame.
[0048] As another example, several video frames can be extracted from the video to be identified as image frames according to a preset time interval.
[0049] As other examples, a fixed number of video frames can be extracted from the video to be identified as image frames according to a preset total number of frames. For example, regardless of the length of the video to be identified, 10 frames can be extracted as image frames.
[0050] In step 130, the video violation category and the image frame are used as input to the video recognition model to obtain the image recognition result of the image frame. The image recognition result includes the violation image location corresponding to the image frame and / or the image violation category to which the violation image location belongs. The video violation category is used to make the features extracted by the video recognition model from the image frame related to the video violation category.
[0051] Here, the video violation category and image frame can be used as input to the video recognition model to obtain the image recognition result of the image frame. This image recognition result includes the location of the violation image corresponding to the image frame and / or the image violation category to which the violation image location belongs. The violation image location refers to the location of the image containing the violation within the image frame, which can be marked using bounding boxes or image coordinates. The image violation category refers to the violation category corresponding to the image at the violation image location, such as categories like "prohibited goods sale," "pornography," "violence," and "vulgarity."
[0052] It should be understood that the image recognition results output by the video recognition model can be set according to the actual situation, and are independent of the image recognition logic of the video recognition model. For example, the output result of the video recognition model can be set to output the location of the violation image, or it can be set to output the location of the violation image and the video violation category.
[0053] It is worth noting that the video violation category reflects the violation problem in the video to be identified. By combining the video violation category, the video recognition model can use the video violation category as prior information to accurately extract features related to the video violation category from the image frame, thereby enabling the determination of the location of the violation image and / or the image violation category related to the video violation category based on the features.
[0054] Figure 2 This is a schematic diagram illustrating an application scenario of a video recognition method based on some embodiments. For example... Figure 2 As shown, in practical applications, the video to be identified can be used as input to the video detection model to obtain the video violation category. If the video violation category indicates that the video to be identified does not have a violation, the video recognition process ends. If the video violation category indicates that the video to be identified has a violation, image frames are extracted from the video to be identified, and these image frames, along with the video violation category, are used as input to the video recognition model to obtain the image recognition result.
[0055] Therefore, by using video violation categories and image frames as input to the video recognition model, image recognition results for the image frames are obtained. Since the video violation category provides prior information for the video recognition model to recognize the image frames, the features extracted by the video recognition model from the image frames are related to the video violation category, thus making the obtained violation image location and / or image violation category more accurate. Furthermore, it ensures that the violation image location and / or image violation category of the obtained image frames are consistent with the video violation category of the video to be recognized. For example, the violation image location and / or image violation category output by the video recognition model can be consistent with the video violation category.
[0056] In some feasible implementations, the video recognition model is configured as follows:
[0057] The text features of the video violation category and the image features of the image frame are fused to obtain fused features, and the image recognition result is obtained based on the fused features.
[0058] Here, the video recognition model can process video violation categories into text features through an embedding layer, and process image frames into image features through Convolutional Neural Networks (CNNs). Then, the text features and image features are fused to obtain fused features. Furthermore, the corresponding image recognition result can be obtained using these fused features.
[0059] For example, a video recognition model can fuse text features with image features by concatenating the text features and image features to obtain fused features.
[0060] Therefore, by configuring the video recognition model to fuse text features of video violation categories and image features of image frames to obtain fused features, the extracted image features and prediction results can be strongly correlated with video violation categories during the feature learning and prediction stages, thereby making the obtained image recognition results more accurate.
[0061] Figure 3 This is a structural schematic diagram of a video recognition model based on some embodiments. For example... Figure 3 As shown, in some feasible implementations, the video recognition model includes a first feature extraction layer, a second feature extraction layer, a fusion layer, a feature learning layer, and a prediction layer. The first feature extraction layer, fusion layer, feature learning layer, and prediction layer are connected sequentially, and the second feature extraction layer is connected to the fusion layer.
[0062] The first feature extraction layer is configured to extract image features from image frames; the second feature extraction layer is configured to extract text features from video violation categories; the fusion layer is configured to receive the image features output by the first feature extraction layer and the text features output by the second feature extraction layer, and fuse the image features and text features to obtain fused features; the feature learning layer is configured to perform vector encoding on the fused features to obtain feature vectors; and the prediction layer is configured to obtain image recognition results based on the feature vectors.
[0063] The first feature extraction layer can be a Convolutional Neural Network (CNN). The second feature extraction layer can be an embedding layer. The fusion layer can concatenate text features and image features to obtain fused features. The feature learning layer can be a Transformer neural network, which uses an attention mechanism to make the learned feature vectors more accurate. The prediction layer can be a Fully Connected Neural Network (FNN).
[0064] Therefore, by processing video violation categories and image frames through a video recognition model, the video violation category can be used as prior information, enabling the feature learning layer of the video recognition model to extract image features related to the video violation category from the image frame, and enabling the prediction layer to obtain more accurate image recognition results.
[0065] In some feasible implementations, the video recognition model may include:
[0066] The third feature extraction layer is configured to extract the image features from the image frame;
[0067] The fourth feature extraction layer is configured to extract the text features from the video violation categories;
[0068] The Transformer neural network is configured to obtain sequence features based on the image features and position encoding, process the sequence features through an encoder to obtain an encoding vector, obtain the fused features based on the text features and learnable position nesting, and process the encoding vector and the fused features through a decoder to obtain a feature vector.
[0069] The prediction layer is configured to obtain the image recognition result based on the feature vector.
[0070] Figure 4 This is a structural schematic diagram of a video recognition model according to some other embodiments. For example... Figure 4As shown, the video recognition model includes a third feature extraction layer 410, a fourth feature extraction layer 420, a Transformer neural network 430, and a prediction layer 440.
[0071] The third feature extraction layer 410 is configured to extract image features from image frames, and this third feature extraction layer 410 can be a convolutional neural network (CNN). The fourth feature extraction layer 420 is configured to extract text features from video violation categories, and this fourth feature extraction layer 420 can be an embedding layer.
[0072] The Transformer neural network 430 includes a first fusion module 431, an encoder 432, a second fusion module 433, and a decoder 434. The Transformer neural network 430 uses the first fusion module 431 to perform vector summation on image features and positional encoding to obtain sequence features. The encoder 432 then processes the sequence features to obtain an encoded vector. The second fusion module 433 fuses text features and learnable positional embeddings (or object queries) to obtain fused features. Finally, the decoder 434 processes the encoded vector output by the encoder 432 and the fused features output by the second fusion module 433 to obtain a feature vector.
[0073] It's worth noting that the decoder uses an attention mechanism to enable each element in the learnable positional nesting to capture object information with different positions and sizes in the original image. By fusing the learnable positional nesting with text features, the decoder can focus on features related to the video violation category in the encoded vector when extracting feature vectors.
[0074] The prediction layer 440 receives the feature vector output by the decoder 434 and obtains the image recognition result based on the feature vector. The prediction layer can be a fully connected neural network (FNN).
[0075] Therefore, by processing video violation categories and image frames through the video recognition model, the Transformer neural network of the video recognition model can extract image features related to video violation categories from image frames, and enable the prediction layer to obtain more accurate image recognition results.
[0076] In some feasible implementations, the video recognition model may include:
[0077] The fifth feature extraction layer is configured to extract the image features from the image frame;
[0078] The sixth feature extraction layer is configured to extract the text features from the video violation categories;
[0079] The Transformer neural network is configured to obtain a first fusion feature based on the image features, the text features, and position encoding; process the first fusion feature through an encoder to obtain an encoding vector; obtain a second fusion feature based on the text features and learnable position nesting; and process the encoding vector and the second fusion feature through a decoder to obtain a feature vector.
[0080] The prediction layer is configured to obtain the image recognition result based on the feature vector.
[0081] Figure 5 This is a structural schematic diagram of a video recognition model according to some other embodiments. For example... Figure 5 As shown, the video recognition model includes a fifth feature extraction layer 510, a sixth feature extraction layer 520, a Transformer neural network 530, and a prediction layer 540.
[0082] The fifth feature extraction layer 510 is configured to extract image features from image frames, and this fifth feature extraction layer 510 can be a convolutional neural network (CNN). The sixth feature extraction layer 520 is configured to extract text features from video violation categories, and this sixth feature extraction layer 520 can be an embedding layer.
[0083] The Transformer neural network 530 includes a first fusion module 531, an encoder 532, a second fusion module 533, and a decoder 534. The Transformer neural network 530 uses the first fusion module 531 to perform vector summation on image features, text features, and positional encoding to obtain sequence features. The encoder 532 processes the sequence features to obtain encoded vectors. The second fusion module 533 fuses text features and learnable positional embeddings (or object queries) to obtain fused features. The decoder 534 processes the encoded vector output by the encoder 532 and the fused features output by the second fusion module 533 to obtain a feature vector.
[0084] It's worth noting that the decoder uses an attention mechanism to enable each element in the learnable positional nesting to capture object information with different positions and sizes in the original image. By fusing the learnable positional nesting with text features, the decoder can focus on features related to the video violation category in the encoded vector when extracting feature vectors.
[0085] The prediction layer 540 receives the feature vector output by the decoder 534 and obtains the image recognition result based on the feature vector. The prediction layer can be a fully connected neural network (FNN).
[0086] Therefore, by processing video violation categories and image frames through the video recognition model, the Transformer neural network of the video recognition model can extract image features related to video violation categories from image frames, and enable the prediction layer to obtain more accurate image recognition results.
[0087] Figure 6 This is a flowchart illustrating a video recognition model training method according to some embodiments. For example... Figure 6 As shown, this disclosure provides a video recognition model training method. This method can be executed by an electronic device, specifically by a video recognition model training device. This device can be implemented in software and / or hardware and configured in the electronic device. Figure 6 As shown, the method may include the following steps.
[0088] In step 610, a training image set is obtained, wherein the training image set includes at least one image sample, the image sample has a first label and a second label, wherein the first label includes a bounding box for marking the location of the violation image in the image sample and the image violation category of the image to which the bounding box belongs, and the second label includes the video violation category corresponding to the image sample.
[0089] Here, the training image set includes at least one image sample, and each image sample includes a first label and a second label. The first label includes a bounding box used to mark the location of a violation in the image sample, and the image violation category to which the bounding box belongs. For example, the location of a violation in the image sample is marked using a bounding box. The image violation category refers to the violation category to which the violation image location belongs, such as categories like "prohibited goods sale," "pornography," "violence," and "vulgarity." The second label includes the video violation category corresponding to the image sample. This video violation category can refer to the video violation category of the video to which the image sample corresponds, meaning the image sample can be a video frame extracted from a video. Alternatively, the video violation category can refer to the overall violation category of the image sample itself.
[0090] In some embodiments, image samples can be used as input to a video detection model to obtain the violation category of the video.
[0091] It should be understood that the video detection model has been described in detail in the above embodiments, and will not be repeated here.
[0092] In step 620, the machine learning model is trained using the training image set to obtain a video recognition model.
[0093] Here, the process of training a machine learning model using a training image set can be as follows: input image samples into the machine learning model to obtain the predicted violation image location and the predicted violation category of the image sample predicted by the machine learning model; calculate the loss value between the predicted violation image location and the predicted violation category and the first label using a loss function; and adjust the parameters of the machine learning model according to the loss value until it converges to the preset conditions, thereby completing the training of the machine learning model and obtaining a video recognition model.
[0094] The machine learning model can be one of the above. Figure 3 The video recognition model shown Figure 4 The video recognition model shown and Figure 5 This is one of the video recognition models shown.
[0095] Therefore, by training the machine learning model using image samples carrying first and second labels, a video recognition model can be obtained, making the location and / or category of the violation images more accurate. Furthermore, it ensures that the location and / or category of the violation images in the obtained image frames remain consistent with the video violation category.
[0096] Figure 7 This is a schematic diagram of the module connections of a video recognition device according to some embodiments. Figure 7 As shown, the video recognition device 700 includes:
[0097] The determination module 701 is configured to determine the video violation category corresponding to the video to be identified.
[0098] Extraction module 702 is configured to extract image frames from the video to be identified;
[0099] The recognition module 703 is configured to take the video violation category and the image frame as input to the video recognition model to obtain the image recognition result of the image frame. The image recognition result includes the violation image location corresponding to the image frame and / or the image violation category to which the violation image location belongs. The video violation category is used to make the features extracted by the video recognition model from the image frame related to the video violation category.
[0100] Optionally, the video recognition model is configured as follows:
[0101] The text features of the video violation category and the image features of the image frame are fused to obtain fused features, and the image recognition result is obtained based on the fused features.
[0102] Optionally, the video recognition model includes:
[0103] The first feature extraction layer is configured to extract the image features from the image frame;
[0104] The second feature extraction layer is configured to extract the text features from the video violation category;
[0105] A fusion layer is configured to fuse the image features and the text features to obtain the fused features;
[0106] The feature learning layer is configured to perform vector encoding on the fused features to obtain feature vectors;
[0107] The prediction layer is configured to obtain the image recognition result based on the feature vector.
[0108] Optionally, the feature learning layer includes a Transformer neural network.
[0109] Optionally, the video recognition model includes:
[0110] The third feature extraction layer is configured to extract the image features from the image frame;
[0111] The fourth feature extraction layer is configured to extract the text features from the video violation categories;
[0112] The Transformer neural network is configured to obtain sequence features based on the image features and position encoding, process the sequence features through an encoder to obtain an encoding vector, obtain the fused features based on the text features and learnable position nesting, and process the encoding vector and the fused features through a decoder to obtain a feature vector.
[0113] The prediction layer is configured to obtain the image recognition result based on the feature vector.
[0114] Optionally, the video recognition model includes:
[0115] The fifth feature extraction layer is configured to extract the image features from the image frame;
[0116] The sixth feature extraction layer is configured to extract the text features from the video violation categories;
[0117] The Transformer neural network is configured to obtain a first fusion feature based on the image features, the text features, and position encoding; process the first fusion feature through an encoder to obtain an encoding vector; obtain a second fusion feature based on the text features and learnable position nesting; and process the encoding vector and the second fusion feature through a decoder to obtain a feature vector.
[0118] The prediction layer is configured to obtain the image recognition result based on the feature vector.
[0119] Optionally, the determining module 701 is specifically configured as follows:
[0120] The video to be identified is used as input to a video detection model to obtain the video violation category. The video detection model is obtained by training a machine learning model using video samples carrying labels of video violation categories.
[0121] Regarding the video recognition device 700 in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0122] Figure 8 This is a schematic diagram of the module connections of a video recognition model training device according to some embodiments. Figure 8 As shown, the video recognition model training device 800 includes:
[0123] The acquisition module 801 is configured to acquire a training image set, wherein the training image set includes at least one image sample, the image sample has a first label and a second label, wherein the first label includes a bounding box for marking the location of the violation image in the image sample and the image violation category of the image to which the bounding box belongs, and the second label includes the video violation category corresponding to the image sample;
[0124] The training module 802 is configured to train the machine learning model using the training image set to obtain a video recognition model.
[0125] Regarding the video recognition device 800 in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0126] The following is for reference. Figure 9 This diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0127] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0128] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0129] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0130] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0131] In some implementations, electronic devices can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0132] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0133] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: determine the video violation category corresponding to the video to be identified; extract image frames from the video to be identified; and use the video violation category and the image frames as input to a video recognition model to obtain an image recognition result for the image frames, wherein the image recognition result includes the violation image location corresponding to the image frame and / or the image violation category to which the violation image location belongs, and the video violation category is used to make the features extracted by the video recognition model from the image frames related to the video violation category.
[0134] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a training image set, wherein the training image set includes at least one image sample, the image sample having a first label and a second label, wherein the first label includes a bounding box for marking the location of a violation image in the image sample and the image violation category to which the bounding box belongs, and the second label includes the video violation category corresponding to the image sample; and train a machine learning model using the training image set to obtain a video recognition model.
[0135] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0137] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0138] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0139] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0140] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0141] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0142] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A video recognition method, characterized in that, include: Determine the video violation category corresponding to the video to be identified; Extract image frames from the video to be identified; The video violation category and the image frame are used as input to the video recognition model to obtain the image recognition result of the image frame. The image recognition result includes the violation image location corresponding to the image frame and / or the image violation category to which the violation image location belongs. The video violation category is used to make the features extracted by the video recognition model from the image frame related to the video violation category. The video recognition model is configured as follows: The text features of the video violation category and the image features of the image frame are fused to obtain fused features, and the image recognition result is obtained based on the fused features.
2. The method according to claim 1, characterized in that, The video recognition model includes: The first feature extraction layer is configured to extract the image features from the image frame; The second feature extraction layer is configured to extract the text features from the video violation category; A fusion layer is configured to fuse the image features and the text features to obtain the fused features; The feature learning layer is configured to perform vector encoding on the fused features to obtain feature vectors; The prediction layer is configured to obtain the image recognition result based on the feature vector.
3. The method according to claim 2, characterized in that, The feature learning layer includes a Transformer neural network.
4. The method according to claim 1, characterized in that, The video recognition model includes: The third feature extraction layer is configured to extract the image features from the image frame; The fourth feature extraction layer is configured to extract the text features from the video violation categories; The Transformer neural network is configured to obtain sequence features based on the image features and position encoding, process the sequence features through an encoder to obtain an encoding vector, obtain the fused features based on the text features and learnable position nesting, and process the encoding vector and the fused features through a decoder to obtain a feature vector. The prediction layer is configured to obtain the image recognition result based on the feature vector.
5. The method according to claim 1, characterized in that, The video recognition model includes: The fifth feature extraction layer is configured to extract the image features from the image frame; The sixth feature extraction layer is configured to extract the text features from the video violation categories; The Transformer neural network is configured to obtain a first fusion feature based on the image features, the text features, and position encoding; process the first fusion feature through an encoder to obtain an encoding vector; obtain a second fusion feature based on the text features and learnable position nesting; and process the encoding vector and the second fusion feature through a decoder to obtain a feature vector. The prediction layer is configured to obtain the image recognition result based on the feature vector.
6. The method according to any one of claims 1-5, characterized in that, Determining the video violation category corresponding to the video to be identified includes: The video to be identified is used as input to a video detection model to obtain the video violation category. The video detection model is obtained by training a machine learning model using video samples carrying labels of video violation categories.
7. A video recognition model training method, characterized in that, include: Obtain a training image set, wherein the training image set includes at least one image sample, the image sample having a first label and a second label, wherein the first label includes a bounding box for marking the location of the violation image in the image sample and the image violation category to which the bounding box belongs, and the second label includes the video violation category corresponding to the image sample; The machine learning model is trained using the training image set to obtain a video recognition model. The video violation category is used to make the features extracted by the video recognition model from the image frame related to the video violation category. The video recognition model is configured to fuse the text features of the video violation category and the image features of the image frame to obtain fused features, and obtain an image recognition result based on the fused features. The image recognition result includes the violation image location corresponding to the image frame and / or the image violation category to which the violation image location belongs.
8. A video recognition device, characterized in that, include: The determination module is configured to determine the video violation category corresponding to the video to be identified; The extraction module is configured to extract image frames from the video to be identified; The recognition module is configured to take the video violation category and the image frame as input to the video recognition model to obtain the image recognition result of the image frame. The image recognition result includes the violation image location corresponding to the image frame and / or the image violation category to which the violation image location belongs. The video violation category is used to make the features extracted by the video recognition model from the image frame related to the video violation category. The video recognition model is configured as follows: The text features of the video violation category and the image features of the image frame are fused to obtain fused features, and the image recognition result is obtained based on the fused features.
9. A video recognition model training device, characterized in that, include: The acquisition module is configured to acquire a training image set, wherein the training image set includes at least one image sample, the image sample has a first label and a second label, wherein the first label includes a bounding box for marking the location of the violation image in the image sample and the image violation category of the image to which the bounding box belongs, and the second label includes the video violation category corresponding to the image sample; The training module is configured to train a machine learning model using the training image set to obtain a video recognition model. The video violation category is used to ensure that the features extracted by the video recognition model from image frames are related to the video violation category. The video recognition model is configured to fuse the text features of the video violation category and the image features of the image frame to obtain fused features, and to obtain an image recognition result based on the fused features. The image recognition result includes the location of the violation image corresponding to the image frame and / or the image violation category to which the violation image location belongs.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the video recognition method according to any one of claims 1 to 6, or the steps of the video recognition model training method according to claim 7.
11. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device is configured to execute the computer program in the storage device to implement the steps of the video recognition method according to any one of claims 1 to 6, or to implement the steps of the video recognition model training method according to claim 7.
Citation Information
Patent Citations
Live broadcast content identification method and device, computer equipment and storage medium
CN114639056A
Video description generation method based on multi-concept knowledge mining and storage medium
CN114743143A