Machine learning based video recognition method, device, server and storage medium

By using machine learning methods to perform content and audio recognition on target and source videos, the problem of inefficient video recognition in existing technologies has been solved, and efficient and accurate recognition of funny dubbed videos has been achieved.

CN115705705BActive Publication Date: 2026-01-16TENCENT TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110925014.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-12
Publication Date
2026-01-16
Estimated Expiration
2041-08-12

AI Technical Summary

Technical Problem

In existing technologies, video platforms mainly rely on manual methods to identify funny dubbed videos, which is inefficient and difficult to guarantee accuracy, and cannot quickly and effectively identify a large number of new and existing videos.

Method used

A machine learning-based video recognition method is used to identify the content and audio type of the target video by comparing the content of the target video and the source video, and combining image, audio and subtitle features to determine whether it is a funny dubbed video.

Benefits of technology

It achieves efficient and accurate recognition of funny dubbed videos, improving video recognition efficiency and enabling accurate identification of derivative funny dubbed videos from a large number of videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115705705B_ABST
    Figure CN115705705B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a video recognition method and device based on machine learning, a server and a storage medium. Embodiments of the present application can obtain a target video; obtain a source video corresponding to the target video, the target video being created by processing the source video; compare the target video and the source video in content to obtain a content type of the target video; when the content type of the target video is a funny content type, perform audio recognition on the target video and the source video to determine an audio type of the target video; when the audio type of the target video is a funny dubbing type, determine the target video as a funny dubbing video, so as to push the funny dubbing video to a user. Embodiments of the present application compare the target video with the source video in dimensions such as content and audio to identify whether the target video is a funny dubbing video created by processing the source video. Thus, the present application can accurately identify a funny dubbing video from a plurality of videos, and improves the efficiency of video recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the computer field, and in particular to a video recognition method and device based on machine learning, a server and a storage medium. BACKGROUND

[0002] In order to achieve a humorous effect, a video author of a funny dubbing video will cut a video segment from an original video and replace the original background music (BGM) of the video segment with other funny audios, such as the video author's own dialect dubbing, funny sound effects, funny music, etc. This kind of funny dubbing video created by the video author has a large number of audiences.

[0003] Before pushing these funny dubbing videos to users, it is necessary to pre-identify this kind of funny dubbing video in a large number of videos. However, at present, the video platform mainly identifies funny dubbing videos through artificial identification and manual classification. This method is not only low in efficiency, high in identification cost, and difficult to guarantee accuracy, but also cannot quickly and effectively identify a large number of incremental and inventory videos. Therefore, the current video recognition method is low in efficiency. SUMMARY

[0004] The embodiments of the present application provide a video recognition method and device based on machine learning, a server and a storage medium, which can improve the efficiency of video recognition.

[0005] The embodiments of the present application provide a video recognition method based on machine learning, comprising:

[0006] obtaining a target video;

[0007] obtaining a source video corresponding to the target video, the target video being created by processing the source video;

[0008] comparing the content of the target video and the source video to obtain the content type of the target video;

[0009] when the content type of the target video is a funny content type, performing audio recognition on the target video and the source video to determine the audio type of the target video;

[0010] when the audio type of the target video is a funny dubbing type, determining the target video as a funny dubbing video, so as to push the funny dubbing video to users.

[0011] The embodiments of the present application also provide a video recognition device based on machine learning, comprising:

[0012] an obtaining unit configured to obtain a target video;

[0013] The source unit is configured to obtain a source video corresponding to a target video, the target video being created by processing the source video;

[0014] The content unit is configured to compare the target video and the source video in terms of content to obtain a content type of the target video.

[0015] The audio unit is configured to perform audio recognition on the target video and the source video to determine an audio type of the target video when the content type of the target video is a funny content type.

[0016] The determining unit is configured to determine the target video as a funny voice-over video when the audio type of the target video is a funny voice-over type, so as to push the funny voice-over video to a user.

[0017] In some embodiments, the content unit comprises:

[0018] The content recognition subunit is configured to perform content recognition on the target video and the source video to obtain a content funny probability of the target video and a content funny probability of the source video.

[0019] The content type subunit is configured to determine the content type of the target video based on the content funny probability of the target video and the content funny probability of the source video.

[0020] In some embodiments, the content type subunit is configured to:

[0021] perform difference processing on the content funny probability of the target video and the content funny probability of the source video to obtain a content funny probability difference.

[0022] determine the content type of the target video as a funny content type when the content funny probability difference is greater than a preset difference threshold and the content funny probability of the target video is greater than a preset content funny probability threshold.

[0023] In some embodiments, the content recognition subunit comprises:

[0024] The model submodule is configured to obtain a content recognition model.

[0025] The content recognition submodule is configured to perform content recognition on the target video by using the content recognition model to obtain the content funny probability of the target video.

[0026] The probability submodule is configured to perform content recognition on the source video by using the content recognition model to obtain the content funny probability of the source video.

[0027] In some embodiments, the model submodule is configured to:

[0028] obtain a preset content recognition model.

[0029] obtain training samples labeled with content types, the content types including but not limited to a humorous content type and a non-humorous content type, the training samples including but not limited to a video segment of a video, a video audio, and a video subtitle;

[0030] train the preset content recognition model by using the training samples labeled with the content types until the preset content recognition model converges, to obtain a content recognition model.

[0031] In some embodiments, the content recognition model includes a feature extraction layer, a feature fusion layer, and an output layer, the feature extraction layer including but not limited to an image feature extraction network, an audio feature extraction network, and a subtitle feature extraction network, and the content recognition sub-module is configured to:

[0032] obtain a video segment, a video audio, and a video subtitle of the target video;

[0033] extract image features of the video segment by using the image feature extraction network;

[0034] extract audio features of the video audio by using the audio feature extraction network;

[0035] extract subtitle features of the video subtitle by using the subtitle feature extraction network;

[0036] perform feature fusion processing on the image features, the audio features, and the subtitle features by using the feature fusion layer, to obtain fused features;

[0037] calculate a content humor probability of the target video based on the fused features by using the output layer.

[0038] In some embodiments, an audio unit includes:

[0039] a speech recognition sub-unit configured to perform speech recognition on the target video and the source video, determine an audio type of the target video, and determine an audio type of the source video, the audio type including but not limited to a non-dialect type and a dialect type;

[0040] an audio type sub-unit configured to determine the audio type of the target video as a humorous dubbing type when the audio type of the target video is the dialect type and the audio type of the source video is the non-dialect type.

[0041] In some embodiments, the speech recognition sub-unit is configured to:

[0042] obtain a speech recognition model;

[0043] The voice recognition model is used to perform funny voice recognition on the target video to obtain a funny voice probability of the target video, and an audio type of the target video is determined based on the funny voice probability of the target video. The audio type includes, but is not limited to, a non-dialect type and a dialect type.

[0044] The voice recognition model is used to perform funny voice recognition on the source video to obtain a funny voice probability of the source video, and an audio type of the source video is determined based on the funny voice probability of the source video.

[0045] In some embodiments, the audio type includes, but is not limited to, a funny background sound type and a non-funny background sound type, and the audio unit is configured to:

[0046] The target video and the source video are subjected to background sound recognition to determine a funny background sound probability of the target video and a funny background sound probability of the source video.

[0047] The funny background sound probability of the target video and the funny background sound probability of the source video are subjected to difference processing to obtain a funny background sound probability difference.

[0048] When the funny background sound probability difference is greater than a preset difference threshold value, and the funny background sound probability of the target video is greater than a preset funny background sound threshold value, the audio type of the target video is determined as a funny dubbing type.

[0049] In some embodiments, the source unit includes:

[0050] The retrieval feature subunit is configured to obtain a retrieval feature of the target video. The retrieval feature includes, but is not limited to, an image feature and a subtitle feature.

[0051] The search subunit is configured to search for a source video corresponding to the target video from a retrieval library based on the retrieval feature.

[0052] In some embodiments, the retrieval feature subunit is configured to:

[0053] The target video is subjected to frame extraction processing to obtain a plurality of image segments.

[0054] For each of the image segments, an image feature is extracted from the image segment.

[0055] For each of the image segments, subtitle recognition is performed on the image segment to obtain a subtitle text, and a subtitle feature is extracted from the subtitle text.

[0056] In some embodiments, the retrieval library includes retrieval features of candidate segments. The candidate segments are video segments of candidate videos. The retrieval feature subunit is configured to:

[0057] For each image segment of the image, based on the search feature corresponding to the image segment and the search feature of the candidate segment, determine the similarity between the image segment and the candidate segment;

[0058] For each image segment of the image, based on the similarity, determine the most similar candidate segment to the image segment in the search library;

[0059] Splice the most similar candidate segment to obtain the source video corresponding to the target video.

[0060] The embodiment of the application also provides a server, comprising a memory storing a plurality of instructions; the processor loads the instructions from the memory to execute the steps in any of the video recognition methods based on machine learning provided by the embodiment of the application.

[0061] The embodiment of the application also provides a computer readable storage medium, the computer readable storage medium stores a plurality of instructions, the instructions are suitable for the processor to load, to execute the steps in any of the video recognition methods based on machine learning provided by the embodiment of the application.

[0062] The embodiment of the application can obtain a target video; obtain a source video corresponding to the target video, the target video is obtained by processing and creating the source video; compare the content of the target video and the source video, to obtain the content type of the target video; when the content type of the target video is a humorous content type, perform audio recognition on the target video and the source video, to determine the audio type of the target video; when the audio type of the target video is a humorous dubbing type, determine the target video as a humorous dubbing video, so as to push the humorous dubbing video to the user.

[0063] The embodiment of the application compares the target video with the source video in the content (such as image content and subtitle content) and audio accompaniment dimensions, to determine whether the target video is more humorous than the source video in content and audio, so as to identify whether the target video is a humorous dubbing video obtained by processing and creating the source video. Therefore, the present application can accurately identify the humorous dubbing video from a plurality of videos, and improves the efficiency of video recognition. BRIEF DESCRIPTION OF DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0065] Figure 1ais a scene schematic diagram of a video recognition method based on machine learning provided by an embodiment of the present application;

[0066] Figure 1b is a flow schematic diagram of a video recognition method based on machine learning provided by an embodiment of the present application;

[0067] Figure 1c is a retrieval schematic diagram of a video recognition method based on machine learning provided by an embodiment of the present application;

[0068] Figure 1d is a model structure schematic diagram of a video recognition method based on machine learning provided by an embodiment of the present application;

[0069] Figure 1e is a model structure schematic diagram of a video recognition method based on machine learning provided by an embodiment of the present application;

[0070] Figure 1f is a model structure schematic diagram of a video recognition method based on machine learning provided by an embodiment of the present application;

[0071] Figure 2a is a flow schematic diagram of a video recognition method based on machine learning applied in a short video recommendation scene provided by an embodiment of the present application;

[0072] Figure 2b is a model structure schematic diagram of a video recognition method based on machine learning provided by an embodiment of the present application;

[0073] Figure 2c is a model structure schematic diagram of a video recognition method based on machine learning provided by an embodiment of the present application;

[0074] Figure 2d is a model structure schematic diagram of a video recognition method based on machine learning provided by an embodiment of the present application;

[0075] Figure 2e is a flow schematic diagram of a video recognition method based on machine learning applied in a short video recommendation scene provided by an embodiment of the present application;

[0076] Figure 3 is a structure schematic diagram of a video recognition device based on machine learning provided by an embodiment of the present application;

[0077] Figure 4 is a structure schematic diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION

[0078] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0079] The embodiments of the present application provide a video recognition method and device based on machine learning, a server and a storage medium.

[0080] The video recognition device based on machine learning can be integrated in an electronic device, which can be a terminal, a server, etc. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a notebook computer, or a personal computer (PC), etc. The server can be a single server or a server cluster composed of multiple servers.

[0081] In some embodiments, the video recognition device based on machine learning can also be integrated in multiple electronic devices, for example, the video recognition device based on machine learning can be integrated in multiple servers, and the multiple servers can be used to implement the video recognition method based on machine learning of the present application.

[0082] In some embodiments, the server can also be implemented in the form of a terminal.

[0083] For example, with reference to Figure 1a The electronic device can be a server, which can obtain a target video and a source video corresponding to the target video from a video database, the target video being created by processing the source video. Then, the server can compare the content of the target video and the source video to obtain the content type of the target video. When the content type of the target video is a funny content type, the server can perform audio recognition on the target video and the source video to determine the audio type of the target video. When the audio type of the target video is a funny dubbing type, the target video is determined as a funny dubbing video, so as to push the funny dubbing video to a user terminal.

[0084] The following will be described in detail. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.

[0085] Artificial Intelligence (AI) is a technology that uses digital computers to simulate human perception of the environment, acquire knowledge, and use knowledge, which can enable machines to have functions similar to human perception, reasoning, and decision-making. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, automatic driving, intelligent transportation, etc.

[0086] Among them, computer vision (CV) is a technology that uses computers to replace human eyes to identify, measure, and further process target images. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, virtual reality, augmented reality, simultaneous localization and mapping, automatic driving, intelligent transportation, etc. It also includes common face recognition, fingerprint recognition, and other biometric identification technologies. For example, image processing technologies such as image coloring and image edge extraction.

[0087] The key technologies of speech technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, and voice has become one of the most promising human-computer interaction methods in the future.

[0088] Natural language processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, i.e., the language used in daily life, so it has a close relationship with the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph, etc.

[0089] Machine Learning (ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It is a specialized study of how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0090] Autonomous driving technology usually includes high-precision map, environment perception, behavior decision, path planning, motion control, etc. Autonomous driving technology has a wide application prospect.

[0091] With the research and progress of artificial intelligence technology, artificial intelligence technology is being researched and applied in many fields, such as common smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned vehicles, autonomous vehicles, drones, robots, smart medical care, smart customer service, Internet of Vehicles, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0092] In this embodiment, a machine learning-based video recognition method involving artificial intelligence is provided, as shown in Figure 1b The specific process of the machine learning-based video recognition method can be as follows:

[0093] 101, obtaining a target video.

[0094] The target video refers to a video to be recognized. The method of obtaining the target video has many kinds, for example, it can be obtained from a video database, it can be obtained from a client, it can also be read in a local memory, etc.

[0095] The target video can be a funny dubbing video or other ordinary video. The funny dubbing video is generally a creative video segment cut from an original long video. The creator replaces the background sound of the original video segment with his own dubbing or funny sound effects, thereby forming a secondary creative video. This kind of video enhances the funny effect of the video through dubbing, and there are a large number of such video audience users on the video platform.

[0096] The target video can be a funny dubbing video or other ordinary video. The funny dubbing video is generally a creative video segment cut from an original long video. The creator replaces the background sound of the original video segment with his own dubbing or funny sound effects, thereby forming a secondary creative video. This kind of video enhances the funny effect of the video through dubbing, and there are a large number of such video audience users on the video platform.

[0097] 102, obtaining a source video corresponding to the target video, the target video being processed and created from the source video.

[0098] The target video corresponding source video can be automatically found by step 102, i.e., the original video segment in the original long video.

[0099] If the target video corresponding source video cannot be found, it can be directly determined that the target video is not a funny dubbing video; if the target video corresponding source video is found, step 103 can be executed to further determine whether the target video is a funny dubbing video.

[0100] In some embodiments, step 102 can include the following steps:

[0101] (1) Obtain the retrieval features of the target video, which include but are not limited to image features and subtitle features;

[0102] (2) Based on the retrieval features, find the target video corresponding source video from the retrieval library.

[0103] The retrieval library stores retrieval features of a plurality of candidate video segments, and the retrieval features of the candidate video segments in the retrieval library can be pre-constructed.

[0104] In some embodiments, the retrieval features include image features and subtitle features, and the retrieval library includes an image frame retrieval library and a subtitle text retrieval library.

[0105] For example, referring to Figure 1c , first, slice the candidate long video according to a preset time length, for example, slice the video according to a time length of 3 seconds to obtain a plurality of candidate video segments, then construct the image features and the subtitle features of each candidate video segment, the image features of the candidate video segments can be stored in the image frame retrieval library, and the subtitle features of the candidate video segments can be stored in the subtitle text retrieval library, when retrieval, the target video corresponding source video can be found from the image frame retrieval library and the subtitle text retrieval library based on the retrieval features of the target video.

[0106] Optionally, whether it is the target video or the source video, the method for constructing the retrieval features is similar, wherein the method for constructing the image features can be frame extraction processing on the target video / source video to extract the image features of each image segment; the method for constructing the subtitle features can be subtitle recognition on each image segment to obtain the subtitle text, and extract the subtitle features from the subtitle text.

[0107] For example, in some embodiments, step (1) of obtaining the retrieval features of the target video can include the following steps:

[0108] Frame extraction processing is performed on the target video to obtain a plurality of image segments;

[0109] For each image segment, extract the image features from the image segment;

[0110] For each image segment, the image segment is subjected to subtitle recognition to obtain a subtitle text, and a subtitle feature is extracted from the subtitle text.

[0111] For each image segment, the method of extracting the image feature from the image segment can be implemented by an image feature extraction network, which can be a convolutional neural network (CNN) such as EfficientNet, for example:

[0112] The image segment depth feature vector is extracted by EfficientNet, i.e., the pixel value of the image segment is input into EfficientNet, and the image feature of the image segment is output.

[0113] For each image segment, the method of performing subtitle recognition on the image segment can be implemented by OCR (Optical Character Recognition) technology, for example, by inputting the image segment into a character recognition network to output the subtitle text.

[0114] The subtitle text can be directly used as a subtitle feature, or a word-sense vector can be extracted from the subtitle text as a subtitle feature.

[0115] The image features and the subtitle features in the retrieval library can be associated with the video segments by indexing. For example, the subtitle features in the subtitle text retrieval library can be constructed by searching the engine to build a subtitle text inverted index, and the image features in the image frame retrieval library can be constructed by searching the engine to build a vector retrieval index.

[0116] In some embodiments, the retrieval library can include retrieval features of candidate segments, and the candidate segments can be video segments of candidate videos. The step (2) of searching for the source video corresponding to the target video based on the retrieval features can include the following steps:

[0117] For each image segment, based on the retrieval feature corresponding to the image segment and the retrieval feature of the candidate segment, the similarity between the image segment and the candidate segment is determined.

[0118] For each image segment, based on the similarity, the most similar candidate segment to the image segment is determined in the retrieval library.

[0119] The most similar candidate segment is spliced to obtain the source video corresponding to the target video.

[0120] For example, based on the image features and the text features of the target video, the image frame retrieval library and the subtitle text retrieval library are queried respectively to find the most similar and continuous candidate segments, and these continuous candidate segments are spliced into a source video as the origin of the current target video.

[0121] Wherein, when determining the most similar candidate segment to the image segment in the retrieval library, the similarity between the image segment and the candidate segment is required to meet a certain threshold. The similarity can be the Cosin distance.

[0122] Wherein, the present scheme can determine the most similar candidate segment to the image segment in the retrieval library based on the similarity through the Jaccard similarity coefficient.

[0123] Wherein, the Jaccard similarity coefficient refers to the proportion of the number of intersection elements of two sets A and B in the union set of A and B. The Jaccard similarity coefficient is an index for measuring the similarity of two sets, and the greater the value, the higher the similarity.

[0124] 103. Comparing the target video and the source video in content to obtain the content type of the target video.

[0125] Wherein, the video content can include the image, audio, subtitle, title and other content of the video, and the content type can include the funny content type and the non-funny content type. The funny content type refers to the humorous content of the target video, and the non-funny content type refers to the non-humorous content of the target video.

[0126] In some embodiments, step 103 can include the following steps:

[0127] (1) performing content recognition on the target video and the source video to obtain the content funny probability of the target video and the content funny probability of the source video;

[0128] (2) determining the content type of the target video based on the content funny probability of the target video and the content funny probability of the source video.

[0129] Since whether a video is humorous or not is a relatively subjective concept, it needs to be judged by people. Therefore, in the embodiments of the present application, a content recognition model can be trained by artificial training samples, so as to determine whether the target video is humorous or not by using the content recognition model.

[0130] Wherein, the training sample can be a network video, which can be manually watched and labeled after watching. The label can include but is not limited to "humorous content" and "non-humorous content".

[0131] Therefore, in some embodiments, step (1) of performing content recognition on the target video and the source video to obtain the content funny probability of the target video and the content funny probability of the source video can include the following steps:

[0132] (1.1) obtaining a content recognition model;

[0133] (1.2) using the content recognition model, content recognition of the target video is performed to obtain a content funny probability of the target video;

[0134] (1.3) using the content recognition model, content recognition of the source video is performed to obtain a content funny probability of the source video.

[0135] In some embodiments, step (1.1) of obtaining the content recognition model can include the following steps:

[0136] obtaining a preset content recognition model;

[0137] obtaining training samples labeled with content types, the content types including but not limited to funny content types and non-funny content types, and the training samples including but not limited to video clips of videos, video audios and video subtitles;

[0138] training the preset content recognition model using the training samples labeled with the content types until the preset content recognition model converges, to obtain the content recognition model.

[0139] The internal structure of the content recognition model will be introduced below:

[0140] In some embodiments, the content recognition model can include a feature extraction layer, a feature fusion layer and an output layer, the feature extraction layer can include but is not limited to an image feature extraction network, an audio feature extraction network and a subtitle feature extraction network, and step (1.2) of using the content recognition model to perform content recognition of the target video to obtain the content funny probability of the target video can include the following steps:

[0141] obtaining a video clip, a video audio and a video subtitle of the target video;

[0142] extracting image features of the video clip through the image feature extraction network;

[0143] extracting audio features of the video audio through the audio feature extraction network;

[0144] extracting subtitle features of the video subtitle through the subtitle feature extraction network;

[0145] performing feature fusion processing on the image features, the audio features and the subtitle features through the feature fusion layer to obtain fused features;

[0146] using the output layer to calculate the content funny probability of the target video based on the fused features.

[0147] For example, refer to Figure 1dThe content recognition model can include a feature extraction layer, a feature fusion layer, and an output layer. The feature extraction layer can include an image feature extraction network, an audio feature extraction network, and a subtitle feature extraction network. The image feature extraction network can include an image feature extraction module and an encoding module, the audio feature extraction network can include an audio feature extraction module and an encoding module, and the subtitle feature extraction network can include a subtitle feature extraction module and an encoding module. These modules can be artificial neural networks. For example, the encoding module can be a Transformer Encoder, the image feature extraction module can be an EfficientNet, the audio feature extraction module can be a VGGish, and the subtitle feature extraction module can be an Albert.

[0148] In some embodiments, the content type of the target video can be determined directly according to the humor probability of the content of the target video, for example:

[0149] When the humor probability of the content of the target video belongs to the preset threshold range, the content type of the target video is determined as a humor content type.

[0150] When the humor probability of the content of the target video does not belong to the preset threshold range, the content type of the target video is determined as a non-humor content type.

[0151] To further improve the judgment accuracy, the content humor probability of the target video and the content humor probability of the source video can be used for judgment. When the content of the target video is sufficiently humorous and the content of the target video is more humorous than the content of the source video, the content type of the target video is determined as a humor content type.

[0152] For example, in some embodiments, step (2) of determining the content type of the target video based on the content humor probability of the target video and the content humor probability of the source video can include the following steps:

[0153] The content humor probability of the target video and the content humor probability of the source video are processed by difference to obtain a content humor probability difference.

[0154] When the content humor probability difference is greater than a preset difference threshold, and the content humor probability of the target video is greater than a preset content humor probability threshold, the content type of the target video is determined as a humor content type.

[0155] For example, assuming that the content humor probability of the target video is P_L1 and the content humor probability of the source video is P_L2, when P_L1 is greater than the content humor probability threshold K_L and P_L1-P_L2 is greater than the content humor probability difference K_Ld, the content type of the target video is determined as the humor content type; otherwise, the content type of the target video is determined as the non-humor content type.

[0156] If the content type of the target video is the non-humor content type, it can be directly determined that the target video is not the humor dubbing video; if the content type of the target video is the humor content type, it can be determined that the content of the target video is humorous enough, but it cannot be determined whether the audio of the target video is the dubbing of secondary creation, and therefore, in step 104, whether the target video is the dubbing video can be identified from the perspective of the audio:

[0157] 104. When the content type of the target video is the humor content type, performing audio identification on the target video and the source video to determine the audio type of the target video.

[0158] Among them, since the dubbing of the humor dubbing video can be dialect dubbing or humor background sound, in some embodiments, whether the target video is the humor dubbing video can be determined by identifying whether there is a dialect in the video audio of the target video, and in other embodiments, whether the target video is the humor dubbing video can also be determined by judging whether the background sound of the target video is more humorous than that of the source video.

[0159] Therefore, the dialect and background sound scenarios will be introduced below respectively:

[0160] (1) Dialect scenario.

[0161] In the dialect identification scenario, the audio type of the target video includes but is not limited to non-dialect type and dialect type. Dialect refers to the local variant of language, which is different from the standard language (such as Mandarin) and is only used in a region.

[0162] For example, dialects can include Sichuanese, Cantonese, Northeastern Mandarin, etc. Non-dialects can include Mandarin, non-human audio, etc.

[0163] In some embodiments, step 104 can include the following steps:

[0164] (1) performing speech recognition on the target video and the source video to determine the audio type of the target video and the audio type of the source video, the audio type including but not limited to non-dialect type and dialect type;

[0165] (2) when the audio type of the target video is the dialect type and the audio type of the source video is the non-dialect type, the audio type of the target video is determined as the humor dubbing type.

[0166] In some embodiments, the step (1) can include the following steps:

[0167] Therefore, in some embodiments, the step (1) can include the following steps:

[0168] obtaining a voice recognition model;

[0169] applying the voice recognition model to the target video to obtain a voice funny probability of the target video, and determining the audio type of the target video based on the voice funny probability of the target video, the audio type including but not limited to non-dialect type and dialect type;

[0170] applying the voice recognition model to the source video to obtain a voice funny probability of the source video, and determining the audio type of the source video based on the voice funny probability of the source video.

[0171] In some embodiments, the internal structure of the voice recognition model can refer to Figure 1e The voice recognition model can include an audio feature extraction network, an encoding network and an output layer. The output layer can output the probability of containing human voice and the probability of containing dialect voice in the audio of the video. According to the numerical value of the probability, it can be determined whether the audio of the video contains human voice and whether it contains dialect voice. For example, as shown in Table 1:

[0172] Table 1

[0173] Video audio name Contains human voice Contains dialect voice Audio 1 Yes Yes Audio 2 No No … … … Audio Y Yes No

[0174] If the target video does not contain dialect voice or human voice, the target video can be directly determined as a non-funny dubbing video. If the target video contains dialect voice, the audio type of the target video can be determined as a funny dubbing type, and the target video is determined as a funny dubbing video.

[0175] (II) Background sound scene

[0176] In some embodiments, the audio type includes but is not limited to funny background sound type and non-funny background sound type. The step 104 can include the following steps:

[0177] applying the voice recognition model to the target video and the source video to obtain a background sound funny probability of the target video and a background sound funny probability of the source video;

[0178] The background sound funny probability difference between the target video and the source video is obtained by subtracting the background sound funny probability of the source video from the background sound funny probability of the target video.

[0179] When the background sound funny probability difference is greater than a preset difference threshold, and the background sound funny probability of the target video is greater than a preset background sound funny threshold, the audio type of the target video is determined as a funny dubbing type.

[0180] The background sound recognition model can be used to recognize the background sound of the target video and the source video. In the training phase of the background sound recognition model, the training samples can be manually labeled as "background sound funny" and "background sound not funny".

[0181] The internal structure of the speech recognition model can refer to Figure 1f The speech recognition model can include an audio feature extraction network, an encoding network, and an output layer. The output layer can output the background sound funny probability of the video audio.

[0182] By comparing the background sound funny probability of the target video and the source video, the audio type of the target video can be determined.

[0183] For example, assuming that the background sound funny probability of the target video is P_B1, and the background sound funny probability of the source video is P_B2, when P_B1 is greater than the background sound funny threshold K_B, and P_B1-P_B2 is greater than the difference threshold K_Bd, the audio type of the target video is determined as a funny dubbing type; otherwise, the content type of the target video is determined as a non-funny dubbing type.

[0184] In some embodiments, in order to make the determination more accurate, in addition to the background sound funny probability difference being greater than a preset difference threshold, and the background sound funny probability of the target video being greater than a preset background sound funny threshold, the audio type of the target video is determined as a funny dubbing type only when the background sound funny probability of the source video is less than another threshold.

[0185] 105、When the audio type of the target video is a funny dubbing type, the target video is determined as a funny dubbing video, so as to push the funny dubbing video to the user.

[0186] In summary, only when all the following conditions are met, the target video can be determined as a funny dubbing video:

[0187] (1) The target video has a corresponding source video;

[0188] (2) The content type of the target video is a funny content type;

[0189] (3) The audio type of the target video is a funny dubbing type;

[0190] Otherwise, as long as any condition is not met, the target video is determined as a non-humor dubbing video.

[0191] As can be seen from the above, the embodiment of the present application can obtain a target video; obtain a source video corresponding to the target video, the target video being processed and created from the source video; perform content comparison on the target video and the source video to obtain a content type of the target video; when the content type of the target video is a humor content type, perform audio recognition on the target video and the source video to determine an audio type of the target video; and when the audio type of the target video is a humor dubbing type, determine the target video as a humor dubbing video, so as to push the humor dubbing video to a user.

[0192] Therefore, the embodiment of the present application can accurately identify a humor dubbing video from numerous videos, and improves the efficiency of video identification.

[0193] The method described in the above embodiment will be further described in detail below.

[0194] The scheme provided by the embodiment of the present application can be applied in various video pushing scenarios. For example, the method of the embodiment of the present application is described in detail below by taking a humor dubbing short video as an example.

[0195] As shown in Figure 2a , a specific process of a video identification method based on machine learning is as follows:

[0196] 201, obtain a target video.

[0197] After obtaining the target video, the source video corresponding to the target video can be searched in a retrieval library.

[0198] A humor dubbing short video is usually a long video with a certain segment being dubbed to enhance the humor effect. Therefore, it is necessary to find a source video (i.e., a certain segment of a long video) for the target video. The embodiment of the present application finds the original source video through image and text features.

[0199] First, a retrieval feature of a long video of a platform needs to be constructed. Each long video is time-slice sliced, for example, the video is sliced according to a time length of 3 seconds, and a retrieval feature is constructed for each slice:

[0200] Image feature: a plurality of image frames are extracted from a video segment, and a deep feature is extracted from the image frames by EfficientNet (a kind of artificial neural network for image feature extraction). The deep feature is constructed into a vector retrieval index by ElasticFaiss (a kind of search engine).

[0201] Text feature: The long video clip is identified by OCR to recognize the subtitles, and the subtitle text is segmented. An inverted index is built for text segmentation, such as building a subtitle text inverted index through ElasticSearch (a search engine).

[0202] First, multiple image frame features are constructed by frame extraction and EfficientNet, and the subtitle text is recognized based on OCR to construct text features. Then, based on image features and text features, the long video image frame feature and subtitle text feature retrieval library is queried to find the most similar continuous long video clips, and multiple continuous clips are spliced into a short video as the source video of the current target video. The similarity between the current short video and the source video meets a certain threshold. In the similarity calculation, the image frame feature can be calculated by Cosin distance, and the text feature can be calculated by Jaccard similarity coefficient.

[0203] 202, if the target video corresponding to the source video is found, go to step 203, otherwise go to step 207.

[0204] Funny dubbing short video requires that the short video has humor, and the short video after being dubbed by the creator has a higher humor effect than the source video found in step 201.

[0205] Therefore, in the embodiments of the present application, the humor of the video can be judged by the model as shown in Figure 2b The model is trained on a pre-constructed humor and non-humor video dataset. The text, video, and audio features of the input video are input into the model, and the difference between the humor probability output by the model and the annotation of the dataset is calculated, and the model parameters are updated to make the model have the ability to input the features of the video and output the humor probability of the video content.

[0206] As shown in Figure 2b The encoding module of the model can be a Transformer-Encoder (a type of encoding network), the image feature extraction module can be an EfficientNet (a type of feature extraction model), the audio feature extraction module can be a VGGish (a type of feature extraction model), and the subtitle feature extraction module can be an Albert (a type of feature extraction model).

[0207] The humor probability P_L1 of the target video content is calculated, and the humor probability P_L2 of the source video content is calculated. The humor probability P_L1 of the target video content meets a certain threshold, and P_L1-P_L2 is greater than a certain threshold. If the above conditions are met, the content type of the target video is determined as a humor content type, and subsequent calculations are continued, otherwise the short video is not a funny dubbing short video.

[0208] 203、If the content type of the target video is a funny content type, go to step 204, otherwise go to step 207.

[0209] There are generally two types of funny dubbing, one is funny dubbing of video voice in other dialects, and the other is background music replaced by other funny background music.

[0210] Among them, the funny dubbing of video voice in other dialects is identified:

[0211] Whether there is voice in the short video is identified, if the original source video is non-dialect voice, and the voice of the target video is dialect voice, it is determined that the current short video is a funny dubbing short video.

[0212] As shown in FIG. 1, a model shown in FIG. 1 can be used to determine whether the video audio contains voice and whether it is a dialect voice type. Figure 2c

[0213] Among them, as shown in FIG. 1, the encoding module of the model can be a Transformer-Encoder, and the audio feature extraction module can be a VGGish. Figure 2c The model is trained on a pre-constructed video audio data set, and the model has the probability of inputting video audio, outputting whether it contains voice, and whether the voice is a dialect.

[0214] Optionally, some short video titles will appear "dialect dubbing", "Sichuan version" and other features, which can be mined by rules on the short video titles to be identified, and combined with the above audio feature model to determine whether the target video is a dialect dubbing.

[0215] Among them, the background music is replaced by other funny background music:

[0216] As shown in FIG. 2, a model shown in FIG. 2 can be used to extract the background sound of the short video and identify the funny background sound. The model is trained on a pre-constructed audio and funny condition data set, and the model has the probability of inputting video background sound and outputting the funny condition of the background sound.

[0217] Figure 2d Among them, as shown in FIG. 2, the encoding module of the model can be a Transformer-Encoder, and the audio feature extraction module can be a VGGish.

[0218] Figure 2d

[0219] ​​​​If the probability of the background sound of the target video being funny meets a certain threshold, the probability of the background sound of the source video being funny is lower than a certain threshold, and the funny conditions of the background sounds of the two videos are quite different, the content type of the target video is determined as a funny content type, and subsequent calculation is continued, otherwise, it is determined that the short video is not a funny dubbing short video.

[0220] 204, if the audio type of the target video is a funny dubbing type, go to step 205, otherwise go to step 207.

[0221] Through the above steps, it is realized to identify whether the short video is a funny dubbing, which can enrich the resource pool of funny dubbing, and provide a data basis for funny dubbing video recommendation and distribution.

[0222] 205, the target video is determined as a funny dubbing video.

[0223] 206, push the funny dubbing video to the user.

[0224] 207, the target video is determined as a non-funny dubbing video.

[0225] Reference Figure 2e The funny dubbing video intelligent identification embodiment proposed in the embodiment can be applied to the field of short videos. For example, by combining the original dubbing of the current short video, the dubbing of the current short video is compared and identified, and at the same time, the funny dubbing recognition is automatically recognized by combining the funny recognition of the video, the recognition efficiency of the funny dubbing video is improved, the intelligent automatic recognition of the large increment and stock video of the video platform is realized, the artificial recognition cost is reduced, the funny dubbing video resource pool is enriched, the recommendation and distribution effect of such videos is enhanced, and the personalized playback demand of platform audience users is met.

[0226] As can be seen from the above, the funny dubbing video can be accurately identified from a large number of videos, and the efficiency of video recognition is improved.

[0227] In order to better implement the above method, the embodiment of the application also provides a video recognition device based on machine learning. The video recognition device based on machine learning can be integrated in an electronic device, which can be a terminal, a server, etc. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a notebook computer, a personal computer, etc. The server can be a single server or a server cluster composed of multiple servers.

[0228] For example, in the embodiment, the video recognition device based on machine learning is integrated in the server, and the method of the embodiment of the application is described in detail.

[0229] For example, as shown in FIG. 1, the video recognition device based on machine learning is integrated in the server.Figure 3 As shown, the machine learning-based video recognition apparatus can include an acquisition unit 301, a source unit 302, a content unit 303, and an audio unit 304, and a determination unit 305, as follows:

[0230] (I) The acquisition unit 301 is configured to acquire a target video.

[0231] (II) The source unit 302 is configured to acquire a source video corresponding to the target video, the target video being created by processing the source video.

[0232] In some embodiments, the source unit 302 can include a retrieval feature subunit and a search subunit, as follows:

[0233] The retrieval feature subunit can be configured to acquire retrieval features of the target video, the retrieval features including but not limited to image features and subtitle features.

[0234] The search subunit can be configured to search for the source video corresponding to the target video from a retrieval library based on the retrieval features.

[0235] In some embodiments, the retrieval feature subunit can be configured to:

[0236] frame the target video to obtain a plurality of image segments;

[0237] extract image features from each image segment;

[0238] perform subtitle recognition on each image segment to obtain subtitle text and extract subtitle features from the subtitle text.

[0239] In some embodiments, the retrieval library can include retrieval features of candidate segments, the candidate segments being video segments of candidate videos, and the retrieval feature subunit can be configured to:

[0240] for each image segment, determine a similarity between the image segment and the candidate segments based on the retrieval features corresponding to the image segment and the retrieval features of the candidate segments;

[0241] for each image segment, determine the most similar candidate segment to the image segment in the retrieval library based on the similarity;

[0242] splice the most similar candidate segment to obtain the source video corresponding to the target video.

[0243] (III) The content unit 303 is configured to compare the content of the target video and the source video to obtain a content type of the target video.

[0244] In some embodiments, the content unit 303 can include a content recognition subunit and a content recognition subunit, as follows:

[0245] The content recognition subunit can be configured to perform content recognition on the target video and the source video to obtain a content humor probability of the target video and a content humor probability of the source video.

[0246] The content recognition subunit can be configured to determine a content type of the target video based on the content humor probability of the target video and the content humor probability of the source video.

[0247] In some embodiments, the content type subunit can be configured to:

[0248] perform difference processing on the content humor probability of the target video and the content humor probability of the source video to obtain a content humor probability difference;

[0249] when the content humor probability difference is greater than a preset difference threshold and the content humor probability of the target video is greater than a preset content humor probability threshold, determine the content type of the target video as a humor content type.

[0250] In some embodiments, the content recognition subunit can include a model submodule, a content recognition submodule, and a probability submodule, as follows:

[0251] The model submodule can be configured to obtain a content recognition model.

[0252] The content recognition submodule can be configured to perform content recognition on the target video using the content recognition model to obtain the content humor probability of the target video.

[0253] The probability submodule can be configured to perform content recognition on the source video using the content recognition model to obtain the content humor probability of the source video.

[0254] In some embodiments, the model submodule can be configured to:

[0255] obtain a preset content recognition model;

[0256] obtain training samples labeled with content types, the content types can include but are not limited to a humor content type and a non-humor content type, and the training samples can include but are not limited to video segments of videos, video audios, and video subtitles;

[0257] train the preset content recognition model using the training samples labeled with the content types until the preset content recognition model converges to obtain the content recognition model.

[0258] In some embodiments, the content recognition model can include a feature extraction layer, a feature fusion layer, and an output layer, the feature extraction layer can include but is not limited to an image feature extraction network, an audio feature extraction network, and a subtitle feature extraction network, and the content recognition submodule can be configured to:

[0259] obtaining a video segment, a video audio and a video subtitle of the target video;

[0260] extracting image features of the video segment through an image feature extraction network;

[0261] extracting audio features of the video audio through an audio feature extraction network;

[0262] extracting subtitle features of the video subtitle through a subtitle feature extraction network;

[0263] performing feature fusion processing on the image features, the audio features and the subtitle features through a feature fusion layer to obtain fusion features;

[0264] adopting an output layer to calculate a funny content probability of the target video based on the fusion features.

[0265] The audio unit 304 is configured to perform audio recognition on the target video and the source video when the content type of the target video is the funny content type, and determine an audio type of the target video.

[0266] In some embodiments, the audio unit 304 can include a speech recognition subunit and an audio type subunit, as follows:

[0267] The speech recognition subunit can be configured to perform speech recognition on the target video and the source video, determine an audio type of the target video, and an audio type of the source video. The audio type can include, but is not limited to, a non-dialect type and a dialect type.

[0268] The audio type subunit can be configured to determine the audio type of the target video as a funny dubbing type when the audio type of the target video is the dialect type and the audio type of the source video is the non-dialect type.

[0269] In some embodiments, the speech recognition subunit can be configured to:

[0270] obtain a speech recognition model;

[0271] adopt the speech recognition model to perform funny speech recognition on the target video, obtain a funny speech probability of the target video, and determine the audio type of the target video based on the funny speech probability of the target video. The audio type can include, but is not limited to, a non-dialect type and a dialect type.

[0272] adopt the speech recognition model to perform funny speech recognition on the source video, obtain a funny speech probability of the source video, and determine the audio type of the source video based on the funny speech probability of the source video.

[0273] In some embodiments, the audio type can include, but is not limited to, a funny background sound type and a non-funny background sound type. The audio unit 304 can be configured to:

[0274] The background sound of the target video and the source video is recognized, the background sound funny probability of the target video is determined, and the background sound funny probability of the source video is determined.

[0275] The background sound funny probability of the target video and the background sound funny probability of the source video are subtracted to obtain a background sound funny probability difference.

[0276] When the background sound funny probability difference is greater than a preset difference threshold, and the background sound funny probability of the target video is greater than a preset background sound funny threshold, the audio type of the target video is determined as a funny dubbing type.

[0277] The determining unit 305 is configured to determine the target video as a funny dubbing video when the audio type of the target video is the funny dubbing type, so as to push the funny dubbing video to the user.

[0278] In specific implementation, the above units can be implemented as independent entities, or can be combined as the same or several entities, and the specific implementation of the above units can be referred to the method embodiments above, which will not be described herein.

[0279] As can be seen from the above, the video recognition apparatus based on machine learning in the embodiment comprises an obtaining unit configured to obtain a target video; a source unit configured to obtain a source video corresponding to the target video, the target video being created by processing the source video; a content unit configured to compare the target video and the source video in content to obtain a content type of the target video; an audio unit configured to perform audio recognition on the target video and the source video to determine an audio type of the target video when the content type of the target video is a funny content type; and a determining unit configured to determine the target video as a funny dubbing video when the audio type of the target video is a funny dubbing type, so as to push the funny dubbing video to the user. Thus, the funny dubbing video can be accurately identified from a plurality of videos, and the efficiency of video recognition is improved.

[0280] The embodiment of the application further provides an electronic device, which can be a terminal, a server, or the like. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a notebook computer, a personal computer, or the like. The server can be a single server or a server cluster composed of multiple servers.

[0281] In some embodiments, the video recognition apparatus based on machine learning can be integrated in multiple electronic devices, for example, the video recognition apparatus based on machine learning can be integrated in multiple servers, and the multiple servers can implement the video recognition method based on machine learning.

[0282] In the embodiment, the electronic device is taken as a server for detailed description, for example, as shown in Figure 4As shown, it shows a structural schematic diagram of a server related to embodiments of the present application, in particular:

[0283] The server can include a processor 401 with one or more processing cores, a memory 402 with one or more computer readable storage media, a power supply 403, an input module 404, a communication module 405, and the like. Those skilled in the art can understand that the server structure shown in the figure does not constitute a limitation on the server, and can include more or fewer components than shown, or combine certain components, or different component arrangements. Among them: Figure 4

[0284] The processor 401 is the control center of the server, which connects various parts of the server through various interfaces and lines, executes various functions of the server and processes data by running or executing software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, thereby overall monitoring the server. In some embodiments, the processor 401 can include one or more processing cores; in some embodiments, the processor 401 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 401.

[0285] The memory 402 can be used to store software programs and modules, and the processor 401 executes various functions and data processing by running the software programs and modules stored in the memory 402. The memory 402 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), and the like; the data storage area can store data created according to the use of the server, etc. In addition, the memory 402 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory 402 can also include a memory controller to provide access for the processor 401 to the memory 402.

[0286] The server also includes a power supply 403 for powering various components, and in some embodiments, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include one or more direct current or alternating current power supplies, a recharging system, a power supply failure detection circuit, a power supply converter or inverter, a power supply state indicator, and the like. Any component.

[0287] ​The server can also include an input module 404, which can be used to receive inputted digital or character information, and to generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0288] The server can also include a communication module 405, which in some embodiments can include a wireless module, through which the server can perform short-range wireless transmission, thereby providing the user with wireless broadband Internet access. For example, the communication module 405 can be used to help the user send and receive emails, browse web pages, and access streaming media, etc.

[0289] Although not shown, the server can also include a display unit, etc., which will not be described here. In particular in the present embodiment, the processor 401 in the server will load one or more executable files corresponding to the processes of the application programs into the memory 402 according to the following instructions, and run the application programs stored in the memory 402 by the processor 401, thereby realizing various functions, such as the following:

[0290] Obtaining a target video;

[0291] Obtaining a source video corresponding to the target video, the target video being created by processing the source video;

[0292] Performing content comparison on the target video and the source video to obtain a content type of the target video;

[0293] When the content type of the target video is a funny content type, performing audio recognition on the target video and the source video to determine an audio type of the target video;

[0294] When the audio type of the target video is a funny dubbing type, determining the target video as a funny dubbing video, so as to push the funny dubbing video to the user.

[0295] The specific implementation of each operation can refer to the previous embodiments, which will not be described here.

[0296] As can be seen from the above, the present scheme can accurately identify a funny dubbing video from numerous videos, thereby improving the efficiency of video recognition.

[0297] Those skilled in the art can understand that all or part of the steps of the various methods of the above embodiments can be completed by instructions, or by instructions controlling related hardware, which can be stored in a computer readable storage medium and loaded and executed by a processor.

[0298] To this end, an embodiment of the present application provides a computer readable storage medium, wherein a plurality of instructions are stored, the instructions being loadable by a processor to execute steps of any of the video recognition methods based on machine learning provided by the embodiments of the present application. For example, the instructions can execute the following steps:

[0299] obtaining a target video;

[0300] obtaining a source video corresponding to the target video, the target video being created by processing the source video;

[0301] performing content comparison on the target video and the source video to obtain a content type of the target video;

[0302] when the content type of the target video is a funny content type, performing audio recognition on the target video and the source video to determine an audio type of the target video;

[0303] when the audio type of the target video is a funny dubbing type, determining the target video as a funny dubbing video, so as to push the funny dubbing video to a user.

[0304] The storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0305] According to an aspect of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the methods provided in the various optional implementations of the video recognition aspect or the video pushing aspect provided in the embodiments described above.

[0306] Since the instructions stored in the storage medium can execute the steps of any of the video recognition methods based on machine learning provided by the embodiments of the present application, the beneficial effects of any of the video recognition methods based on machine learning provided by the embodiments of the present application can be achieved, which are described in detail in the foregoing embodiments and will not be described here again.

[0307] The above describes in detail the video recognition method, device, server and computer readable storage medium provided by the embodiment of the application based on machine learning. The principle and implementation mode of the application are described by applying specific examples. The above embodiment is only used to help understand the method of the application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the application, the specific implementation mode and application range will be changed. In conclusion, the content of the specification should not be understood as a limitation of the application.

Claims

1. A method for video recognition based on machine learning, characterized in that, The method comprises the following steps: acquiring a target video; acquiring a source video corresponding to the target video, the target video being created by processing the source video; comparing the content of the target video and the source video to obtain the content type of the target video; when the content type of the target video is a funny content type, performing audio recognition on the target video and the source video to determine the audio type of the target video; when the audio type of the target video is a funny voice type, determining the target video as a funny voice video, so as to push the funny voice video to a user. 2.The machine learning based video recognition method of claim 1, wherein, The step of comparing the content of the target video and the source video to obtain the content type of the target video comprises the following steps: performing content recognition on the target video and the source video to obtain the content funny probability of the target video and the content funny probability of the source video; determining the content type of the target video based on the content funny probability of the target video and the content funny probability of the source video. 3.The machine learning based video recognition method of claim 2, wherein, The step of determining the content type of the target video based on the content funny probability of the target video and the content funny probability of the source video comprises the following steps: performing difference processing on the content funny probability of the target video and the content funny probability of the source video to obtain a content funny probability difference; when the content funny probability difference is greater than a preset difference threshold value and the content funny probability of the target video is greater than a preset content funny probability threshold value, determining the content type of the target video as a funny content type. 4.The method of claim 2, wherein, The step of performing content recognition on the target video and the source video to obtain the content funny probability of the target video and the content funny probability of the source video comprises the following steps: acquiring a content recognition model; performing content recognition on the target video by using the content recognition model to obtain the content funny probability of the target video; performing content recognition on the source video by using the content recognition model to obtain the content funny probability of the source video.

5. The machine learning based video recognition method of claim 4, wherein, The step of acquiring a content recognition model comprises the following steps: acquiring a preset content recognition model; acquiring training samples labeled with a content type, the content type comprising but not limited to a funny content type and a non-funny content type, the training samples comprising but not limited to a video segment of a video, a video audio and a video subtitle; training the preset content recognition model by using the training samples labeled with the content type until the preset content recognition model converges to obtain a content recognition model. 6.The method of claim 4, wherein, The content recognition model comprises a feature extraction layer, a feature fusion layer and an output layer, the feature extraction layer comprising but not limited to an image feature extraction network, an audio feature extraction network and a subtitle feature extraction network, and the step of performing content recognition on the target video by using the content recognition model to obtain the content funny probability of the target video comprises the following steps: acquiring a video segment, a video audio and a video subtitle of the target video; extracting image features of the video segment by using the image feature extraction network; extracting audio features of the video audio by using the audio feature extraction network; extracting subtitle features of the video subtitle by using the subtitle feature extraction network; The image features, audio features and subtitle features are subjected to feature fusion processing by the feature fusion layer to obtain fused features; The output layer is used to calculate the funny content probability of the target video based on the fused features.

7. The machine learning based video recognition method of claim 1, wherein, The audio recognition of the target video and the source video is performed to determine the audio type of the target video, which includes: The audio recognition of the target video and the source video is performed to determine the audio type of the target video and the audio type of the source video, which includes but is not limited to non-dialect type and dialect type; When the audio type of the target video is dialect type and the audio type of the source video is non-dialect type, the audio type of the target video is determined as funny dubbing type.

8. The machine learning based video recognition method of claim 7, wherein, The audio recognition of the target video and the source video is performed to determine the audio type of the target video and the audio type of the source video, which includes: A speech recognition model is obtained; The speech recognition model is used to perform funny speech recognition on the target video to obtain the speech funny probability of the target video, and the audio type of the target video is determined based on the speech funny probability of the target video, which includes but is not limited to non-dialect type and dialect type; The speech recognition model is used to perform funny speech recognition on the source video to obtain the speech funny probability of the source video, and the audio type of the source video is determined based on the speech funny probability of the source video. 9.The machine learning based video recognition method of claim 1, wherein, The audio type includes but is not limited to funny background sound type and non-funny background sound type, and the audio recognition of the target video and the source video to determine the audio type of the target video includes: The background sound recognition of the target video and the source video is performed to determine the background sound funny probability of the target video and the background sound funny probability of the source video; The background sound funny probability difference is obtained by performing difference processing on the background sound funny probability of the target video and the background sound funny probability of the source video; When the background sound funny probability difference is greater than a preset difference threshold and the background sound funny probability of the target video is greater than a preset background sound funny threshold, the audio type of the target video is determined as funny dubbing type. 10.The machine learning based video recognition method of claim 1, wherein, The source video corresponding to the target video is obtained, which includes: The retrieval features of the target video are obtained, which include but are not limited to image features and subtitle features; The source video corresponding to the target video is searched from a retrieval library based on the retrieval features.

11. The machine learning based video recognition method of claim 10, wherein, The retrieval features of the target video are obtained, which include: The target video is subjected to frame extraction processing to obtain a plurality of image segments; For each image segment, image features are extracted from the image segment; For each image segment, subtitle recognition is performed on the image segment to obtain subtitle text, and subtitle features are extracted from the subtitle text.

12. The machine learning based video recognition method of claim 11, wherein, The retrieval library includes retrieval features of candidate segments, the candidate segments are video segments of candidate videos, and the source video corresponding to the target video is searched from a retrieval library based on the retrieval features, which includes: For each frame of the image segment, based on the search feature corresponding to the image segment and the search feature of the candidate segment, determine the similarity between the image segment and the candidate segment; For each frame of the image segment, based on the similarity, determine the most similar candidate segment to the image segment in the search library; Splicing the most similar candidate segment to obtain the source video corresponding to the target video. 13.A machine learning based video recognition apparatus, characterized in that, Comprise: An acquisition unit is used for acquiring a target video; A source unit is used for acquiring a source video corresponding to the target video, and the target video is obtained by processing and creating the source video; A content unit is used for comparing the content of the target video and the source video to obtain the content type of the target video; An audio unit is used for when the content type of the target video is a funny content type, performing audio recognition on the target video and the source video to determine the audio type of the target video; A determination unit is used for when the audio type of the target video is a funny voice type, determining the target video as a funny voice video, so as to push the funny voice video to the user.

14. A server, characterized by Comprise a processor and a memory, the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps in the video recognition method based on machine learning as claimed in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a plurality of instructions, the instructions are suitable for the processor to load, to execute the steps in the video recognition method based on machine learning as claimed in any one of claims 1-12.

Citation Information

Patent Citations

  • Video classification method and device and server

    CN111428088A

  • Teaching video auditing method and device, equipment and medium

    CN112860943A