Video Processing Method, Apparatus, Video Processing Device, and Storage Medium

By using the target video processing model for feature extraction, classification and label identification in video processing, video data identification information that is both robust and separable is generated, and the shortcomings of identification information in the prior art are solved.

CN113822127BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110712104.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-25
Publication Date
2025-06-27
Estimated Expiration
2041-06-25

AI Technical Summary

Technical Problem

The existing video processing methods cannot have the robustness and separability of video data identification information, resulting in insufficient identification information in terms of robustness and separability.

Method used

By calling the target video processing model, feature extraction, classification processing and label identification processing are performed on the video data, and identification information based on classification information and label information is generated, making it both robust and divisible.

Benefits of technology

The robustness and separability of video data identification information are realized, the risk of overfitting of video processing models is reduced, and the identification information can comprehensively describe video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113822127B_ABST
    Figure CN113822127B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of video processing, and in particular, to a video processing method, apparatus, video processing device, and storage medium. Among them, the video processing method includes: calling a target video processing model to extract features from target video data to obtain video features of the target video data; performing classification processing on the target video data based on the video features to obtain classification information of the target video data; performing label recognition processing on the target video data based on the video features to obtain label information of the target video data; determining identification information of the target video data according to the classification information and the label information. This identification information has both robustness and separability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technologies, and in particular, to a video processing method, apparatus, video processing device, and storage medium. Background Art

[0002] With the rapid popularization of intelligent mobile terminals and the development of multimedia technologies, videos have gradually become carriers of information dissemination. In recent years, short videos have rapidly emerged, and videos have become a major way for people to entertain. Therefore, the field of video processing technologies has become a popular research direction. In the field of video processing technologies, the identification information of video data can be determined according to the video features of the video data. However, the existing video processing methods for determining identification information through video features cannot have both robustness and separability. Therefore, a video processing method that makes the identification information have both robustness and separability is an important research topic in the field of video processing technologies. Summary of the Invention

[0003] Embodiments of this application provide a video processing method, apparatus, video processing device, and storage medium, which can determine the classification information of target video data through a target video processing model, and determine the label information of the target video data through the target video processing model, so that the identification information determined based on the classification information and the label information has robustness and separability.

[0004] On the one hand, embodiments of this application provide a video processing method, which includes:

[0005] Invoking a target video processing model to perform feature extraction on target video data to obtain the video features of the target video data;

[0006] Performing classification processing on the target video data based on the video features to obtain the classification information of the target video data;

[0007] Performing label recognition processing on the target video data based on the video features to obtain the label information of the target video data;

[0008] Determining the identification information of the target video data according to the classification information and the label information.

[0009] On the other hand, embodiments of this application provide a video processing apparatus, which includes:

[0010] A feature extraction unit, configured to invoke a target video processing model to perform feature extraction on target video data to obtain the video features of the target video data;

[0011] A processing unit, configured to perform classification processing on the target video data based on the video features to obtain the classification information of the target video data;

[0012] The processing unit is further configured to perform label recognition processing on the target video data based on the video features to obtain the label information of the target video data;

[0013] The determination unit is configured to determine the identification information of the target video data according to the classification information and the label information.

[0014] On the other hand, an embodiment of the present application provides a video processing device, which includes an input interface and an output interface. The video processing device further includes:

[0015] A processor adapted to implement one or more instructions; and,

[0016] A computer storage medium storing one or more instructions, where the one or more instructions are adapted to be loaded and executed by the processor to perform the following steps:

[0017] Call the target video processing model to extract features from the target video data to obtain the video features of the target video data;

[0018] Perform classification processing on the target video data based on the video features to obtain the classification information of the target video data;

[0019] Perform label recognition processing on the target video data based on the video features to obtain the label information of the target video data;

[0020] Determine the identification information of the target video data according to the classification information and the label information.

[0021] On the other hand, an embodiment of the present application provides a computer storage medium storing one or more instructions, where the one or more instructions are adapted to be loaded and executed by the processor to perform the following steps:

[0022] Call the target video processing model to extract features from the target video data to obtain the video features of the target video data;

[0023] Perform classification processing on the target video data based on the video features to obtain the classification information of the target video data;

[0024] Perform label recognition processing on the target video data based on the video features to obtain the label information of the target video data;

[0025] Determine the identification information of the target video data according to the classification information and the label information.

[0026] In an embodiment of the present application, when target video data is obtained, a video processing device may call a target video processing model to classify the target video data based on the video features of the target video data to obtain classification information of the target video data; and perform label recognition processing on the target video data based on the video features of the target video data to obtain label information of the target video data; and determine identification information of the target video data according to the classification information and the label information. Since the classification processing in the target video processing model has high accuracy, few categories, and coarse granularity, the classification information obtained through classification processing is more robust; and the label recognition processing is more specific and has finer granularity, and the label information obtained through label recognition processing has better distinguishability, so that the identification information obtained based on the classification information and the label information has both the robustness of the classification information and the distinguishability of the label information. The target video processing model will not overfit, avoiding overfitting to the classification information, resulting in insufficient separability of the identification information, and also avoiding overfitting to the label information, resulting in poor robustness of the identification information, reducing the overfitting risk of the target video processing model. At the same time, the identification information is obtained based on the classification information and the label information, and the identification information can meet both robustness and separability, and the identification information can comprehensively describe the target video data. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0028] Figure 1a is the application of the video processing method provided by the embodiment of the present application in the video deduplication task;

[0029] Figure 1b is the application of the video processing method provided by the embodiment of the present application in the video deduplication task;

[0030] Figure 2a is the application of the video processing method provided by the embodiment of the present application in the video recommendation task;

[0031] Figure 2b is the application of the video processing method provided by the embodiment of the present application in the video recommendation task;

[0032] Figure 3 is a schematic flowchart of a video processing method provided by an embodiment of the present application;

[0033] Figure 4 is a schematic flowchart of a method for obtaining multimodal features provided by an embodiment of the present application;

[0034] Figure 5 It is a schematic flow chart of a video processing model provided by an embodiment of the present application;

[0035] Figure 6 It is a schematic structural diagram of a blockchain provided by an embodiment of the present application;

[0036] Figure 7 It is a schematic flow chart of another video processing method provided by an embodiment of the present application;

[0037] Figure 8 It is a schematic flow chart of a video processing model with a cascade structure provided by an embodiment of the present application;

[0038] Figure 9 It is a schematic flow chart of the training process of a video processing model provided by an embodiment of the present application;

[0039] Figure 10 It is a schematic structural diagram of a video processing device provided by an embodiment of the present application;

[0040] Figure 11 It is a schematic structural diagram of a video processing device provided by an embodiment of the present application. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0042] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. Among them, Machine Learning (ML) is a multi-disciplinary subject that involves several disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance.

[0043] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. In addition, artificial intelligence technology can also be applied in other fields. For example, machine learning in artificial intelligence technology can be applied to representation learning in the video technology field. Among them, representation learning refers to converting video data into a form that can be effectively developed, that is, removing invalid or redundant information in the video data, refining the valid information, and forming identification information so that the identification information of the video data can be applied to various downstream tasks. Among them, the embodiments of the present application propose a video processing method based on machine learning, enabling the video processing device to use multi-task learning in machine learning to construct a video processing model, and can call the video processing model to perform classification processing to obtain classification information of the video data; and call the video processing model to perform label recognition processing to obtain label information of the video data, so that the identification information of the video data can be obtained based on the classification information and label information of the video data, making the identification information obtained by the video processing model have both the generalization of the classification information and the specificity of the label information.

[0044] In specific implementation, the video processing method can be executed by a video processing device. The video processing device mentioned here can refer to any device with data computing capabilities, such as a terminal device or a server. Among them, the terminal device can include, but is not limited to: smart phones, tablets, laptop computers, wearable devices, desktop computers, etc.; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, Content Delivery Network (CDN), middleware services, domain name services, security services, and big data and artificial intelligence platforms, etc.

[0045] The video processing device can use this video processing method to perform representation learning processing on the video data collected in various scenarios to obtain the identification information of the video data, so that the identification information of the video data can be used as the underlying feature of the video data for downstream tasks. Among them, the downstream tasks can include, but are not limited to, video recommendation ranking, video recall scattering, and video deduplication, etc.

[0046] In one embodiment, the downstream task may be a video deduplication task. When the video processing device obtains the target video data to be processed, it may use the trained target video processing model to obtain the identification information of the target video data, and then use the indexing tool to retrieve in the first preset video library. When there is original video data in the first preset video library whose similarity with the identification information of the target video data is greater than the preset threshold, obtain the first user identification for publishing the target video data and the second user identification for publishing the original video data; if the users indicated by the first user identification and the second user identification are different, determine the target video data as plagiarized video data. At this time, when the video processing device determines the target video data as plagiarized video data, it may reduce the weight of the target video data, hide the target video data, reduce the exposure of the target video data, and improve the user experience of the video product.

[0047] Among them, the first preset video library includes multiple original video data. The original video data may refer to video data with unique personality in content or form, and is video data independently created by users. For example, video data obtained by users using video capture devices to shoot natural environments.

[0048] Among them, the indexing tool may be a tool related to the identification information of video data. For example, when the identification information of video data is a vector, the indexing tool may be a vector indexing tool (such as faiss). When the indexing tool is a vector indexing tool and the identification information is a vector, the indexing tool can judge the vector distance between the identification information of each original video data in the first preset video library and the identification information of the target video data. When there is original video data in the first preset video library with a vector distance less than the preset threshold, the first user identification for publishing the target video data and the second user identification for publishing the original video data can be obtained; among them, the user identification may be a user name or a user ID, etc. If the users indicated by the first user identification and the second user identification are different, determine the target video data as plagiarized video data. If the users indicated by the first user identification and the second user identification are the same, determine the target video data as normal video data.

[0049] For example, for Figure 1a the target video data shown above, the video processing device may retrieve in the first preset video library and obtain Figure 1a original video data whose similarity with the identification information of the target video data shown below is greater than the preset threshold. The video processing device can obtain that the first user identification for publishing the Figure 1a target video data shown above is "Drama Highlights", and for publishing Figure 1aThe second user identifier of the original video data shown on the lower side is "XX Video Culture". Since the first user identifier and the second user identifier indicate different users, the video processing device may Figure 1a determine the target video data shown on the upper side as copied video data. For example, for Figure 1b the target video data shown on the upper side, the video processing device may retrieve in the first preset video library and obtain Figure 1b the original video data whose similarity between the identifier information shown on the lower side and the identifier information of the target video data is greater than the preset threshold. The video processing device may obtain that Figure 1b the first user identifier of the target video data shown on the upper side is "XX Entertainment", and the Figure 1b second user identifier of the original video data shown on the lower side is "XX Video Culture". Since the first user identifier and the second user identifier indicate the same user, namely user "XX", the video processing device may Figure 1b determine the target video data shown on the upper side as normal video data.

[0050] In another embodiment, the downstream task may also be a video recommendation task. On the user side of the video recommendation task, the identifier information of the video data may be introduced, and the identifier information of the video data may be used as a continuous feature, and the user profile information of the target user may be used as a sparse feature. Based on the identifier information of the target video data and the user profile information, the user feature information of the target user may be determined. Specifically, when the video processing device obtains the target video data to be processed, it may use the trained target video processing model to obtain the identifier information of the target video data, and determine the user feature information of the target user based on the identifier information of the target video data and the user profile information. In one embodiment, the video processing device may obtain the user profile information of the target user accessing the target video data, and perform embedding processing on the user profile information of the target user through an embedding layer to obtain the processed user profile information. Then, through a dense embedding layer, the identifier information of the target video data and the processed user profile information are concatenated to obtain the user feature information of the target user, such as Figure 2aAs shown in the left figure. Correspondingly, on the item side, the identification information of each candidate video data in the second preset video library can also be used as a continuity feature. And based on the identification information of the target video data and the identification information of each candidate video data in the second preset video library, the candidate video feature information of each candidate video data in the second preset video library is determined. In one embodiment, the video processing device can use the video features of the target video data as discrete features, and perform embedding processing on the video features of the target video data through an embedding layer to obtain the identification information of the target video features. Then, through a dense embedding layer, the identification information of the target video data is concatenated with the identification information of each candidate video data in the second preset video library to obtain the candidate video feature information of each candidate video data in the second preset video library. As Figure 2a shown in the right figure. In another embodiment, the video processing device can also directly obtain the identification information of the target video data through the target video processing model, and then through a dense embedding layer, the identification information of the target video data is concatenated with the identification information of each candidate video data in the second preset video library to obtain the candidate video feature information of each candidate video data in the second preset video library. The user feature information obtained on the user side and the candidate video feature information of each candidate video data in the second preset video library obtained on the item side can be applied to the upper network. For example, candidate video feature information matching the user feature information can be searched for in the second preset video library as the target candidate video feature information, and the target candidate video data corresponding to the target candidate video feature information is used as the recommended video data for the target user.

[0051] Among them, the user who accesses the video data can refer to the user who has performed a user operation on the video data. For example, the user who clicks on the video data or browses the video data.

[0052] Among them, the user portrait information can refer to the information used to describe the user's characteristics, such as name, nickname, gender, age, and so on.

[0053] It should be noted that the collection and processing of relevant data (such as user portrait information and other data) in the embodiments of the present application should strictly meet the requirements of relevant laws and regulations. Obtaining personal information requires the informed consent of the individual subject (or having a legal basis for information acquisition), and subsequent data use and processing behaviors should be carried out within the scope authorized by laws and regulations and the individual information subject. For example, when the embodiments of the present application are applied to specific products or technologies, when obtaining user portrait information, it is necessary to inform the user through an interactive interface and obtain the permission or consent of the object, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant region.

[0054] In one embodiment, the video recommendation task can be a model with a two-tower structure, such asFigure 2b As shown Figure 2b shows a schematic structural diagram of a video recommendation task. As Figure 2b shown in the left figure, the video recommendation task of the embodiment of the present application introduces the identification information of the target video data, takes the identification information of the target video data as a continuous feature, and splices it with the discrete feature of the user portrait information to obtain user feature information. As Figure 2b shown in the right figure, the video recommendation task of the embodiment of the present application introduces the identification information of the candidate video data, takes the identification information of each candidate video data in the second preset video library as a continuous feature, and splices it with the identification information of the target video data to obtain the candidate video feature information of each candidate video data. Then, the candidate video feature information matching the user feature information can be searched in the second preset video library as the target candidate video feature information, and the target candidate video data corresponding to the target candidate video feature information is used as the recommended video data of the target user. In the video recommendation task of the embodiment of the present application, the user feature information of the target user is obtained by splicing the identification information of the target video data and the user portrait information of the target user. This user feature information can more comprehensively represent the target user accessing the target video data. And the candidate video feature information is obtained by splicing the identification information of the target video data and the identification information of the candidate video data. This candidate video feature information can more comprehensively represent the features of the candidate video data in the second preset video library. Therefore, the target candidate video data corresponding to the target candidate video feature information found in the second preset video library is more accurate, that is, the recommended video data of the target user is more accurate. The accuracy of the video recommendation task is improved, and the user experience is improved.

[0055] Based on the above description, an embodiment of the present application proposes a video processing method; this video processing method can be executed by the above-mentioned video processing device. See Figure 3 shown, this video processing method may include the following steps S301 - S304:

[0056] S301: Invoke the target video processing model to extract features from the target video data to obtain the video features of the target video data.

[0057] Among them, the video data may include any type of video data. For example, the video data may be film and television video data, short video data, real-time shared video data, and so on. Among them, the short video can also be called a short film video, which is generally a video with a playback duration within N minutes (such as 4 minutes, 5 minutes, etc.) transmitted on the Internet new media. The real-time shared video data may include, but is not limited to: live video data, network conference video data, and so on.

[0058] Among them, video features may include features obtained by describing video data from any angle. For example, title features obtained by describing from the perspective of the video data title; for another example, video stream features obtained by describing from the perspective of the video data content; and for still another example, audio features obtained by describing from the audio perspective of the video data. In one embodiment, the video features may include features from one angle of the video data. In another embodiment, the video features may include features from multiple angles of the video data, that is, multi-modal features. The multi-modal features may include any combination of title features obtained by describing from the perspective of the video data title, video stream features obtained by describing from the perspective of the video data content, and audio features obtained by describing from the audio perspective of the video data. It should be understood that the more description angles the multi-modal features contain, the more accurate the description of the video data by the multi-modal features.

[0059] Specifically, the video processing device can perform feature extraction on the video data to obtain the video features of the video data. In one embodiment, the video features refer to the multi-modal features combined by the video stream features from the perspective of the video data content and the title features from the perspective of the video data title. The video processing device can obtain the video information describing the video data content angle and the title information describing the video data title angle, and then obtain the video stream features of the video data from the video information through the video processing model, and obtain the title features of the video data from the title information through the video processing model. The video stream features of the video data and the title features of the video data are fused to obtain the multi-modal features of the video data.

[0060] Among them, the video information describing the video data content angle may be the video stream in the video data. Among them, the video processing model may include a video feature extraction module, and the video feature extraction module can be used to obtain the video stream features of the video data. The video feature extraction module may include a sampling module, an image feature extraction module, a frame feature aggregation module, and a feature enhancement module. Specifically, the sampling module can be used to perform global and / or sparse sampling on the video stream to obtain a set of frame images. The image feature extraction module can be used to extract the image features of each frame image in the set of frame images, the frame feature aggregation module can be used to aggregate the image features of each frame image to obtain initial video features, and the feature enhancement module can be used to perform feature enhancement on the initial video features to obtain the video stream features of the video data. Among them, the image feature extraction module can be any image feature extraction network (such as InceptionResNetV2, ResNet, and EfficienNet, etc.), among which, the frame feature aggregation module can be any image aggregation network (such as NeXtVLAD), and among which, the feature enhancement module can be any image enhancement network (such as SENet).

[0061] Among them, the title information describing the title angle of the video data can be the title text of the video data. Among them, the video processing model can include a title feature extraction module, which can be used to obtain the title features of the video data. Among them, the title feature extraction module can include a word segmentation module, a word embedding module, a hybrid deep neural network model, and a pooling layer. Specifically, the word segmentation module can be used to segment the title text to obtain segmented text, the word embedding module can be used to perform embedding processing on the segmented text to obtain word vectors, the hybrid deep neural network model is used to extract initial features from the word vectors, and the pooling layer is used to perform pooling processing on the initial features to obtain title features.

[0062] See Figure 4 , Figure 4 shows a schematic flowchart of obtaining multimodal features. As Figure 4 shown in the upper diagram of, the target video data includes a video stream 401 during a boxing match. The sampling module can sample the video stream 401 to obtain a set of frame images shown as 402, and then use the image feature extraction network InceptionResNetV2 to extract the image features of each frame image in the set of frame images. The frame feature aggregation module NeXtVLAD can aggregate the image features of each frame image to obtain initial video features, and the feature enhancement module can perform feature enhancement on the initial video features to obtain the video stream features of the video data. As Figure 4 shown in the lower diagram of, the title text 403 in the target video data is "UFC e-sports: BB participates in the competition, smashes the fighting champion into a coma angrily, so domineering". The word segmentation module can segment the title text 403 to obtain segmented text, the word embedding module performs embedding processing on the segmented text to obtain word vectors, the hybrid deep neural network model can extract initial features from the word vectors, and the pooling layer performs pooling processing on the initial features to obtain title features. Finally, the video features of the video data and the title features of the video data can be fused through a feature fusion module (such as the GateMultimodal Unit structure) to obtain the multimodal features of the video data.

[0063] S302: Classify the target video data based on the video features to obtain the classification information of the target video data.

[0064] Among them, the classification information of the video data can include the category to which the video data belongs. The video processing device can call the target video processing model to classify the target video data based on the video features to determine the category to which the target video data belongs. Specifically, the video processing device can call the target video processing model to classify the target video data based on the video features to determine the probability of the target video data under each category, and determine the category corresponding to the maximum probability as the category to which the target video data belongs.

[0065] In one embodiment, the video processing model of the embodiments of the present application may include multiple tasks. For example, it may include a classification task and a tagging task. It should be noted that with the development of the business, the video processing model may also include other tasks. For example, the video processing model may also include an account task, etc., which is not limited in the present application.

[0066] Among them, the classification task can be used to determine the classification information of video data. In one embodiment, the classification task can be a multi-classification task. The video processing device can call the classification task in the target video processing model to determine the probability that the target video data belongs to each category based on the video features, and determine the category with the highest probability as the category to which the target video data belongs.

[0067] Please refer to Figure 5 as shown Figure 5 which shows a schematic flowchart of a video processing model. Among them, Figure 5 the upper diagram shows a schematic flowchart of the classification task. The video processing device can call the classification task in the target video processing model to determine the probability that the target video data belongs to each category based on the video features. Among them, Figure 5 continuing Figure 4 the example shown, the probabilities of each category can be shown as Figure 5 shown in the upper diagram, that is, the probability that the target video data belongs to the "funny" category, the probability that the target video data belongs to the "movie" category, the probability that the target video data belongs to the "TV drama" category, the probability that the target video data belongs to the "variety show" category, the probability that the target video data belongs to the "entertainment" category, the probability that the target video data belongs to the "game" category, the probability that the target video data belongs to the "internet celebrity" category, the probability that the target video data belongs to the "music" category, the probability that the target video data belongs to the "quyi" category, and the probability that the target video data belongs to the "animation" category. It can be seen that the probability of the "game" category is the highest. Therefore, the video processing device calls the classification task in the target video processing model to determine that the category to which the target video data belongs is "game".

[0068] S303: Perform label recognition processing on the target video data based on the video features to obtain the label information of the target video data.

[0069] Optionally, the label information may include the labels to which the video data belongs. Since there may be multiple labels for the video data, it is necessary to determine whether the video data contains each label separately. Specifically, the video processing device can call the target video processing model to determine the probability that the target video data belongs to each label based on the video features, and determine the labels with probabilities greater than the probability threshold as the labels contained in the target video data.

[0070] Among them, the tagging task can be used to determine the tagging information of video data. In one embodiment, the tagging task can be a combination of multiple binary classification tasks, and one binary classification task can be used to determine whether the video data contains the tag corresponding to this binary classification task. Specifically, the video processing device can use the target binary classification task to determine the probability that the target video data contains the target tag. If the probability that the target video data contains the target tag is greater than the probability threshold, it is determined that the target video data contains the target tag; if the probability that the target video data contains the target tag is less than or equal to the probability threshold, it is determined that the target video data does not contain the target tag.

[0071] Among them, Figure 5 Continuing Figure 4 the example shown, the probabilities of each tag can be shown as Figure 5 shown in the following figure. That is, the probability that the target video data contains the tag "Romance of the Three Kingdoms", the probability that the target video data contains the tag "Shuangyashan", the probability that the target video data contains the tag "Fighting Game", the probability that the target video data contains the tag "AA", the probability that the target video data contains the tag "BB", the probability that the target video data contains the tag "Program Commentary", the probability that the target video data contains the tag "Snow Leopard", the probability that the target video data contains the tag "UFC", and the probability that the target video data contains the tag "Honor of Kings". It can be seen that the probabilities of the tags "Fighting Game", "BB", "Program Commentary", and "UFC" are greater than the probability threshold. Therefore, the video processing device invokes the tagging task in the target video processing model to determine that the tags contained in the target video data based on the video features are: "Fighting Game", "BB", "Program Commentary", and "UFC".

[0072] It should be noted that S302 and S303 can be parallel steps. In the embodiment of the present application, the step shown in S302 is executed first, and then the step shown in S303 is executed. In other embodiments, the step shown in S303 can also be executed first, and then the step shown in S302 is executed, that is, first invoke the tagging task in the target video processing model to obtain the tagging information of the target video data; then invoke the classification task in the target video processing model to obtain the classification information of the target video data.

[0073] S304: Determine the identification information of the target video data according to the classification information and the tagging information.

[0074] In one embodiment, the video processing device may perform attention processing on the classification information and the label information through the attention mechanism of the target video processing model to obtain the identification information of the target video data. Among them, the attention mechanism refers to the ability to focus attention on actually important features through attention weights. For example, when the target video processing model pays more attention to the classification information, the attention weight of the classification information can be set to be greater than the attention weight of the label information. Another example is that when the target video processing model pays more attention to the label information, the attention weight of the label information can be set to be greater than the attention weight of the classification information.

[0075] In another embodiment, the label information may also be filtered based on the classification information, and the filtered label information is used as the identification information of the target video data. The filtered label information matches the category indicated by the classification information. Due to the generalization of the label information, the label information may include labels under multiple categories at the same time. For example, the label information may include the "single-player game" label under the "game" category and the "second dimension" label under the "animation" category. If the classification information indicates that the category to which the target video data belongs is the "game" category, then the "single-player game" label under the "game" category and the "second dimension" label under the "animation" category included in the label information can be filtered based on the category indicated by the classification information. The filtered "single-player game" label matches the "game" category, and the "single-player game" label is used as the identification information of the target video data.

[0076] In a feasible implementation manner, in order to facilitate the downstream task to call the identification information of the target video data, the blockchain technology can be used to write the identification information of the target video data into the blockchain. Specifically, the video processing device may encapsulate the identification information of the target video data into a block and store the block on the blockchain.

[0077] Among them, the blockchain is a chain data structure formed by combining data blocks in sequence according to the time sequence, and a distributed ledger that ensures the data cannot be tampered with and forged by cryptographic means. Multiple independent distributed nodes store the same records. The blockchain technology has achieved decentralization and has become the cornerstone of trusted digital asset storage, transfer, and transaction.

[0078] For Figure 6Taking the schematic structural diagram of the blockchain shown as an example, when writing the identification information of the target video data into the blockchain, the identification information of the target video data can be encapsulated into a block and added to the end of the existing blockchain. The consensus algorithm ensures that the newly added blocks at each node are exactly the same. Each block records several pieces of identification information and also contains the hash value of the previous block. All blocks save the hash value of the previous block in this way and are connected in sequence to form the blockchain. The hash value of the previous block is stored in the block header of the next block in the blockchain. When the identification information in the previous block changes, the hash value of this block will also change accordingly. Therefore, it is difficult to tamper with the identification information uploaded to the blockchain, improving the reliability of the data.

[0079] In one embodiment, in a subsequent time period, the video processing device can directly obtain the identification information of the target video data in the blockchain to perform downstream tasks without having to obtain the identification information of the target video data again, improving timeliness and accuracy.

[0080] In the embodiment of the present application, when the target video data is obtained, the video processing device can call the target video processing model to classify the target video data based on the video features of the target video data to obtain the classification information of the target video data; and perform label recognition processing on the target video data based on the video features of the target video data to obtain the label information of the target video data; and obtain the identification information of the target video data according to the classification information and the label information. Since the classification processing in the target video processing model has high accuracy, few categories, and coarse granularity, the classification information obtained through classification processing is more robust. And the label recognition processing is more specific and has finer granularity, and the label information obtained through label recognition processing has better distinguishability. The identification information obtained based on the classification information and the label information has both the robustness of the classification information and the distinguishability of the label information. The target video processing model will not overfit, avoiding overfitting to the classification information, resulting in insufficient separability of the identification information, and also avoiding overfitting to the label information, resulting in poor robustness of the identification information, reducing the overfitting risk of the target video processing model. At the same time, the identification information is obtained based on the classification information and the label information, and the identification information can satisfy both the robustness of the classification information and the separability of the label information, and this identification information can more comprehensively describe the target video data.

[0081] Referring to the above Figure 3 description of the relevant method embodiments shown, Figure 3The video processing method shown can obtain the classification information of the target video data through classification processing and obtain the label information of the target video data through label recognition processing. Since classification processing is much simpler than label recognition processing, the accuracy rate of classification processing is relatively high. Although label recognition processing is more specific, due to problems such as an excessively large number of label information, fine granularity, and severe long-tail distribution in label recognition processing, label recognition processing does not have the same robust characteristics as classification processing. After introducing classification information, the label information obtained through label recognition processing is more robust. Therefore, in this solution, classification processing can also be used as an auxiliary process for label recognition processing. Based on this, the embodiments of the present application provide another video processing method. Refer to Figure 7 as shown, this video processing method may include the following steps S701-S704:

[0082] S701: Invoke the target video processing model to extract features from the target video data to obtain the video features of the target video data.

[0083] S702: Based on the video features, perform classification processing on the target video data to obtain the classification information of the target video data.

[0084] It should be noted that the specific implementation manners of steps S701-S702 can refer to Figure 3 the specific descriptions of the relevant embodiments in, and will not be elaborated here.

[0085] S703: Based on the classification information and the video features, perform label recognition processing on the target video data to obtain the label information of the target video data.

[0086] Specifically, perform label recognition processing on the target video data according to the category indicated by the classification information and the video features to obtain the label information of the target video data, and the label information matches the category indicated by the classification information. That is to say, the classification task and the label task can be set as a cascaded structure, and the classification task is an auxiliary task of the label task.

[0087] Please refer to Figure 8 as shown, Figure 8 shows a schematic flow diagram of the video processing model of the cascaded structure. Among them, Figure 8 the upper side diagram shows the schematic flow diagram of the classification task. The video processing device can determine the classification information of the target video data based on the video features, that is, determine the probability that the target video data belongs to each category, and determine the category to which the target video data belongs based on the probabilities of each category. Among them, Figure 8 the lower side diagram shows the schematic flow diagram of the label task. In the label task, the label information of the target video data can be determined based on the classification information determined by the classification task and the video features of the target video data. As Figure 8As shown, the classification information corresponding to the classification task and the video features of the target video data are both input into the tagging task in the target video processing model. At this time, the video processing device can call the tagging task in the target video processing model to perform tagging recognition processing on the target video data according to the category indicated by the classification information and the video features, and obtain the tagging information of the target video data, where the tagging information matches the category indicated by the classification information.

[0088] This tagging information matches the category indicated by the classification information. For example, if the category indicated by the classification information is "game", then each tag in the tagging information will match "game", and the tagging information will include tags related to game categories such as "fighting", "single-player game", "game commentary", etc. Another example, if the category indicated by the classification information is "animation", then each tag in the tagging information will match "animation", and the tagging information will include tags related to animation categories such as "Japanese animation", "two-dimensional", "animation adaptation", etc.

[0089] In this case, the tagging information matches the category indicated by the classification information, and the situation of mutually exclusive tags in the tagging information is greatly reduced. It avoids the situation where the tagging information includes tags under multiple categories in a separate tagging task. For example, in a separate tagging task, the tagging information may simultaneously include the tag "single-player game" under the "game" category and the tag "two-dimensional" under the "animation" category, and the accuracy of the tagging information is relatively low. Since the classification task is much simpler than the tagging task, the accuracy of the classification task is relatively high. When using the category indicated by the classification information obtained from the classification task as a benchmark, there will be no tags under multiple categories, and the accuracy of the tagging information is higher.

[0090] S704: Determine the identification information of the target video data according to the classification information and the tagging information.

[0091] It should be noted that the specific implementation method of step S704 can refer to Figure 3 the specific description of the relevant embodiments therein, which will not be elaborated here.

[0092] In the embodiments of the present application, tagging recognition processing is performed on the target video data according to the category indicated by the classification information and the video features, and the tagging information of the target video data is obtained. This tagging information matches the category indicated by the classification information. Since the classification task is much simpler than the tagging task, the accuracy of the classification task is relatively high. When using the category indicated by the classification information obtained from the classification task as a benchmark, it avoids the appearance of tags under multiple categories and improves the accuracy of the tagging information.

[0093] Further, to verify the beneficial effects of the video processing method according to the embodiments of the present application, the identification information of the same video stream can be obtained by using a single-label task and a cascade structure, and compared through evaluation metrics such as MAP@10, accuracy, recall, and F1. The specific evaluation metric results are shown in Table 1 as follows:

[0094] Table 1 Evaluation Metric Results

[0095] Task MAP@10 Accuracy Recall F1 Single-label task 0.7072 0.7903 0.5865 0.6733 Cascade structure 0.7161 0.7659 0.6066 0.6770

[0096] As can be seen from Table 1, the MAP@10, recall, and F1 of the cascade structure are all better than those of the single-label task, which proves that the identification information obtained through the cascade structure is better.

[0097] Referring to the relevant descriptions of the above Figure 3 or Figure 7 shown method embodiments, it can be known that Figure 3 or Figure 7 shown video processing method can call the trained target video processing model to obtain the identification information of the target video data. Then, before calling the trained target video processing model, the target video processing model needs to be trained. Based on this, referring to Figure 9 , Figure 9 shows a schematic flowchart of another video processing method, which may include S901 - S907:

[0098] S901: Obtain training samples, where the training samples include sample video data, the reference classification of the sample video data, and the reference label of the sample video data.

[0099] Among them, the sample video data can be labeled to obtain the reference classification and the reference label of the sample video.

[0100] S902: Extract features from the sample video data through the initial video processing model to obtain the video features of the sample video data.

[0101] S903: Classify the sample video data based on the video features of the sample video data to obtain the classification information of the sample video data.

[0102] S904: Perform label recognition processing on the sample video data based on the video features of the sample video data to obtain the label information of the sample video data.

[0103] S905: Determine the first loss value according to the reference classification and the classification information of the sample video data, and determine the second loss value according to the reference label and the label information of the sample video data.

[0104] Among them, the first loss value can be the loss value corresponding to the classification task. As described above, the classification task can be a multi-classification task. Then, the first loss value determined by the video processing device according to the reference classification and classification information of the sample video data can be the multi-class cross-entropy loss value.

[0105] Among them, the second loss value can be the loss value corresponding to the tagging task. As described above, the tagging task is a combination of multiple binary classification tasks. Then, the second loss value determined by the video processing device according to the reference tag and tag information of the sample video data can include multiple binary cross-entropy loss values.

[0106] S906: Based on the first loss value and the second loss value, obtain the loss value of the initial video processing model.

[0107] In one embodiment, the video processing device can obtain the weight factor corresponding to the classification task and the weight factor corresponding to the tagging task, and process the first loss value and the second loss value based on the weight factor corresponding to the classification task and the weight factor corresponding to the tagging task to obtain the loss value of the initial video processing model.

[0108] Optionally, the sum of the weight factor corresponding to the classification task and the weight factor corresponding to the tagging task can be equal to the reference value. Then, the video processing device can perform weighted summation on the first loss value and the second loss value based on the weight factor corresponding to the classification task and the weight factor corresponding to the tagging task to obtain the loss value of the initial video processing model.

[0109] Optionally, both the weight factor corresponding to the classification task and the weight factor corresponding to the tagging task can be the reference value. The video processing device can directly add the first loss value and the second loss value to obtain the loss value of the initial video processing model.

[0110] S907: Train the initial video processing model according to the loss value of the initial video processing model to obtain the target video processing model.

[0111] Specifically, the video processing device can perform a derivative calculation on the loss value of the initial video processing model to obtain the update parameters of the initial video processing model, and perform gradient backpropagation on the classification task execution module in the initial video processing model and the tagging task execution module in the initial video processing model based on the update parameters of the initial video processing model until the converged target video processing model is obtained.

[0112] When training a video processing model in an embodiment of the present application, the magnitude of the loss value of the initial video processing model is affected by both the first loss value and the second loss value. Therefore, the direction of parameter update cannot be biased towards only one direction, but rather satisfies that the loss values of both tasks are reduced. This enables the trained target video processing model to focus on features that are important for both tasks, reducing the risk of the target video processing model overfitting to a certain task. At the same time, the identification information learned by the target video processing model meets the requirements of both the classification task and the tagging task, ensuring generalization ability.

[0113] Based on the description of the above video processing method embodiment, an embodiment of the present application also discloses a video processing apparatus 100. The video processing apparatus 100 may be a computer program (including program code) running in one of the above-mentioned video processing devices. The video processing apparatus 100 can execute Figure 3 、 Figure 7 or Figure 9 the methods shown. Please refer to Figure 10 and the video processing apparatus 100 can operate the following units:

[0114] A feature extraction unit 1001, configured to call a target video processing model to perform feature extraction on target video data to obtain video features of the target video data;

[0115] A processing unit 1002, configured to perform classification processing on the target video data based on the video features to obtain classification information of the target video data;

[0116] The processing unit 1002 is further configured to perform tagging recognition processing on the target video data based on the video features to obtain tagging information of the target video data;

[0117] A determination unit 1003, configured to determine identification information of the target video data according to the classification information and the tagging information.

[0118] In one implementation, the processing unit 1002 is configured to perform tagging recognition processing on the target video data based on the video features to obtain tagging information of the target video data, including:

[0119] Performing tagging recognition processing on the target video data based on the classification information and the video features to obtain tagging information of the target video data.

[0120] In another implementation, the processing unit 1002 is configured to perform tagging recognition processing on the target video data based on the classification information and the video features to obtain tagging information of the target video data, including:

[0121] Perform label recognition processing on the target video data according to the category indicated by the classification information and the video features, and obtain the label information of the target video data, where the label information matches the category indicated by the classification information.

[0122] In another implementation, the processing unit 1002 is further configured to:

[0123] Obtain training samples, where the training samples include sample video data, the reference classification of the sample video data, and the reference label of the sample video data;

[0124] Call the initial video processing model to extract features from the sample video data to obtain the video features of the sample video data;

[0125] Perform classification processing on the sample video data based on the video features of the sample video data to obtain the classification information of the sample video data;

[0126] Perform label recognition processing on the sample video data based on the video features of the sample video data to obtain the label information of the sample video data;

[0127] Determine a first loss value according to the reference classification and the classification information of the sample video data, and determine a second loss value according to the reference label and the label information of the sample video data;

[0128] Based on the first loss value and the second loss value, obtain the loss value of the initial video processing model;

[0129] Train the initial video processing model according to the loss value of the initial video processing model to obtain the target video processing model.

[0130] In another implementation, the feature extraction unit 1001 is configured to call the target video processing model to extract features from the target video data to obtain the video features of the target video data, including:

[0131] Obtain the video stream features of the target video data through the target video processing model, and obtain the title features of the target video data through the target video processing model;

[0132] Fuse the video stream features of the target video data and the title features of the target video data to obtain the multi-modal features of the target video data.

[0133] In another implementation, the processing unit 1002 is further configured to:

[0134] When there is original video data in the first preset video library whose similarity between the identification information and the identification information of the target video data is greater than a preset threshold, obtain the first user identification for publishing the target video data and the second user identification for publishing the original video data;

[0135] If the users indicated by the first user identifier and the second user identifier are different, the target video data is determined to be transported video data.

[0136] In yet another implementation manner, the processing unit 1002 is further configured to:

[0137] Obtaining user portrait information of a target user who accesses the target video data, and determining user feature information of the target user based on identification information of the target video data and the user portrait information;

[0138] Acquire identification information of each candidate video data in the second preset video library, and determine candidate video feature information of each candidate video data in the second preset video library based on the identification information of the target video data and the identification information of each candidate video data;

[0139] The candidate video feature information matching the user feature information is searched in the second preset video library as the target candidate video feature information, and the target candidate video data corresponding to the target candidate video feature information is used as the recommended video data for the target user.

[0140] According to one embodiment of the present application, Figure 3 , Figure 7 or Figure 9 Each step involved in the method shown can be performed by Figure 10 The various units in the video processing device 100 shown in FIG. Figure 3 Step S301 shown is composed of Figure 10 The feature extraction unit 1001 shown in FIG. 1 is used to perform the steps S302-S303. Figure 10 The processing unit 1002 shown in FIG. 1 is executed, and step S304 is performed by Figure 10 The determination unit 1003 shown in FIG. 1 is used to perform the above operation.

[0141] According to another embodiment of the present application, Figure 10 The various units in the video processing device 100 shown can be respectively or all merged into one or several other units to constitute, or one (some) of the units can also be split into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In practical applications, the function of a unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, other units can also be included based on the video processing device 100. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.

[0142] According to another embodiment of the present application, it may include processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), and a Read-Only Memory (ROM). For example, a computer program (including program code) that can execute the steps involved in the corresponding methods shown in Figure 3 , Figure 7 or Figure 9 runs on a general computing device such as a computer to construct a video processing device 100 as shown in Figure 10 and to implement the video processing method of the embodiments of the present application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above video processing device through the computer-readable recording medium, and run therein.

[0143] In the embodiments of the present application, when the target video data is acquired, the video processing device can call the target video processing model to classify the target video data based on the video features of the target video data to obtain the classification information of the target video data; and perform label recognition processing on the target video data based on the video features of the target video data to obtain the label information of the target video data; and determine the identification information of the target video data according to the classification information and the label information. Since the classification processing in the target video processing model has high accuracy, few categories, and coarse granularity, the classification information obtained through classification processing is more robust; and the label recognition processing is more specific and has finer granularity, and the label information obtained through label recognition processing has better distinguishability, so that the identification information obtained based on the classification information and the label information has both the robustness of the classification information and the distinguishability of the label information. The target video processing model will not overfit, avoiding overfitting to the classification information, resulting in insufficient separability of the identification information, and also avoiding overfitting to the label information, resulting in poor robustness of the identification information, reducing the overfitting risk of the target video processing model. At the same time, the identification information is obtained based on the classification information and the label information, and the identification information can satisfy both the robustness of the classification information and the separability of the label information, and the identification information can comprehensively describe the target video data.

[0144] Based on the description of the embodiments of the above video processing method, the embodiments of the present application also disclose a video processing device 110. Please refer to Figure 11 . The video processing device 110 at least includes a processor 1101, an input interface 1102, an output interface 1103, and a computer storage medium 1104, which can be connected through a bus or other means.

[0145] The computer storage medium 1104 is a memory device in the video processing device 110 for storing programs and data. It can be understood that the computer storage medium 1104 here can include both the built-in storage medium of the video processing device 110 and, of course, the extended storage medium supported by the video processing device 110. The computer storage medium 1104 provides a storage space, and the operating system of the video processing device 110 is stored in this storage space. Moreover, one or more instructions suitable for being loaded and executed by the processor 1101 are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium 1104 here can be a high-speed RAM memory; optionally, it can also be at least one computer storage medium far from the aforementioned processor. This processor can be called a Central Processing Unit (CPU), which is the core and control center of the video processing device 110 and is suitable for implementing one or more instructions, specifically loading and executing one or more instructions to implement the corresponding method processes or functions.

[0146] In one embodiment, one or more instructions stored in the computer storage medium 1104 can be loaded and executed by the processor 1101 to implement the execution of Figure 3 , Figure 7 or Figure 9 each step involved in the corresponding method shown in. Specifically, in the implementation, one or more instructions in the computer storage medium 1104 are loaded and executed by the processor 1101 to perform the following steps:

[0147] Call the target video processing model to extract features from the target video data to obtain the video features of the target video data;

[0148] Perform classification processing on the target video data based on the video features to obtain the classification information of the target video data;

[0149] Perform label recognition processing on the target video data based on the video features to obtain the label information of the target video data;

[0150] Determine the identification information of the target video data according to the classification information and the label information.

[0151] In one implementation manner, the processor 1101 is used to perform label recognition processing on the target video data based on the video features to obtain the label information of the target video data, including:

[0152] Perform label recognition processing on the target video data based on the classification information and the video features to obtain the label information of the target video data.

[0153] In another implementation, the processor 1101 is configured to perform label recognition processing on the target video data based on the classification information and the video features to obtain the label information of the target video data, including:

[0154] Perform label recognition processing on the target video data according to the category indicated by the classification information and the video features to obtain the label information of the target video data, and the label information matches the category indicated by the classification information.

[0155] In another implementation, the processor 1101 is further configured to: obtain training samples, where the training samples include sample video data, the reference classification of the sample video data, and the reference label of the sample video data;

[0156] Extract features from the sample video data through the initial video processing model to obtain the video features of the sample video data;

[0157] Perform classification processing on the sample video data based on the video features of the sample video data to obtain the classification information of the sample video data;

[0158] Perform label recognition processing on the sample video data based on the video features of the sample video data to obtain the label information of the sample video data;

[0159] Determine the first loss value according to the reference classification and the classification information of the sample video data, and determine the second loss value according to the reference label and the label information of the sample video data;

[0160] Based on the first loss value and the second loss value, obtain the loss value of the initial video processing model;

[0161] Train the initial video processing model according to the loss value of the initial video processing model to obtain the target video processing model.

[0162] In another implementation, the processor 1101 is configured to extract features from the target video data through the target video processing model to obtain the video features of the target video data, including:

[0163] Obtain the video stream features of the target video data through the target video processing model, and obtain the title features of the target video data through the target video processing model;

[0164] Fuse the video stream features of the target video data and the title features of the target video data to obtain the multi-modal features of the target video data.

[0165] In another implementation, the processor 1101 is further configured to:

[0166] When there is original video data in the first preset video library whose similarity between the identification information and the identification information of the target video data is greater than a preset threshold, obtain the first user identification of the user who publishes the target video data and the second user identification of the user who publishes the original video data;

[0167] If the users indicated by the first user identification and the second user identification are different, determine the target video data as copied video data.

[0168] In another implementation manner, the processor 1101 is further configured to:

[0169] Obtain the user profile information of the target user who accesses the target video data, and determine the user feature information of the target user based on the identification information of the target video data and the user profile information;

[0170] Obtain the identification information of each candidate video data in the second preset video library, and determine the candidate video feature information of each candidate video data in the second preset video library based on the identification information of the target video data and the identification information of each candidate video data;

[0171] Search in the second preset video library for candidate video feature information that matches the user feature information as the target candidate video feature information, and use the target candidate video data corresponding to the target candidate video feature information as the recommended video data for the target user.

[0172] In the embodiments of the present application, when the target video data is obtained, the video processing device may call the target video processing model to perform classification processing on the target video data based on the video features of the target video data to obtain the classification information of the target video data; and perform label recognition processing on the target video data based on the video features of the target video data to obtain the label information of the target video data; and determine the identification information of the target video data according to the classification information and the label information. Since the classification processing in the target video processing model has high accuracy, few categories, and coarse granularity, the classification information obtained through classification processing is more robust; and the label recognition processing is more specific and has finer granularity, and the label information obtained through label recognition processing has better distinguishability, so that the identification information obtained based on the classification information and the label information has both the robustness of the classification information and the distinguishability of the label information. The target video processing model will not overfit, avoiding overfitting to the classification information, resulting in insufficient separability of the identification information, and also avoiding overfitting to the label information, resulting in poor robustness of the identification information, reducing the overfitting risk of the target video processing model. At the same time, the identification information is obtained based on the classification information and the label information, and the identification information can satisfy both the robustness of the classification information and the separability of the label information, and this identification information can comprehensively describe the target video data.

[0173] It should be noted that the embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the video processing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the video processing device executes the steps performed in the above-mentioned embodiments of the video processing method. Figure 3 , Figure 7 or Figure 9 the steps performed in.

[0174] The foregoing disclosure is only a preferred embodiment of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the application.

Claims

1. A video processing method, characterized in that, Including: Invoking a target video processing model to extract features from target video data to obtain video features of the target video data; Performing classification processing on the target video data based on the video features to obtain classification information of the target video data; Performing label recognition processing on the target video data based on the video features to obtain label information of the target video data; Determining identification information of the target video data according to the classification information and the label information; Performing a downstream task according to the identification information, where the downstream task includes any one or two of a video recommendation task and a video deduplication task; Wherein, when the downstream task is a video deduplication task, the performing the downstream task according to the identification information includes: when there is original video data in a first preset video library whose similarity with the identification information of the target video data is greater than a preset threshold, obtaining a first user identifier of the user who publishes the target video data and a second user identifier of the user who publishes the original video data; if the users indicated by the first user identifier and the second user identifier are different, determining the target video data as copied video data; When the downstream task is a video recommendation task, the performing the downstream task according to the identification information includes: obtaining user portrait information of a target user who accesses the target video data, and determining user feature information of the target user based on the identification information of the target video data and the user portrait information; obtaining identification information of each candidate video data in a second preset video library, and determining candidate video feature information of each candidate video data in the second preset video library based on the identification information of the target video data and the identification information of each candidate video data; searching in the second preset video library for candidate video feature information that matches the user feature information as target candidate video feature information, and using the target candidate video data corresponding to the target candidate video feature information as recommended video data for the target user.

2. The method according to claim 1, wherein The performing the label recognition processing on the target video data based on the video features to obtain the label information of the target video data includes: Performing label recognition processing on the target video data based on the classification information and the video features to obtain the label information of the target video data.

3. The method according to claim 2, wherein The performing the label recognition processing on the target video data based on the classification information and the video features to obtain the label information of the target video data includes: Performing label recognition processing on the target video data according to the category indicated by the classification information and the video features to obtain the label information of the target video data, and the label information matches the category indicated by the classification information.

4. The method according to claim 1, wherein The method further includes: Obtaining training samples, where the training samples include sample video data, a reference classification of the sample video data, and a reference label of the sample video data; Extracting features from the sample video data through an initial video processing model to obtain video features of the sample video data; Classify the sample video data based on the video features of the sample video data to obtain the classification information of the sample video data; Perform label recognition processing on the sample video data based on the video features of the sample video data to obtain the label information of the sample video data; Determine a first loss value according to the reference classification and classification information of the sample video data, and determine a second loss value according to the reference label and label information of the sample video data; Based on the first loss value and the second loss value, obtain the loss value of the initial video processing model; Train the initial video processing model according to the loss value of the initial video processing model to obtain the target video processing model.

5. The method according to claim 1, wherein The step of calling the target video processing model to extract features from the target video data to obtain the video features of the target video data includes: Obtain the video stream features of the target video data through the target video processing model, and obtain the title features of the target video data through the target video processing model; Fuse the video stream features of the target video data and the title features of the target video data to obtain the multi-modal features of the target video data.

6. A video processing device, characterized in that, It includes: A feature extraction unit for calling the target video processing model to extract features from the target video data to obtain the video features of the target video data; A processing unit for classifying the target video data based on the video features to obtain the classification information of the target video data; The processing unit is further configured to perform label recognition processing on the target video data based on the video features to obtain the label information of the target video data; A determination unit for determining the identification information of the target video data according to the classification information and the label information; The processing unit is further configured to execute a downstream task according to the identification information, and the downstream task includes any one or two of a video recommendation task and a video deduplication task; Wherein, when the downstream task is a video deduplication task, the processing unit is configured to, when there is original video data in the first preset video library whose similarity with the identification information of the target video data is greater than a preset threshold, obtain the first user identifier for publishing the target video data and the second user identifier for publishing the original video data; if the users indicated by the first user identifier and the second user identifier are different, determine the target video data as copied video data; When the downstream task is a video recommendation task, the processing unit is configured to obtain the user profile information of the target user who accesses the target video data, and determine the user feature information of the target user based on the identification information of the target video data and the user profile information; obtain the identification information of each candidate video data in the second preset video library, and determine the candidate video feature information of each candidate video data in the second preset video library based on the identification information of the target video data and the identification information of each candidate video data; search in the second preset video library for candidate video feature information that matches the user feature information as the target candidate video feature information, and use the target candidate video data corresponding to the target candidate video feature information as the recommended video data for the target user.

7. A video processing device, comprising an input interface and an output interface, characterized in that, Further comprising: a processor adapted to implement one or more instructions; and, a computer storage medium storing one or more instructions, the one or more instructions being adapted to be loaded and executed by the processor to perform the video processing method according to any one of claims 1-5.

8. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, the one or more instructions being adapted to be loaded and executed by the processor to perform the video processing method according to any one of claims 1-5.

9. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the video processing method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Label data processing method and device and computer readable storage medium

    CN111711869A