Video content recognition method and apparatus, storage medium, and electronic device
By fusing multidimensional video features and user behavior data, a multimodal video feature vector is generated. Combined with a graded recognition parameter set, this solves the problem of low accuracy in video content recognition and achieves more efficient low-end video recognition and quality control.
Patent Information
- Application Number
- CN202110176761.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-09
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2041-02-09
AI Technical Summary
In existing technologies, the accuracy of video content recognition is relatively low, making it difficult to effectively identify and distribute low-end videos, which affects the overall quality of video platforms and user experience.
By extracting multidimensional features from the video to be identified and fusing them into a multimodal video feature vector, and using the first identification label generated by the level definition and the second identification label generated by the user playback behavior coefficient, the first and second level identification parameter sets are obtained, and the target content quality level of the video is determined by combining the two.
It improves the coverage and accuracy of low-end video recognition, reduces the cost of manual recognition and annotation, and enhances the overall quality of video platforms and user experience.
Smart Images

Figure CN113569610B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computers, and more specifically, to a video content recognition method and apparatus, storage medium, and electronic device. Background Technology
[0002] Nowadays, more and more users are uploading and publishing their own short videos or micro-videos to content distribution platforms through their personal accounts. The distribution and recommendation strategies provided by these technologies typically set recommendation criteria based on the data traffic of each client, such as setting an upper limit on playback traffic (i.e., rate limiting), to control the exposure of videos uploaded and published by each client.
[0003] However, the quality of videos posted by these users varies greatly. Some low-quality videos do not actually meet the recommendation criteria for sharing with other users through content distribution platforms. For example, videos with poor image clarity and low playback completion rates are not suitable for multiple distributions and recommendations.
[0004] Currently, content distribution platforms often rely on manual identification and labeling for these low-quality videos. This can easily lead to the omission of videos that do not meet the recommendation criteria, resulting in low accuracy in video content identification.
[0005] There is currently no effective solution to the above problems. Summary of the Invention
[0006] This invention provides a video content recognition method and apparatus, storage medium and electronic device, to at least solve the technical problem of low accuracy in video content recognition.
[0007] According to one aspect of the present invention, a video content recognition method is provided, comprising: fusing multidimensional features extracted from the video content of an object video to be recognized to obtain a multimodal video feature vector corresponding to the object video; obtaining a first-level recognition parameter set based on the multimodal video feature vector and a first weight set determined based on a first recognition label, wherein the first recognition label is a level recognition label generated according to a level definition, and the first-level recognition parameter set is used to indicate the probability that the object video is classified into various content quality levels according to the first recognition label; obtaining a second-level recognition parameter set based on the multimodal video feature vector and a second weight set determined based on a second recognition label, wherein the second recognition label is a level recognition label generated according to a user playback behavior coefficient, and the second-level recognition parameter set is used to indicate the probability that the object video is classified into various content quality levels according to the second recognition label; and determining a target content quality level matched by the object video based on the first-level recognition parameter set and the second-level recognition parameter set.
[0008] According to another aspect of the present invention, a video content recognition apparatus is also provided, comprising: a fusion unit, configured to fuse multidimensional features extracted from the video content of an object video to be recognized, to obtain a multimodal video feature vector corresponding to the object video; a first acquisition unit, configured to acquire a first level recognition parameter set based on the multimodal video feature vector and a first weight set determined based on a first recognition label, wherein the first recognition label is a level recognition label generated according to a level definition, and the first level recognition parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first recognition label; a second acquisition unit, configured to acquire a second level recognition parameter set based on the multimodal video feature vector and a second weight set determined based on a second recognition label, wherein the second recognition label is a level recognition label generated according to a user playback behavior coefficient, and the second level recognition parameter set is used to indicate the probability that the object video is classified into each content quality level according to the second recognition label; and a determination unit, configured to determine the target content quality level matched by the object video based on the first level recognition parameter set and the second level recognition parameter set.
[0009] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, wherein the computer program is configured to execute the video content recognition method described above when running.
[0010] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the video content recognition method described above through the computer program.
[0011] In this embodiment of the invention, multi-dimensional features extracted from the video content of the object video to be identified are fused to obtain a multimodal video feature vector corresponding to the object video; based on the multimodal video feature vector and a first weight set determined based on a first identification label, a first level identification parameter set is obtained, wherein the first identification label is a level identification label generated according to the level definition, and the first level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label; based on the multimodal video feature vector and a second weight set determined based on a second identification label, a second level identification parameter set is obtained, wherein the second identification label is a level identification label generated according to the user playback behavior coefficient, and the second level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label. The probability of classifying content quality levels according to the second identification label is determined; the target content quality level of the object video is determined based on the first and second level identification parameter sets. This is achieved by fusing multi-dimensional features extracted from the video content of the object video to obtain a multimodal video feature vector corresponding to the object video. The first and second level identification parameter sets are then obtained based on the multimodal video feature vector to determine the target content quality level of the object video. This approach aims to improve the low-end video recognition capability, thereby enhancing the coverage and accuracy of low-end recognition, reducing manual identification and annotation costs, improving the overall video quality of the platform, and enhancing the user's viewing experience of the platform's videos. Ultimately, this solves the technical problem of low accuracy in video content recognition. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0013] Figure 1 This is a schematic diagram of an application environment for an optional video content recognition method according to an embodiment of the present invention;
[0014] Figure 2 This is a schematic diagram of an application environment for another optional video content recognition method according to an embodiment of the present invention;
[0015] Figure 3 This is a flowchart of an optional video content recognition method according to an embodiment of the present invention;
[0016] Figure 4 This is a schematic diagram of a low-level video recognition architecture for an optional video content recognition method according to an embodiment of the present invention;
[0017] Figure 5 This is a schematic diagram of a low-end video content recognition model structure based on multi-dimensional video content recognition method according to an embodiment of the present invention.
[0018] Figure 6 This is a schematic diagram of a low-end recognition process based on multi-dimensional video content, according to an embodiment of the present invention;
[0019] Figure 7 This is a schematic diagram of a low-end identification model structure based on video distribution user behavior, which is an optional video content recognition method according to an embodiment of the present invention.
[0020] Figure 8 This is a flowchart of a user behavior-based low-end video content recognition method according to an embodiment of the present invention;
[0021] Figure 9 This is a schematic diagram of the structure of an optional video content recognition device according to an embodiment of the present invention;
[0022] Figure 10 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] According to one aspect of the present invention, a video content recognition method is provided. Optionally, as an optional implementation, the above-described video content recognition method may be applied to, but is not limited to, [examples of applications]. Figure 1 The application environment shown includes: a terminal device 102 for human-computer interaction, a network 104, and a server 106. User 108 can interact with terminal device 102, which runs a video content recognition application client. Terminal device 102 includes a human-computer interaction screen 1022, a processor 1024, and a memory 1026. The human-computer interaction screen 1022 presents the video content of the object video to be recognized. The processor 1024 fuses the multi-dimensional features extracted from the video content of the object video to obtain a multimodal video feature vector corresponding to the object video. The memory 1026 stores the video content of the object video to be recognized and the multimodal video feature vector corresponding to the object video.
[0026] Furthermore, server 106 includes database 1062 and processing engine 1064. Database 1062 stores multimodal video feature vectors corresponding to the object video, and stores a first-level recognition parameter set and a second-level recognition parameter set; it also stores the target content quality level for matching the object video. Processing engine 1064 is used to obtain the first-level recognition parameter set based on the multimodal video feature vectors and a first weight set determined based on the first recognition label; obtain the second-level recognition parameter set based on the multimodal video feature vectors and a second weight set determined based on the second recognition label; and determine the target content quality level for matching the object video based on the first-level and second-level recognition parameter sets.
[0027] The specific process is as follows: assuming... Figure 1The terminal device 102 shown runs a video content recognition application client. The user 108 operates the human-computer interaction screen 1022 to manage and operate the video content. For example, in step S102, the multi-dimensional features extracted from the video content of the object video to be recognized are fused to obtain the multi-modal video feature vector corresponding to the object video. Then, step S104 is executed to send the multi-modal video feature vector to the server 106 through the network 104. After receiving the request, server 106 executes steps S106-S108. Based on the multimodal video feature vector and the first weight set determined based on the first identification label, it obtains a first-level identification parameter set. The first identification label is a level identification label generated according to the level definition, and the first-level identification parameter set indicates the probability that the object video will be classified into each content quality level according to the first identification label. Based on the multimodal video feature vector and the second weight set determined based on the second identification label, it obtains a second-level identification parameter set. The second identification label is a level identification label generated according to the user playback behavior coefficient, and the second-level identification parameter set indicates the probability that the object video will be classified into each content quality level according to the second identification label. Based on the first-level identification parameter set and the second-level identification parameter set, it determines the target content quality level matching the object video. Then, as in step S112, it notifies terminal device 102 via network 104 to return the target content quality level matching the object video.
[0028] As another optional implementation, the video content recognition method described above in this application can be applied to... Figure 2 The application environment shown. For example... Figure 2 As shown, user 202 and user equipment 204 can interact. User equipment 204 includes a memory 206 and a processor 208. In this embodiment, user equipment 204 can, but is not limited to, referencing and executing the operations performed by terminal device 102 to obtain a target route matching the target route.
[0029] Optionally, the terminal device 102 and user device 204 may be, but are not limited to, mobile phones, tablets, laptops, PCs, etc., and the network 104 may be, but is not limited to, wireless networks or wired networks. The wireless network includes Wi-Fi and other networks that enable wireless communication. The wired network may include, but is not limited to, wide area networks (WANs), metropolitan area networks (MANs), and local area networks (LANs). The server 106 may include, but is not limited to, any hardware device capable of computation. The above is merely an example, and no limitations are imposed in this embodiment.
[0030] Low-quality videos negatively impact the overall content quality of video platforms. These platforms typically need to identify and mitigate the impact of low-quality videos. The technologies used are primarily based on text mining and image recognition, without adequately integrating multi-dimensional video content with user behavior in video recommendation and distribution.
[0031] To address the aforementioned technical problems, alternatively, as an optional implementation method, such as Figure 3 As shown, the above video content recognition method includes:
[0032] S302, fuse the multi-dimensional features extracted from the video content of the object video to be identified to obtain the multimodal video feature vector corresponding to the object video;
[0033] S304, based on the multimodal video feature vector and the first weight set determined based on the first identification label, obtain the first level identification parameter set, wherein the first identification label is a level identification label generated according to the level definition, and the first level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label;
[0034] S306, Based on the multimodal video feature vector and the second weight set determined based on the second identification label, obtain the second level identification parameter set, wherein the second identification label is a level identification label generated according to the user playback behavior coefficient, and the second level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the second identification label;
[0035] S308, determine the target content quality level for object video matching based on the first-level recognition parameter set and the second-level recognition parameter set.
[0036] In step S302, in practical applications, the video to be identified can include, but is not limited to, movies, TV series, and various long and short videos on any video platform. Here, the multi-dimensional features can include, but are not limited to, video text features, image features, or audio features; no limitation is made here. The multimodal video feature vector includes, but is not limited to, the following: text features include word vector sequences obtained by segmenting and vectorizing the text, and the resulting vector after encoding the word vector sequences; image features include features obtained by inputting each subject's keyframe into an image recognition model with time-series fusion capabilities; and audio features include features obtained by inputting each audio frame into an audio recognition model with time-series fusion capabilities.
[0037] In step S304, in practical applications, the first identification label may include, but is not limited to, preset different levels. The first level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label. For example, the first identification label may be divided into 5 levels, from 1 to 5. The probability of the low-end level corresponding to these 5 levels is [0.12, 0.52, 0.36, 0.08, 0.19]. That is, the probability of the low-end level corresponding to the level 1 identification label is 0.12, the probability of the low-end level corresponding to the level 2 identification label is 0.52, the probability of the low-end level corresponding to the level 3 identification label is 0.36, the probability of the low-end level corresponding to the level 4 identification label is 0.08, and the probability of the low-end level corresponding to the level 5 identification label is 0.19.
[0038] In step S306, in practical applications, the second identification label may include, but is not limited to, defining a level identification label generated according to the user's playback behavior coefficient by statistically analyzing the video playback rate (number of plays / number of exposures) and playback completion rate (total playback duration / duration viewed by the user) of the videos in the recommendation pool.
[0039] Here, c1*play rate + c2*play completion rate is defined as the video distribution behavior score, where c1 and c2 are weights, and c1 + c2 = 1. The video behavior score is divided into K low-end level intervals, such as [0, 0.2] for low-end K level, meaning the probability range from 0 to 0.2 corresponds to low-end K level; [0.8, 1.0] for low-end 1 level; and from 0.8 to 1 for low-end 1 level. The level identification label generated based on the user playback behavior coefficient, the second level identification parameter set contains the set of different probabilities from low-end 1 level to low-end K level.
[0040] In step S308, in practical applications, the target content quality level of object video matching can be achieved using methods including but not limited to the following: Video low-end level probability = x1 * low-end probability based on video multi-dimensional content recognition model + x2 * low-end probability based on user behavior recognition model, where x1 + x2 = 1, and the fused video low-end level probability is taken as the final low-end level of the video. Here, the low-end probability based on video multi-dimensional content recognition model can include but is not limited to the first-level recognition parameter set, and the low-end probability based on user behavior recognition model can include but is not limited to the second-level recognition parameter set.
[0041] In this embodiment of the invention, multi-dimensional features extracted from the video content of the object video to be identified are fused to obtain a multimodal video feature vector corresponding to the object video; based on the multimodal video feature vector and a first weight set determined based on a first identification label, a first level identification parameter set is obtained, wherein the first identification label is a level identification label generated according to the level definition, and the first level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label; based on the multimodal video feature vector and a second weight set determined based on a second identification label, a second level identification parameter set is obtained, wherein the second identification label is a level identification label generated according to the user playback behavior coefficient, and the second level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label. The probability of classifying content quality levels according to the second identification label is determined; the target content quality level of the object video is determined based on the first and second level identification parameter sets. This is achieved by fusing multi-dimensional features extracted from the video content of the object video to obtain a multimodal video feature vector corresponding to the object video. The first and second level identification parameter sets are then obtained based on the multimodal video feature vector to determine the target content quality level of the object video. This approach aims to improve the low-end video recognition capability, thereby enhancing the coverage and accuracy of low-end recognition, reducing manual identification and annotation costs, improving the overall video quality of the platform, and enhancing the user's viewing experience of the platform's videos. Ultimately, this solves the technical problem of low accuracy in video content recognition.
[0042] In one embodiment, step S302 includes: obtaining the weighted sum of the multimodal video feature vector and each weight value in the first weight set to obtain a first-level recognition parameter set; here, the text features, image features, or audio features corresponding to the multimodal video feature vector can be weighted according to their respective weights to obtain the weighted sum of each weight value as the first-level recognition parameter set; for example, the feature vector value corresponding to the text feature in the video content of video A is 2, and the weight is set to 0.3; the feature vector value corresponding to the image feature is set to 5, and the weight is set to 0.4; the feature vector value corresponding to the audio feature is set to 3, and the weight is set to 0.3. The first-level recognition parameter corresponding to video A is 2*0.3+5*0.4+3*0.3=3.5; the feature vector value corresponding to the text feature in the video content of the current object to be processed, video B, is 3, and the weight is set to 0.3; the feature vector value corresponding to the image feature is set to 4, and the weight is set to 0.4; the feature vector value corresponding to the audio feature is set to 2, and the weight is set to 0.3. The first-level recognition parameter corresponding to video B is 3*0.3+4*0.4+2*0.3=3.1; then the first-level recognition parameter set can be [3.5, 3.1]. In addition, this parameter set can be normalized to obtain the first-level recognition parameter set as [0.35, 0.31]. Here, the process of obtaining the first-level recognition parameter set is only an example and is not limited here.
[0043] Step S304 includes: obtaining the weighted sum of the multimodal video feature vector and each weight value in the second weight set to obtain the second-level recognition parameter set; for example, it may include, but is not limited to, calculating the video playback rate (number of plays / number of exposures) and playback completion rate (total playback duration / duration viewed by the user) of videos in the recommendation pool. Here, c1*playback rate + c2*play completion rate is defined as the second-level recognition parameter set, where c1 and c2 are weights, and c1 + c2 = 1.
[0044] For example, video A in the video recommendation pool has been played 3000 times, exposed 5000 times, has a total playback time of 6000 hours, and has been viewed by users for 8000 hours. Therefore, video A has a playback rate of 0.6 and a completion rate of 0.75. Here, when c1 is 0.4 and c2 is 0.6, the second-level recognition parameter for video A can be 0.4*0.6 + 0.75*0.6 = 0.67; video B in the video recommendation pool has a playback rate of 0.6. With 2000 plays, 4000 exposures, a total playback time of 3000 hours, and 4000 hours of user viewing time, video B has a playback rate of 0.5 and a completion rate of 0.75. Here, when c1 is 0.4 and c2 is 0.6, the second-level recognition parameter for video B can be 0.4*0.5 + 0.6*0.75 = 0.65; therefore, the second-level recognition parameter set can be [0.67, 0.65]. The process of obtaining the second-level recognition parameter set is merely an example and is not limited here.
[0045] In one embodiment, before fusing the multidimensional features extracted from the video content of the object video to be identified to obtain the multimodal video feature vector corresponding to the object video, the method further includes: obtaining a first sample video set; configuring a first identification label for each first sample video in the first sample video set according to a level definition; inputting the first sample video set and the corresponding first identification label into an initialized content level recognition model for training, and obtaining a training output result, wherein, in each training process of the content level recognition model, the first sample content quality level corresponding to the first sample video is determined based on the multidimensional features extracted from the video content of the first sample video; and when the training output result indicates that a first convergence condition has been met, a target content level recognition model for obtaining the first level recognition parameter set is determined, wherein the first convergence condition is used to indicate that the difference between the determined first sample content quality level and the content quality level indicated by the first identification label is less than or equal to a first threshold.
[0046] Here, the first sample video set can be video files of different types extracted from the video recommendation pool of the video platform. A first recognition label is configured for each first sample video in the first sample video set according to the level definition. The first sample video set and its corresponding first recognition label are input into the initialized content level recognition model for training to obtain the training output. Alternatively, videos from the video recommendation pool can be input into the initialized content level recognition model according to text features, image features, and audio features to obtain the training output.
[0047] The text features corresponding to the model training set can be obtained by jointly using video titles, subtitles, and dialogue text. Subtitles can be extracted using a general Optical Character Recognition (OCR) model, such as Google Tesseract, while dialogue text can be recognized using a general Automatic Speech Recognition (ASR) model. The ASR model can identify the text corresponding to the dialogue segments in the video. The video title, dialogue, and subtitles are concatenated to form the video text. Then, word segmentation is performed on the video text, and the word vector for each word is retrieved. This word vector sequence is then input into the ALBERT Encoder model, and the model's output serves as the text representation of the video. The first vector output by ALBERT can be used to represent the overall input text. Then, a word segmentation algorithm is constructed, the output word vectors are processed, and the distance or similarity between two texts is calculated using the ALBERT Encoder.
[0048] Here, the video image processing process, which may include but is not limited to, involves extracting keyframes related to the video's theme to represent the video. Keyframe extraction is a sequence labeling model, where each frame in the video is labeled with 0 or 1, with 1 indicating that the frame is a keyframe. A training dataset is constructed by manually labeling each frame with 0s and 1s. The model is then trained on this dataset to output a sequence of keyframes from a given video. Each keyframe is input into a pre-trained EfficientNet model, and the output of the last hidden layer before the classification layer (e.g., a 1024-dimensional floating-point vector) is used as the frame's representation. After obtaining the keyframe representation, each keyframe is sequentially input into a model layer with time-series fusion capabilities, such as the NetXVlad model, to construct the video image-side representation. The NetXVlad model algorithm can be divided into the following steps: 1. Extract SIFT descriptors from the image; 2. Train a codebook using the extracted SIFT descriptors (so the SIFT of the training image), the training method can be K-means; 3. Assign all SIFT descriptors of an image to the codebook according to the nearest neighbor principle (that is, assign them to K cluster centers); 4. Perform residual summation for each cluster center (that is, sum the SIFTs belonging to the current cluster center minus the cluster center); 5. Perform L2 normalization on this residual, and then concatenate them into a long vector of K*128, where 128 is the length of a single SIFT.
[0049] The audio feature representation of video is similar to that of images. First, the VGGish model is used to model the audio frames to obtain the audio frame representation. Then, NetXVlad is used to perform temporal fusion of multiple audio frame representations to obtain the audio-side representation of the video.
[0050] In each training process of the content quality recognition model, the content quality level of the first sample video is determined based on the multidimensional features extracted from the video content of the first sample video; for example, different content quality levels can be obtained and divided into levels 1-K.
[0051] If the training output indicates that the first convergence condition has been met, a target content level recognition model for obtaining the first level recognition parameter set is determined. The first convergence condition indicates that the difference between the determined content quality level of the first sample and the content quality level indicated by the first recognition label is less than or equal to a first threshold. In other words, when the difference between the content quality level obtained through model training and the content quality level indicated by the first label is within a preset range, the convergence condition is met, and the training process of the content level recognition model stops.
[0052] In one embodiment, before fusing the multidimensional features extracted from the video content of the object video to be identified to obtain the multimodal video feature vector corresponding to the object video, the method further includes: obtaining a second sample video set; configuring a second identification label for each second sample video in the second sample video set according to the user playback behavior coefficient; inputting the second sample video set and the corresponding second identification label into an initialized behavior level recognition model for training, and obtaining a training output result, wherein, in each training process of the behavior level recognition model, the second sample content quality level corresponding to the second sample video is determined based on the multidimensional features extracted from the video content of the second sample video and the user playback behavior coefficient corresponding to the second sample video; and when the training output result indicates that a second convergence condition has been met, a target behavior level recognition model for obtaining the second level recognition parameter set is determined, wherein the second convergence condition is used to indicate that the difference between the determined second sample content quality level and the content quality level indicated by the second identification label is less than or equal to a second threshold.
[0053] Here, for example, training a low-end video recognition model based on behavior low-end levels can be done by statistically analyzing the video playback rate (number of plays / number of exposures) and playback completion rate (total playback duration / duration viewed by users) of videos in the recommendation pool. The video distribution behavior score is defined as c1*playback rate + c2*play completion rate, where c1 and c2 are weights, and c1 + c2 = 1. The video behavior score is then divided into K low-end behavior level intervals, such as [0, 0.2] for low-end level K and [0.8, 1.0] for low-end level 1. Then, based on the videos in the recommendation pool and their corresponding low-end behavior levels, a low-end video recognition model is trained. The model's input features are the multi-dimensional features of the video, and the classification output target is the low-end distribution behavior level of the video. The multi-dimensional features can include the video's text, image frames, audio frames, and other multi-dimensional features. After the above model is applied, the video's low-end level predicted based on user behavior and the corresponding level probability are output.
[0054] If the training output indicates that the second convergence condition has been met, a target behavior level recognition model for obtaining the second level recognition parameter set is determined. The second convergence condition indicates that the difference between the determined second sample content quality level and the content quality level indicated by the second recognition label is less than or equal to a second threshold. In other words, when the difference between the content quality level obtained through model training and the content quality level indicated by the second label is within a preset range, the convergence condition is met, and the training process of the content level recognition model stops.
[0055] In one embodiment, configuring the second identification tag for each second sample video in the second sample video set according to the user playback behavior coefficient includes: taking each second sample video in the second sample video set as the current sample video in turn, and performing the following operations: calculating the playback rate and playback completion rate of the current sample video, wherein the playback rate is used to indicate the ratio between the number of times the current sample video is actually played on the playback client and the number of exposures, and the playback completion rate is used to indicate the ratio between the duration of the current sample video being actually played on the playback client and the total playback duration of the current sample video; determining the current user playback behavior coefficient matched by the current sample video based on the playback rate and the playback completion rate; determining the current content quality level corresponding to the current user playback behavior coefficient according to the level interval divided for the user playback behavior coefficient; and configuring the second identification tag corresponding to the current content quality level for the current sample video.
[0056] For example, in the current sample video, video A has been played 3000 times, exposed 5000 times, has a total playback duration of 6000 hours, and has been viewed by users for 8000 hours. Therefore, video A has a playback rate of 0.6 and a completion rate of 0.75. Here, when c1 is 0.4 and c2 is 0.6, the second-level recognition parameter for video A can be 0.4*0.6 + 0.75*0.6 = 0.67; the playback rate of video B in the video recommendation pool... If video B is played 2000 times, exposed 4000 times, has a total playback time of 3000 hours, and has been viewed by users for 4000 hours, then its playback rate is 0.5 and its completion rate is 0.75. Here, when c1 is 0.4 and c2 is 0.6, the second-level recognition parameter for video B can be 0.4*0.5 + 0.6*0.75 = 0.65; therefore, the second-level recognition parameter set can be [0.67, 0.65]. 0.65 corresponds to level 1, and 0.65 corresponds to level 2. Therefore, the level recognition label generated for video A based on the user playback behavior coefficient is level 1, and the level recognition label generated for video A based on the user playback behavior coefficient is level 2. The process of obtaining the second-level recognition label is only an example and is not limited here.
[0057] In this embodiment, step S308 includes: traversing each content quality level, taking each content quality level as the current content quality level, and sequentially performing the following operations: obtaining the first-level identification parameter corresponding to the current content quality level from the first-level identification parameter set, and obtaining the second-level identification parameter corresponding to the current content quality level from the second-level identification parameter set; performing a weighted summation of the first-level identification parameter and the second-level identification parameter to obtain the current-level identification parameter corresponding to the current content quality level; and, given that the level identification parameters corresponding to each content quality level are obtained, determining the largest level identification parameter value, and determining the content quality level corresponding to the largest level identification parameter value as the target content quality level. Here, a video-based multi-dimensional content recognition model and a user behavior-based low-end identification model can be used in combination. The probability of a video's low-end level is calculated as x1 * the low-end probability of the video-based multi-dimensional content recognition model + x2 * the low-end probability of the user behavior-based low-end identification model, where x1 + x2 = 1. The fused video low-end level probability is taken as the target content quality level. When the low-end level exceeds a certain threshold, the video may not be distributed.
[0058] In one embodiment, step S302 includes: extracting at least one of the following features from the video content of the object video: text features, image features, and audio features; concatenating the text information contained in the object video to obtain the object text to be processed corresponding to the object video; performing word segmentation and vector transformation processing on the object text to obtain a word vector sequence; encoding the word vector sequence to obtain the text features; inputting the keyframes of each theme contained in the object video into an image recognition model with time series fusion capability to obtain the image features; and inputting the audio frames contained in the object video into an audio recognition model with time series fusion capability to obtain the audio features.
[0059] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0060] Based on the above embodiments, in one application embodiment, the video content recognition method further includes:
[0061] A low-level video recognition method based on video multimodal content and user behavior, with the following overall architecture: Figure 4 As shown: Step S402, acquire the video to be identified; Step S404, perform low-end identification based on video content from the video to be identified; Step S406, perform low-end prediction of user behavior; Step S408, put the processed video into the video recommendation pool; Step S410, distribute exposure, clicks, etc. through video distribution; Step S412, perform model learning; Step S414, obtain the low-end identification model based on user behavior; Step S416, after labeling the low-end videos generated by the low-end prediction of user behavior, generate a low-end video library; then proceed to Step S418, perform model learning; Step S420, obtain the low-end identification model based on video content.
[0062] Based on the above embodiments, in one application embodiment, the process of low-level recognition based on multi-dimensional video content includes the following:
[0063] The video platform predefines K low-end levels, where K is the number of low-end categories, with categories 1 to K representing progressively higher low-end levels. A large number of videos in the video library are labeled with these low-end level categories to construct a low-end dataset. A low-end recognition model based on multimodal video content is then trained on this dataset. The model structure is as follows: Figure 5 As shown, after a user uploads a video, the video platform uses a multimodal deep learning model to identify low-quality video content. The low-quality video content identification model based on multimodal video content is as follows:
[0064] The video text representation is achieved by jointly using the video title, subtitles, and dialogue text. Subtitles can be extracted using a general OCR model, such as Google Tesseract, while dialogue text can be recognized using a general ASR recognition model. The video text is constructed by concatenating the title, dialogue, and subtitles. This text is then segmented into words, and the word vector for each word is retrieved. The word vector sequence is then input into the ALBERT Encoder model, and the model's output serves as the video's text representation.
[0065] On the video image side, keyframes related to the video's theme are extracted and used to represent the video. Keyframe extraction is a sequence labeling model, where each frame in the video is labeled with 0s and 1s, with 1 indicating a keyframe. A training dataset is built by manually labeling each frame with 0s and 1s. The model is then trained on this dataset to output a sequence of keyframes from a given video. Each keyframe is input into a pre-trained EfficientNet model, and the output of the last hidden layer before the classification layer (e.g., a 1024-dimensional floating-point vector) is used as the frame's representation. After obtaining the keyframe representation, each keyframe is sequentially input into a model layer with time-series fusion capabilities, such as NetXVlad, to construct the video image-side representation. The video audio-side representation is similar to the image method. First, the VGGish model is used to model the audio frames, obtaining their representations. Then, NetXVlad performs time-series fusion of multiple audio frame representations to obtain the video audio-side representation. A low-level classification and recognition model is constructed by fusing multi-dimensional video features. The video text, image, and audio features constructed above are concatenated and then passed through a fully connected network for multi-dimensional feature fusion representation. A low-level classification output layer is then constructed based on the multi-dimensional fusion representation of the video to classify the low-level video. By training the above model on a pre-labeled content-based low-level training set, the model is able to output the low-level video.
[0066] When performing low-level identification on user-uploaded videos, such as Figure 6As shown, the process includes the following steps: Step S602, acquiring the video to be identified; Step S604, performing low-end identification based on the content of the video to be identified, proceeding to Step S606, labeling and confirming the identified content to obtain a low-end video library; then proceeding to Step S608, performing model learning; Step S610, obtaining a low-end identification model based on video content. In this embodiment, text, image frame, and audio frame features are extracted based on the above scheme, and then the model outputs the low-end level of the video and the corresponding level probability. For example, if the number of low-end levels K is 5, the probability of the model outputting each low-end level from 1 to K for the current video is [0.05432093, 0.53563935, 0.18928528, 0.13303354, 0.0877209]. If judged only from the low-end level identification model based on video content, the low-end level of this video is level 2, with a corresponding probability of 0.53563935.
[0067] Based on the above embodiments, in one application embodiment, the video content recognition method further includes: the low-level recognition process based on video distribution user behavior may include the following:
[0068] If classification is based solely on video content, the boundaries between video content may be blurry, making it difficult for models and humans to accurately identify low-end levels. If the initial judgment of a video's low-end level based on content meets the recommendation and distribution threshold, and the video has a high play rate and completion rate after being recommended and distributed multiple times, it indicates that the video is indeed not a low-end video. Conversely, the video may have potential low-end risk.
[0069] By statistically analyzing the play rate (number of plays / number of exposures) and completion rate (total playback time / time viewed by users) of videos in the recommendation pool, the distribution behavior score of a video is defined as c1*play rate + c2*completion rate, where c1 and c2 are weights, and c1 + c2 = 1. The behavior score of a video is divided into K low-end behavior level intervals, such as [0, 0.2] for low-end level K and [0.8, 1.0] for low-end level 1. Then, a low-end video recognition model is trained based on the videos in the recommendation pool and their corresponding low-end behavior levels. The input features of the model are the multi-dimensional features of the video, and the classification output target is the low-end distribution behavior level of the video. The input corpus to the model consists of the video and its corresponding score level; the output target of the model is the probability that the output video is classified as a low-end video.
[0070] In this embodiment, as Figure 7As shown, the probability of low-end video classification can be obtained through the following steps: Step S702, input video data; Step S704, extract multi-dimensional features of the video; Step S706, divide the video into K low-end behavior levels through a fully connected layer; Step S708, obtain the video multimodal vector; Step S710, then obtain the low-end classification of the distribution user behavior video.
[0071] When performing low-level identification on user-uploaded videos, multi-dimensional features such as text, image frames, and audio frames can be extracted. Then, the model outputs the predicted low-level video based on user behavior and the corresponding probability. For example, if the number of low-level levels K is 5, the model outputs the probability of each low-level level from 1 to K for the current video as [0.06821877999999992, 0.08823278, 0.09661109, 0.19159416, 0.55534319]. If the low-level identification model is based solely on user behavior, the low-level of this video is 5, with a corresponding probability of 0.55534319. In one embodiment, when performing low-level identification on user-uploaded videos, such as... Figure 8 As shown, the process can be completed through the following steps: Step S802, obtain videos from the video recommendation pool; Step S804, distribute, expose, and click on the videos; then proceed to Step S806, perform model learning; Step S808, obtain a low-end recognition model based on video content; Step S810, obtain the video to be recognized; Step S812, perform low-end prediction of user behavior in the video to be recognized based on the low-end recognition model of video content; Step S814, after labeling and confirming the predicted videos, obtain a low-end video library.
[0072] Based on the above embodiments, in one application embodiment, the video content recognition method further includes:
[0073] The low-end rating identification of videos, which integrates multi-dimensional video content and user behavior, includes the following: To improve the coverage and accuracy of low-end rating identification, a combination of a video multi-dimensional content recognition model and a user behavior-based low-end rating identification model can be used. The low-end rating probability of a video is calculated as follows: x1 * low-end rating probability based on the video multi-dimensional content recognition model + x2 * low-end rating probability based on the user behavior recognition model, where x1 + x2 = 1. The fused low-end rating probability is taken as the final low-end rating of the video. When the low-end rating exceeds a certain level, the video may not be distributed.
[0074] Based on multi-dimensional content recognition of videos, this invention constructs a low-level recognition model by combining video distribution information. Through user behavior prediction and analysis, it improves the low-level video recognition capability, reduces the cost of manual recognition and annotation, and further enhances the overall video quality of the video platform.
[0075] According to another aspect of the present invention, a video content recognition apparatus for implementing the above-described video content recognition method is also provided. For example... Figure 9 As shown, the device includes:
[0076] The fusion unit 902 is used to fuse the multi-dimensional features extracted from the video content of the object video to be identified, so as to obtain the multimodal video feature vector corresponding to the object video.
[0077] The first acquisition unit 904 is used to acquire a first level recognition parameter set based on the multimodal video feature vector and a first weight set determined based on the first recognition label, wherein the first recognition label is a level recognition label generated according to the level definition, and the first level recognition parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first recognition label.
[0078] The second acquisition unit 906 is used to acquire a second level identification parameter set based on the multimodal video feature vector and the second weight set determined based on the second identification label, wherein the second identification label is a level identification label generated according to the user playback behavior coefficient, and the second level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the second identification label.
[0079] The determining unit 908 is used to determine the target content quality level of the object video matching based on the first level identification parameter set and the second level identification parameter set.
[0080] In this embodiment of the invention, the video to be identified may include, but is not limited to, movies, TV series, and various long and short videos on any video platform. The multi-dimensional features may include, but are not limited to, text features, image features, or audio features of the video; no limitation is made here. The multimodal video feature vector includes, but is not limited to, the following: text features include word vector sequences obtained by segmenting and vectorizing the text, and the resulting vector after encoding the word vector sequences; image features include features obtained by inputting each subject's keyframe into an image recognition model with time-series fusion capabilities; and audio features include features obtained by inputting each audio frame into an audio recognition model with time-series fusion capabilities.
[0081] In this embodiment of the invention, the first identification label may include, but is not limited to, preset different levels. The first level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label. For example, the first identification label may be divided into 5 levels, from 1 to 5. The probability of the low-end level corresponding to the 5 levels is [0.12, 0.52, 0.36, 0.08, 0.19]. That is, the probability of the low-end level corresponding to the level 1 identification label is 0.12, the probability of the low-end level corresponding to the level 2 identification label is 0.52, the probability of the low-end level corresponding to the level 3 identification label is 0.36, the probability of the low-end level corresponding to the level 4 identification label is 0.08, and the probability of the low-end level corresponding to the level 5 identification label is 0.19.
[0082] In this embodiment of the invention, the second identification label may include, but is not limited to, defining a level identification label generated according to the user's playback behavior coefficient by statistically analyzing the video playback rate (number of plays / number of exposures) and playback completion rate (total playback duration / duration viewed by the user) of videos in the recommendation pool. Here, c1*playback rate + c2*play completion rate is defined as the video distribution behavior score, where c1 and c2 are weights, c1+c2=1, and the video behavior score is divided into K low-end behavior level intervals, such as [0, 0.2] being the low-end K level, that is, from 0 to 0.2 is the probability range corresponding to the low-end K level, [0.8, 1.0] is the low-end 1 level, and from 0.8 to 1 is the probability range corresponding to the low-end 1 level. The level identification label generated according to the user's playback behavior coefficient includes a set of different probabilities from the low-end 1 level to the low-end K level.
[0083] In this embodiment of the invention, the target content quality level of object video matching may be achieved using methods including but not limited to the following: Video low-end level probability = x1 * low-end probability based on video multi-dimensional content recognition model + x2 * low-end probability based on user behavior recognition model, where x1 + x2 = 1, and the fused video low-end level probability is taken as the final low-end level of the video. Here, the low-end probability based on video multi-dimensional content recognition model may include but is not limited to a first-level recognition parameter set, and the low-end probability based on user behavior recognition model may include but is not limited to a second-level recognition parameter set.
[0084] In this embodiment of the invention, multi-dimensional features extracted from the video content of the object video to be identified are fused to obtain a multimodal video feature vector corresponding to the object video; based on the multimodal video feature vector and a first weight set determined based on a first identification label, a first level identification parameter set is obtained, wherein the first identification label is a level identification label generated according to the level definition, and the first level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label; based on the multimodal video feature vector and a second weight set determined based on a second identification label, a second level identification parameter set is obtained, wherein the second identification label is a level identification label generated according to the user playback behavior coefficient, and the second level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label. The probability of classifying content quality levels according to the second identification label is determined; the target content quality level of the object video is determined based on the first and second level identification parameter sets. This is achieved by fusing multi-dimensional features extracted from the video content of the object video to obtain a multimodal video feature vector corresponding to the object video. The first and second level identification parameter sets are then obtained based on the multimodal video feature vector to determine the target content quality level of the object video. This approach aims to improve the low-end video recognition capability, thereby enhancing the coverage and accuracy of low-end recognition, reducing manual identification and annotation costs, improving the overall video quality of the platform, and enhancing the user's viewing experience of the platform's videos. Ultimately, this solves the technical problem of low accuracy in video content recognition.
[0085] According to another aspect of the present invention, an electronic device for implementing the above-described video content recognition method is also provided. This electronic device may be... Figure 1 The terminal device or server shown is illustrated in this embodiment. This example uses this electronic device for illustration. Figure 10 As shown, the electronic device includes a memory 1002 and a processor 1004. The memory 1002 stores a computer program, and the processor 1004 is configured to execute the steps of any of the above method embodiments via the computer program.
[0086] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0087] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0088] S1, fuse the multi-dimensional features extracted from the video content of the object video to be identified to obtain the multimodal video feature vector corresponding to the object video;
[0089] S2, based on the multimodal video feature vector and the first weight set determined based on the first identification label, obtain the first level identification parameter set, wherein the first identification label is a level identification label generated according to the level definition, and the first level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label;
[0090] S3. Based on the multimodal video feature vector and the second weight set determined based on the second identification label, obtain the second level identification parameter set. The second identification label is a level identification label generated according to the user playback behavior coefficient. The second level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the second identification label.
[0091] S4. Determine the target content quality level for object video matching based on the first-level recognition parameter set and the second-level recognition parameter set.
[0092] Alternatively, as those skilled in the art will understand, Figure 10 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 10 This does not limit the structure of the aforementioned electronic devices or electronic equipment. For example, electronic devices or electronic equipment may also include components that are more... Figure 10 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 10 The different configurations shown.
[0093] The memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the video content recognition method and apparatus in this embodiment of the invention. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, thereby realizing the aforementioned video content recognition method. The memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1002 may further include memory remotely located relative to the processor 1004, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1002 may be used, but is not limited to, to store information such as multimodal video feature vectors corresponding to the object video. As an example, such as Figure 10As shown, the memory 1002 may include, but is not limited to, the fusion unit 902, the first acquisition unit 904, the second acquisition unit 906, and the determination unit 908 in the video content recognition device. Furthermore, it may include, but is not limited to, other module units in the video content recognition device, which will not be elaborated upon in this example.
[0094] Optionally, the transmission device 1006 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1006 includes a Network Interface Controller (NIC), which can be connected to other network devices and routers via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1006 is a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0095] In addition, the aforementioned electronic device also includes: a display 1008 for displaying the aforementioned multimodal video feature vector information; and a connection bus 1010 for connecting the various module components in the aforementioned electronic device.
[0096] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.
[0097] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned method for displaying a simulated ground surface. The computer program is configured to execute the steps of any of the above method embodiments during runtime.
[0098] Optionally, in this embodiment, the computer-readable storage medium described above may be configured to store a computer program for performing the following steps:
[0099] S1, fuse the multi-dimensional features extracted from the video content of the object video to be identified to obtain the multimodal video feature vector corresponding to the object video;
[0100] S2, based on the multimodal video feature vector and the first weight set determined based on the first identification label, obtain the first level identification parameter set, wherein the first identification label is a level identification label generated according to the level definition, and the first level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the first identification label;
[0101] S3. Based on the multimodal video feature vector and the second weight set determined based on the second identification label, obtain the second level identification parameter set. The second identification label is a level identification label generated according to the user playback behavior coefficient. The second level identification parameter set is used to indicate the probability that the object video is classified into each content quality level according to the second identification label.
[0102] S4. Determine the target content quality level for object video matching based on the first-level recognition parameter set and the second-level recognition parameter set.
[0103] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0104] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0105] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0106] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0107] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.
[0108] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0109] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0110] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A video content recognition method, characterized in that, include: Based on the user playback behavior coefficient that matches the object video in the recommendation pool, and the level range divided for the user playback behavior coefficient, the current content quality level of the object video is determined. The user playback behavior coefficient is determined based on the playback index value of the object video, according to the current content quality level and the corresponding second identification tag is configured for the object video. The multidimensional features extracted from the video content of the object video are fused to obtain a multimodal video feature vector; A first weight set is determined based on the first identification label. A first level identification parameter set is determined based on the weighted sum of the first feature vector value corresponding to each feature vector in the multimodal video feature vector and the first weight value corresponding to each feature vector in the first weight set. The first identification label is a level identification label generated according to the level definition. Based on the second identification label, a second weight set is determined. According to the playback index values of multiple playback indicators that match the object video, and the weighted sum of the second weight values corresponding to the multiple playback indicators in the second weight set, a second level identification parameter set for indicating the second probability of the object video being classified into each content quality level is determined. From the first level identification parameter set and the second level identification parameter set, respectively obtain the first probability and the second probability that match each of the content quality levels; Calculate the weighted sum of the first probability and the second probability under each of the content quality levels, and determine the content quality level corresponding to the largest weighted sum as the target content quality level of the object video.
2. The method according to claim 1, characterized in that, Obtaining the first level recognition parameter set based on the multimodal video feature vector and the first weight set determined based on the first recognition label includes: obtaining the weighted summation result between the multimodal video feature vector and each weight value in the first weight set to obtain the first level recognition parameter set; Obtaining the second-level recognition parameter set based on the multimodal video feature vector and the second weight set determined based on the second recognition label includes: obtaining the weighted summation result between the multimodal video feature vector and each weight value in the second weight set to obtain the second-level recognition parameter set.
3. The method according to claim 2, characterized in that, Before fusing the multidimensional features extracted from the video content of the object video to obtain a multimodal video feature vector, the method further includes: Obtain the first sample video set; Configure the first identification label for each first sample video in the first sample video set according to the level definition; The first sample video set and the corresponding first identification label are input into the initialized content level recognition model for training to obtain the training output result. In each training process of the content level recognition model, the first sample content quality level corresponding to the first sample video is determined based on the multi-dimensional features extracted from the video content of the first sample video. When the training output indicates that the first convergence condition has been met, a target content level recognition model for obtaining the first level recognition parameter set is determined, wherein the first convergence condition is used to indicate that the difference between the determined first sample content quality level and the content quality level indicated by the first recognition label is less than or equal to a first threshold.
4. The method according to claim 2, characterized in that, Before fusing the multidimensional features extracted from the video content of the object video to obtain a multimodal video feature vector, the method further includes: Obtain the second set of sample videos; Configure the second identification label for each second sample video in the second sample video set according to the user playback behavior coefficient; The second sample video set and the corresponding second identification label are input into the initialized behavior level recognition model for training to obtain the training output result. In each training process of the behavior level recognition model, the second sample content quality level corresponding to the second sample video is determined based on the multi-dimensional features extracted from the video content of the second sample video and the user playback behavior coefficient corresponding to the second sample video. When the training output indicates that the second convergence condition has been met, a target behavior level recognition model for obtaining the second level recognition parameter set is determined, wherein the second convergence condition is used to indicate that the difference between the determined second sample content quality level and the content quality level indicated by the second recognition label is less than or equal to a second threshold.
5. The method according to claim 4, characterized in that, The step of configuring the second identification label for each second sample video in the second sample video set according to the user playback behavior coefficient includes: Take each second sample video in the second sample video set as the current sample video and perform the following operations: The playback rate and playback completion rate of the current sample video are calculated. The playback rate is used to indicate the ratio between the number of times the current sample video is actually played on the playback client and the number of times it is exposed. The playback completion rate is used to indicate the ratio between the actual playback duration of the current sample video on the playback client and the total playback duration of the current sample video. The current user playback behavior coefficient matching the current sample video is determined based on the playback rate and the playback completion rate. The current content quality level corresponding to the current user playback behavior coefficient is determined according to the level range divided by the user playback behavior coefficient; Configure the second identification tag corresponding to the current content quality level for the current sample video.
6. The method according to any one of claims 1 to 5, characterized in that, Determining the target content quality level for the object video matching based on the first level recognition parameter set and the second level recognition parameter set includes: Iterate through each content quality level, taking each content quality level as the current content quality level, and perform the following operations in sequence: obtain the first level identification parameter corresponding to the current content quality level from the first level identification parameter set, and obtain the second level identification parameter corresponding to the current content quality level from the second level identification parameter set; perform a weighted summation of the first level identification parameter and the second level identification parameter to obtain the current level identification parameter corresponding to the current content quality level; Having obtained the level identification parameters corresponding to each of the content quality levels, the largest level identification parameter value is determined, and the content quality level corresponding to the largest level identification parameter value is determined as the target content quality level.
7. The method according to any one of claims 1 to 5, characterized in that, The process of fusing the multidimensional features extracted from the video content of the object video to obtain a multimodal video feature vector includes: Extract at least one of the following features from the video content of the object video: text features, image features, and audio features; Upon identifying the various text information contained in the object video, the various text information are concatenated to obtain the object text to be processed corresponding to the object video; the object text is then processed by word segmentation and vector transformation to obtain a word vector sequence; the word vector sequence is then encoded to obtain the text features; Once the keyframes of each theme contained in the object video are identified, the keyframes of each theme are input into an image recognition model with time series fusion capability to obtain the image features. If each audio frame contained in the object video is identified, the audio frames are input into an audio recognition model with time series fusion capability to obtain the audio features.
8. A video content recognition device, characterized in that, include: The fusion unit is used to determine the current content quality level of the object video based on the user playback behavior coefficient that matches the object video in the recommendation pool and the level range divided for the user playback behavior coefficient; and to configure a corresponding second identification tag for the object video according to the current content quality level, wherein the user playback behavior coefficient is determined based on the playback index value of the object video; The multidimensional features extracted from the video content of the object video are fused to obtain a multimodal video feature vector; The first acquisition unit is configured to determine a first weight set based on a first identification label, and to determine a first level identification parameter set for indicating the first probability that the object video is classified into each content quality level, based on the first feature vector value corresponding to each feature vector in the multimodal video feature vector and the weighted sum of the first weight values corresponding to each feature vector in the first weight set. The first identification label is a level identification label generated according to the level definition. The second acquisition unit is used to determine a second weight set based on the second identification label, and to determine a second level identification parameter set for indicating the second probability that the object video is classified into each content quality level by weighted summation of the playback index values of each of the multiple playback indicators that match the object video and the second weight values corresponding to the multiple playback indicators in the second weight set. The determining unit is configured to obtain, from the first level identification parameter set and the second level identification parameter set, the first probability and the second probability that match each of the content quality levels respectively; calculate the weighted sum of the first probability and the second probability under each of the content quality levels; and determine the content quality level corresponding to the largest weighted sum as the target content quality level of the object video.
9. The apparatus according to claim 8, characterized in that, The first acquisition unit includes: a first acquisition module, used to acquire the weighted summation result between the multimodal video feature vector and each weight value in the first weight set, to obtain the first level recognition parameter set; The second acquisition unit includes a second acquisition module, used to acquire the weighted summation result between the multimodal video feature vector and each weight value in the second weight set, to obtain the second level recognition parameter set.
10. The apparatus according to claim 9, characterized in that, The device further includes: The third acquisition unit is used to acquire the first sample video set; The first configuration unit is used to configure the first identification label for each first sample video in the first sample video set according to the level definition; The first training unit is used to input the first sample video set and the corresponding first recognition label into the initialized content level recognition model for training, and obtain training output results. In each training process of the content level recognition model, the first sample content quality level corresponding to the first sample video is determined based on the multi-dimensional features extracted from the video content of the first sample video. The first determining unit is configured to determine a target content level recognition model for obtaining the first level recognition parameter set when the training output result indicates that a first convergence condition has been met, wherein the first convergence condition indicates that the difference between the determined first sample content quality level and the content quality level indicated by the first recognition label is less than or equal to a first threshold.
11. The apparatus according to claim 9, characterized in that, Before fusing the multidimensional features extracted from the video content of the object video to obtain a multimodal video feature vector, the method further includes: The fourth acquisition unit is used to acquire the second sample video set; The second configuration unit is used to configure the second identification label for each second sample video in the second sample video set according to the user playback behavior coefficient; The second training unit is used to train the behavior level recognition model by inputting the second sample video set and the corresponding second recognition label into the initialization, and to obtain the training output result. In each training process of the behavior level recognition model, the second sample content quality level corresponding to the second sample video is determined based on the multi-dimensional features extracted from the video content of the second sample video and the user playback behavior coefficient corresponding to the second sample video. The second determining unit, when the training output result indicates that the second convergence condition has been met, determines the target behavior level recognition model for obtaining the second level recognition parameter set, wherein the second convergence condition is used to indicate that the difference between the determined second sample content quality level and the content quality level indicated by the second recognition label is less than or equal to a second threshold.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method described in any one of claims 1 to 7.
13. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 7 through the computer program.
Citation Information
Patent Citations
Video content rating method and device
CN104486649A
A low-quality video identification method and device
CN109684513A
Video category identification method and related device
CN110826545A