A video processing method, apparatus, device, and medium
By extracting keyframes from videos and using a multi-dimensional classification model for processing, the problem of incomplete video tags in the existing technology is solved, the comprehensive and accurate generation of video tags is achieved, and the accuracy of video recommendations is improved.
Patent Information
- Application Number
- CN202010658845.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-07-09
AI Technical Summary
In the prior art, video tag generation is not comprehensive enough, the accuracy is not high, it is difficult to comprehensively summarize the video content, and the single-dimensional classification method is difficult to weigh the subject and background.
Keyframes are extracted from the target video to form a frame sequence, and a multi-dimensional classification model is used to classify the frame sequence, and video tags are generated through repeated semantic filtering, including analysis of object, scene and content dimensions.
It improves the comprehensiveness and accuracy of video tags, so that video tags can reflect video content and scene information more comprehensively, and improves the accuracy of video recommendations.
Smart Images

Figure CN111783712B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly relates to a video processing method, a video processing device, a video processing equipment and a computer-readable storage medium. Background Art
[0002] With the progress of computer technology, the number of videos collected in video platforms is increasing. At present, video platforms usually adopt an information flow interaction mode to recommend videos to users. This interaction mode relies on video tags to achieve, which requires pre-processing of videos to generate video tags. It is found in practice that the video tags generated by video processing in the prior art are difficult to comprehensively summarize the content of the video and have low accuracy. Summary of the Invention
[0003] Embodiments of the present invention provide a video processing method, device, equipment and computer-readable storage medium, which can generate comprehensive and accurate video tags for a target video.
[0004] On the one hand, embodiments of the present application provide a video processing method, which includes:
[0005] Obtain a target video to be processed;
[0006] Extract a frame sequence from the target video, where the frame sequence includes key frames of the target video;
[0007] Call a multi-dimensional classification model to classify the frame sequence to obtain a candidate tag set of the target video, where the candidate tag set contains classification tags of the target video in at least two dimensions;
[0008] Perform repeated semantic screening on the candidate tag set to obtain a video tag set of the target video.
[0009] On the one hand, the present application provides a video processing device, which includes:
[0010] An obtaining unit, configured to obtain a target video to be processed;
[0011] A processing unit, configured to extract a frame sequence from the target video, where the frame sequence includes key frames of the target video; call a multi-dimensional classification model to classify the frame sequence to obtain a candidate tag set of the target video, where the candidate tag set contains classification tags of the target video in at least two dimensions; perform repeated semantic screening on the candidate tag set to obtain a video tag set of the target video.
[0012] In one embodiment, the number of dimensions is denoted as P, and the multi-dimensional classification model includes P classification sub-models; the i-th classification sub-model is used to classify the frame sequence in the i-th dimension; P is an integer greater than 1, and i is an integer greater than 1 and i ≤ P.
[0013] In one embodiment, the processing unit is further configured to extract a frame sequence from the target video, specifically:
[0014] Determine the frame extraction frequency according to the frame density required by the P classification sub-models;
[0015] Perform frame extraction processing on the target video according to the frame extraction frequency to obtain a frame sequence.
[0016] In one embodiment, the processing unit is further configured to determine the frame extraction frequency according to the frame density required by the P classification sub-models, specifically:
[0017] Obtain the frame density required by each classification sub-model among the P classification sub-models;
[0018] Select the maximum frame density from the P frame densities and determine it as the frame extraction frequency.
[0019] In one embodiment, the processing unit is further configured to call the multi-dimensional classification model to classify the frame sequence to obtain a candidate label set of the target video, specifically:
[0020] Call the P classification sub-models respectively to classify the frame sequence to obtain the classification labels of the target video in the P dimensions;
[0021] Add the classification labels of the target video in the P dimensions to the candidate label set of the target video.
[0022] In one embodiment, before calling the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension, the processing unit is further configured to:
[0023] Detect whether the frame density required by the i-th classification sub-model matches the frame extraction frequency of the frame sequence;
[0024] If the frame density required by the i-th classification sub-model matches the frame extraction frequency of the frame sequence, then perform the step of calling the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension;
[0025] If the frame density required by the i-th classification sub-model does not match the frame extraction frequency of the frame sequence, then perform frame extraction processing on the frame sequence according to the frame density required by the i-th classification sub-model, and call the i-th classification sub-model to classify the frame sequence after frame extraction processing to obtain the classification label of the target video in the i-th dimension.
[0026] In one embodiment, the processing unit is further configured to perform repeated semantic screening on the candidate label set to obtain the video label set of the target video, specifically:
[0027] Perform repeated semantic mapping on each classification label in the candidate label set to obtain a standard category label set, which includes multiple standard categories and multiple classification labels under each standard category;
[0028] Count the number N of classification labels belonging to the target standard category, and count the number M of times the P classification sub-models perform classification processing on the frame sequence; the target standard category is any one of the standard categories in the standard category label set, and N and M are positive integers;
[0029] If the ratio between N and M is greater than or equal to the threshold, add the target standard category to the video label set of the target video.
[0030] In one embodiment, the i-th dimension is the object dimension, and the i-th classification sub-model includes an identification network; the processing unit is further configured to call the i-th classification sub-model to perform classification processing on the frame sequence to obtain the classification label of the target video in the i-th dimension, specifically:
[0031] Call the identification network of the i-th classification sub-model to identify the frame sequence to obtain the features of the objects included in each video frame at at least two granularities;
[0032] Determine the classification label of the target video in the object dimension according to the features of the objects included in each video frame at at least two granularities.
[0033] In one embodiment, the i-th dimension is the scene dimension, and the i-th classification sub-model includes a residual network; the processing unit is further configured to call the i-th classification sub-model to perform classification processing on the frame sequence to obtain the classification label of the target video in the i-th dimension, specifically:
[0034] Call the residual network of the i-th classification sub-model to perform weighted processing on each video frame in the frame sequence to obtain the weighted features of each video frame at at least two granularities;
[0035] Determine the classification label of the target video in the scene dimension according to the weighted features of each video frame at at least two granularities.
[0036] In one embodiment, the frame sequence is divided into at least one group, each group of frame sequences includes at least two video frames, the i-th dimension is the content dimension, and the i-th classification sub-model includes a temporal convolutional network and a spatial convolutional network; the processing unit is further configured to call the i-th classification sub-model to perform classification processing on the frame sequence to obtain the classification label of the target video in the i-th dimension, specifically:
[0037] Call the spatial domain convolutional network of the i-th classification sub-model to extract the features of the key frames in each group of frame sequences;
[0038] Call the temporal domain convolutional network of the i-th classification sub-model to extract the features of the data optical flow in each group of frame sequences, and the data optical flow is generated according to the inter-frame difference between adjacent frames in the same group of video frame sequences;
[0039] Determine the classification label of the target video in the content dimension according to the features of the key frames and the features of the data optical flow in each group of frame sequences.
[0040] In one embodiment, the processing unit is further configured to:
[0041] In response to the video service request of the target user, display a video service page;
[0042] Obtain the preference label set of the target user, and the preference label set contains at least one preference label;
[0043] If there is a classification label in the video label set of the target video that matches the preference label in the preference label set, recommend the target video on the video service page.
[0044] In one embodiment, a recommendation list is displayed on the video service page, and the recommendation list includes multiple recommended videos, and the target video is any one in the recommendation list; the processing unit is further configured to recommend the target video on the video service page, specifically:
[0045] Sort the recommendation list in descending order of the relevance of each video in the recommendation list to the preferences of the target user;
[0046] Display the videos in the recommendation list that are arranged before the recommended position on the video service page according to the sorting result;
[0047] Among them, the relevance of the target video to the preferences of the target user is determined according to the number of classification labels in the video label set that match the preference labels in the preference label set.
[0048] On the one hand, the present application provides a video processing device, and the device includes:
[0049] A processor, adapted to execute a computer program;
[0050] A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, the above-mentioned video processing method is implemented.
[0051] On the one hand, the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor to perform the above video processing method.
[0052] On the one hand, the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above video processing method.
[0053] In an embodiment of the present application, a frame sequence is extracted from a target video, and the frame sequence contains key frames of the target video. Since key frames usually have the characteristics of high picture quality and complete picture information, using this frame sequence as the object of video processing to generate a video tag of the target video can make the video tag more comprehensively reflect the content and scene information of the target video, and improve the accuracy of the video tag; in addition, a multi-dimensional classification model is used to classify the frame sequence of the video from at least two dimensions, so as to obtain classification tags of the video in at least two dimensions, and a video tag set of the video is obtained by performing repeated semantic screening on the classification tags. By performing semantic analysis and classification on the content of the video from at least two dimensions through the multi-dimensional classification model, the comprehensiveness and accuracy of the video tag are further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0055] Figure 1a Shows an architecture diagram of a video processing system provided by an exemplary embodiment of the present application;
[0056] Figure 1b Shows a video processing flow chart provided by an exemplary embodiment of the present application;
[0057] Figure 1c Shows another video processing flow chart provided by an exemplary embodiment of the present application;
[0058] Figure 2 Shows a flow chart of a video processing method provided by an exemplary embodiment of the present application;
[0059] Figure 3Shows a flowchart of a frame sequence extraction provided by an exemplary embodiment of the present application;
[0060] Figure 4 Shows a flowchart of another video processing method provided by an exemplary embodiment of the present application;
[0061] Figure 5a Shows an object dimension classification sub-model provided by an exemplary embodiment of the present application;
[0062] Figure 5b Shows a scene dimension classification sub-model provided by an exemplary embodiment of the present application;
[0063] Figure 5c Shows a content dimension classification sub-model provided by an exemplary embodiment of the present application;
[0064] Figure 5d Shows a schematic diagram of a standard category label set provided by an exemplary embodiment of the present application;
[0065] Figure 5e Shows a flowchart of processing a video file in three dimensions provided by an exemplary embodiment of the present application;
[0066] Figure 6 Shows a flowchart of another video processing method provided by an exemplary embodiment of the present application;
[0067] Figure 7a Shows a video service page diagram provided by an exemplary embodiment of the present application;
[0068] Figure 7b Shows another video service page diagram provided by an exemplary embodiment of the present application;
[0069] Figure 8 Shows a schematic structural diagram of a video processing device provided by an exemplary embodiment of the present application;
[0070] Figure 9 Shows a schematic structural diagram of a video processing device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0071] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings.
[0072] Embodiments of this application relate to Artificial Intelligence (AI), Natural Language Processing (NLP), and Machine Learning (ML). By combining AI, NLP, and ML, it is possible to discover hidden and potentially valuable information in videos, enabling devices to more accurately predict and identify objects, scenes, content, etc. in videos, thereby generating video tags corresponding to the videos. Among them, AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0073] AI technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. Basic AI technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, processing technologies for large applications, operating / interactive systems, and mechatronics. AI software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0074] NLP is an important direction in the field of computer science and the field of AI. It studies various theories and methods that can enable effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. NLP technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0075] ML is an interdisciplinary subject that crosses multiple fields and involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0076] Video tags generally refer to the high-level semantic descriptions of video content. As can be seen from the description of the background art above, through practice, it is found that in the prior art, tags are usually added to videos from a single dimension of the main content of the video. This makes the video tags insufficiently comprehensive and inaccurate. In addition, it is difficult to balance the main body and the background with a single-dimensional classification method, which further reflects the deficiencies of the existing video tags. Based on this, the embodiments of the present application propose a video processing solution that can generate more comprehensive and accurate video tags for the target video. This solution has the following characteristics: (1) Extract a frame sequence containing key frames from the target video as the processing object. Since key frames usually have high picture quality and complete picture information, this enables the video tags to more comprehensively reflect the content and scene information of the target video and improve the accuracy of the video tags; (2) Classify the frame sequence from multiple dimensions (such as object dimension, content dimension, scene dimension), which enables the video tags to comprehensively summarize the high-level semantics of the target video; (3) Obtain the video tag set of the target video by performing repeated semantic screening on the classified tags. Through repeated semantic screening, the finally obtained video tags are more accurate in expressing the target video.
[0077] Figure 1a FIG. shows an architecture diagram of a video processing system provided by an exemplary embodiment of the present application. As Figure 1a shown, the video processing system may include one or more terminal devices 101 and one or more servers 102. Figure 1a The numbers of terminal devices and servers in the shown video processing system are only examples. For example, the numbers of terminal devices and servers can be multiple, and the present application does not limit the numbers of terminal devices and servers.
[0078] The terminal device 101 is the device used by the user. The terminal device 101 may include, but is not limited to: smart phones (such as Android phones, iOS phones, etc.), tablet computers, portable personal computers, mobile Internet devices (Mobile Internet Devices, abbreviated as MID), etc. The embodiments of the present invention do not make limitations. The terminal device 101 includes at least one video client, and the video client can be used to provide video services for the user, including but not limited to: video playback service, video search service, video recommendation service, etc. Specifically, the video client in the terminal device 101 provides a video service page 103, such as Figure 1aAn exemplary interface diagram of the video service page 103 is shown; the video client can provide video services to users through the video service page 103. The server 102 refers to a background device that can provide technical support for video services to the terminal device 101; in one embodiment, the server 102 can be the background server of the video client in the terminal device 101. The server 102 can include, but is not limited to, a cluster server.
[0079] Figure 1a In the shown video processing system, in order to better provide video services, the terminal device 101 or the server 102 needs to pre-execute a video processing process to generate video tags for each video in the video library of the video processing system. The video processing process mainly includes the following steps ① to ③: ① Obtain the target video to be processed, and the target video can be any video in the video library of the video processing system; and extract a frame sequence from the target video (such as extracting the key frame sequence of the target video); ② Call a multi-dimensional classification model to perform classification processing on the frame sequence to obtain a candidate tag set for the target video (such as calling a multi-dimensional classification model to perform classification processing on the frame sequence, and obtaining the candidate tag for video 1 as "football" in the first dimension and "playing football" in the second dimension, then the candidate tag set includes "football" and "playing football"); ③ Perform duplicate semantic screening on the candidate tag set to obtain the video tag set for the target video (such as performing duplicate semantic screening on "football" and "playing football", and since "football" includes "playing football", "football" is added to the video tag set of the target video).
[0080] In one embodiment, the terminal device 101 can include a multi-dimensional classification model. Figure 1b Shows a video processing flow chart provided by an exemplary embodiment of the present application. As Figure 1b shown, the above steps ① to ③ can be executed by the terminal device 101. On the basis of steps ① to ③, the video processing process can further include the following steps ④ to ⑥: ④ When the video client on the terminal device 101 is triggered by the target user (for example, the target user opens the video client), the terminal device 101 displays the video service page; ⑤ The terminal device 101 obtains the preference tag set of the target user (such as generating the preference tag set of the target user according to the search keywords of the target user, or the historical browsing record of the target user, etc.); ⑥ The terminal device 101 matches the video tag set of the target video with the preference tag set of the target user. If there is a classification tag in the video tag set that matches the preference tag in the preference tag set, the target video is recommended on the video service page (such as both the video tag set and the preference tag set of video 1 include "football", then video 1 is recommended on the video service page).
[0081] In another embodiment, the server 102 may also include a multi-dimensional classification model. Figure 1c FIG. shows another video processing flow chart provided by an exemplary embodiment of the present application. As Figure 1c shown, the above steps ① to ③ may also be executed by the server 102. On the basis of steps ① to ③, the video processing flow may further include the following steps ⑦ to ⑦ When the video client on the terminal device 101 is triggered by the target user (for example, the target user opens the video client), the terminal device 101 displays a video service page; ⑧ The terminal device 101 obtains a set of preference tags of the target user (such as generating a set of preference tags of the target user according to the search keywords of the target user, or the historical browsing records of the target user, etc.); ⑨ The terminal device 101 requests the server 102 to obtain a video and sends the user preference set to the server 102 together; ⑩ The server 102 matches the set of video tags of the target video with the set of preference tags of the target user. If there is a classification tag in the set of video tags that matches the preference tag in the set of preference tags, the server 102 returns the target video to the terminal device 101; The terminal device 101 then recommends the target video on the video service page.
[0082] In the embodiments of the present application, a multi-dimensional classification model is used to classify the frame sequence of the video from at least two dimensions, so as to obtain classification tags of the video in at least two dimensions, and a set of video tags of the video is obtained by performing repeated semantic screening on the classification tags. It can be seen that calling the multi-dimensional classification model to classify the video can semantically describe the content of the video from different dimensions, making the video tags of the video more comprehensive and accurate. In addition, by detecting the set of preference tags of the user and the set of video tags of the target video, it is determined whether the target video is the content that the user is interested in. It can be seen that for different users, the recommended videos are also different, ensuring that the recommended videos seen by each user are content related to their own preferences (that is, interested), and improving the user experience.
[0083] Figure 2 FIG. shows a flow chart of a video processing method provided by an exemplary embodiment of the present application. The video processing method may be executed by the video processing device proposed in the embodiments of the present application, and the video processing device may be Figure 1a the terminal device 101 or the server 102 shown; as Figure 2 shown, the video processing method includes but is not limited to the following steps 201 to 204. The video processing method provided by the embodiments of the present application will be introduced in detail below:
[0084] 201. The video processing device obtains the target video to be processed.
[0085] The target video can be a video already published on the network, such as an educational video on a learning website, a funny video on an entertainment website, a news video on a news website, etc.; it can also be a video uploaded by a user to the server through a terminal device (i.e., a video that has not been made public), such as user A uploading video 1 to the server after shooting it with a terminal device.
[0086] 202. The video processing device extracts a frame sequence from the target video, and the frame sequence includes the key frames of the target video.
[0087] The frame sequence is obtained by extracting video frames of the target video according to the frame extraction frequency. Figure 3 Fig. shows a flowchart of frame sequence extraction provided by an exemplary embodiment of the present application. As Figure 3 shown, the video source of the target video is input into a decoder to obtain the video frame data stream of the target video. The video frame data stream contains multiple Groups of Pictures (GOPs). GOP represents the distance between two I-frames. An I-frame refers to the first frame in each group of pictures, that is, the key frame. Each GOP contains a set of consecutive pictures. When there are drastic changes in the video picture, the value of GOP will become smaller to ensure the video picture quality. Frame extraction processing is performed on the video frame data stream according to the key frame extraction rule (i.e., the frame extraction frequency) to obtain the video frame sequence. For example, assume that the video frame data stream of video 1 contains 10 GOPs, each GOP contains 6 frames of images, and the frame extraction frequency is to extract one frame every 3 frames of images. Then the number of video frames in the frame sequence of video 1 obtained is 20, and the frame sequence contains 10 key frames in 10 GOPs.
[0088] It should be noted that since the picture quality of the key frame is relatively high, and the positions where there are drastic changes in the video picture (i.e., the content of the video has changed) are usually the positions where the key frames are located, therefore, extracting key frames during frame extraction is beneficial to improving the classification accuracy of the multi-dimensional classification model.
[0089] 203. The video processing device calls a multi-dimensional classification model to perform classification processing on the frame sequence to obtain a candidate label set of the target video, and the candidate label set contains classification labels of the target video in at least two dimensions.
[0090] In one embodiment, the video processing device calls a multi-dimensional classification model to extract features from each frame image in the frame sequence in different dimensions, generates corresponding classification labels according to the extracted features, and then adds the classification labels to the candidate label set of the target video. For example, the content of Video 1 is playing football. The video processing device calls the multi-dimensional classification model to classify the frame sequence of Video 1, and obtains the labels of Video 1 in the object detection dimension as "athlete", "football", and the label in the scene dimension as "football field". Then, the candidate label set of Video 1 includes "athlete", "football", and "football field".
[0091] 204. The video processing device performs repeated semantic screening on the candidate label set to obtain the video label set of the target video.
[0092] In one embodiment, the video processing device screens the labels in the candidate label set that have the same semantics or an inclusion relationship or an association relationship, and adds the screened labels to the video label set of the target video. For example, the candidate label set contains two labels, "football" and "playing football". Since "football" includes "playing football", "football" is added to the video label set of the target video.
[0093] In the embodiments of the present application, a frame sequence is extracted from the target video, and the frame sequence contains the key frames of the target video. Since the key frames usually have the characteristics of high picture quality and complete picture information, using this frame sequence as the object of video processing to generate the video labels of the target video can make the video labels more comprehensively reflect the content and scene information of the target video and improve the accuracy of the video labels; in addition, a multi-dimensional classification model is used to classify the frame sequence of the video from at least two dimensions, so as to obtain the classification labels of the video in at least two dimensions, and the video label set of the video is obtained by performing repeated semantic screening on the classification labels. By using the multi-dimensional classification model to perform semantic analysis and classification on the content of the video from at least two dimensions, the comprehensiveness and accuracy of the video labels are further improved.
[0094] Figure 4 The flowchart of another video processing method provided by an exemplary embodiment of the present application is shown. This video processing method can be executed by the video processing device proposed in the embodiments of the present application, and the video processing device can be Figure 1a the terminal device 101 or the server 102 shown in the figure; as Figure 4 shown in the figure, the video processing method includes but is not limited to the following steps 401 to 407. The following provides a detailed introduction to a video processing method provided by the embodiments of the present application:
[0095] 401. The video processing device obtains the target video to be processed.
[0096] For the specific implementation of step 401, reference can be made to Figure 2 the implementation of step 201 in
[0097] 402. The video processing device determines the frame extraction frequency according to the frame density required by the i-th classification sub-model.
[0098] The frame density is used to measure the number of video frames in a frame sequence. It can be understood that the larger the number of video frames in the frame sequence, the greater the frame density; correspondingly, the smaller the number of video frames in the frame sequence, the smaller the frame density. The frame extraction frequency is calculated based on the number of video frames in the video frame data stream of the target video and the frame density required by the i-th classification sub-model. The number of dimensions is P, that is, the multi-dimensional classification model includes P classification sub-models. The i-th classification sub-model is used to classify the frame sequence in the i-th dimension. P is an integer greater than 1, and i is an integer greater than 1 and i ≤ P.
[0099] In one implementation, when each classification sub-model processes the frame sequence, the required frame densities are different. The i-th classification sub-model refers to the sub-model with the largest required frame density among the P classification sub-models. For example, assume that the number of dimensions is 3, that is, the multi-dimensional classification model includes 3 classification sub-models. The required frame density of the first classification sub-model is 3, that is, the number of video frames in the frame sequence is 3; the required frame density of the second classification sub-model is 6; the required frame density of the third classification sub-model is 36; the number of video frames in the video frame data stream of the target video is 108. Then the video processing device determines the frame extraction frequency as extracting 1 frame every 3 frames according to the frame density required by the third classification sub-model.
[0100] In another implementation, when each classification sub-model processes the frame sequence, the required frame densities are the same, then the frame extraction frequency is determined according to the frame density required by the i-th classification sub-model. At this time, the i-th classification sub-model can refer to any sub-model among the P classification sub-models.
[0101] 403. The video processing device extracts a frame sequence from the target video according to the frame extraction frequency. The frame sequence includes the key frames of the target video.
[0102] For the specific implementation of step 403, reference can be made to Figure 2 the implementation of step 202 in
[0103] 404. The video processing device detects whether the frame density required by the i-th classification sub-model matches the frame extraction frequency of the frame sequence.
[0104] In one embodiment, the i-th classification sub-model may refer to any one of the P classification sub-models. If the frame density required by the i-th classification sub-model matches the frame extraction frequency of the frame sequence, that is, the frame density of the frame sequence obtained by extracting frames according to the frame extraction frequency is the same as the frame density required by the i-th classification sub-model, then step 405 is continued. If the frame density required by the i-th classification sub-model does not match the frame extraction frequency of the frame sequence, that is, the frame density of the frame sequence obtained by extracting frames according to the frame extraction frequency is different from the frame density required by the i-th classification sub-model, then the frame sequence is extracted according to the frame density required by the i-th classification sub-model to obtain a processed frame sequence. For example, assume that the frame density required by the second classification sub-model is 6, and the frame density of the frame sequence obtained by extracting frames according to the frame extraction frequency is 18. Then, the frame sequence is extracted according to the frame density of 6 required by the second classification sub-model (extract 1 frame every 3 frames) to obtain a processed frame sequence. At this time, the frame density of the frame sequence is 6.
[0105] 405. The video processing device respectively calls the P classification sub-models to classify the frame sequence, and obtains the classification labels of the target video in P dimensions.
[0106] In one embodiment, the i-th dimension is the object dimension, and the i-th classification sub-model includes an identification network, and the identification network is used to extract and fuse the features of the video frame at at least two granularities. The i-th classification sub-model generates corresponding classification labels according to the features of the object included in each video frame output by the identification network at at least two granularities. Figure 5a FIG. shows an object dimension classification sub-model provided by an exemplary embodiment of the present application. As Figure 5a shown, the object dimension classification sub-model is constructed based on the YOLOv3 network framework. The object dimension classification sub-model includes a residual block, an upsampling layer, a detection layer, and a progressive layer. In the object dimension classification sub-model, the identification network fuses the features of the video frame at 3 granularities. It should be noted that the object dimension classification sub-model may also be other network models based on multi-granularity prediction and multi-granularity fusion, such as Fast Region-based Convolutional Neural Networks (Fast R-CNN), Single Shot MultiBox Detector (SSD), etc.
[0107] In another embodiment, the i-th dimension is the scene dimension, and the i-th classification sub-model includes a residual network, and the residual network is used to extract and fuse the features of the video frame at at least two granularities. The i-th classification sub-model generates corresponding classification labels according to the features of the scene included in each video frame output by the residual network at at least two granularities. Figure 5bShows a scene dimension classification sub-model provided by an exemplary embodiment of the present application. As Figure 5b shown, this scene dimension classification sub-model is constructed based on the Residual Network 34 (ResNet34). The scene dimension classification sub-model contains 34 convolutional layers. Among them, 3x3 represents the filter in the convolutional layer, and 64-256 represents the granularity size into which the video frames are divided in the current convolutional layer. It should be noted that the object dimension classification sub-model can also be constructed based on other residual networks, such as ResNet18, ResNet101, etc.
[0108] In another implementation, the frame sequence is divided into at least one GOP. Each GOP includes at least two video frames. The i-th dimension is the content dimension. The i-th classification sub-model includes a temporal convolutional network and a spatial convolutional network. The spatial convolutional network is used to extract the features of the key frames in each GOP, and the temporal convolutional network is used to extract the features of the data optical flow in each GOP. Among them, the data optical flow is generated according to the inter-frame difference between adjacent video frames in the same GOP. The i-th classification sub-model generates corresponding classification labels according to the features of the content contained in each video frame output by the temporal convolutional network and the spatial convolutional network in the temporal and spatial domains. Figure 5c Shows a content dimension classification sub-model provided by an exemplary embodiment of the present application. As Figure 5c shown, this content dimension classification sub-model is constructed based on the time-sensitive network (TSN). Each GOP contains 3 video frames. The temporal convolutional network and the spatial convolutional network are respectively used to extract features and classify each GOP. Then, the results in the two dimensions are combined and sent to the Softmax layer to predict the probability that each GOP belongs to a certain category. Finally, the predicted values of each GOP are fused by weighted average to finally obtain the probability value of the target video in each category. It should be noted that the content dimension classification sub-model can also be other network models based on the temporal convolutional network and the spatial convolutional network. For example, the content dimension classification sub-model can also be constructed based on the Temporal Relation Network (TRN), etc.
[0109] It can be understood that the multi-dimension classification model can include one or more of the above 3 dimension classification sub-models, or can also include classification sub-models of other dimensions.
[0110] 406. The video processing device adds the classification labels in the P dimensions to the candidate label set of the target video.
[0111] For example, assume that the classification labels under the first dimension are "football" and "athlete", the classification label under the second dimension is "outdoor sport", and the classification label under the third dimension is "football field". Then, the candidate label set of the target video includes "football", "athlete", "outdoor sport", and "football field".
[0112] 407. The video processing device performs duplicate semantic screening on the candidate label set to obtain the video label set of the target video.
[0113] In one implementation, the video processing device maps the labels with duplicate (identical) semantics in the candidate label set to obtain standard category labels, and adds the standard category labels to the standard category label set. For example, the candidate label set contains two labels, "pop music" and "folk music". Since both "pop music" and "folk music" belong to "music", "music" is added to the standard category label set as the standard category label. Figure 5d FIG. shows a schematic diagram of a standard category label set provided by an exemplary embodiment of the present application. As Figure 5d shown, the standard category label set includes multiple standard categories, and each standard category includes multiple classification labels.
[0114] Statistically count the number N of classification labels belonging to the target standard category, and statistically count the number M of times that P classification sub-models perform classification processing on the frame sequence. Calculate the ratio of N to M. If the ratio between N and M is greater than or equal to the threshold, add the target standard category to the video label set of the target video, where the target standard category is any one of the standard categories in the standard category label set. For example, assume that the number of classification labels belonging to the "music" category in the standard category label set of video 1 is 87, the multi-dimensional classification model includes 3 classification sub-models, the number of times that the first classification sub-model and the second classification sub-model perform classification processing on the frame sequence is 40 each, the number of times that the third classification sub-model performs classification processing on the frame sequence is 20, and the threshold is 0.8. Then, the value of N is 87, M = 40 + 40 + 20 = 100, and the ratio of N to M is 0.87 > 0.8. Therefore, "music" is added to the video label set of video 1 (i.e., "music" is determined as a video label of video 1). Correspondingly, if the ratio between N and M is less than the threshold, the target standard category is discarded.
[0115] Figure 5e FIG. shows a flowchart of processing a video file in three dimensions according to an exemplary embodiment of the present application. As Figure 5eAs shown in the figure, after obtaining the video file, first determine the video frame extraction frequency (i.e., the frame extraction strategy) according to the frame density required by the object dimension classification sub-model, the scene dimension classification sub-model, and the content dimension classification sub-model. Assume that the number of video frames in the video frame data stream of the video file is 150, the video frame sequences required by the object dimension classification sub-model and the scene dimension classification sub-model are the key frame sequences of the video file (the frame density is 10), and the frame density required by the content dimension classification sub-model is 30. Then determine the frame extraction frequency to extract 1 frame every 5 frames. Perform frame extraction processing on the video frame data stream of the video file according to the frame extraction frequency to obtain a video frame sequence, and the density of the obtained frame sequence is 30. Then adapt the frame sequence according to the frame density required by each classification sub-model. Since the frame density required by the object dimension classification sub-model and the scene dimension classification sub-model is 10, it is necessary to perform frame extraction processing on the frame sequence (extract 1 frame every 3 frames) to obtain the adapted frame sequence, and call the object dimension classification sub-model and the scene dimension classification sub-model to perform classification processing on the adapted frame sequence. The frame density required by the content dimension classification sub-model is the same as the density of the frame sequence, which is 30. Therefore, directly call the content dimension classification sub-model to perform classification processing on the frame sequence. After the 3 classification sub-models complete the classification processing of the corresponding frame sequences, a candidate label set of the video file can be obtained, and repeated semantic screening is performed on the candidate label set to obtain the video label set of the target video (i.e., the video multi-label description).
[0116] In the embodiment of the present application, a frame sequence is extracted from the target video, and the frame sequence contains the key frames of the target video. Since the key frames usually have the characteristics of high picture quality and complete picture information, using this frame sequence as the object of video processing to generate the video label of the target video can make the video label more comprehensively reflect the content and scene information of the target video and improve the accuracy of the video label; in addition, a multi-dimensional classification model is used to classify the frame sequence of the video from at least two dimensions, so as to obtain classification labels of the video in at least two dimensions, and the video label set of the video is obtained by performing repeated semantic screening on the classification labels. By performing semantic analysis and classification on the content of the video from at least two dimensions through the multi-dimensional classification model, the comprehensiveness and accuracy of the video label are further improved.
[0117] Figure 6 The flowchart of another video processing method provided by an exemplary embodiment of the present application is shown. This video processing method can be executed by the video processing device proposed in the embodiment of the present application, and the video processing device can be Figure 1a the terminal device 101 shown in the figure; as Figure 6 shown, the video processing method includes but is not limited to the following steps 601 to step 603. The following provides a detailed introduction to a video processing method provided by an embodiment of the present application:
[0118] 601. In response to a video service request from a target user, the video processing device displays a video service page.
[0119] In one embodiment, when the video processing device detects that the target user opens the video client, the video processing device displays a video service page.
[0120] 602. The video processing device obtains a set of preference tags of the target user, and the set of preference tags contains at least one preference tag.
[0121] The set of preference tags of the target user can be obtained according to keywords input by the user, or can be generated based on the historical browsing records of the target user. The set of preference tags includes one or more preference tags. For example, when user A opens a video recommendation software, the video processing device obtains that the videos browsed by user A in the past week are mainly related to music and pets. Then the preference tags included in the set of preference tags are "music" and "pets". Then it is detected that user A enters the keyword "football" in the search bar. At this time, the preference tags included in the set of preference tags are "football".
[0122] 603. If there is a classification tag in the video tag set of the target video that matches the preference tag in the set of preference tags, the video processing device recommends the target video on the video service page.
[0123] In one embodiment, the video processing device obtains and compares the classification tags in the video tag set of the target video with the preference tags in the obtained set of preference tags. If there is a classification tag in the video tag set of the target video that matches the preference tag in the set of preference tags, the video processing device recommends the target video on the video service page. Among them, the video tag set of the target video is obtained through the above Figure 2 or Figure 4 video processing method. For example, the video tag set of video 1 includes "music" and "concert", and the set of preference tags includes "music" and "pets". Since both the video tag set and the set of preference tags include the "music" tag, the video processing device recommends video 1 on the service page. Figure 7a Shows a video service page diagram provided by an exemplary embodiment of the present application.
[0124] Further, the video processing device recommends videos to the target user by displaying a recommended list on the service page. The recommended list includes multiple recommended videos, and the recommended videos in the recommended list are arranged in descending order of relevance to the preferences of the target user. When displaying, the video processing device displays the recommended videos in the recommended list that are before the recommended position on the video service page according to the sorting result. Among them, the relevance of the recommended video to the preferences of the target user is determined according to the number of classification labels in the video label set that match the preference labels in the preference label set. The more the number of classification labels in the video label set that match the preference labels in the preference label set, the higher the relevance of the recommended video to the preferences of the target user. For example, assume that the preference label set obtained by the video processing device and the video label sets of recommended videos 1 to 3 are shown in Table 1:
[0125] Table 1
[0126] Set of preference labels "Football", "Funny", "Outdoor", "Pet" Set of video tags for recommended video 1 "Football", "Outdoor", "Pet" Set of video tags for recommended video 2 "Pet", "Training" Set of video tags for recommended video 3 "Football", "Outdoor"
[0127] As can be seen from Table 1, the number of classification labels in the video label set of recommended video 1 that match the preference labels in the preference label set is 3, the number of classification labels in the video label set of recommended video 2 that match the preference labels in the preference label set is 1, and the number of classification labels in the video label set of recommended video 3 that match the preference labels in the preference label set is 2. Therefore, the sorting result of recommended videos 1 to 3 in descending order of relevance to the preferences of the target user is: recommended video 1 → recommended video 3 → recommended video 2. If the recommended position is 2 (that is, the first two videos in the recommended list are recommended), the video processing device displays recommended video 1 and recommended video 3 on the service page. Figure 7b Fig. shows another video service page diagram provided by an exemplary embodiment of the present application.
[0128] In another implementation, the video processing device sends a recommended video acquisition request to the server. The recommended video acquisition request includes the preference label set of the target user. The server determines the recommended videos according to the preference label set of the target user and the video label set of the target video, and sends them to the video processing device. After the video processing device acquires the recommended videos, it displays the recommended videos on the service page. The specific implementation of the server determining the recommended target videos according to the preference label set of the target user and the video label set of the target video can refer to the previous implementation, and will not be elaborated here.
[0129] In the embodiments of the present application, by detecting the set of preference tags of the user and the set of video tags of the target video, it is determined whether the target video is the content that the user is interested in. It can be seen that for different users, the recommended videos are also different, ensuring that the recommended videos seen by each user are all content related to their own preferences (i.e., interested), thus improving the user experience.
[0130] The method of the embodiments of the present application is elaborated in detail above. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, the device of the embodiments of the present application is provided below.
[0131] Please refer to Figure 8 , Figure 8 which shows a schematic structural diagram of a video processing device provided by an exemplary embodiment of the present application. This video processing device can be mounted on the video processing device in the above method embodiment. This video processing device can be an application program in the video processing device (for example: a video application program); Figure 8 The shown video processing device can be used to execute some or all of the functions in the above Figure 2 , Figure 4 and Figure 6 described method embodiments. Among them, the detailed descriptions of each unit are as follows:
[0132] An obtaining unit 801, configured to obtain a target video to be processed;
[0133] A processing unit 802, configured to extract a frame sequence from the target video, where the frame sequence includes key frames of the target video;
[0134] Call a multi-dimensional classification model to perform classification processing on the frame sequence to obtain a candidate tag set of the target video, where the candidate tag set contains classification tags of the target video in at least two dimensions;
[0135] Perform duplicate semantic screening on the candidate tag set to obtain a video tag set of the target video.
[0136] In one implementation, the number of dimensions is denoted as P, and the multi-dimensional classification model includes P classification sub-models; the i-th classification sub-model is used to perform classification processing on the frame sequence in the i-th dimension; P is an integer greater than 1, and i is an integer greater than 1 and i ≤ P.
[0137] In one implementation, the processing unit 802 is further configured to extract a frame sequence from the target video, specifically:
[0138] Determine the frame extraction frequency according to the frame density required by the P classification sub-models;
[0139] Perform frame extraction processing on the target video according to the frame extraction frequency to obtain a frame sequence.
[0140] In one embodiment, the processing unit 802 is further configured to determine the frame extraction frequency according to the frame densities required by P classification sub-models, specifically:
[0141] Obtain the frame densities respectively required by each of the P classification sub-models;
[0142] Select the maximum frame density from the P frame densities and determine it as the frame extraction frequency.
[0143] In one embodiment, the processing unit 802 is further configured to call a multi-dimensional classification model to classify the frame sequence, and obtain a candidate label set of the target video, specifically:
[0144] Call the P classification sub-models respectively to classify the frame sequence, and obtain the classification labels of the target video in P dimensions;
[0145] Add the classification labels of the target video in P dimensions to the candidate label set of the target video.
[0146] In one embodiment, before calling the i-th classification sub-model to classify the frame sequence and obtain the classification label of the target video in the i-th dimension, the processing unit 802 is further configured to:
[0147] Detect whether the frame density required by the i-th classification sub-model matches the frame extraction frequency of the frame sequence;
[0148] If the frame density required by the i-th classification sub-model matches the frame extraction frequency of the frame sequence, then perform the step of calling the i-th classification sub-model to classify the frame sequence and obtain the classification label of the target video in the i-th dimension;
[0149] If the frame density required by the i-th classification sub-model does not match the frame extraction frequency of the frame sequence, then perform frame extraction on the frame sequence according to the frame density required by the i-th classification sub-model, and call the i-th classification sub-model to classify the frame sequence after frame extraction, and obtain the classification label of the target video in the i-th dimension.
[0150] In one embodiment, the processing unit 802 is further configured to perform duplicate semantic screening on the candidate label set to obtain a video label set of the target video, specifically:
[0151] Perform duplicate semantic mapping on each classification label in the candidate label set to obtain a standard category label set, where the standard category label set includes multiple standard categories and multiple classification labels under each standard category;
[0152] Count the number N of classification labels that belong to the target standard category, and count the number M of times the P classification sub-models classify the frame sequence; the target standard category is any one of the standard category label sets, and N and M are positive integers;
[0153] If the ratio between N and M is greater than or equal to the threshold, add the target standard category to the video label set of the target video.
[0154] In one implementation, the i-th dimension is the object dimension, and the i-th classification sub-model includes an identification network; the processing unit 802 is further configured to call the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension, specifically:
[0155] Call the identification network of the i-th classification sub-model to identify the frame sequence to obtain the features of the objects included in each video frame at at least two granularities;
[0156] Determine the classification label of the target video in the object dimension according to the features of the objects included in each video frame at at least two granularities.
[0157] In one implementation, the i-th dimension is the scene dimension, and the i-th classification sub-model includes a residual network; the processing unit 802 is further configured to call the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension, specifically:
[0158] Call the residual network of the i-th classification sub-model to perform weighted processing on each video frame in the frame sequence to obtain the weighted features of each video frame at at least two granularities;
[0159] Determine the classification label of the target video in the scene dimension according to the weighted features of each video frame at at least two granularities.
[0160] In one implementation, the frame sequence is divided into at least one group, each group of frame sequences includes at least two video frames, the i-th dimension is the content dimension, and the i-th classification sub-model includes a temporal convolutional network and a spatial convolutional network; the processing unit 802 is further configured to call the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension, specifically:
[0161] Call the spatial convolutional network of the i-th classification sub-model to extract the features of the key frames in each group of frame sequences;
[0162] Call the temporal convolutional network of the i-th classification sub-model to extract the features of the data optical flow in each group of frame sequences, and the data optical flow is generated according to the inter-frame difference between adjacent frames in the same group of video frame sequences;
[0163] Determine the classification label of the target video in the content dimension according to the features of the key frames in each group of frame sequences and the features of the data optical flow.
[0164] In one implementation, the processing unit 802 is further configured to:
[0165] In response to the video service request of the target user, display a video service page;
[0166] Obtain the preference label set of the target user, where the preference label set contains at least one preference label;
[0167] If there is a classification label in the video label set of the target video that matches the preference label in the preference label set, recommend the target video on the video service page.
[0168] In one implementation, a recommendation list is displayed on the video service page, and the recommendation list includes multiple recommended videos. The target video is any one of the videos in the recommendation list; the processing unit 802 is further configured to recommend the target video on the video service page, specifically:
[0169] Sort the recommendation list in descending order according to the relevance of each video in the recommendation list to the preferences of the target user;
[0170] Display the videos in the recommendation list that are arranged before the recommended position on the video service page according to the sorting result;
[0171] Among them, the relevance of the target video to the preferences of the target user is determined according to the number of classification labels in the video label set that match the preference labels in the preference label set.
[0172] According to an embodiment of the present application, Figure 2 , Figure 4 and Figure 6 Some of the steps involved in the video processing method shown can be executed by each unit in the Figure 8 shown video processing device. For example, Figure 2 the step 201 shown in Figure 8 can be executed by the obtaining unit 801 shown, and the steps 202-step 204 can be executed by the Figure 8 shown processing unit 802. Figure 4 The step 401 shown in Figure 8 can be executed by the obtaining unit 801 shown, and the steps 402-step 407 can be executed by the Figure 8 shown processing unit 802. Figure 6 The step 602 shown in Figure 8 can be executed by the obtaining unit 801 shown, and the steps 601 and 603 can be executed by the Figure 8 shown processing unit 802.Figure 8 Each unit in the video processing device shown can be separately or all combined into one or several other units to form, or some of them can be further split into multiple smaller units with more specific functions to form. This can achieve the same operations without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the video processing device may also include other units. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.
[0173] According to another embodiment of this application, it can be achieved by running a computer program (including program code) that can execute the respective steps involved in the corresponding methods shown in Figure 2 , Figure 4 and Figure 6 on a general computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct a video processing device as shown in Figure 8 and to implement the video processing method of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0174] Based on the same inventive concept, the principle of solving problems and the beneficial effects of the video processing device provided in the embodiments of this application are similar to those of the video processing method in the method embodiments of this application. For the principle and beneficial effects of the method implementation, reference can be made, and for the sake of brevity, they will not be elaborated here.
[0175] Please refer to Figure 9 , Figure 9 which shows a schematic structural diagram of a video processing device provided by an exemplary embodiment of this application. The video processing device can be Figure 1aThe terminal device 101 or the server 102 in the system shown; the video processing device at least includes a processor 901, a communication interface 902, and a memory 903. Among them, the processor 901, the communication interface 902, and the memory 903 can be connected through a bus or other means. In the embodiments of the present application, the connection through a bus is taken as an example. Among them, the processor 901 (or the Central Processing Unit (CPU)) is the computing core and control core of the video processing device, which can parse various instructions in the terminal device and process various data of the terminal device. For example, the CPU can be used to parse the power-on and power-off instructions sent by the user to the terminal device and control the terminal device to perform power-on and power-off operations; for another example, the CPU can transmit various interactive data between the internal structures of the terminal device, and so on. The communication interface 902 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and under the control of the processor 901, it can be used to send and receive data; the communication interface 902 can also be used for the transmission and interaction of internal data of the terminal device. The memory 903 (Memory) is the memory device in the terminal device, used to store programs and data. It can be understood that the memory 903 here can include both the built-in memory of the terminal device and, of course, the extended memory supported by the terminal device. The memory 903 provides a storage space, and this storage space stores the operating system of the terminal device, which can include but is not limited to: Android system, iOS system, Windows Phone system, etc. The present application does not make any limitations in this regard.
[0176] In one embodiment, the video processing device may refer to a terminal device or a server, for example Figure 1a the terminal device 101 or the server 102 shown. In this case, by running the executable program code in the memory 903, the processor 901 performs the following operations:
[0177] Obtain the target video to be processed through the communication interface 902;
[0178] Extract a frame sequence from the target video, and the frame sequence includes the key frames of the target video;
[0179] Call a multi-dimensional classification model to classify the frame sequence, and obtain a candidate label set of the target video. The candidate label set contains classification labels of the target video in at least two dimensions;
[0180] Perform repeated semantic screening on the candidate label set to obtain the video label set of the target video.
[0181] As an alternative implementation, the number of dimensions is denoted as P, and the multi-dimensional classification model includes P classification sub-models; the i-th classification sub-model is used to classify the frame sequence in the i-th dimension; P is an integer greater than 1, and i is an integer greater than 1 and i ≤ P.
[0182] As an alternative implementation, the specific implementation of the processor 901 extracting the frame sequence from the target video is as follows:
[0183] Determine the frame extraction frequency according to the frame density required by the P classification sub-models;
[0184] Perform frame extraction on the target video according to the frame extraction frequency to obtain a frame sequence.
[0185] As an alternative implementation, the specific implementation of the processor 901 determining the frame extraction frequency according to the frame density required by the P classification sub-models is as follows:
[0186] Obtain the frame density required by each classification sub-model among the P classification sub-models;
[0187] Select the maximum frame density from the P frame densities and determine it as the frame extraction frequency.
[0188] As an alternative implementation, the specific implementation of the processor 901 calling the multi-dimensional classification model to classify the frame sequence to obtain the candidate label set of the target video is as follows:
[0189] Call the P classification sub-models respectively to classify the frame sequence to obtain the classification labels of the target video in the P dimensions;
[0190] Add the classification labels of the target video in the P dimensions to the candidate label set of the target video.
[0191] As an alternative implementation, before calling the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension, the processor 901 also performs the following operations by running the executable program code in the memory 903:
[0192] Detect whether the frame density required by the i-th classification sub-model matches the frame extraction frequency of the frame sequence;
[0193] If the frame density required by the i-th classification sub-model matches the frame extraction frequency of the frame sequence, then perform the step of calling the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension;
[0194] If the frame density required by the i-th classification sub-model does not match the frame extraction frequency of the frame sequence, the frame sequence is frame-extracted according to the frame density required by the i-th classification sub-model, and the i-th classification sub-model is called to classify the frame sequence after frame extraction to obtain the classification label of the target video in the i-th dimension.
[0195] As an alternative implementation, the specific implementation of the processor 901 performing repeated semantic screening on the candidate label set to obtain the video label set of the target video is as follows:
[0196] Perform repeated semantic mapping on each classification label in the candidate label set to obtain a standard category label set, where the standard category label set includes multiple standard categories and multiple classification labels under each standard category;
[0197] Count the number N of classification labels belonging to the target standard category, and count the number M of times the P classification sub-models classify the frame sequence; the target standard category is any standard category in the standard category label set, and N and M are positive integers;
[0198] If the ratio between N and M is greater than or equal to the threshold, add the target standard category to the video label set of the target video.
[0199] As an alternative implementation, the i-th dimension is the object dimension, and the i-th classification sub-model includes an identification network; the specific implementation of the processor 901 calling the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension is as follows:
[0200] Call the identification network of the i-th classification sub-model to identify the frame sequence to obtain the features of the objects included in each video frame at at least two granularities;
[0201] Determine the classification label of the target video in the object dimension according to the features of the objects included in each video frame at at least two granularities.
[0202] As an alternative implementation, the i-th dimension is the scene dimension, and the i-th classification sub-model includes a residual network; the specific implementation of the processor 901 calling the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension is as follows:
[0203] Call the residual network of the i-th classification sub-model to perform weighted processing on each video frame in the frame sequence to obtain the weighted features of each video frame at at least two granularities;
[0204] Determine the classification label of the target video in the scene dimension according to the weighted features of each video frame at at least two granularities.
[0205] As an alternative implementation, the frame sequence is divided into at least one group, each group of frame sequences includes at least two video frames, the i-th dimension is the content dimension, and the i-th classification sub-model includes a temporal convolutional network and a spatial convolutional network; the specific implementation of the processor 901 calling the i-th classification sub-model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension is as follows:
[0206] Call the spatial convolutional network of the i-th classification sub-model to extract the features of the key frames in each group of frame sequences;
[0207] Call the temporal convolutional network of the i-th classification sub-model to extract the features of the data optical flow in each group of frame sequences, and the data optical flow is generated according to the inter-frame difference between adjacent frames in the same group of video frame sequences;
[0208] Determine the classification label of the target video in the content dimension according to the features of the key frames and the features of the data optical flow in each group of frame sequences.
[0209] As an alternative implementation, the processor 901 also performs the following operations by running the executable program code in the memory 903:
[0210] In response to the video service request of the target user, display the video service page;
[0211] Obtain the preference label set of the target user, and the preference label set contains at least one preference label;
[0212] If there is a classification label in the video label set of the target video that matches the preference label in the preference label set, recommend the target video on the video service page.
[0213] As an alternative implementation, a recommendation list is displayed on the video service page, and the recommendation list includes multiple recommended videos, and the target video is any one in the recommendation list; the specific implementation of the processor 901 recommending the target video on the video service page is as follows:
[0214] Sort the recommendation list in descending order according to the relevance of each video in the recommendation list to the preferences of the target user;
[0215] Display the videos in the recommendation list that are arranged before the recommended position on the video service page according to the sorting result;
[0216] Among them, the relevance of the target video to the preferences of the target user is determined according to the number of classification labels in the video label set that match the preference labels in the preference label set.
[0217] Based on the same inventive concept, the principle and beneficial effects of the video processing device provided in the embodiments of the present application for solving problems are similar to those of the video processing method in the method embodiments of the present application. The principle and beneficial effects of the method implementation can be referred to. For the sake of concise description, they will not be elaborated here.
[0218] The embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored. The computer program is adapted to be loaded and executed by a processor to perform the video processing method in the above method embodiments.
[0219] The embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above video processing method.
[0220] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0221] The steps in the method embodiments of the present application can be adjusted, combined, and deleted according to actual needs.
[0222] The modules in the device embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0223] Those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the readable storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0224] The above-disclosed is only a preferred embodiment of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand the entire or partial processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.
Claims
1. A video processing method, characterized in that, The method includes: Obtaining a target video to be processed; Determining a frame extraction frequency according to the frame densities required by P classification sub-models included in the multi-dimensional classification model; wherein, the determined frame extraction frequency refers to: the frame extraction frequency corresponding to the maximum frame density among the frame densities required by each of the P classification sub-models; when each classification sub-model processes a frame sequence, the required frame densities are different; Performing frame extraction processing on the target video according to the frame extraction frequency to obtain a frame sequence, where the frame sequence includes key frames of the target video; Detecting whether the frame density required by the i-th classification sub-model among the P classification sub-models matches the frame extraction frequency of the frame sequence; P is an integer greater than 1, and i is an integer greater than 1 and i ≤ P; the i-th classification sub-model among the P classification sub-models is used to classify the frame sequence in the i-th dimension; If they match, calling the i-th classification sub-model in the multi-dimensional classification model to classify the frame sequence to obtain a classification label of the target video in the i-th dimension; If they do not match, performing frame extraction processing on the frame sequence according to the frame density required by the i-th classification sub-model, and calling the i-th classification sub-model to classify the frame sequence after frame extraction processing to obtain a classification label of the target video in the i-th dimension; According to the classification labels respectively obtained by the P classification sub-models in the multi-dimensional classification model, obtaining a candidate label set of the target video; Performing repeated semantic screening on the candidate label set to obtain a video label set of the target video.
2. The method according to claim 1, wherein The calling the multi-dimensional classification model to classify the frame sequence to obtain a candidate label set of the target video includes: Respectively calling the P classification sub-models to classify the frame sequence to obtain classification labels of the target video in P dimensions; Adding the classification labels of the target video in P dimensions to the candidate label set of the target video.
3. The method according to claim 1, wherein The performing repeated semantic screening on the candidate label set to obtain a video label set of the target video includes: Performing repeated semantic mapping on each classification label in the candidate label set to obtain a standard category label set, where the standard category label set includes multiple standard categories and multiple classification labels under each standard category; Counting the number N of classification labels belonging to the target standard category, and counting the number M of times the P classification sub-models classify the frame sequence; the target standard category is any one of the standard categories in the standard category label set, and N and M are positive integers; If the ratio between N and M is greater than or equal to a threshold, adding the target standard category to the video label set of the target video.
4. The method according to claim 2, characterized in that The i-th dimension is an object dimension, and the i-th classification sub-model includes an identification network; the calling the i-th classification sub-model in the multi-dimensional classification model to classify the frame sequence to obtain a classification label of the target video in the i-th dimension includes: Invoke the recognition network of the i-th classification sub-model to recognize the frame sequence, and obtain the features of the objects included in each video frame at at least two granularities; Determine the classification label of the target video in the object dimension according to the features of the objects included in each video frame at at least two granularities.
5. The method according to claim 2, characterized in that, The i-th dimension is the scene dimension, and the i-th classification sub-model includes a residual network; the step of invoking the i-th classification sub-model in the multi-dimensional classification model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension includes: Invoke the residual network of the i-th classification sub-model to perform weighted processing on each video frame in the frame sequence, and obtain the weighted features of each video frame at at least two granularities; Determine the classification label of the target video in the scene dimension according to the weighted features of each video frame at at least two granularities.
6. The method according to claim 2, wherein The frame sequence is divided into at least one group, each group of frame sequences includes at least two video frames, the i-th dimension is the content dimension, and the i-th classification sub-model includes a temporal convolutional network and a spatial convolutional network; the step of invoking the i-th classification sub-model in the multi-dimensional classification model to classify the frame sequence to obtain the classification label of the target video in the i-th dimension includes: Invoke the spatial convolutional network of the i-th classification sub-model to extract the features of the key frames in each group of frame sequences; Invoke the temporal convolutional network of the i-th classification sub-model to extract the features of the data optical flow in each group of frame sequences, where the data optical flow is generated according to the inter-frame difference between adjacent frames in the same group of video frame sequences; Determine the classification label of the target video in the content dimension according to the features of the key frames and the features of the data optical flow in each group of frame sequences.
7. The method according to claim 1, wherein The method further includes: In response to a video service request from a target user, display a video service page; Obtain the set of preference labels of the target user, where the set of preference labels contains at least one preference label; If there is a classification label in the video label set of the target video that matches the preference label in the set of preference labels, recommend the target video on the video service page.
8. The method according to claim 7, characterized in that A recommendation list is displayed on the video service page, and the recommendation list includes multiple recommended videos, and the target video is any one of the videos in the recommendation list; The step of recommending the target video on the video service page includes: Sort the recommendation list in descending order according to the relevance of each video in the recommendation list to the preferences of the target user; Display the videos in the recommendation list that are arranged before the recommended position on the video service page according to the sorting result; Wherein, the relevance of the target video to the preferences of the target user is determined according to the number of classification labels in the video label set that match the preference labels in the set of preference labels.
9. A video processing device, characterized in that, including: An acquisition unit, configured to acquire a target video to be processed; A processing unit for determining a frame extraction frequency according to the frame densities required by P classification sub-models included in a multi-dimensional classification model; wherein, the determined frame extraction frequency means: the frame extraction frequency corresponding to the maximum frame density among the frame densities required by each of the P classification sub-models; the frame densities required by each classification sub-model are different when processing a frame sequence; the target video is frame-extracted according to the frame extraction frequency to obtain a frame sequence, and the frame sequence includes key frames of the target video; detecting whether the frame density required by the i-th classification sub-model among the P classification sub-models matches the frame extraction frequency of the frame sequence; P is an integer greater than 1, and i is an integer greater than 1 and i ≤ P; the i-th classification sub-model among the P classification sub-models is used to classify the frame sequence in the i-th dimension. If they match, the i-th classification sub-model in the multi-dimensional classification model is called to classify the frame sequence to obtain a classification label of the target video in the i-th dimension; if they do not match, the frame sequence is frame-extracted according to the frame density required by the i-th classification sub-model, and the i-th classification sub-model is called to classify the frame sequence after frame extraction to obtain a classification label of the target video in the i-th dimension; according to the classification labels respectively obtained by the P classification sub-models in the multi-dimensional classification model, a candidate label set of the target video is obtained; the candidate label set is subjected to repeated semantic screening to obtain a video label set of the target video.
10. A video processing device, characterized in that, Comprising: A processor adapted to execute a computer program; A computer-readable storage medium storing a computer program, which when executed by the processor, implements the video processing method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the video processing method according to any one of claims 1-8.
12. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the video processing method according to any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Method and apparatus for generating information
CN109325148A
Game acquisition method and device, computer equipment and storage medium
CN111277859A
Method for recognizing video action, and device and storage medium thereof
US20220130146A1