Identification and replacement methods of video elements and video recommendation methods
By extracting the constituent element characteristics, short-term and long-term characteristics of the video frame sequence and performing feature fusion, the problem of low accuracy in recognition of constituent elements in the video is solved, and comprehensive and accurate identification of constituent elements in the video is achieved.
Patent Information
- Application Number
- CN202210674399.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-15
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-06-15
AI Technical Summary
In the prior art, the accuracy of the identification of constituent elements in video files is low, and it is difficult to fully and accurately identify constituent elements in video.
By obtaining the constituent element characteristics, short-term characteristics and long-term characteristics of the video frame sequence, the constituent elements in the video are identified by using the feature fusion method, and combining pixel difference analysis and element category information, fine-grained feature extraction and recognition of video frames are achieved.
It improves the accuracy of the recognition of components in videos, can fully identify single-frame, short-term and long-term components, and enhances the recognition effect of components in videos.
Smart Images

Figure CN115115979B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer equipment, storage medium, and computer program product for identifying constituent elements in a video, as well as a method for replacing constituent elements in a video and a method for recommending videos based on constituent elements. Background Art
[0002] With the continuous development of the internet, people's interactions are becoming more diverse. In addition to text-based and voice-based interactions, video has become a crucial way for people to communicate and share knowledge. Compared to other media files like voice and text, videos provide more useful information, and their content is more vivid, visual, and intuitive. Furthermore, with the rapid development of video technology, various elements such as text, animation, and backgrounds can be added to videos, creating even richer videos.
[0003] For synthesized video files, in order to reversely identify the component elements contained in the video file, an image classification network is generally used to predict the component elements frame by frame, but this processing method has the problem of low accuracy. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, computer equipment, computer-readable storage medium and computer program product for identifying component elements in a video that can improve recognition accuracy in response to the above technical problems.
[0005] In a first aspect, the present application provides a method for identifying component elements in a video. The method comprises:
[0006] Acquire a video frame sequence consisting of at least a portion of video frames of a target video, wherein the video frames in the video frame sequence are arranged according to a time sequence in the target video;
[0007] Extracting component element features of each video frame in the video frame sequence;
[0008] Extracting short temporal features and long temporal features of the video frame sequence respectively; a time span of a video frame matched by the short temporal features in the video frame sequence is smaller than a time span of a video frame matched by the long temporal features in the video frame sequence;
[0009] Based on a feature fusion result obtained by fusing the component element features, the short time sequence features, and the long time sequence features, the component elements in the target video are identified.
[0010] In one embodiment, the comparing by pixel unit to obtain the pixel difference between frames includes:
[0011] Obtain pixel data of each pixel unit to be compared, wherein the pixel data includes at least one type of data among RGB data, brightness data, and grayscale data; compare the pixel data of the pixel unit at the same position in the two images to obtain the pixel difference between the frames.
[0012] In a second aspect, the present application further provides a device for identifying elements in a video. The device comprises:
[0013] A sequence acquisition module is used to acquire a video frame sequence consisting of at least a portion of video frames of a target video, wherein each video frame in the video frame sequence is arranged according to a time sequence in the target video;
[0014] A first feature extraction module is used to extract the component element features of each video frame in the video frame sequence;
[0015] A second feature extraction module is used to extract short-term temporal features of the video frame sequence;
[0016] A third feature extraction module is configured to extract long temporal features of the video frame sequence; the time span of the video frames matched by the short temporal features in the video frame sequence is smaller than the time span of the video frames matched by the long temporal features in the video frame sequence;
[0017] A feature fusion module is used to identify the component elements in the target video based on a feature fusion result obtained by fusing the component element features, the short-term features and the long-term features.
[0018] In a third aspect, the present application also provides a method for replacing elements in a video. The method comprises:
[0019] Obtain a target video, and obtain a replacement element for a target component element in the target video;
[0020] Based on the above-mentioned method for identifying component elements in a video, identifying the component elements in the target video;
[0021] In a case where it is identified that the component elements include the target component element, determining element filling data of the target component element in the target video;
[0022] Based on the element filling data, the target component element in the target video is replaced with the replacement element.
[0023] In a fourth aspect, the present application further provides a device for replacing elements in a video. The device includes:
[0024] A target video acquisition module is used to acquire a target video and acquire a replacement element for a target component element in the target video;
[0025] The recognition result acquisition module is used to obtain the component elements obtained by performing component element recognition on the target video from the recognition device of the component elements in the video;
[0026] an element filling data determining module, configured to, when identifying that the component elements include the target component element, determine the element filling data of the target component element in the target video;
[0027] A component element replacement module is used to replace the target component element in the target video with the replacement element based on the element filling data.
[0028] In a fifth aspect, the present application also provides a video recommendation method based on component elements. The method includes:
[0029] Based on the above-mentioned method for identifying component elements in a video, identifying component elements in a target video;
[0030] Perform tag matching on the element tags of the constituent elements and the interest tags of the video viewing objects, and filter out target objects with successful tag matching from the video viewing objects;
[0031] Push the target video to the target object.
[0032] In a sixth aspect, the present application further provides a video recommendation device based on component elements. The device comprises:
[0033] The recognition result acquisition module is used to obtain the component elements obtained by performing component element recognition on the target video from the recognition device of the component elements in the video;
[0034] a tag matching module, configured to perform tag matching on the element tags of the constituent elements and the interest tags of the video viewing objects, and filter out target objects with successful tag matching from the video viewing objects;
[0035] The video push module is used to push the target video to the target object.
[0036] In a seventh aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above methods when executing the computer program.
[0037] In an eighth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above methods when executed by a processor.
[0038] In a ninth aspect, the present application further provides a computer program product, which includes a computer program that implements the steps of the above methods when executed by a processor.
[0039] The above-mentioned method, apparatus, computer equipment, storage medium and computer program product for identifying component elements in a video take a video frame sequence composed of at least a portion of the video frames of a target video as the analysis object, ensure that each video frame in the video frame sequence retains the time sequence in the target video, accurately extract short-time sequence features and long-time sequence features representing different time spans in the video frame sequence, and based on the feature fusion of the component element features, short-time sequence features and long-time sequence features of each video frame in the video frame sequence, can identify single-frame component elements, short-time component elements and long-time component elements with different time sequence information in the target video, and then comprehensively and accurately identify the various component elements filled in the target video, thereby improving the accuracy of the component element identification results in the target video.
[0040] Furthermore, the above-mentioned method, apparatus, computer equipment, storage medium and computer program product for replacing constituent elements in the video utilize constituent elements accurately identified from the target video, and when the constituent elements include target constituent elements, determine the element filling data of the target constituent elements. By accurately identifying the target constituent elements and determining the element filling data of the target constituent elements to perform element replacement, the accuracy of replacing the target constituent elements in the target video can be effectively improved.
[0041] Furthermore, the above-mentioned component-based video recommendation method, device, computer equipment, storage medium and computer program product use the component elements accurately identified from the target video to match the video viewing object based on the element tags of the component elements, which can accurately locate the recommendation object and achieve effective recommendation of the target video. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 FIG. 1 is an application environment diagram of a method for identifying component elements in a video according to an embodiment;
[0043] Figure 2 FIG2 is an application environment diagram of a method for identifying component elements in a video according to another embodiment;
[0044] Figure 3 1 is a flow chart of a method for identifying component elements in a video according to an embodiment;
[0045] Figure 4 A schematic diagram of a video frame in which elements are filled in the left and right areas of a video in one embodiment;
[0046] Figure 5A schematic diagram of a video frame in which elements are filled in the upper and lower areas of a video in one embodiment;
[0047] Figure 6 A schematic diagram of dividing a video frame filled in the left and right areas into a 3*3 grid in one embodiment;
[0048] Figure 7 A schematic diagram of dividing a video frame filled in upper and lower areas into a 3*3 grid in one embodiment;
[0049] Figure 8 A schematic diagram of a sample image of a video frame with left and right area filling according to an embodiment;
[0050] Figure 9 A schematic diagram of a sampled image of a video frame with upper and lower area filling according to an embodiment;
[0051] Figure 10 A schematic diagram of the model structure of a scene transition recognition model in one embodiment;
[0052] Figure 11 is a schematic structural diagram of a network unit in another scene transition recognition model in one embodiment;
[0053] Figure 12 1 is a flow chart of a method for replacing component elements in a video according to an embodiment;
[0054] Figure 13 1 is a flow chart of a video recommendation method based on component elements in one embodiment;
[0055] Figure 14 1 is a flow chart of a method for identifying component elements in a video according to an embodiment;
[0056] Figure 15 is a structural block diagram of a device for identifying component elements in a video in one embodiment;
[0057] Figure 16 is a structural block diagram of a device for replacing component elements in a video according to an embodiment;
[0058] Figure 17 is a structural block diagram of a video recommendation device based on component elements in one embodiment;
[0059] Figure 18 is a diagram of the internal structure of a computer device in one embodiment;
[0060] Figure 19 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0062] The method for identifying elements in a video provided by the embodiment of the present application can be applied to Figure 1 In the application environment shown. The terminal 102 communicates with the server 104 via a network. The terminal 102 sends the target video to the server 104, and the server 104 obtains a video frame sequence composed of at least a portion of the video frames of the target video. The video frames in the video frame sequence are arranged in a time sequence in the target video. The server 104 extracts the constituent element features of each video frame in the video frame sequence, and extracts the short-time sequence features and long-time sequence features of the video frame sequence respectively, wherein the time span of the video frame matched by the short-time sequence features in the video frame sequence is shorter than the time span of the video frame matched by the long-time sequence features in the video frame sequence. The server 104 identifies the constituent elements in the target video based on the feature fusion result obtained by fusing the constituent element features, the short-time sequence features, and the long-time sequence features, and feeds back the identification result of the constituent elements to the terminal 102.
[0063] The method for identifying elements in a video provided by the embodiment of the present application can be applied to Figure 2 In the application environment shown. In response to a component element recognition event triggered for a target video, the terminal 200 obtains a video frame sequence composed of at least a portion of video frames of the target video, wherein the video frames in the video frame sequence are arranged in a time sequence in the target video. The terminal 200 extracts component element features of each video frame in the video frame sequence, and respectively extracts short-time sequence features and long-time sequence features of the video frame sequence, wherein the time span of the video frame matched by the short-time sequence feature in the video frame sequence is shorter than the time span of the video frame matched by the long-time sequence feature in the video frame sequence. The terminal 200 identifies the component elements in the target video based on a feature fusion result obtained by fusing the component element features, the short-time sequence features, and the long-time sequence features, and displays the recognition result of the component elements.
[0064] Terminals 102 and 200 may be, but are not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. IoT devices may include intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, and the like. Portable wearable devices may include smart watches, smart bracelets, head-mounted devices, and the like. Server 104 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0065] It should be noted that the methods for identifying component elements in videos in some embodiments of the present application utilize artificial intelligence technology. For example, the component element features of each video frame in a video frame sequence are extracted, and the short-term temporal features and long-term temporal features of the video frame sequence are extracted separately. Specifically, artificial intelligence technology can be used to train a feature extraction model, and feature extraction is performed based on the feature extraction model.
[0066] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0067] Artificial intelligence technology is a comprehensive discipline covering a wide range of fields, encompassing both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technologies, operating / interactive systems, and mechatronics. Artificial intelligence software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning. It is understood that the prediction models used in some of the embodiments of this application are equivalent to neural network models trained using machine learning techniques.
[0068] Computer vision (CV) is the science of making machines "see." Specifically, it refers to using cameras and computers to replace the human eye in identifying and measuring objects, and then further processing the images to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, and smart transportation. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0069] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0070] Deep learning (DL) involves learning the inherent patterns and representational hierarchies of sample data. The information gained from this learning process is highly helpful in interpreting data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to acquire the same analytical and learning capabilities as humans, enabling them to recognize data such as text, images, and sounds.
[0071] The method for identifying constituent elements in a video in the embodiments of the present application can be applied to any video understanding scenario, such as a constituent element replacement scenario, and can be applied to video classification, video recommendation, etc. It is even possible to extract constituent elements based on the method for identifying constituent elements in a video of the present application, thereby realizing tasks such as video clustering and video retrieval.
[0072] In one embodiment, Figure 3 As shown, a method for identifying elements in a video is provided, and the method is described by taking the application of the method to a computer device as an example. The computer device may be Figure 1 The server in Figure 2The terminal in the , including the following steps:
[0073] Step 302: Obtain a video frame sequence consisting of at least a portion of video frames of the target video, wherein the video frames in the video frame sequence are arranged in a time sequence in the target video.
[0074] The target video is a video composed of multiple components. Specifically, it can be an unmodified original video directly obtained, or a composite video obtained by synthesizing the original video and the components. The original video can be a video directly acquired by a video acquisition device. The video acquisition device can be a camera, a mobile phone, or other device with video acquisition function. The video acquisition method can be shooting or screen recording. For example, the original video can be a video shot by a mobile phone, a video recorded by a mobile phone, etc. The components can be the original components of the target video, or one of the original components or filler elements added to the original video. The components can be static elements or dynamic elements. Static elements can be static images, such as stickers, text, etc. Dynamic elements can be changing elements, such as animated images, transition animations, audio, etc. Dynamic elements can be automatically played elements, or interactive controls that can be triggered to change the display state, such as red envelope controls.
[0075] The video frame sequence may include all video frames of the target video. Subsequent feature extraction and feature fusion processing of all video frames can make the component element identification result obtained more accurate. The video frame sequence may also include a portion of the video frames of the target video. Compared with the processing method of adding all video frames to the video frame sequence for data processing, extracting a portion of the video frames of the target video for data processing can reduce the subsequent data processing volume and improve data processing efficiency. The total number of frames of the partial video frames extracted from the target video frames can be determined according to the duration of the target video. The number of extracted video frames can be positively correlated with the duration of the target video. The longer the duration of the target video, the more video frames are extracted.
[0076] Each video frame in the target video carries a corresponding timestamp. Each timestamp is sequential and can be used to indicate the position of the video frame in the target video. The video frames arranged in a time sequence within the target video can specifically be arranged in chronological order based on the timestamps they carry.
[0077] Specifically, for the case where the video frame sequence includes a portion of video frames of the target video, the computer device extracts a portion of video frames from the target video to construct a video frame sequence. Extracting a portion of video frames from the target video can be performed by screening out a portion of video frames from all video frames of the target video at the same time interval. By arranging the screened video frames in a time sequence in the target video, it can be ensured that the video frame sequence is arranged at equal time intervals, thereby improving the accuracy of the short-term temporal features and long-term temporal features subsequently extracted. It will be understood that in other embodiments, the computer device extracts a portion of video frames from the target video to construct a video frame sequence, or it can be performed by randomly selecting a portion of video frames from all video frames of the target video.
[0078] In the case where the frame sequence includes all video frames of the target video, the computer device can directly perform frame processing on the target video to obtain a video frame sequence consisting of video frames arranged in a time sequence in the target video.
[0079] Step 304: extract the component element features of each video frame in the video frame sequence.
[0080] The component element features are features used to describe the component element identification results of a video frame. A computer device performs component element identification processing on a video frame, and the results obtained are features used to describe the component element identification results of the video frame. The component element features can be identified using a neural network model capable of identifying image component elements, or can be determined by matching a video frame with an image component element template. The neural network model capable of identifying image component elements can be a model trained using filled images labeled with component elements.
[0081] Specifically, the component element feature extraction processes for each video frame in the video frame sequence are independent of each other and do not interfere with each other. The computer device can synchronously extract component element features for each video frame in the video frame sequence through multiple data processing processes that are the same as the number of video frames in the video frame sequence. The computer device can also extract component element features for a portion of the video frames in the video frame sequence each time based on at least two data processing processes, wherein the number of data processing processes is less than the number of video frames in the video frame sequence. The computer device can also sequentially extract component element features for each video frame in the video frame sequence through a single data processing process. The number of data processing processes executed simultaneously can be determined based on the computing power that the computer device can provide, or can be determined by pre-set processing parameters.
[0082] In a specific application, a computer device uses each video frame in a video frame sequence as input data for a neural network model with image component element recognition function, so that the neural network model extracts the component element features of each video frame respectively. Based on the powerful data processing capabilities of the trained neural network model, the component element features of each video frame in the video frame sequence can be obtained quickly and accurately.
[0083] Step 306 , extracting short temporal features and long temporal features of the video frame sequence respectively; the time span of the video frames matched by the short temporal features in the video frame sequence is smaller than the time span of the video frames matched by the long temporal features in the video frame sequence.
[0084] Among them, short-term features are the results of feature extraction processing on a smaller amount of time-series data, and correspondingly, long-term features are the results of feature extraction processing on a larger amount of time-series data. Short-term features are used to characterize objects with short-term characteristics in the target video, such as transition animations and inserted red envelopes in the target video, while long-term features are used to characterize objects with long-term characteristics in the target video, such as background fills in the target video.
[0085] In a video frame sequence, the data used to extract short-term temporal features and the data used to extract long-term temporal features are both continuous video frames in the video frame sequence. The video frames in a video frame sequence are arranged in time sequence. The more continuous video frames there are, the longer the time span is. The time span of the video frames matched by the short-term temporal features in the video frame sequence is smaller than the time span of the video frames matched by the long-term temporal features in the video frame sequence. For example, the temporal features corresponding to two continuous video frames in a video frame sequence are short-term temporal features, and the temporal features corresponding to ten continuous video frames in the video frame sequence are long-term temporal features.
[0086] Specifically, short-time series features and long-time series features can be realized based on different extraction methods. The features obtained based on the short-time series feature extraction method are short-time series features, and the features obtained based on the long-time series feature extraction method are long-time series features. The short-time series feature extraction method includes:
[0087] To extract short-term temporal features, the computer device can group a preset number of consecutive video frames in a video frame sequence into groups, perform feature extraction on each group, and then fuse the feature extraction results for each group to obtain the short-term temporal features of the video frame sequence. To extract long-term temporal features, the computer device can identify the scene composition of the target video from the video frame sequence and identify the long-term temporal features of the video frame sequence based on the distribution characteristics of each scene in the target video or video frame sequence.
[0088] Step 308 : Identify the component elements in the target video based on the feature fusion result obtained by fusing the component element features, the short-term features, and the long-term features.
[0089] The number of component element features is the same as the number of video frames in the video frame sequence, or is a multiple of the number of video frames in the video frame sequence. Specifically, the number of component element features of each video frame in the video frame sequence is the same, and each video frame has at least one component element feature. In a specific application, a computer device performs image depth feature extraction and component element prediction processing on a video frame to obtain image depth features and predicted features. The component element features can be a feature combination composed of image depth features and predicted features, or can be a feature obtained by fusion of image depth features and predicted features.
[0090] Specifically, the computer device can align the features of each component element, the short-time features and the long-time features, and then perform feature fusion processing on the aligned feature data. Since the features of each component element can represent the components contained in each video frame, the short-time features can represent the objects with short-time distribution characteristics in the target video, and the long-time features can represent the objects with long-time distribution characteristics in the target video, the feature fusion result obtained based on the combination of the three can be used to identify single-frame component elements that appear in a single frame in the target video, short-time component elements that appear in fewer video frames, and long-time component elements that appear continuously in more video frames, thereby comprehensively and accurately identifying the various component elements filled in the target video.
[0091] In one embodiment, the computer device can normalize the feature dimensions of each component element feature, short time series feature, and long time series feature to obtain component element features, short time series features, and long time series features with the same feature dimensions, and then perform feature fusion on the component element features, short time series features, and long time series features with the same feature dimensions to obtain a feature fusion result, thereby improving the accuracy of the feature fusion result.
[0092] The above-mentioned method for identifying component elements in a video takes a video frame sequence composed of at least a part of the video frames of the target video as the analysis object, ensures that each video frame in the video frame sequence retains the time sequence in the target video, accurately extracts short-time sequence features and long-time sequence features representing different time spans in the video frame sequence, and based on the feature fusion of the component element features, short-time sequence features and long-time sequence features of each video frame in the video frame sequence, it can ensure the identification of single-frame component elements, short-time component elements and long-time component elements with different time sequence information in the target video, and then comprehensively and accurately identify the various component elements filled in the target video, thereby improving the accuracy of the component element identification results in the target video.
[0093] In one embodiment, extracting short-term temporal features contained in a video frame sequence includes:
[0094] Image difference data of each group of adjacent video frames in the video frame sequence is obtained; based on the image difference data, element features of the adjacent video frames are extracted to obtain short-term features composed of the element features of each group of adjacent video frames.
[0095] Adjacent video frames refer to two adjacent video frames in the video frame sequence, sorted by time in the target video. For example, if there are 10 video frames arranged in time, numbered 1-10, then video frame 1 and video frame 2 form a pair of adjacent video frames, video frame 2 and video frame 3 form a pair of adjacent video frames, and so on, for a total of 9 pairs of adjacent video frames.
[0096] Image difference data refers to the similarity between two video frames in a group of adjacent video frames, and the image difference data includes the differences between different positions in the two video frames. Image difference data can be obtained by image block level difference comparison or by pixel level difference comparison. Image block level difference comparison refers to directly dividing the entire image into multiple image blocks, and comparing the image data based on the image blocks. The image data can be at least one of the image clarity, image brightness and other data. Correspondingly, pixel level difference comparison refers to comparing the pixel data of pixel units at the same position in the two images. The pixel data can be at least one of the pixel three primary color values, grayscale values and other data. Since the time interval between the two video frames in each group of adjacent video frames is small, the element features of each group of adjacent video frames can be used as short-term features of the video frame sequence.
[0097] Specifically, the computer device groups the video frames in the video frame sequence according to adjacent video frames to obtain multiple groups of adjacent video frames. For the two video frames contained in each group of adjacent video frames, the image difference data is calculated, and then based on the image difference data of each group of adjacent video frames, the elements represented by the image data whose differences meet the element change conditions in the adjacent video frames are identified to obtain the element features of the adjacent video frames. The element features of the adjacent video frames can represent the elements that appear in the target video for a short time. The element features of the adjacent video frames are a short-term feature of the video frame sequence. The short-term feature of the video frame sequence includes the element features of each group of adjacent video frames in the video frame sequence.
[0098] In this embodiment, the computer device obtains image difference data for each group of adjacent video frames in the video frame sequence. Through the image difference data, it can accurately identify the element features of short-time appearing elements existing in at least two consecutive video frames of the video frame sequence, and use the element features of the adjacent video frames as part of the short-time features of the video frame sequence, thereby achieving accurate expression of the short-time features of the video frame sequence.
[0099] Furthermore, in the case where the video frames in the video frame sequence are filtered from the target video at a certain time interval, there are still video frames that have not been filtered out between adjacent video frames in the video frame sequence, which means that there is a certain time interval between adjacent video frames in the video frame sequence. The image data whose differences in adjacent video frames meet the element change conditions represents an element in the target video, thereby realizing accurate identification of the element and obtaining accurate element features.
[0100] In one embodiment, obtaining image difference data of each group of adjacent video frames in a video frame sequence includes: for each group of adjacent video frames in the video frame sequence, comparing two video frames in the adjacent video frames in pixel units to obtain inter-frame pixel differences; and determining image difference data of the adjacent video frames based on the inter-frame pixel differences.
[0101] The pixel unit is a component of the image, which can be a pixel point or a pixel block. In a specific application, the comparison can be performed by pixel unit, by pixel point, by pixel block, or by a combination of pixel point and pixel block comparison. Comparing two video frames by pixel unit means comparing the pixel data of the pixel units at the same position in the two images. The pixel data includes RGB (optical primary colors, R represents red, G represents green, and B represents blue) data, such as at least one type of data from an RGB histogram, brightness data, and grayscale data. The comparison result of the RGB data of the pixel unit is used to indicate a change in the displayed color; the comparison result of the brightness data of the pixel unit is used to indicate a change in the brightness of the image; and the comparison result of the grayscale data of the pixel unit is used to indicate a change in the grayscale of the image.
[0102] Specifically, the computer device obtains pixel data of each pixel unit in each video frame contained in adjacent video frames, compares similar pixel data of pixel units at the same position in the two video frames, and determines the pixel difference between adjacent video frames based on the difference distribution of similar pixel data.
[0103] In this embodiment, by performing pixel data comparison on a pixel basis and utilizing the pixel unit as the smallest constituent unit of an image, fine-grained data comparison is achieved, and inter-frame pixel differences between adjacent video frames can be accurately obtained.
[0104] In one embodiment, two video frames in adjacent video frames are compared in pixel units to obtain pixel differences between frames, including: based on the same sampling parameters, local image sampling is performed on the two video frames in the adjacent video frames to obtain captured images contained in each of the two video frames; the sampled images at the same position in the two video frames are compared in pixel units to obtain pixel differences between frames.
[0105] Among them, the sampling parameters are used to characterize the specific image sampling method used for the image. Using the same sampling method for adjacent video frames can ensure that each sampled image obtained by sampling has an object for pixel unit comparison, avoiding redundant sampled images and causing waste of data processing resources. Local image sampling is the sampling image collected after sampling the video frame, which constitutes a part of the video frame. The local image sampling process of each video frame in adjacent video frames is independent of each other and does not interfere with each other. The computer device can synchronously perform local image sampling on two video frames in adjacent video frames through two local image sampling processes. The computer device can also perform local image sampling on two video frames in adjacent video frames in sequence through a single local image sampling process. The specific paradigm of local image sampling can be determined based on the computing power that the computer device can provide, or it can be determined by pre-configured sampling parameters.
[0106] Specifically, the computer device may compare pixel data of sampled images at the same location in two video frames on a pixel-by-pixel basis. The pixel data may include at least one type of data: RGB values, brightness values, and grayscale values. In one specific application, the computer device performs local image sampling on adjacent video frames to obtain the sampled images contained therein, determines the pixel data of each pixel unit in the sampled pattern, treats the sampled images at the same location in the two video frames as a comparison group, compares similar pixel data of pixel units at the same location in the same group of sampled images, and determines the pixel difference between the adjacent video frames based on the difference distribution of the similar pixel data.
[0107] In this embodiment, on the one hand, local image extraction of adjacent video frames can be achieved through local image sampling, which can reduce the number of pixel units involved in the comparison and improve data processing efficiency; on the other hand, fine-grained data comparison is achieved by comparing pixel data by pixel unit, and the pixel differences between adjacent video frames can be accurately obtained.
[0108] In one embodiment, the method for identifying component elements in a video further includes: identifying static elements and dynamic elements in adjacent video frames to obtain element category information.
[0109] The components of a video are classified into two categories: static elements and dynamic elements. Static elements, such as single stickers or text, are elements whose display remains constant across multiple frames of the target video. Dynamic elements, such as animated graphics, transitions, audio, and interactive components, are elements whose display varies across multiple frames of the target video. Specifically, for video frames in a video frame sequence, the differences in their display within the video frames can be used to distinguish between static and dynamic elements within the video frames, thereby obtaining element category information. These differences in the display between static and dynamic elements within the video frames include, for example, differences in clarity or changes in optical flow. Specifically, for static and dynamic elements added to the target video, the clarity of the static elements within the video frames will be higher than that of the dynamic elements, and the optical flow changes of the static elements within the video frames will be smaller than those of the dynamic elements.
[0110] Furthermore, based on the image difference data, element features of adjacent video frames are extracted to obtain short-term features composed of the element features of each group of adjacent video frames, including: based on the image difference data of adjacent video frames and the element category information of adjacent video frames, element features of adjacent video frames are extracted to obtain short-term features composed of the element features of each group of adjacent video frames.
[0111] For static elements, the specific elements they represent are located in areas where there is no image difference data in adjacent video frames. For dynamic elements, the specific elements they represent are located in areas where there is image difference data in adjacent video frames.
[0112] Specifically, the computer device determines the distribution data of various elements in the video frames based on the positions of the image difference data of adjacent video frames in the video frames and the data changes corresponding to the element category information of the adjacent video frames, and extracts element features of the adjacent video frames based on the element distribution data to obtain short-term features composed of the element features of each group of adjacent video frames.
[0113] In this embodiment, by combining the image difference data and element category information of adjacent video frames, the differences in the display results of static elements and dynamic elements in the video frames can be taken into account, and the element features of the elements in adjacent video frames can be effectively extracted to improve the accuracy of the element features.
[0114] In one embodiment, based on image difference data of adjacent video frames and element category information of adjacent video frames, element features of adjacent video frames are extracted to obtain short-term features composed of element features of each group of adjacent video frames, including: obtaining a target element category obtained by performing element category matching on image difference data of adjacent video frames; extracting initial element features of adjacent videos based on the image difference data; and when the target element category successfully matches the element category information, performing feature fusion on the category features of the target element category and the initial element features to obtain short-term features composed of element features of adjacent video frames.
[0115] Specifically, for regions consisting of pixel units that do not contain image difference data in adjacent video frames, the target element category matched by the target element located in that region is a static element; for regions consisting of pixel units that contain image difference data in adjacent video frames, the target element category matched by the target element located in that region is a dynamic element. For the same region in a video frame, when the matched target element category is the same as the element category represented by the element category information, the category features of the target element category are fused with the initial element features to obtain the element features of the adjacent video frame.
[0116] In this embodiment, by performing element category matching on image difference data, if the target element category is successfully matched, the accuracy of the description of the element features of adjacent video frames can be improved by fusing the category features with the initial element features. In one embodiment, static and dynamic elements are identified in adjacent video frames to obtain element category information, including: for any of the adjacent video frames, obtaining element category information based on the difference in image clarity of each region in the video frame.
[0117] Among them, the element categories represented by the element category information include static elements and dynamic elements. Dynamic elements will change in multiple video frames of the target video, while static elements include elements that only appear in one video frame, or elements that are stable and unchanged in multiple video frames. The stability of static elements is stronger than that of dynamic elements. From the perspective of image clarity of video frames, the clarity difference between static elements and dynamic elements is that the image clarity of static elements is greater than that of dynamic elements.
[0118] Specifically, if the element categories of two adjacent video frames are the same, the element category information of one of the adjacent video frames can be directly obtained based on the element category analysis result of the adjacent video frames. The computer device performs image similarity difference analysis on one of the adjacent video frames to obtain the element category information of the video frame.
[0119] Each region in the video frame is divided based on the region composition represented by the target video's screen boundary filling method. The purpose of filling the target boundary is to make the initial video be in the middle of the target video. Specifically, the screen boundary filling method can also be used to fill the left and right areas of the video, such as Figure 4 As shown, the video is divided into three areas, left, middle and right. The screen boundary filling method can be to fill the upper and lower areas of the video screen, such as Figure 5 As shown, the video is divided into three regions: upper, middle, and lower. Generally speaking, the elements in the padded region are static, while the elements in the region containing the original video are dynamic. Within each region, the element category information can be determined based on the image clarity differences by identifying the image clarity of each region. The element category information for adjacent video frames is the sum of the element category information for each region.
[0120] In this embodiment, since in the video frame sequence, except for the video frames at the first and last positions, each of the remaining video frames has two adjacent video frames, therefore, except for the video frames at the first and last positions, each of the remaining video frames belongs to two different groups of adjacent video frames. By comparing the image clarity difference of the previous video frame or the next video frame in each group of adjacent video frames according to each area in the video frame, on the one hand, the clarity difference between static elements and dynamic elements can be utilized to quickly and accurately obtain element category information. On the other hand, only one video frame in the adjacent video frames is processed, which can effectively reduce the data processing amount and improve data processing efficiency.
[0121] In one embodiment, element category information is determined based on the image clarity differences of each image region in a video frame, including: obtaining clarity evaluation data of the regional image of each region in the video frame; and determining element category information based on the differences in the clarity evaluation data of each regional image.
[0122] The clarity evaluation data refers to the evaluation result obtained by calculating the clarity of an image. Specifically, clarity evaluation can be implemented based on any of the Tenengrad (a clarity evaluation function) gradient method, the Laplacian gradient method, and the variance method. The Tenengrad gradient method uses the Sobel operator to calculate the gradients in the horizontal and vertical directions. For the same image, the higher the gradient value, the clearer the image. The Tenengrad gradient method differs from the Laplacian gradient method in the operator used. The Laplacian gradient method uses the Laplacian operator. Variance is a measure used in probability theory to examine the degree of dispersion between a set of discrete data and the expected value (i.e., the mean value). A larger variance indicates greater deviation between the data set, with some data within the set being larger and some smaller, resulting in an uneven distribution. A smaller variance indicates less deviation between the data set, with the data within the set being evenly distributed and of similar size. When applied to the evaluation of image clarity, the grayscale difference between data of a clear image is greater than that of a blurred image, that is, the variance will be larger. The clarity of the image can be measured by the variance of the image grayscale data. The larger the variance, the better the clarity.
[0123] Specifically, the computer device extracts the regional image of each area in the video frame according to the regional division of the video frame; performs image gradient calculation on each regional image to obtain clarity evaluation data of each regional image; compares the clarity evaluation data of each regional image to obtain the difference in the clarity evaluation data of each regional image; determines the element category of the elements in the area where the regional image with relatively smaller clarity is located as a dynamic type element, and determines the element category of the elements in the area where the regional image with relatively larger clarity is located as a static type element, thereby obtaining accurate element category information.
[0124] In this embodiment, clarity evaluation data is obtained by evaluating the clarity of the regional image of each region in the video frame, which can achieve rapid and accurate judgment of the clarity of each region in the video frame. Furthermore, image gradient calculation based on the Tenengrad gradient method or the Laplace gradient method can be used to quickly generate clarity evaluation data for each regional image, ensuring the accuracy of the clarity evaluation data. Then, by comparing the differences in the clarity evaluation data of each region, element category information of the video frame can be quickly and accurately obtained.
[0125] In a specific embodiment, static elements and dynamic elements are identified in adjacent video frames to obtain element category information, including: for any video frame in the adjacent video frames, according to the area division of the video frame, extracting the regional image of each area in the video frame; performing image gradient calculation on each regional image to obtain clarity evaluation data of each regional image; based on the difference in clarity evaluation data of each regional image, determining the element category information according to the clarity difference characteristics between the static elements and the dynamic elements.
[0126] By comparing the image clarity differences of each area image in the video frame, on the one hand, we can use the clarity difference between static elements and dynamic elements to quickly and accurately obtain element category information. On the other hand, only processing one video frame among the adjacent video frames can effectively reduce the data processing amount and improve data processing efficiency.
[0127] In one embodiment, the processing of determining element category information includes: gridding the video frame to obtain multiple grid images contained in the video frame; grouping the grid images in the same area according to the area division of the video frame; calculating a first clarity difference between the grid images in the same group and a second clarity difference between the grid images in different groups; and determining the element category information of the video frame based on the first clarity difference and the second clarity difference and according to the clarity difference characteristics between the static type elements and the dynamic type elements.
[0128] Specifically, the video frame is divided into N*M (N, M are both positive integers) grids. The values of N and M can be determined according to the allowed splicing methods in the process of fusing the initial video components to obtain the target video. For example, when the target video allows three-segment filling (top, middle, bottom or left, middle, right), the video frame can be divided into 9 3*3 grids. Figure 6 and Figure 7 As shown, the upper three grids can be divided into the same group, the middle three grids can be divided into the same group, and the lower three grids can be divided into the same group. Then, the first clarity difference between the grid images in the same group and the second clarity difference between the grid images in different groups are calculated respectively. Then, based on the larger difference between the first clarity difference and the second clarity difference, the element category information of each area is determined.
[0129] Specifically, if the grouping is correct, the first clarity difference between grid images in the same group will be significantly smaller than the second clarity difference between grid images in different groups. In this case, the element category information for each region can be determined based on the second clarity difference. When the first clarity difference between grid images in the same group is significantly larger than the second clarity difference between grid images in different groups, it indicates a grouping error. In this case, the correct grouping method should be to divide the three grids on the left, the three grids in the middle, and the three grids on the right into the same group. The first clarity difference is the difference between the top three grids, representing the image clarity difference between the topmost grids according to the left, middle, and right division. Therefore, the element category information for each region can be determined directly based on the first clarity difference.
[0130] In this embodiment, the clarity differences of various areas in the video frame are determined by grid division and grouping. Based on the characteristic that the clarity differences of each grid image in the same area are small, the area division method of the video frame can be accurately identified, and the element category information of each area corresponding to the area division method can be obtained, thereby improving the object representation accuracy of the element category information.
[0131] In another embodiment, the processing procedure for determining element category information includes: performing local image sampling on a video frame to obtain multiple sampled images contained in the video frame; dividing the sampled images in the same area into a group according to the area division of the video frame; calculating a first clarity difference between the sampled images in the same group and a second clarity difference between the sampled images in different groups; and determining the element category information of the video frame based on the first clarity difference and the second clarity difference and according to the clarity difference characteristics between static elements and dynamic elements.
[0132] Partial image sampling involves sampling a video frame, and the resulting sampled image forms part of the video frame. The sampling parameters used for partial image sampling for different video frames can be the same or different. These parameters characterize the specific sampling method used for the image. Using the same sampling method for all video frames simplifies the sampling process and improves sampling efficiency.
[0133] Performing local image sampling on the video frame can collect local images from different positions of the video frame. Local image sampling can be random sampling or matrix sampling. Taking matrix sampling as an example, Figure 8 and Figure 9As shown, 9 3*3 sampling images can be collected from the video frame through 3*3 matrix sampling parameters, and then each sampling image is grouped, and the first clarity difference between the grid images in the same group and the second clarity difference between the grid images in different groups are calculated respectively. Then, based on the larger difference between the first clarity difference and the second clarity difference, the element category information of each area is determined.
[0134] Specifically, when grouping is correct, the first clarity difference between the sample images in the same group will be significantly smaller than the second clarity difference between the sample images in different groups. In this case, the element category information of each region can be determined based on the second clarity difference. When the first clarity difference between the sample images in the same group is significantly larger than the second clarity difference between the sample images in different groups, it indicates a grouping error. In this case, the correct grouping method should be to divide the three sample images on the left, the three sample images in the middle, and the three sample images on the right into the same group. The first clarity difference is the difference between the top three sample images, representing the image clarity difference between the topmost sample images according to the left, middle, and right division method. Therefore, the element category information of each region can be determined directly based on the first clarity difference.
[0135] In this embodiment, by sampling and grouping local images, the clarity differences between regions within a video frame are determined. Based on the relatively small clarity differences between sampled images within the same region, the correct grouping method and the corresponding video frame region division method are determined. Element category information corresponding to the region division method is then obtained, improving the accuracy of object representation in this element category information. Furthermore, compared to grid-based analysis, analysis of sampled images obtained through local image sampling can reduce the number of pixel units that comprise the image, minimizing the amount of data processing and improving data processing efficiency.
[0136] In one embodiment, extracting long-term temporal features contained in a video frame sequence includes:
[0137] Identify scene transition features in video frame sequences, perform feature conversion on the scene transition features, and obtain long-term temporal features contained in the video frame sequences.
[0138] The scene transition feature characterizes the distribution of scene boundary frames within a video frame sequence. Scene boundary frames are the video frames corresponding to scene transitions between multiple scenes within the target video. The distribution of scene boundary frames within a video frame sequence can characterize the starting positions of the video segments corresponding to each of the multiple scenes in the target video, as well as the scene transition patterns between these segments.
[0139] Specifically, when two scenes are converted, the number of scene boundary frames that are matched can be one or more. When the scene boundary frame is 1 frame, the scene splicing method used to represent the two scenes before and after the conversion is hard splicing, that is, directly connecting the video clip of the first scene with the video clip of the second scene; when the scene boundary frame is multiple frames, the scene splicing method used to represent the two scenes before and after the conversion is soft splicing, which is also called transition splicing. The transition animation used in the transition splicing can be realized through component elements. By identifying the scene conversion features in the video frame sequence, the component elements that may exist in the target video and the starting position of the video clip corresponding to each scene can be identified. The computer device obtains the long-term features contained in the video frame sequence by performing feature conversion on the scene conversion features.
[0140] In this embodiment, by identifying the scene transition features in the video frame sequence, the distribution of scene boundary frames in the video frame sequence can be determined, the starting positions of the video clips corresponding to multiple scenes in the target video, and the scene transition methods between the video clips can be identified. Based on these, the scene transition features are subjected to feature conversion, and the long-term temporal features contained in the video frame sequence can be accurately obtained.
[0141] In one embodiment, extracting long-term temporal features contained in a video frame sequence includes:
[0142] Based on the scene transition recognition model, scene transition features are extracted from the video frame sequence to obtain the scene transition features in the video frame sequence; the scene transition recognition model is trained based on transition videos carrying scene transition labels; feature conversion is performed on the scene transition features to obtain the long-term temporal features contained in the video frame sequence.
[0143] The scene transition recognition model is a deep learning model, which can be a scene transition recognition model trained based on transition videos carrying scene transition labels, such as a TransNet model or a TransNet V2 model.
[0144] Specifically, the input data of the TransNet model is a video frame sequence of length N, and the size of each video frame is 48x27x3. The video frame sequence will pass through four dilated convolutional layers with different expansion rates in the time dimension to reduce the number of training parameters. The output features of the four dilated convolutional networks are then fully connected, or fully connected and max pooled. Finally, after two layers of full connection and logistic regression layers, an N×2N feature vector is output. The output feature vector can be used to characterize whether each input video frame is a scene boundary frame.
[0145] The TransNet V2 model is improved on the basis of TransNet. The overall network structure of the TransNet V2 model is as follows: Figure 10 As shown, it includes 6 network units, where the structure of the network unit is as follows Figure 11 As shown, each network unit is constructed through a series of spatial and temporal convolutions to reduce overfitting. The TransNet V2 model input is a video frame sequence, and the model compresses each video frame to a uniform small size of 48x27x3. TransNetV2 integrates RGB histogram features into its features to enhance their expressiveness. TransNetV2 adds a learnable similarity module after three mean pooling operations. By concatenating RGB histogram features with learnable similarity features, and then performing fully connected feature transformations based on the concatenated features, it obtains the long-term temporal features contained in the video frame sequence, suitable for identifying video transitions.
[0146] In this embodiment, based on the scene transition recognition model obtained by training the transition video with scene transition labels, scene transition features are extracted from the video frame sequence, which can effectively utilize the efficient data processing capabilities of the scene transition recognition model to quickly and accurately obtain the scene transition features in the video frame sequence.
[0147] In one embodiment, feature conversion is performed on scene transition features to obtain long-term temporal features contained in a video frame sequence, including: performing maximum pooling processing on the scene transition features of the video frame sequence to obtain pooled features; performing feature connection on the pooled features based on a fully connected layer with a feature discarding processing function to obtain fully connected features; and performing dimensionality compression and feature fuzzy transformation on the fully connected features to obtain long-term temporal features contained in the video frame sequence.
[0148] Specifically, when the scene transition recognition model performs feature conversion on the scene transition features, the corresponding specific processing levels are as follows: Figure 10 As shown in the box in the lower left corner, based on the maximum pooling layer, the N-dimensional scene conversion features outputted previously are subjected to maximum pooling processing to obtain pooled features. Then, based on the dense layer (Dense+Dropout) including feature discarding processing, some features of the pooled features are discarded and fully connected to obtain fully connected features. Based on the compression layer (Flatten), the fully connected features are dimensionally compressed to obtain one-dimensional compressed features. Based on the feature fuzzy transformation layer (Feature-dim-transform), the compressed features are subjected to feature fuzzy transformation to obtain long-term features.
[0149] In this embodiment, the scene transition recognition model described above can accurately extract scene transition features. By adjusting the model hierarchical structure, the scene transition features can be further extracted and converted to obtain long-term features for feature fusion, thereby achieving accurate extraction of long-term features.
[0150] In one embodiment, extracting the component element features of each video frame in the video frame sequence includes: performing image feature extraction and component element prediction on each video frame in the video frame sequence to obtain the component element features of each video frame.
[0151] Among them, image features refer to the vectorized expression of the results obtained by extracting the information contained in the image, and component element prediction refers to the processing process of predicting the component elements contained in the image based on image features. The result obtained by component element prediction can be component element features, and component element features include category feature data of the elements in the video frame being a specific component element and position feature data in the video frame.
[0152] Furthermore, deep feature extraction and component element prediction are performed on each video frame in the video frame sequence to obtain the component element features of each video frame, which can be obtained based on an image processing model trained with a filled image carrying the component element identifier. The image processing model can specifically be an image classification model or an image depth model. The image depth model includes any one of the BiT (Big Transfer, a pre-trained residual network that serves as the starting point for any visual task) model, the ViT (vision transformer, a model based on the attention mechanism) model, the EfficientNet (a model based on dilated convolution) model, or the Swin model. The image depth model can extract the feature information of the deeper level components in the video frame, thereby improving the accuracy of the extracted component element features.
[0153] In a specific application, taking the image depth model as the BiT model as an example, the training process of the image depth model includes: obtaining a BiT model that has been pre-trained based on a large-scale pre-training corpus; adjusting the parameters of the BiT model based on a filled image carrying component element identifiers to obtain an image depth model.
[0154] Among them, compared with other image depth models, the main feature of the BiT model is that it is optimized for pre-training, using a larger pre-training corpus, and replacing the batch normalization (BN) in the residual network model (ResNet) with group normalization (GN) and weight standardization in the pre-training stage, reducing the impact of the batch size of the residual network model on the pre-training process, and based on the hyper rule mechanism (HyperRule) to reduce the parameter adjustment work in the actual application field, that is, the fine-tuning stage, so that the representation ability of the BiT model is greatly improved through pre-training optimization. Among them, the hyper rule mechanism is a mechanism used to make the BiT model perform hyperparameter adjustment based only on high-level dataset features, such as image resolution and the number of labeled samples, which effectively reduces the cost of task adaptation.
[0155] When applied to scenarios involving component element features, only a relatively small number of labeled samples with component element identifiers are needed for fine-tuning model hyperparameters to achieve good results. Specifically, the number of padded images with component element identifiers used as training samples for scenarios involving component element features is significantly different from the number of pre-training corpora used in BiT model pre-training. The number of pre-training corpora can be several orders of magnitude larger than the padded images used as training samples.
[0156] In this embodiment, by using the trained BiT model to extract the features of the constituent elements, on the one hand, it is possible to quickly train based on the pre-trained BiT model and obtain a BiT model that can be applied to the scene of the constituent element features at a lower cost. On the other hand, the BiT model, as an image depth model, can extract deeper feature information in the video frame, thereby improving the accuracy of the extracted constituent element features.
[0157] In a specific application scenario, such as Figure 12 As shown, a method for replacing component elements in a video is also provided. The replacement method is implemented based on a method for identifying component elements, and specifically includes:
[0158] Step 1202: Obtain a target video and obtain a replacement element for a target component element in the target video;
[0159] Step 1204: obtaining a video frame sequence consisting of at least a portion of the video frames of the target video, wherein the video frames in the video frame sequence are arranged in a time sequence in the target video;
[0160] Step 1206, extracting the component element features of each video frame in the video frame sequence;
[0161] Step 1208: extracting short-term temporal features and long-term temporal features of the video frame sequence respectively; the time span of the video frames matched by the short-term temporal features in the video frame sequence is smaller than the time span of the video frames matched by the long-term temporal features in the video frame sequence;
[0162] Step 1210: identifying the component elements in the target video based on the feature fusion result obtained by fusing the component element features, the short-term features, and the long-term features;
[0163] Step 1212: when the identified component elements include the target component element, determining element filling data of the target component element in the target video;
[0164] Step 1214: Based on the element filling data, the target component element in the target video is replaced with the replacement element.
[0165] The target component element in the target video is the component element that the user can directly identify by playing the target video and wants to replace with other components. The replacement element is used to replace the target component element. To ensure the successful replacement of the component element, the element attributes of the replacement element should be the same as the element attributes of the target component element in the target video. For example, when the element attribute of the target video element is a static element, the element attribute of the replacement element should also be a static element. When the element attribute of the target video element is a transition animation, the element attribute of the replacement element should also be a transition animation.
[0166] Steps 1204 to 1210 are steps for implementing the method for identifying component elements in a video. The specific implementation process can be implemented using the various embodiments of the method for identifying component elements in a video described above, thereby obtaining an identification result of the component elements in the target video. Specifically, the identification result of the component elements in the target video can include information about all component elements in the target video, including the category of the component elements, the position of the component elements in the target video, and the position of the component elements in the target video. The filling position includes the video frame to be filled and the position within the filled video frame.
[0167] In the case where the constituent elements include the target constituent elements, the computer device can obtain information such as the element category of the target constituent elements and the filling position of the target constituent elements in the target video from the constituent element information. Based on the element category and filling position of the target constituent elements, the element filling data of the target constituent elements in the target video can be obtained. Based on the element filling data, the target constituent elements can be separated from the target video, and then the replacement elements can be filled into the target video according to the inverse process of the separation process of the target constituent elements, so as to realize the replacement of the target constituent elements in the target video and obtain the target video filled with the replacement elements.
[0168] In a specific application, taking the target video as an advertising video as an example, advertising videos often consist of multiple scenes. The scene composition can be identified and split in different dimensions, and the identified advertising elements can be expanded and replaced. For example, if it is recognized that the advertising elements filled in the advertising video include a promotion page, the video frame that hits the promotion page can be replaced with the new promotion page when creating a new advertisement; or after hitting the curtain tag, the original curtain background can be directly replaced. For users, they only need to upload the scene video or promotion page, or curtain background and then replace it to complete the production of the new advertising video.
[0169] In this embodiment, based on the replacement method of component elements in a video, the component elements of the target video can be accurately identified from the target video. When the component elements include the target component elements, the element filling data of the target component elements is determined. By accurately identifying the target component elements and determining the element filling data of the target component elements to perform element replacement, the accuracy of replacing the target component elements in the target video can be effectively improved.
[0170] In a specific application scenario, such as Figure 13 As shown, a video recommendation method based on component elements is also provided. The recommendation method is implemented based on a method for identifying component elements in a video, and specifically includes:
[0171] Step 1302: obtaining a video frame sequence consisting of at least a portion of video frames of a target video, wherein the video frames in the video frame sequence are arranged in a time sequence in the target video;
[0172] Step 1304, extracting the component element features of each video frame in the video frame sequence;
[0173] Step 1306: extracting short-term temporal features and long-term temporal features of the video frame sequence respectively; the time span of the video frames matched by the short-term temporal features in the video frame sequence is smaller than the time span of the video frames matched by the long-term temporal features in the video frame sequence;
[0174] Step 1308: identifying the component elements in the target video based on the feature fusion result obtained by fusing the component element features, the short-term features, and the long-term features;
[0175] Step 1310: performing tag matching on the element tags of the constituent elements and the interest tags of the video viewing objects, and filtering out target objects with successful tag matching from the video viewing objects;
[0176] Step 1312: Push the target video to the target object.
[0177] Steps 1302 to 1308 are steps for implementing the method for identifying component elements in a video. The specific implementation process can be implemented using the various embodiments of the method for identifying component elements in a video described above, thereby obtaining an identification result of the component elements in the target video. Specifically, the identification result of the component elements in the target video can include element labels corresponding to all component elements in the target video. The element labels can include element categories, element names, etc.
[0178] The video viewing object can be a user account in the video platform. The interest tag of the video viewing object can be selected from the candidate interest tags based on the user account, or can be obtained based on the analysis of the user's historical viewing videos with the user's authorization. A successful tag match can specifically be a case where the element tag of the constituent element hits the interest tag of the video viewing object. When the element tag of the constituent element successfully matches the interest tag of the video viewing object, the target objects with successful tag matching are filtered out from the video viewing object. The number of target objects can be a single user or a class of users including multiple users.
[0179] In a specific application, taking the replacement of advertising elements in an advertising video as an example, after the computer device obtains the uploaded advertising video, the advertising feature platform identifies the advertising elements in the advertising video and extracts the advertising element tags; at the same time, it can also obtain the interest tags entered by the user for tag matching. If the advertising video material hits the advertising element tag and matches the interest tag, it will be weighted in the advertising recommendation for this user or this type of user to increase the recommendation probability of the advertising video.
[0180] In this embodiment, based on the replacement method of the constituent elements in the video, the constituent elements of the target video can be accurately identified from the target video, the element tags of the constituent elements are matched with the interest tags of the video viewing objects, and the target objects with successful label matching are screened out from the video viewing objects to achieve precise matching of the target objects, which can effectively improve the accuracy of recommending the target meta-video.
[0181] In a specific application scenario, the method for identifying the components of a video in an advertising scenario can be applied as follows: Figure 14 This is achieved using the advertising element recognition model shown in the figure that integrates multi-frame information. The overall structure of the advertising element recognition model mainly consists of four sub-models: an image depth sub-model (such as the BiT model), an ensemble learning sub-model (such as the GBDT model, a Gradient Boosting Decision Tree, a boosted tree algorithm model), a visual feature extraction sub-model for extracting short-term temporal features, and a scene transition recognition sub-model (such as TransNetV2) for extracting long-term temporal features.
[0182] First, the computer equipment cuts the advertising material video into frames. After cutting, each video frame is predicted through the image depth sub-model to generate a prediction vector or image feature; at the same time, the adjacent video frames are processed through the visual feature extraction sub-model to obtain the inter-frame pixel difference at time t and time t-1, as well as the Laplace gradient of each frame to determine the short-term visual features; at the same time, the scene transition recognition sub-model is used to process the video frame sequence to obtain the long-term video features; finally, the image features, short-term visual features, and long-term scene features are integrated through the integrated learning module to obtain the final multi-label prediction result.
[0183] Specifically, the BiT model is optimized for the pre-training process and uses a larger pre-training corpus. In the pre-training phase, group normalization replaces batch normalization in the residual network model to reduce the impact of the residual network model's batch size on the pre-training process. Furthermore, based on the super-rule mechanism, the parameter adjustment work in the actual application field, i.e., the fine-tuning phase, is reduced. The feature representation capability of the BiT model has been greatly improved through pre-training optimization. In the scenario of advertising label recognition, only a small number of labeled samples are needed for fine-tuning to achieve good results.
[0184] The integrated learning sub-model can specifically be XGBoost (eXtreme Gradient Boosting). The Boosting algorithm is an effective and widely used model training algorithm. XGBoost is implemented based on the Boosting algorithm. The idea of the Boosting algorithm is to continuously improve and enhance weak classifiers, and integrate these classifiers together to form a strong classifier. The XGBoost algorithm is an integrated boosting algorithm, which is a powerful model formed by integrating many basic models. The basic model here can be a classification and regression decision tree (CART) or a linear model. XGBoost is an improvement on the gradient boosting algorithm. The Newton method is used to solve the extreme value of the loss function, and the loss function is Taylor expanded to the second order. In addition, a regularization term is added to the loss function. The objective function during training consists of two parts, the first part is the gradient boosting algorithm loss, and the second part is the regularization term. The loss function of XGBoost is defined as:
[0185] Ω(f)=γT+1 / 2λ‖w‖ 2
[0186] γ and λ are manually set parameters, w is the vector formed by the values of all leaf nodes of the XGBoost decision tree, and T is the number of leaf nodes.
[0187] The visual feature extraction sub-model is used to extract short-term temporal features, which can be elemental features from adjacent video frames. It leverages the sharpness relationship between static and dynamic elements, meaning that images have higher sharpness than videos, and static images have higher sharpness than dynamic images. Specifically, a grid-based sampling method is used to reduce redundant information interference and computational complexity. Nine regions are sampled, and the pixel difference between the video frame's X and Y axes and the previous moment is calculated, as well as the Laplace gradient difference between the center and the edges of the X and Y axes.
[0188] Taking the scene transition recognition sub-model, TransNet, as an example, its input data is a sequence of video frames of length N, each of size 48x27x3. The sequence frames are passed through four dilated convolutional layers with different dilation rates in the temporal dimension to reduce the number of training parameters. The output features of the four dilated convolutional networks are then concatenated or concatenated and max-pooled. Finally, they pass through two fully connected layers and a logistic regression layer to output an N×2N feature vector indicating whether each input video frame is a scene boundary frame. The scene transition recognition sub-model takes a sequence of video frames as input, and compresses each frame to a uniform small size of 48x27x3. Compared to TransNet, TransNetV2 incorporates RGB histogram features to enhance feature representation. TransNetV2 adds a learnable similarity module that concatenates RGB histogram features with learnable similarity features. This concatenation is then used to perform fully connected feature transformations and other operations to obtain long-term temporal features contained in the video frame sequence, which are suitable for scene recognition based on the video's component elements.
[0189] In this embodiment, the feature information of shorter or longer time series is taken into consideration, and the time series information of different lengths is effectively fused, making full use of the image change information between frames, which plays a very important role in identifying element information of different time series. Specifically, the long and short time series information are fused by combining single-frame images with multi-frame features, and the BiT sub-module is used to extract the features and prediction results of single-frame images. The multi-frame prediction results and features are fused through an integrated learning method, so that the model has the ability to recognize long time series information without losing the single-frame prediction results. This method has scalability and flexibility that other methods cannot have. And the visual feature extraction sub-model and the TransNetV2 sub-model are used to expand the multi-frame information. In addition to the integrated learning of single-frame image results, visual features such as Laplace gradient and pixel difference are introduced, and the video features extracted by the deep model TransNetV2 are used as extended information to identify elements existing in multiple frames, while increasing the accuracy of identifying elements in a single frame.
[0190] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0191] Based on the same inventive concept, an embodiment of the present application also provides an identification device for implementing the above-mentioned method for identifying constituent elements in a video; a replacement device for implementing the above-mentioned method for replacing constituent elements in a video; and a video recommendation device for implementing the above-mentioned method for recommending videos based on constituent elements.
[0192] The implementation solutions to the problems provided by the above-mentioned devices are similar to the implementation solutions described in the above-mentioned methods. Therefore, the specific limitations in the embodiments of the device for identifying constituent elements in one or more videos provided below can refer to the limitations on the method for identifying constituent elements in videos described above; the specific limitations in the embodiments of the device for replacing constituent elements in one or more videos provided below can refer to the limitations on the method for replacing constituent elements in videos described above; the specific limitations in the embodiments of the device for recommending videos based on constituent elements provided below can refer to the limitations on the method for recommending videos based on constituent elements described above, and will not be repeated here.
[0193] In one embodiment, Figure 15 As shown, a device 1500 for identifying component elements in a video is provided, comprising: a sequence acquisition module 1502, a first feature extraction module 1504, a second feature extraction module 1506, a third feature extraction module 1508 and a feature fusion module 1510, wherein:
[0194] A sequence acquisition module 1502 is configured to acquire a video frame sequence consisting of at least a portion of video frames of a target video, wherein the video frames in the video frame sequence are arranged in a time sequence in the target video;
[0195] A first feature extraction module 1504 is used to extract the component element features of each video frame in the video frame sequence;
[0196] A second feature extraction module 1506 is used to extract short-term temporal features of the video frame sequence;
[0197] A third feature extraction module 1508 is configured to extract long temporal features of the video frame sequence; the time span of the video frame matched by the short temporal features in the video frame sequence is smaller than the time span of the video frame matched by the long temporal features in the video frame sequence;
[0198] The feature fusion module 1510 is used to identify the component elements in the target video based on the feature fusion results obtained by fusing the component element features, short-term features and long-term features.
[0199] In one embodiment, the second feature extraction module is further used to obtain image difference data of each group of adjacent video frames in the video frame sequence; based on the image difference data, element features of the adjacent video frames are extracted to obtain short-term features composed of the element features of each group of adjacent video frames.
[0200] In one embodiment, the second feature extraction module is further used to compare two video frames in each group of adjacent video frames in the video frame sequence in pixel units to obtain inter-frame pixel differences; and determine image difference data of adjacent video frames based on the inter-frame pixel differences.
[0201] In one embodiment, the second feature extraction module is further used to perform local image sampling on two adjacent video frames based on the same sampling parameters to obtain the captured images contained in each of the two video frames; and compare the sampled images at the same position in the two video frames in pixel units to obtain the pixel difference between the frames.
[0202] In one embodiment, the second feature extraction module is also used to obtain pixel data of each pixel unit to be compared, and the pixel data includes at least one type of data among RGB value, brightness value, and grayscale value; and compare the pixel data of the pixel unit at the same position in the two images to obtain the pixel difference between the frames.
[0203] In one embodiment, the second feature extraction module is also used to identify static elements and dynamic elements in adjacent video frames to obtain element category information; based on the image difference data of adjacent video frames and the element category information of adjacent video frames, element features of adjacent video frames are extracted to obtain short-term features composed of the element features of each group of adjacent video frames.
[0204] In one embodiment, the second feature extraction module is also used to obtain the target element category obtained by performing element category matching on the image difference data of adjacent video frames; when the target element category successfully matches the element category information, element features are extracted from the adjacent video frames to obtain short-term features composed of the element features of each group of adjacent video frames.
[0205] In one embodiment, the second feature extraction module is further configured to determine, for any adjacent video frame, element category information according to image definition differences between regions in the video frame.
[0206] In one embodiment, the second feature extraction module is further configured to obtain clarity evaluation data of a regional image of each region in the video frame; and determine element category information based on differences in the clarity evaluation data of the regional images.
[0207] In one embodiment, the second feature extraction module is further used to extract a regional image of each region in the video frame according to the region division of the video frame; perform image gradient calculation on each regional image to obtain clarity evaluation data of each regional image.
[0208] In one embodiment, the second feature extraction module is further used to perform grid division on the video frame to obtain multiple grid images contained in the video frame; divide the grid images in the same area into a group according to the area division of the video frame; calculate the first clarity difference between the grid images in the same group and the second clarity difference between the grid images in different groups; and determine the element category information of the video frame based on the first clarity difference and the second clarity difference.
[0209] In one embodiment, the second feature extraction module is further used to perform local image sampling on the video frame to obtain multiple sampled images contained in the video frame; divide the sampled images in the same area into a group according to the area division of the video frame; calculate the first clarity difference between the sampled images in the same group and the second clarity difference between the sampled images in different groups; and determine the element category information of the video frame based on the first clarity difference and the second clarity difference.
[0210] In one embodiment, the second feature extraction module is further used to identify scene transition features in the video frame sequence, where the scene transition features are used to characterize the distribution of scene boundary frames in the video frame sequence; and feature conversion is performed on the scene transition features to obtain long-term temporal features contained in the video frame sequence.
[0211] In one embodiment, the second feature extraction module is further used to extract scene transition features from the video frame sequence based on a scene transition recognition model to obtain scene transition features in the video frame sequence; the scene transition recognition model is obtained based on training of transition videos carrying scene transition labels.
[0212] In one embodiment, the second feature extraction module is further used to perform maximum pooling processing on the scene transition features of the video frame sequence to obtain pooled features; perform feature connection on the pooled features based on a fully connected layer with a feature discarding processing function to obtain fully connected features; and perform dimension compression and feature fuzzy transformation on the fully connected features to obtain long-term time series features contained in the video frame sequence.
[0213] In one embodiment, the first feature extraction module is further configured to perform image feature extraction and component element prediction on each video frame in the video frame sequence to obtain component element features of each video frame.
[0214] In one embodiment, the first feature extraction module is also used to perform depth feature extraction and component element prediction on each video frame in the video frame sequence based on the image depth model to obtain the component element features of each video frame; the image depth model is obtained by training based on a filled image carrying a component element identifier.
[0215] In one embodiment, the training process of the image depth model includes: obtaining a BiT model that has been pre-trained based on a large-scale pre-training corpus; and adjusting parameters of the BiT model based on a filled image carrying component element identifiers to obtain an image depth model.
[0216] In one embodiment, the device for identifying component elements in a video also includes a feature fusion module, which is used to normalize the feature dimensions of each component element feature, short-time sequence feature, and long-time sequence feature to obtain component element features, short-time sequence features, and long-time sequence features with the same feature dimensions; and to fuse the component element features, short-time sequence features, and long-time sequence features with the same feature dimensions to obtain a feature fusion result.
[0217] In one embodiment, Figure 16 As shown, a device 1600 for replacing component elements in a video is provided, comprising: a target video acquisition module 1602, a component element acquisition module 1604, an element filling data determination module 1606, and a component element replacement module 1608, wherein:
[0218] A target video acquisition module 1602 is configured to acquire a target video and a replacement element for a target component element in the target video;
[0219] The recognition result acquisition module 1604 is used to obtain the component elements obtained by performing component element recognition on the target video from the recognition device for component elements in the video;
[0220] An element filling data determining module 1606 is configured to determine element filling data of the target component element in the target video when the component elements include the target component element;
[0221] The component element replacement module 1608 is configured to replace the target component element in the target video with a replacement element based on the element filling data.
[0222] In one embodiment, Figure 17 As shown, a video recommendation device 1700 based on component elements is provided, comprising: a component element acquisition module 1702, a tag matching module 1704 and a video push module 1706, wherein:
[0223] The recognition result acquisition module 1702 is used to obtain the component elements obtained by performing component element recognition on the target video from the component element recognition device in the video;
[0224] The tag matching module 1704 is used to perform tag matching between the element tags of the constituent elements and the interest tags of the video viewing objects, and filter out target objects with successful tag matching from the video viewing objects;
[0225] The video pushing module 1706 is used to push the target video to the target object.
[0226] Each module in the aforementioned apparatus for identifying component elements in a video, apparatus for replacing component elements in a video, and apparatus for recommending video components based on component elements may be implemented in whole or in part via software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0227] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 18 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store model data for feature extraction. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for identifying component elements in a video, a method for replacing component elements in a video, and a video recommendation method based on component elements.
[0228] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 19 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, it implements a method for identifying component elements in a video, a method for replacing component elements in a video, and a method for recommending videos based on component elements. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse, etc.
[0229] Those skilled in the art will understand that Figure 18 and Figure 19 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0230] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0231] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0232] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0233] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0234] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0235] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0236] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for identifying component elements in a video, characterized in that: The method comprises: Acquire a video frame sequence consisting of at least a portion of video frames of a target video, wherein the video frames in the video frame sequence are arranged according to a time sequence in the target video; Extracting component element features of each video frame in the video frame sequence; Extracting long-term temporal features of the video frame sequence; obtaining image difference data of each group of adjacent video frames in the video frame sequence; performing element feature extraction on the adjacent video frames based on the image difference data to obtain short-term temporal features composed of the element features of each group of adjacent video frames; wherein the time span of the video frames matched by the short-term temporal features in the video frame sequence is shorter than the time span of the video frames matched by the long-term temporal features in the video frame sequence; Based on a feature fusion result obtained by fusing the component element features, the short time sequence features, and the long time sequence features, the component elements in the target video are identified.
2. The method according to claim 1, characterized in that The obtaining of image difference data of each group of adjacent video frames in the video frame sequence includes: For each group of adjacent video frames in the video frame sequence, comparing two video frames in the adjacent video frames in pixel units to obtain a pixel difference between the frames; Image difference data of the adjacent video frames is determined based on the inter-frame pixel differences.
3. The method according to claim 2, characterized in that Comparing two adjacent video frames in pixel units to obtain inter-frame pixel differences includes: Based on the same sampling parameters, local image sampling is performed on two video frames in the adjacent video frames to obtain captured images respectively contained in the two video frames; The sampled images at the same position in the two video frames are compared in pixel units to obtain the pixel difference between the frames.
4. The method according to claim 1, wherein The method further comprises: Identifying static elements and dynamic elements in the adjacent video frames to obtain element category information; The extracting element features of the adjacent video frames based on the image difference data to obtain short-term features composed of the element features of each group of adjacent video frames includes: Based on the image difference data of the adjacent video frames and the element category information of the adjacent video frames, element features of the adjacent video frames are extracted to obtain short-term features composed of the element features of each group of adjacent video frames.
5. The method according to claim 4, characterized in that The step of extracting element features from the adjacent video frames based on the image difference data of the adjacent video frames and the element category information of the adjacent video frames to obtain short-term features composed of the element features of each group of adjacent video frames includes: Obtaining a target element category obtained by performing element category matching on image difference data of adjacent video frames; When the target element category successfully matches the element category information, element features are extracted from the adjacent video frames to obtain short-term features composed of the element features of each group of adjacent video frames.
6. The method according to claim 4, characterized in that The identifying of static elements and dynamic elements in the adjacent video frames to obtain element category information includes: For any video frame in the adjacent video frames, extracting a regional image of each region in the video frame according to the region division of the video frame; Performing image gradient calculation on each of the regional images to obtain clarity evaluation data of each of the regional images; Based on the difference in the clarity evaluation data of the images of the respective regions, the element category information is determined according to the clarity difference characteristics between the static elements and the dynamic elements.
7. The method according to claim 4, characterized in that The identifying of static elements and dynamic elements in the adjacent video frames to obtain element category information includes: Dividing the video frames into grids respectively to obtain a plurality of grid images contained in the video frames; According to the regional division of the video frame, the grid images in the same region are divided into a group; calculating a first definition difference between the same group of grid images and a second definition difference between different groups of grid images; Based on the first definition difference and the second definition difference, and in accordance with definition difference characteristics between static elements and dynamic elements, element category information of the video frame is determined.
8. The method according to claim 4, characterized in that The identifying of static elements and dynamic elements in the adjacent video frames to obtain element category information includes: Performing local image sampling on the video frame to obtain a plurality of sampled images contained in the video frame; According to the region division of the video frame, the sampled images in the same region are divided into a group; Calculating a first definition difference between the same group of sample images and a second definition difference between different groups of sample images; Based on the first definition difference and the second definition difference, and in accordance with definition difference characteristics between static elements and dynamic elements, element category information of the video frame is determined.
9. The method according to claim 1, characterized in that The extracting of the long-term temporal features contained in the video frame sequence includes: Based on a scene transition recognition model, scene transition features are extracted from the video frame sequence to obtain scene transition features in the video frame sequence; wherein the scene transition features are used to characterize the distribution of scene boundary frames in the video frame sequence; and the scene transition recognition model is obtained by training based on transition videos carrying scene transition labels; Feature conversion is performed on the scene transition features to obtain long-term temporal features contained in the video frame sequence.
10. The method according to claim 9, characterized in that The performing feature conversion on the scene transition feature to obtain the long-term temporal features contained in the video frame sequence includes: Performing maximum pooling processing on the scene transition features of the video frame sequence to obtain pooled features; Performing feature connection on the pooled features based on a fully connected layer with a feature discarding processing function to obtain a fully connected feature; The fully connected features are subjected to dimension compression and feature fuzzy transformation to obtain long-term temporal features contained in the video frame sequence.
11. The method according to claim 1, wherein The extracting the component element features of each video frame in the video frame sequence includes: Based on the image depth model, performing depth feature extraction and component element prediction on each video frame in the video frame sequence to obtain component element features of each video frame; The image depth model is obtained by training based on a filled image carrying identifiers of constituent elements.
12. The method according to claim 11, characterized in that The training process of the image depth model includes: Obtain the pre-trained BiT model based on the pre-training corpus; Based on the filled image carrying the component element identification, the parameters of the BiT model are adjusted to obtain the image depth model.
13. The method according to any one of claims 1 to 12, characterized in that The method further comprises: Normalizing the feature dimensions of each of the component element features, the short time series features, and the long time series features to obtain component element features, short time series features, and long time series features with the same feature dimensions; The component element features, short time series features and long time series features with the same feature dimension are fused to obtain the feature fusion result.
14. A method for replacing elements in a video, characterized in that: The method comprises: Obtain a target video, and obtain a replacement element for a target component element in the target video; Identify the component elements in the target video based on the method for identifying component elements in a video according to any one of claims 1 to 13; In the case of identifying that the component elements include the target component element, determining element filling data of the target component element in the target video; Based on the element filling data, the target component element in the target video is replaced with the replacement element.
15. A video recommendation method based on component elements, characterized in that: The method comprises: Identify the component elements in the target video based on the method for identifying component elements in a video according to any one of claims 1 to 13; Perform tag matching on the element tags of the constituent elements and the interest tags of the video viewing objects, and filter out target objects with successful tag matching from the video viewing objects; Push the target video to the target object.
16. A device for identifying elements in a video, characterized in that: The device comprises: A sequence acquisition module is used to acquire a video frame sequence consisting of at least a portion of video frames of a target video, wherein each video frame in the video frame sequence is arranged according to a time sequence in the target video; A first feature extraction module is used to extract the component element features of each video frame in the video frame sequence; a second feature extraction module configured to obtain image difference data of each group of adjacent video frames in the video frame sequence; and extract element features of the adjacent video frames based on the image difference data to obtain a short-term feature composed of the element features of each group of adjacent video frames; A third feature extraction module is configured to extract long-term temporal features of the video frame sequence; the time span of the video frames matched by the short-term temporal features in the video frame sequence is smaller than the time span of the video frames matched by the long-term temporal features in the video frame sequence; A feature fusion module is used to identify the component elements in the target video based on a feature fusion result obtained by fusing the component element features, the short-term features and the long-term features.
17. The device according to claim 16, characterized in that The second feature extraction module is further used to compare two video frames in each group of adjacent video frames in the video frame sequence in pixel units to obtain inter-frame pixel differences; and determine image difference data of the adjacent video frames based on the inter-frame pixel differences.
18. The device according to claim 17, characterized in that The second feature extraction module is also used to perform local image sampling on two video frames in the adjacent video frames based on the same sampling parameters to obtain the captured images contained in each of the two video frames; and compare the sampled images at the same position in the two video frames in pixel units to obtain the pixel difference between the frames.
19. The device according to claim 16, characterized in that The second feature extraction module is also used to identify static elements and dynamic elements in the adjacent video frames to obtain element category information; based on the image difference data of the adjacent video frames and the element category information of the adjacent video frames, element features are extracted from the adjacent video frames to obtain short-term features composed of the element features of each group of adjacent video frames.
20. The device according to claim 19, characterized in that The second feature extraction module is also used to obtain the target element category obtained by performing element category matching on the image difference data of the adjacent video frames; when the target element category successfully matches the element category information, element feature extraction is performed on the adjacent video frames to obtain a short-term feature composed of the element features of each group of adjacent video frames.
21. The device according to claim 19, characterized in that The second feature extraction module is also used to extract the regional image of each area in any video frame among the adjacent video frames according to the area division of the video frame; perform image gradient calculation on each of the regional images to obtain clarity evaluation data of each of the regional images; and determine element category information based on the difference in clarity evaluation data of each of the regional images and the clarity difference characteristics between static elements and dynamic elements.
22. The device according to claim 19, characterized in that The second feature extraction module is further configured to perform grid division on the video frame to obtain a plurality of grid images contained in the video frame; group the grid images in the same region according to the region division of the video frame; and calculate a first definition difference between the grid images in the same group and a second definition difference between the grid images in different groups; Based on the first definition difference and the second definition difference, and in accordance with definition difference characteristics between static elements and dynamic elements, element category information of the video frame is determined.
23. The device according to claim 19, characterized in that The second feature extraction module is further configured to perform local image sampling on the video frame to obtain a plurality of sampled images contained in the video frame; group the sampled images in the same region according to the region division of the video frame; and calculate a first definition difference between the sampled images in the same group and a second definition difference between the sampled images in different groups; Based on the first definition difference and the second definition difference, and in accordance with definition difference characteristics between static elements and dynamic elements, element category information of the video frame is determined.
24. The device according to claim 16, characterized in that The third feature extraction module is further used to extract scene transition features from the video frame sequence based on a scene transition recognition model to obtain scene transition features in the video frame sequence; wherein the scene transition features are used to characterize the distribution of scene boundary frames in the video frame sequence; the scene transition recognition model is obtained based on training of transition videos carrying scene transition labels; and feature conversion is performed on the scene transition features to obtain long-term temporal features contained in the video frame sequence.
25. The device according to claim 24, characterized in that The third feature extraction module is also used to perform maximum pooling processing on the scene transition features of the video frame sequence to obtain pooled features; perform feature connection on the pooled features based on a fully connected layer with a feature discarding processing function to obtain fully connected features; and perform dimensionality compression and feature fuzzy transformation on the fully connected features to obtain long-term time series features contained in the video frame sequence.
26. The device according to claim 16, characterized in that The first feature extraction module is also used to perform depth feature extraction and component element prediction on each video frame in the video frame sequence based on an image depth model to obtain the component element features of each video frame; the image depth model is obtained by training based on a filled image carrying a component element identifier.
27. The device according to claim 26, characterized in that The training process of the image depth model includes: obtaining a BiT model that has been pre-trained based on pre-training corpus; and adjusting parameters of the BiT model based on a filled image carrying component element identifiers to obtain the image depth model.
28. The device according to any one of claims 16 to 27, characterized in that The device also includes a feature fusion module; The feature fusion module is used to normalize the feature dimensions of each of the component element features, the short time series features, and the long time series features to obtain component element features, short time series features, and long time series features with the same feature dimensions; The component element features, short time series features and long time series features with the same feature dimension are fused to obtain the feature fusion result.
29. A device for replacing elements in a video, characterized in that: The device comprises: A target video acquisition module is used to acquire a target video and obtain a replacement element for a target component element in the target video; a recognition result acquisition module, configured to acquire, from the device for recognizing component elements in a video according to any one of claims 16 to 28, component elements obtained by performing component element recognition on the target video; an element filling data determining module, configured to, when identifying that the component elements include the target component element, determine the element filling data of the target component element in the target video; A component element replacement module is used to replace the target component element in the target video with the replacement element based on the element filling data.
30. A video recommendation device based on component elements, characterized in that: The device comprises: a recognition result acquisition module, configured to acquire, from the device for recognizing component elements in a video according to any one of claims 16 to 28, component elements obtained by performing component element recognition on a target video; a tag matching module, configured to perform tag matching on the element tags of the constituent elements and the interest tags of the video viewing objects, and filter out target objects with successful tag matching from the video viewing objects; The video push module is used to push the target video to the target object.
31. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 15 are implemented.
32. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.
33. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
Video classification method, device, electronic equipment and storage medium
CN113010736A
Video processing method, device and equipment, computer program product and storage medium
CN114513653A