Video processing method and device, data training method, device and system
By training the video processing model, focusing on background and motion feature information, and using background loss function and motion loss function to optimize the model, the problem of inaccurate video content representation by the video representation model is solved, and a more accurate video content representation is achieved.
Patent Information
- Application Number
- CN202110179387.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-02-09
AI Technical Summary
Existing video representation models pay too much attention to the background features of the video when extracting features, and pay less attention to motion feature information, which causes the output video representation information of the video representation model to be easily affected by the background of the video image, resulting in inaccurate representation of the video content.
By training the initial model, a video processing model is generated. The supervision task is related to the background feature information and motion feature information of the sample data. The background feature information and motion feature information of the video to be processed are extracted. The model is optimized using the background loss function and the motion loss function to enhance the focus on motion features.
The accuracy of the video representation model in representing video content is improved, the influence of background on feature extraction is avoided, and a more accurate video content representation is achieved.
Smart Images

Figure CN114913444B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing, and in particular to a video processing method and device, and a data training method, device and system. Background Art
[0002] There are a large number of video resources on the Internet. Video representation models can learn from video content on the Internet to detect and label video content. For example, on e-commerce platforms, by identifying short video content of products and adding product tags to the videos, users can quickly find products of interest through search. However, existing video representation models pay too much attention to the background features of the video when extracting features, and pay less attention to the motion feature information in the foreground of the video. As a result, the output video representation information of the video representation model is easily affected by the background of the video image, which leads to inaccurate representation of the video content. For example, if the background of a video contains object A and the foreground contains a moving object B, the video representation model, because it pays more attention to the background features, may regard object A in the background as the subject of the video image and output the features of object A as the video representation information, resulting in a deviation between the representation results of the video representation model and the actual content of the video.
[0003] Currently, no effective solution has been proposed to the technical problem that the video representation model in the above-mentioned prior art does not accurately represent the video content. Summary of the Invention
[0004] The embodiments of the present invention provide a video processing method and device, and a data training method, device, and system to at least solve the technical problem in the prior art that video self-supervised learning is easily affected by image background.
[0005] According to one aspect of an embodiment of the present invention, a video processing method is provided, comprising: receiving a video to be processed; performing feature extraction on the video to be processed through a video representation model to obtain video representation information of the video to be processed, wherein the video representation model is obtained by training an initial model, the initial model is a model obtained by training with sample data, and the training task is related to background feature information and motion feature information of the sample data; and outputting the video representation information of the video to be processed, wherein the video representation information includes background feature information of the video to be processed and motion feature information of the video to be processed.
[0006] According to another aspect of an embodiment of the present invention, a video processing method is provided, comprising: receiving a video to be processed; processing the video to be processed by a video processing model to obtain a video label of the video to be processed, wherein the video processing model is obtained by training an initial model, the initial model is a model obtained by training sample data, and the training task is related to background feature information and motion feature information of the sample data; and displaying the video label of the video to be processed.
[0007] According to another aspect of an embodiment of the present invention, a video processing method is provided, comprising: receiving a live video; processing the live video through a video processing model to obtain a video tag of the live video, wherein the video tag is used to represent the product type of a target object in the live video, and the video processing model is obtained by training an initial model, wherein the initial model is a model obtained by training sample data, and the training task is related to background feature information and motion feature information of the sample data; and displaying the video tag of the live video.
[0008] According to another aspect of an embodiment of the present invention, a data training method is provided, including: obtaining first feature information obtained by performing feature extraction on a sample video clip by a model to be trained; determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information; determining a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and optimizing the model to be trained according to the background loss function and the motion loss function.
[0009] According to another aspect of an embodiment of the present invention, a video processing device is also provided, including: a first receiving module for receiving a video to be processed; a feature extraction module for extracting features from the video to be processed through a video representation model to obtain video representation information of the video to be processed, wherein the video processing model is obtained by training an initial model, and the initial model is a model obtained by training sample data, and the training task is related to the background feature information and motion feature information of the sample data; an output module for outputting the video representation information of the video to be processed, wherein the video representation information includes the background feature information of the video to be processed and the motion feature information of the video to be processed.
[0010] According to another aspect of the embodiments of the present application, a video processing apparatus is also provided, which comprises: a second receiving module configured to receive a video to be processed; a first processing module configured to process the video to be processed by using a video processing model to obtain a video label of the video to be processed, wherein the video processing model is obtained by training an initial model, and the initial model is a model trained by using sample data, and a supervised task is related to background feature information and motion feature information of the sample data; and a first display module configured to display the video label of the video to be processed.
[0011] According to another aspect of the embodiments of the present application, a video processing apparatus is also provided, which comprises: a third receiving module configured to receive a live video; a second processing module configured to process the live video by using a video processing model to obtain a video label of the live video, wherein the video label is used to indicate a product type of a target object in the live video, the video processing model is obtained by training an initial model, and the initial model is a model trained by using sample data, and a training task is related to background feature information and motion feature information of the sample data; and a second display module configured to display the video label of the live video.
[0012] According to another aspect of the embodiments of the present application, a data training apparatus is also provided, which comprises: an obtaining module configured to obtain first feature information extracted by a to-be-trained model from a sample video segment; a first determining module configured to determine a background loss function based on the first feature information and second feature information, wherein the second feature information comprises background feature information of an image in the sample video segment, and the background loss function is used to represent a difference degree between the first feature information and the background feature information; a second determining module configured to determine a motion loss function based on the first feature information and third feature information, wherein the third feature information comprises first motion feature information of an image after the sample video segment, and the motion loss function is used to represent a difference degree between second motion feature information predicted based on the first feature information and the first motion feature information; and an optimization module configured to optimize the to-be-trained model according to the background loss function and the motion loss function.
[0013] According to another aspect of the embodiments of the present application, a storage medium is also provided, which comprises a stored program, wherein the program, when executed, controls a device where the storage medium is located to perform any of the video processing methods.
[0014] According to another aspect of the embodiments of the present application, a processor is also provided, which is used to execute a program, wherein the program, when executed, performs any of the video processing methods.
[0015] According to another aspect of an embodiment of the present invention, a data training system is also provided, including: a processor; and a memory, connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining first feature information obtained by extracting features of a sample video clip by a model to be trained; determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information; determining a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and optimizing the model to be trained according to the background loss function and the motion loss function.
[0016] In an embodiment of the present invention, features of the video to be processed are extracted by using a video processing model obtained by training an initial model, wherein the initial model is a model obtained by training sample data, and the supervision task is related to the background feature information and motion feature information of the sample data, so that the video processing model can extract the background feature information and motion feature information of the video to be processed, and the video representation model pays more attention to the motion feature information during feature extraction, so that the video representation information can contain more motion feature information, thereby avoiding the problem that the video representation model is easily affected by the background in the video to be processed during feature extraction, improving the accuracy of the video representation model in representing the video content, and solving the problem of inaccurate video representation model in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0018] Figure 1 is a schematic diagram of a hardware structure block diagram of a computing device (or mobile device) for implementing a data training method;
[0019] Figure 2 is a flow chart of a video processing method according to an embodiment of the present invention;
[0020] Figure 3 is a flow chart of a data training method according to an embodiment of the present invention;
[0021] Figure 4 is a schematic diagram of a framework of an optional data training method according to an embodiment of the present invention;
[0022] Figure 5 is a schematic diagram of optional compressed video formats according to an embodiment of the present invention;
[0023] Figure 6 is a schematic diagram of a data processing method according to an embodiment of the present invention;
[0024] Figure 7 is a schematic diagram of a data training device according to an embodiment of the present invention;
[0025] Figure 8 is a schematic diagram of a data processing device according to an embodiment of the present invention;
[0026] Figure 9 is a structural block diagram of a computer terminal according to an embodiment of the present invention;
[0027] Figure 10 is a flowchart of a video processing method according to an embodiment of the present invention;
[0028] Figure 11 Schematic diagram of a video self-supervised learning method based on contrastive learning;
[0029] Figure 12 is a schematic diagram of a video processing device according to an embodiment of the present invention;
[0030] Figure 13 is a schematic diagram of a video processing device according to an embodiment of the present invention;
[0031] Figure 14 is a flowchart of a video processing method according to an embodiment of the present invention;
[0032] Figure 15 2 is a schematic diagram of a video processing device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] Example 1
[0036] An embodiment of the present invention provides an embodiment of a video processing method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0037] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computing device or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computing device (or mobile device) for implementing a video processing method. Figure 1 As shown, the computing device 10 (or mobile device 10) may include one or more (illustrated as 102a, 102b, ..., 102n) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0038] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computing device 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0039] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data training method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the vulnerability detection method of the above-mentioned application. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computing device 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0040] The transmission module 106 is configured to receive or send data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computing device 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0041] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of computing device 10 (or mobile device).
[0042] Under the above operating environment, this application provides Figure 2 The processing method of the video shown. Figure 2 FIG. 1 is a flowchart of video processing according to the first embodiment of the present invention. Figure 2 As shown, the method includes the following steps:
[0043] Step S201: receiving a video to be processed.
[0044] The above-mentioned videos to be processed are videos that need to be represented for downstream tasks. The downstream tasks can be classification, recognition, object detection, object tracking, and video labeling based on video representation.
[0045] The video to be processed can be a video or video clip obtained from the internet. For example, the video to be processed could be a live e-commerce video, and the corresponding downstream task could be object detection and object tracking to identify products in the live video and make accurate recommendations. Another example is a live entertainment video, and the corresponding downstream task could be video tagging to determine the video type (sports, beauty, baby girls, movies, etc.) through video tagging, and then make accurate recommendations for entertainment videos.
[0046] The above-mentioned videos to be processed can also be videos in non-entertainment fields such as education or medical fields. For example, the videos to be processed can be videos of distance education courses. For another example, the videos to be processed can also be images used for medical diagnosis, etc. Examples are not given here one by one. Step S202, feature extraction is performed on the video to be processed using a video representation model to obtain video representation information of the video to be processed, wherein the video representation model is obtained by training an initial model, and the initial model is a model obtained by training using sample data, and the supervision task is related to the background feature information and motion feature information of the sample data.
[0047] The above-mentioned initial model can be a feature extraction model obtained after learning and training through sample data. The sample data used by the initial model is a video clip. During the training process, the supervision information includes background feature information and motion feature information, that is, the supervision task is related to both the background feature information and the motion feature information, so that the initial model learns both the background feature information and the motion feature information, avoiding the background bias problem that occurs in the learning of sample data in the prior art. In an optional embodiment, a video clip can be obtained as sample data, and the foreground and background of a frame image containing the background in the video clip are separated to obtain background information and motion information, and the features in the background information are extracted to obtain background feature information, and the features in the motion information are extracted to obtain motion feature information. The background feature information and the motion feature information are used to train the initial model respectively, so that the initial model can learn both background information and motion information. The above-mentioned initial model can also be trained using the method in Example 3 of the present application.
[0048] The video representation model is a feature extraction model used to extract features from the video being processed. The video representation model can be a neural network model with a convolutional layer structure. The video representation model can be obtained by further training the initial model through supervised learning. Specifically, the initial model can be further trained using sample data with preset topic labels, so that the video representation model can extract features of the preset topic content.
[0049] In an optional embodiment, when the video characterization model is used to extract features in a video, label information is added to the background object A and the foreground moving person B in the sample video. After the video characterization model is trained using the sample video with the added label information, the video characterization model can extract the background feature information of the background object A and the motion feature information of the foreground moving person B in the video to be processed.
[0050] Step S203: outputting video characterization information of the video to be processed, wherein the video characterization information includes background feature information of the video to be processed and motion feature information of the video to be processed.
[0051] The above-mentioned video representation information includes background feature information and motion feature information, so that the video representation information can more accurately represent the content of the video. For example, the video to be processed is a live video. The live video contains the background information of the live broadcast room where the anchor is located, and also contains the anchor who is dancing in the foreground. The video representation model is used to extract features from the video to be processed, and the output video representation information contains the background feature information of the live broadcast room in the background, as well as the motion feature information of the anchor. Furthermore, performing downstream tasks based on this video representation information can obtain more accurate results. For example, live videos can be labeled according to the type of dance performed by the anchor.
[0052] In this embodiment, by receiving a video to be processed, feature extraction is performed on the video to be processed through a video representation model to obtain video representation information of the video to be processed. The video processing model is obtained by training an initial model, and the initial model is a model obtained through self-supervised training. The supervision information of the self-supervised training includes background feature information and motion feature information of the sample data, and the video representation information of the video to be processed is output, wherein the video representation information includes background feature information of the video to be processed and motion feature information of the video to be processed, so that the video representation model can extract background feature information and motion feature information in the video to be processed, and the video representation model pays more attention to motion feature information during feature extraction, so that the video representation information can contain more motion feature information, thereby avoiding the problem that the video representation model is easily affected by the background in the video to be processed during feature extraction, improving the accuracy of the video representation model in representing the video content, and solving the problem of inaccurate representation of the video content by the video representation model in the prior art.
[0053] As an optional embodiment, after outputting the video representation information of the video to be processed, the method also includes at least one of the following: performing video classification on the video to be processed based on the video representation information to obtain a video label of the video to be processed; performing object detection on the video to be processed based on the video representation information to obtain a target object in the video to be processed; performing object tracking on the video to be processed based on the video representation information to obtain the position of the target object in each frame image of the video to be processed.
[0054] After the above-mentioned video representation model, a head network model can be added to classify, detect, and recognize the background feature information and motion feature information extracted by the video representation model.
[0055] In an optional implementation, a classification network can be added to the output of the video representation model. The classification network can classify the video to be processed based on the video representation information. For example, the video to be processed is a live broadcast video on an e-commerce platform, which includes a video of the host demonstrating the usage of product C in the live broadcast room. The trained video representation model can extract the background feature information of the live broadcast room and the motion feature information of the host demonstrating product C. By outputting the background feature information and motion feature information to the classification network, a video label "host demonstrating product C" can be obtained. Users can quickly find the demonstration video of product C by searching on the e-commerce platform.
[0056] In another optional implementation, a head network can be added to the output of the above-mentioned video representation model, and the head network can realize object detection in the video to be processed based on the video representation information. For example, on an e-commerce platform, the target object can be a certain product that needs to be removed from the shelves in a centralized manner, and the video to be processed can be a small video of the product on the e-commerce platform. The small video of the product contains the background of the live broadcast room and the demonstration of the product in the foreground. By using sample images with the product labeling information to train the video representation model and the head network, the video representation model can extract the background feature information and motion feature information of the small video of the product on the e-commerce platform separately. The head network further detects the background feature information and motion feature information of the small video of the product, and determines whether the foreground of the small video of the product contains the demonstration information of the product, thereby accurately detecting the video containing the target product from a large number of small videos of the product, and avoiding the influence of the background in the small video of the product on the detection (for example, in the small video of the product, the above-mentioned product to be removed from the shelves may be on the shelf in the background of the live broadcast room, but the host in the foreground is demonstrating another product. The existing feature extraction model may extract the features of the other product in the background as the theme of the video, thereby removing the wrong small video of the product).
[0057] In another optional implementation, a regression network can be added to the output of the video representation model. The regression network can use the video representation information to select the target object in the video to be processed, thereby achieving object tracking in the video to be processed. For example, in a live video, after detecting the target object, a bounding box is used to select the target object in each frame image to achieve tracking of the target object. By tracking objects in the video, it is easier to identify objects in the video and thus determine the subject in the video.
[0058] As an optional embodiment, the above method also includes: obtaining an initial model, and the step of obtaining the initial model includes: obtaining the target loss function of the model to be trained, wherein the target loss function is composed of background feature information and motion feature information; optimizing the model to be trained by solving the target loss function to obtain the initial model.
[0059] Specifically, the model to be trained is a feature extraction model that needs to be optimized, the initial model is the feature extraction model after optimizing the model to be trained, and the feature extraction model may be a three-dimensional video neural network model.
[0060] Since the target loss function is composed of background feature information and motion feature information, the target loss function can be used to train and learn the training model in terms of both background information and motion information, so that the initial model obtained after training can pay more attention to the motion feature information in the video, and then the features obtained after the initial model extracts features from the video can contain more motion feature information. Extracting the motion feature information as a video label can more accurately represent the content of the video.
[0061] As an optional embodiment, obtaining the target loss function of the model to be trained includes: obtaining first feature information obtained by performing feature extraction on a sample video clip by the model to be trained; determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information; determining a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and determining the sum of the background loss function and the motion loss function as the target loss function.
[0062] The sample video clip is a video clip with multiple frames of images, which is used for training the training model. Specifically, the complete video corresponding to the sample video clip can be divided into multiple video clips, each of which contains multiple consecutive frames of images. Any video clip can be selected as the sample video clip.
[0063] In an optional implementation, background feature information can be obtained by separating the foreground and background in a frame image containing the background and extracting features from the background information. It can also be extracted by inputting a frame image containing the background in a sample video clip into an image feature extraction model.
[0064] The above-mentioned background loss function is used to perform contrastive learning on the first feature information and the second feature information. Specifically, the contrastive learning method is a video self-supervised learning method. By performing contrastive learning on unlabeled image features, different segments of the same video are brought closer in the feature space, while different segments from different videos are pushed further away in the feature space, thereby achieving self-supervised learning of the video. In an optional embodiment, the above-mentioned sample video segment is input into a three-dimensional video feature extraction model to extract the first feature information, and any key frame is input into a two-dimensional image feature extraction model to obtain the second feature information. Based on the background loss function, the first feature information and the second feature information are contrastively learned. The first feature information extracted by the three-dimensional video feature extraction model and the second feature information extracted by the two-dimensional image feature extraction model of the background image of the same video can be brought closer, while the first feature information and the second feature information from different videos are pushed further away, thereby enabling the to-be-trained model to learn the background information. The to-be-trained model is trained based on the background loss function to improve the to-be-trained model's ability to recognize the background information of the video.
[0065] The first motion feature information can be extracted by inputting the images following the sample video clip into a three-dimensional feature extraction model. The sample video clip is followed by another video clip located after the sample video clip. For example, if the sample video clip is the images from frames 1 to 5 of the video, the images used to extract the first motion feature information are the images from frames 6 to 10 of the same video. The second motion feature information can be predicted by inputting the first feature information into a neural network with an encoder-decoder structure.
[0066] The above-mentioned motion loss function is used to compare and learn the first feature information and the third feature information. According to the motion loss function, it can be determined that the first motion feature information and the second motion feature information correspond to each other in time and space positions as positive samples, and the first motion feature information and the second motion feature information are inconsistent in any one of time and space positions as negative samples, thereby enabling the model to be trained to learn the motion information.
[0067] It should be noted that, unlike static, coarse-grained background feature information, motion feature information is fine-grained, position-related feature information. Therefore, the sample video clip is predicted through a three-dimensional neural network to obtain the second motion feature information of the image after the sample video clip, and the first motion feature information is obtained by extracting the corresponding image after the real sample video clip through a three-dimensional neural network. By comparing and learning the first motion feature information and the second motion feature information, the model to be trained can learn fine-grained motion information.
[0068] The background loss function and motion loss function can be used to train the training model in terms of background information and motion information. By decoupling the background information and motion information in the sample video clips, the training model can learn both background information and motion information. The sum of the background loss function and the motion loss function is determined as the target loss function, so that the target loss function can compare and learn the training model in terms of background information and motion information, thereby improving the initial model's recognition accuracy of the image's background information and the foreground motion information.
[0069] Example 2
[0070] According to an embodiment of the present invention, an embodiment of a video processing method is also provided. Figure 10 Flowchart of video processing according to embodiment 2 of the present invention, Figure 10 As shown, the method includes the following steps:
[0071] Step S1001: Receive a video to be processed.
[0072] The above-mentioned videos to be processed are videos that need to be represented for use in downstream tasks. The downstream tasks may be detection and recognition based on the video representation and video labeling.
[0073] The video to be processed can be a video or video clip obtained from the internet. For example, the video to be processed could be a live e-commerce video, and the corresponding downstream task could be object detection and object tracking to identify products in the live video and make accurate recommendations. Another example is a live entertainment video, and the corresponding downstream task could be video tagging to determine the video type (sports, beauty, baby girls, movies, etc.) through video tagging, and then make accurate recommendations for entertainment videos.
[0074] The videos to be processed can also be videos in non-entertainment fields such as education or medical fields. For example, the videos to be processed can be videos of distance education courses. For another example, the videos to be processed can also be images from medical diagnosis. Examples are not given here one by one.
[0075] Step S1002: Process the video to be processed through a video processing model to obtain a video label for the video to be processed, wherein the video processing model is obtained by training an initial model, and the initial model is a model obtained by training sample data, and the supervision task is related to the background feature information and motion feature information of the sample data.
[0076] The video representation model is a feature extraction model used for feature extraction of the to-be-processed video. The video representation model can be a neural network model with a convolutional layer structure. The video label is used to represent the theme of the to-be-processed video or the target object in the to-be-processed video. For example, a shopping video of a commodity can use the name of the commodity as the video label. The video representation model can be further trained in a supervised learning manner on the basis of the initial model. Specifically, the initial model can be further trained by using sample data with a preset theme label, so that the video representation model can extract features of the preset theme content and take the preset theme as the video label.
[0077] It should be noted that, since the video representation model can pay more attention to the motion information features in the video, the video label determined according to the features extracted and output by the video representation model can accurately represent the content of the video. For example, the to-be-processed video is a live video, which contains background information of a live room where a host is located and also contains the host dancing in the foreground. The video representation model can extract features of the to-be-processed video, and take the host dancing as the video label. Further, the dancing type of the host can also be taken as the video label of the live video, so that the video label can more accurately represent the content of the live video.
[0078] The initial model can be a feature extraction model obtained by learning and training sample data. The sample data used by the initial model is a video segment. In the training process, the supervision information includes background feature information and motion feature information, that is, the supervision task is related to both the background feature information and the motion feature information, so that the initial model learns both the background feature information and the motion feature information, thereby avoiding the background bias problem in the learning of the sample data in the prior art. In an optional embodiment, a video segment can be obtained as sample data. The foreground and the background in a frame of image containing the background in the video segment are separated to obtain the background information and the motion information. The features in the background information are extracted to obtain the background feature information, and the features in the motion information are extracted to obtain the motion feature information. The initial model is trained by using the background feature information and the motion feature information, so that the initial model can learn both the background information and the motion information. The initial model can also be trained in the manner of the embodiment 3 of the present application.
[0079] In an optional embodiment, the video to be processed is a live shopping video on an e-commerce platform. The live shopping video includes the live broadcast room background, the merchandise shelves in the background, and the products being explained by the host. The preset themes can be the host's explanation of product D and the host's explanation of product E. By extracting features from the live shopping video using the video representation model, the video of the host explaining product D and the video of the host explaining product E can be identified, and the video labels "Product D" and "Product E" can be obtained, respectively.
[0080] It should be noted that since the video processing model can extract background feature information and motion feature information of the video to be processed, it can obtain a video label based on the content of the motion feature information, thereby enabling the video processing model to more accurately represent the video. For example, in the embodiment of the shopping live broadcast video on the above-mentioned e-commerce platform, the video processing model can extract background feature information representing the background of the live broadcast room and the product shelves, as well as motion feature information representing product D that the anchor is explaining, and then obtain "Product D" as the label of the video, thereby avoiding the video processing model mistakenly extracting the features of Product X on the product shelves in the background and using Product X as the label of the video.
[0081] The same video to be processed may include multiple video tags. For example, in the embodiment of the shopping live video on the above-mentioned e-commerce platform, when the host explains the products D and E that are used together at the same time, "Product D" and "Product E" can be used as video tags at the same time.
[0082] Step S1003: Display the video tag of the video to be processed.
[0083] After obtaining the video tag of the video to be processed, the video tag can be displayed at the corresponding position of the target object in the video to be processed. For example, in the embodiment of the shopping live video on the above-mentioned e-commerce platform, the video tag "Product D" can be displayed on the product D explained by the anchor.
[0084] In this embodiment, by receiving a to-be-processed video, the to-be-processed video is processed by a video processing model to obtain a video label of the to-be-processed video, wherein the video processing model is obtained by training an initial model, and the initial model is a model obtained by training sample data, and a supervision task is related to background feature information and motion feature information of the sample data. The video representation model can extract the background feature information and the motion feature information in the to-be-processed video, and the video representation model pays more attention to the motion feature information when extracting features, so that the video representation information can contain more motion feature information, thereby avoiding the problem that the video representation model is easily affected by the background in the to-be-processed video when extracting features, improving the accuracy of the video representation model in representing video content, and solving the problem that the video representation model in the prior art is not accurate in representing video content.
[0085] As an optional embodiment, after displaying the video label of the to-be-processed video, the above method further includes at least one of the following: recommending the to-be-processed video based on the label of the to-be-processed video; displaying the video label of the to-be-processed video, receiving correction information of the video label, and modifying the video label based on the correction information.
[0086] In the embodiment of the shopping live video on the e-commerce platform, different live videos are labeled according to the above scheme, so that different live videos have video labels of corresponding commodities, and further, according to the shopping habits of the user, the video related to the shopping habits of the user can be recommended to the user according to the content of the video label. Since the video label can accurately represent the content of the video, the user is not recommended to watch irrelevant shopping videos, and the user experience is improved.
[0087] In an optional embodiment, the above video label can be further corrected by manual means. In the embodiment of the shopping live video on the e-commerce platform, the consistency of the video label and the content of the live video can be corrected, and when the video label does not match the content of the live video, the accurate content of the video label can be used as correction information, and the original video label can be modified.
[0088] Embodiment 3
[0089] According to the embodiment of the present application, a data training method embodiment is also provided. Figure 3 is a flowchart of the processing of the video according to the embodiment of the present application, as shown in Figure 3 The method comprises the following steps:
[0090] In step S301, first feature information obtained by a to-be-trained model for extracting features from a sample video segment is acquired.
[0091] The model to be trained is a feature extraction model that needs to be optimized, such as a 3D video neural network. The sample video clip is a video clip containing multiple frames of images, which is used to train the model to be trained. Specifically, the complete video corresponding to the sample video clip can be divided into multiple video clips, each of which contains multiple consecutive frames of images. Any of these video clips can be selected as the sample video clip.
[0092] Step S302 : determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information.
[0093] In an optional implementation, background feature information can be obtained by separating the foreground and background in a frame image containing the background and extracting features from the background information. It can also be extracted by inputting a frame image containing the background in a sample video clip into an image feature extraction model.
[0094] The above-mentioned background loss function is used to perform contrastive learning on the first feature information and the second feature information. Specifically, the contrastive learning method is a video self-supervised learning method. By performing contrastive learning on unlabeled image features, different clips of the same video are brought closer in the feature space, and different clips from different videos are pushed further away in the feature space, thereby realizing self-supervised learning of the video. Figure 11 A schematic diagram of a video self-supervised learning method based on contrastive learning is shown in Figure 11 As shown, video 1 and video 2 are two different videos, and the videos are divided into multiple video segments. Segments 11 and 12 are sampled and enhanced from the video segment of video 1, and segment 21 is sampled and enhanced from the video segment of video 2. Segments 11, 12, and 21 are respectively input into the three-dimensional deep neural network to extract features. Image feature information 13 is extracted from segment 11 through the three-dimensional deep neural network Φ1, image feature information 14 is extracted from segment 12 through the three-dimensional deep neural network Φ2, and image feature information 22 is extracted from segment 21 through the three-dimensional deep neural network Φ3. Based on the comparative learning method, the image feature information 13 and image feature information 14 from the same video 1 can be brought closer in the feature space, while the image feature information 22 from different videos 2 can be pulled farther apart in the feature space.
[0095] In an optional embodiment, the sample video clip is input into a three-dimensional video feature extraction model to extract first feature information, and any key frame is input into a two-dimensional image feature extraction model to obtain second feature information. The first feature information and the second feature information are compared and learned based on a background loss function. This allows the first feature information extracted by the three-dimensional video feature extraction model and the second feature information extracted by the two-dimensional image feature extraction model for the same video to be brought closer together, while the first feature information and the second feature information from different videos are pushed further apart, thereby enabling the model to be trained to learn background information. The training of the model to be trained based on the background loss function improves the model's ability to recognize background information in videos.
[0096] Step S303: determine a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes the first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information.
[0097] The first motion feature information can be extracted by inputting the images following the sample video clip into a three-dimensional feature extraction model. The sample video clip is followed by another video clip located after the sample video clip. For example, if the sample video clip is the images from frames 1 to 5 of the video, the images used to extract the first motion feature information are the images from frames 6 to 10 of the same video. The second motion feature information can be predicted by inputting the first feature information into a neural network with an encoder-decoder structure.
[0098] The above-mentioned motion loss function is used to compare and learn the first feature information and the third feature information. According to the motion loss function, it can be determined that the first motion feature information and the second motion feature information correspond to each other in time and space positions as positive samples, and the first motion feature information and the second motion feature information are inconsistent in any one of time and space positions as negative samples, thereby enabling the model to be trained to learn the motion information.
[0099] It should be noted that, unlike static, coarse-grained background feature information, motion feature information is fine-grained, position-related feature information. Therefore, the sample video clip is predicted through a three-dimensional neural network to obtain the second motion feature information of the image after the sample video clip, and the first motion feature information is obtained by extracting the corresponding image after the real sample video clip through a three-dimensional neural network. By comparing and learning the first motion feature information and the second motion feature information, the model to be trained can learn fine-grained motion information.
[0100] Step S304: Optimize the model to be trained according to the background loss function and the motion loss function.
[0101] The background loss function and the motion loss function can be used to train and learn the to-be-trained model in two aspects of background information and motion information, and the to-be-trained model learns the information in two aspects of background information and motion information by decoupling the background information and the motion information in the sample video segment.
[0102] Specifically, the to-be-trained model can be optimized by using the obtained background loss function to obtain a first to-be-trained model that has learned the background information, and the to-be-trained model can be optimized by using the obtained motion loss function to obtain a second to-be-trained model that has learned the motion information, and a to-be-trained model that has comprehensively learned the background information and the motion information is obtained based on the first to-be-trained model and the second to-be-trained model, and the optimization of the to-be-trained model is realized.
[0103] The background loss function and the motion loss function can also be weighted to obtain a loss function used for optimizing the to-be-trained model, and the optimized loss function can be used to optimize the to-be-trained model, so that the to-be-trained model learns the information in two aspects of background information and motion information.
[0104] The above scheme uses background loss function and motion loss function to compare and learn the background information and motion information of the training model, thereby improving the recognition accuracy of the training model for the background information of the image and the motion information as the foreground. For example, when the existing self-supervised learning method learns a video of swimming, the image features obtained through feature extraction may be the swimming pool in the background of the image, or the athlete in the foreground of the image. Since the video itself is not labeled, the existing self-supervised learning method may use the swimming pool in the background of the image as a feature to identify the swimming video. After learning, all images containing the swimming pool may be judged as swimming. However, it is obvious that swimming does not necessarily exist in images containing the swimming pool. The background bias in the related technology leads to inaccurate self-supervised learning. In addition, for fine-grained motion information (such as whether the athlete's swimming posture is butterfly stroke or breaststroke), the existing self-supervised learning cannot perform fine learning. By using the data training method of the present embodiment, the background information of the swimming pool and the swimming information of the athlete as the foreground can be separately extracted from the video of the swimming sport, and the background loss function and the motion loss function can be determined to optimize the model to be trained, so that the model to be trained can identify the image containing the swimming pool as the background information and the motion information of the foreground athlete. Based on the background information of the swimming pool and the motion information of the athlete, it is further determined whether the image is the swimming sport that needs to be identified, thereby avoiding the background bias problem in the prior art. In addition, in the present embodiment, the motion loss function is obtained based on the motion vector of the image, and the model to be trained can perform refined learning of the motion features, thereby improving the recognition accuracy of the motion features. For example, the athlete's movement posture can be identified to identify the specific type of butterfly stroke or breaststroke.
[0105] In this embodiment, first feature information is obtained by extracting features from a sample video clip by a model to be trained; a background loss function is determined based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip; a motion loss function is determined based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip; and the model to be trained is optimized according to the background loss function and the motion loss function. By using the background loss function and the motion loss function to optimize the model to be trained, the model to be trained can learn both background information and motion information through comparative learning, thus avoiding the situation in the prior art where the training model can only learn background information but does not pay attention to background deviation caused by motion information, thereby improving the effect of video self-supervised learning and solving the technical problem in the prior art that video self-supervised learning is easily affected by image background.
[0106] As an optional embodiment, obtaining the first feature information obtained by extracting features from a sample video clip by the model to be trained includes: obtaining a target video and randomly extracting a video clip from the target video to obtain a sample video clip; inputting the sample video clip into the model to be trained to obtain the first feature information output by the model to be trained, wherein the model to be trained is a three-dimensional feature extraction model.
[0107] The target video can be the complete video used to extract the sample video clip. The target video contains multiple frames, and the sample video clip is a continuous multiple-frame image extracted arbitrarily from the multiple frames of the target video. For example, if the target video is 5 seconds long and has a frame rate of 20 bps, the target video can be cut into 10 video clips, each of which has 10 frames. The sample video clip can be any of the 10 video clips, that is, the images from frames 1 to 10, or the images from frames 11 to 20.
[0108] It should be noted that when the target video is a complete video, according to the above-mentioned comparative learning method, since there is no need to manually label each frame of the image, the complete video can be downloaded directly from the Internet, making it more convenient to obtain sample video clips and effectively utilizing massive Internet resources.
[0109] If the model to be trained is a three-dimensional feature extraction model, then the first feature information is three-dimensional feature information. For example, the sample video clip v i Input into the 3D feature extraction model to obtain the first 3D feature information x i , the first feature information x i for:
[0110]
[0111] Among them, C1 is the number of channels of the image, T1 is the time of the image, H1 is the height of the image, and W1 is the width of the image. The three-dimensional feature extraction model can be a C3D network, a ResNet-(2+1)D-26 network (i.e., a three-dimensional 26-layer ResNet network split by time and space convolution), etc.
[0112] As an optional embodiment, before determining the background loss function based on the first feature information and the second feature information, the method also includes: obtaining the second feature information, wherein the step of obtaining the second feature information includes: obtaining compressed data corresponding to the target video; extracting key frames corresponding to the sample video clips from the compressed data; extracting feature information of the key frames through a background feature extraction model to obtain the second feature information, wherein the background feature extraction model is a two-dimensional feature extraction model.
[0113] By converting the target video into compressed data, the data volume of the target video can be reduced, thereby reducing the data computational complexity of the multiple feature extraction models used in this application to extract the first feature information, the second feature information, the third feature information, and the first motion feature information. The compressed data may be in H.264 format or MPEG-4 encoding format. In an optional embodiment, the compressed data is in MPEG-4 encoding format. Figure 5 FIG. 5 is a schematic diagram of a video 50 in MPEG-4 encoding format, such as Figure 5 As described above, a video in the MPEG-4 encoding format includes key frames 501, P / B frames 502, and residual frames 503. During video decoding, the original video image can be restored based on the key frames 501, P / B frames 502, and residual frames 503. The key frames 501 represent static, coarse-grained video background information. The key frames 501 are extracted from an image containing video background information in the target video, and do not represent all images containing video background information in the target video. The P / B frames 502 contain motion vectors extracted from an image containing motion information in the target video, and can represent dynamic, fine-grained motion information. In the MPEG-4 encoding format, the key frames 501 representing video background information and the P / B frames 502 representing motion information have been decoupled. Therefore, there is no need to further decouple the sample video clips in the MPEG-4 encoding format to obtain separate video background information and motion information. The key frames 501 in the MPEG-4 encoding format can directly extract the above-mentioned second feature information through a two-dimensional feature extraction model.
[0114] Specifically, according to the time corresponding to the sample video clip in the time axis of the target video, the key frame within the event time range of the sample video clip is extracted from the MPEG-4 encoded format video. If there are multiple key frames, one of the key frames is selected and input into the two-dimensional background feature extraction model to extract the second feature information z i , the second feature information z i is two-dimensional feature information, for example, the second feature information z i for:
[0115]
[0116] Among them, C2 is the number of channels of the image, H2 is the height of the image, and W2 is the width of the image. Compared with the above three-dimensional first feature information x i , the second feature information z i Reduces the dimension of time.
[0117] The above-mentioned background feature extraction model can be a ResNet2D-10 feature extraction model (i.e., a two-dimensional 10-layer ResNet network).
[0118] As an optional embodiment, extracting key frames corresponding to the sample video clip from compressed data includes: determining the key frames of the target video clip from the compressed data, wherein the start frame of the target video clip is earlier than the start frame of the sample video clip by a first preset number of frames, and the end frame of the target video clip is later than the end frame of the sample video clip by a second preset number of frames; and extracting any one frame from the key frames of the target video as the key frame corresponding to the sample video clip.
[0119] The target video segment is the video segment used for key frame extraction. Since the sample video segment is compressed, the number of key frames containing background information is reduced, and the sample video segment may not contain key frames. Therefore, the video segment used for key frame extraction has more images than the sample video segment. For example, the first preset number of frames may be 10 frames, the second preset number of frames may be 5 frames, and the sample video segment is the 20th to 30th frames of the complete video. Then, the starting frame of the target video segment is the 10th frame of the complete video, and the ending frame of the target video segment is the 35th frame of the complete video. The target video segment is determined to be the images from the 10th to 35th frames.
[0120] The target video clip may contain one key frame or multiple key frames. If the target video clip contains only one key frame, the key frame is determined as the key frame corresponding to the sample video clip and can be used to extract the second feature information. If the target video clip contains multiple key frames, any one of the multiple key frames is selected as the key frame corresponding to the sample video clip. For example, if the target video clip is an image from the 10th to the 35th frame, where the 13th, 15th, and 20th frames are all key frames, the key frame corresponding to the sample video clip can be the 13th frame, or the 15th frame, and the 20th frame. If the target video clip does not contain a key frame, the number of frames of the sample video clip is expanded by adjusting the data range of the first preset number of frames and the second preset number of frames to obtain the key frame.
[0121] It should be noted that the first preset frame number and the second preset frame number can be determined according to the number of image frames in the compressed data. The first preset frame number and the second preset frame number can be the same or different, which is not limited here.
[0122] As an optional embodiment, a background loss function is determined based on the first feature information and the second feature information, including: performing global average pooling processing on the three-dimensional first feature information and the two-dimensional second feature information, respectively, to obtain one-dimensional first feature information and one-dimensional second feature information; mapping the one-dimensional first feature information and the one-dimensional second feature information into a single output to obtain the mapped one-dimensional first feature information and one-dimensional second feature information; determining a first noise contrast estimation loss function based on the mapped one-dimensional first feature information and the one-dimensional second feature information, and determining the first noise contrast estimation loss function as the background loss function.
[0123] Global average pooling can reduce the dimensionality of feature information through the pooling layer. The mapping of the one-dimensional first feature information and the one-dimensional second feature information to a single output can be achieved through an MLP network (Multi-Layer Perceptron). The multi-layer perceptron is an artificial neural network (ANN). Its structure includes an input layer, one or more hidden layers, and an output layer. The output of each hidden layer is transformed by an activation function.
[0124] In an optional embodiment, based on Figure 4 The data processing method framework shown in FIG. 4 is to input a sample video clip 402 into a three-dimensional V-network 405 to extract first feature information 408. The first three-dimensional feature information x i for:
[0125]
[0126] The first feature information 408 is input to the pooling layer 410 for global average pooling to obtain the one-dimensional first feature information 413 after dimensionality reduction. The expression of the one-dimensional first feature information is:
[0127]
[0128] The background image 401 is input into the two-dimensional I-network 404 to extract the second feature information 407. The two-dimensional second feature information z i The expression is:
[0129] The second feature information 407 is subjected to global average pooling processing by the pooling layer 409 to obtain the one-dimensional second feature information 412 after dimensionality reduction. The expression is:
[0130]
[0131] Input the one-dimensional first feature information and the one-dimensional second feature information into the MLP network respectively and MLP network Get the one-dimensional first feature information after mapping:
[0132]
[0133] And the one-dimensional second feature information after mapping:
[0134]
[0135] The one-dimensional first feature information and one-dimensional second feature information Substituting the first noise contrast estimation loss function (i.e., InfoNCE loss function) into the background loss function J I The expression can be:
[0136]
[0137] in, is the one-dimensional second feature information after mapping (i.e. z i The corresponding one-dimensional feature vector, The one-dimensional first feature information after mapping (i.e. the one-dimensional feature vector corresponding to xi), B represents the batch size, Represents the feature vector and eigenvectors The cosine similarity measure between:
[0138]
[0139] As an optional embodiment, before determining the motion loss function based on the first feature information and the third feature information, the method also includes: obtaining the third feature information, wherein the step of obtaining the third feature information includes: extracting the motion vector of multiple consecutive frames of images after the sample video clip; determining the first motion feature information according to the motion vector based on the motion feature extraction model; and determining the first motion feature information as the third feature information.
[0140] Motion vectors can be extracted by separating the foreground and background of multiple consecutive image frames following the sample video clip. In an optional embodiment, when the sample video clip is compressed data in the MPEG-4 encoding format, the P / B frames in the MPEG-4 encoding format represent the motion vectors extracted from the sample video clip. Extracting the P / B frames from the compressed sample video clip can obtain motion vectors for the multiple consecutive image frames.
[0141] In an optional embodiment, based on Figure 4 The data processing method shown in the framework, the motion vector v i403 is extracted from multiple frames of images after the sample video clip, and M-network 406 is a motion feature extraction model that extracts the motion vector v of the continuous multiple frames of images. i Input M-network 406 to extract the three-dimensional first motion feature information v i
[0142] v i ∈(C3×T3×H3×W3);
[0143] The three-dimensional first motion feature information v i This is the third feature information 415.
[0144] As an optional embodiment, determining the motion loss function based on the first feature information and the third feature information includes: predicting the future second motion feature information based on the first feature information of the sample video clip; and determining the motion loss function based on the second motion feature information and the third feature information.
[0145] The future second motion feature information can be predicted by inputting the first feature information of the sample video clip into a neural network having an encoder-decoder structure. The neural network having an encoder-decoder structure can be a ConvGRU or Transformer neural network. The time corresponding to the second motion feature information should coincide with the time on the video timeline when the image used to extract the first motion feature information was generated.
[0146] In an optional embodiment, based on Figure 4 The data processing method shown in the framework is as follows: the sample video clip is a multi-frame image with a time period of 1-T seconds on the time axis of the video. The sample video clip is input into the three-dimensional video neural network to extract the first feature information x i , the first three-dimensional feature information x i for:
[0147]
[0148] The three-dimensional first feature information xi is input into the encoder-decoder network 411 to predict the second motion feature information corresponding to the image at time T+1 to T+S seconds on the time axis of the video.
[0149] Here, S can be equal to T-1.
[0150] The images from T+1 to T+S seconds after the sample video clip in the video are input into the three-dimensional feature extraction model to extract the corresponding three-dimensional first motion feature information v i :
[0151] v i ∈(C3×T3×H3×W3);
[0152] A motion loss function is determined based on the second motion feature information and the first motion feature information, and comparative learning is performed on the model to be trained. If the first motion feature information and the second motion feature information correspond in time and space, they are determined to be positive samples. If the first motion feature information and the second motion feature information are inconsistent in time and space, they are determined to be negative samples. Through comparative learning, the model to be trained can learn fine-grained motion information.
[0153] As an optional embodiment, a motion loss function is determined based on the second motion feature information and the third feature information, including: performing a single output mapping on the second motion feature information and the third feature information to obtain the mapped second motion feature information and the third feature information; determining a second noise contrast estimation loss function based on the mapped second motion feature information and the third feature information, and determining the second noise contrast estimation loss function as the motion loss function.
[0154] The above single output mapping of the second motion feature information and the third feature information can be achieved through the MLP network. Figure 4 The data processing method shown in the framework is as follows: the third feature information 415 and the second motion feature information 414 are input into the MLP network and MLP network The second motion feature information is mapped to obtain mapped second motion feature information:
[0155]
[0156] The third feature information 415 is the three-dimensional first motion feature information v i , the first motion feature information v i Mapping is performed to obtain the mapped first motion feature information:
[0157]
[0158] Substituting the mapped second motion feature information and the third feature information into the InfoNCE loss function, the second noise contrast estimation loss function is obtained as follows:
[0159]
[0160] in, represent The jth column of for , B represents a batch size, N=T3xH3xW3(T is a time corresponding to the first feature information, H is a height corresponding to the first feature information, and W is a width corresponding to the first feature information), i=1-B, k=1-B, j=1-N, and l=1-N.
[0161] As an optional embodiment, the model to be trained is optimized according to the background loss function and the motion loss function, including: constructing a target loss function by the background loss function and the motion loss function based on preset hyperparameters; and solving the target loss function to optimize the model to be trained.
[0162] The target loss function is a loss function used to optimize the model to be trained, and the target loss function used to optimize the model to be trained is obtained by weighting the background loss function and the motion loss function, so that the model to be trained learns information in both the background information and the motion information.
[0163] For example, the target loss function J can be: J=(1-α)J I +αJ M , where J I is the background loss function, J M is the motion loss function, and α is a preset hyperparameter (i.e., a weight coefficient of motion prediction).
[0164] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0165] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and a necessary general hardware platform, and of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk), and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device) to execute the methods described in the embodiments of the present application.
[0166] Embodiment 4
[0167] According to the embodiments of the present application, a data processing method embodiment is also provided, Figure 6 is a flow chart of a data processing method according to an embodiment 4 of the present application, as shown in the figure, the method comprises the following steps: Figure 6
[0168] Step S601, a sample video segment is extracted from a target video.
[0169] The sample video segment is a video segment with multiple frames of images, used for training learning of a to-be-trained model.
[0170] The target video can be a complete video used for extracting the sample video segment, the target video contains multiple frames of images, and the sample video segment is multiple continuous frames of images randomly extracted from the multiple frames of images of the target video. For example, the target video is a video with a time length of 5s and a frame rate of 20bpf, the target video can be cut into 10 video segments, each of the 10 video segments has 10 frames of images, and the sample video segment can be any one of the 10 video segments, i.e., the images from the 1st frame to the 10th frame or the images from the 11th frame to the 20th frame.
[0171] It should be noted that when the target video is a complete video, according to the above-mentioned contrast learning method, since manual labeling of each frame of image is not required, the complete video can be directly downloaded from the Internet, and the acquisition of the sample video segment is more convenient, and the vast amount of Internet resources is effectively utilized.
[0172] Step S602, first feature information obtained by performing feature extraction on the sample video segment by a to-be-trained model is acquired.
[0173] The to-be-trained model is a feature extraction model that needs to be optimized, for example, a three-dimensional video neural network.
[0174] Step S603, background feature information of images in the sample video segment is acquired.
[0175] The background feature information can be obtained by separating the foreground and the background in a frame of image containing the background and extracting the features in the background information, or by inputting a frame of image containing the background in the sample video segment into an image feature extraction model.
[0176] The above-mentioned background loss function is used for contrast learning (contrastive learning) of the first feature information and the second feature information. Specifically, the method of contrast learning is a video self-supervised learning method, which realizes the self-supervised learning of the video by performing contrast learning on the unlabeled image features, and realizes the approximation of different segments of the same video in the feature space and the distancing of different segments from different videos in the feature space. Figure 11 is a schematic diagram of a video self-supervised learning method based on contrast learning, as shown in the figure, Figure 11 As shown, video 1 and video 2 are two different videos, the video is divided into multiple video segments, segment 11 and segment 12 are sampled and enhanced from the video segments of video 1, segment 21 is sampled and enhanced from the video segments of video 2, segment 11, segment 12 and segment 21 are respectively input into a three-dimensional deep neural network to extract features, segment 11 is extracted to image feature information 13 through three-dimensional deep neural network Φ1, segment 12 is extracted to image feature information 14 through three-dimensional deep neural network Φ2, segment 21 is extracted to image feature information 22 through three-dimensional deep neural network Φ3, based on the method of contrast learning, the image feature information 13 and the image feature information 14 from the same video 1 can be pulled closer in the feature space, while the image feature information 22 from different video 2 is pulled away in the feature space.
[0177] In an optional embodiment, the above-mentioned sample video segment is input into a three-dimensional video neural network to extract first feature information, and any key frame is input into a two-dimensional image neural network to obtain second feature information. Based on the background loss function, the contrast learning of the first feature information and the second feature information can pull the first feature information extracted by the three-dimensional video neural network and the second feature information extracted by the two-dimensional image neural network closer, while the first feature information and the second feature information from different videos are pushed away, so as to make the to-be-trained model learn the background information. Based on the background loss function, the training of the to-be-trained model improves the recognition ability of the to-be-trained model to the background information of the video.
[0178] Step S604, obtaining the first motion feature information of the image after the sample video segment.
[0179] The first motion feature information can be extracted by inputting the image containing motion information after the sample video segment into a three-dimensional feature extraction model. The image after the sample video segment can be understood as another video segment located after the sample video segment in the time axis of the video, for example, the sample video segment is the image of the first frame to the fifth frame in the video, and the multiple frames of images in the time axis of the video are between 1-10 seconds, and the image for extracting the first motion feature information is the image of the sixth frame to the tenth frame in the same video, and the multiple frames of images in the time axis of the video are between 11-15 seconds. The second motion feature information can be predicted by inputting the first feature information into a neural network with an encoding-decoding structure (such as ConvGRU or Transformer neural network).
[0180] The above-mentioned motion loss function is used to compare and learn the first feature information and the third feature information. According to the motion loss function, it can be determined that the first motion feature information and the second motion feature information that correspond to each other in time and space are positive samples, and the first motion feature information and the second motion feature information that are inconsistent in time and space are negative samples, thereby enabling the model to be trained to learn the motion information.
[0181] It should be noted that, unlike static, coarse-grained background feature information, motion feature information is fine-grained, position-related feature information. Therefore, the sample video clip is predicted through a three-dimensional neural network to obtain the second motion feature information of the image after the sample video clip, and the first motion feature information is obtained by extracting the corresponding image after the real sample video clip through a three-dimensional neural network. By comparing and learning the first motion feature information and the second motion feature information, the model to be trained can learn fine-grained motion information.
[0182] Step S605 : training the to-be-trained model based on the first feature information, the background feature information, and the first motion feature information.
[0183] The background loss function and the motion loss function can be used to train the model to be trained in terms of background information and motion information, so that the model to be trained can learn both background information and motion information.
[0184] Specifically, the obtained background loss function can be used to optimize the model to be trained to obtain a first model to be trained that has learned background information, and the obtained motion loss function can be used to optimize the model to be trained to obtain a second model to be trained that has learned motion information. Based on the first model to be trained and the second model to be trained, a model to be trained that has comprehensively learned background information and motion information is obtained, thereby achieving optimization of the model to be trained. The background loss function and the motion loss function can also be weighted to obtain a loss function for optimizing the model to be trained. The optimized loss function can be used to optimize the model to be trained so that the model to be trained learns both background information and motion information.
[0185] In this embodiment, the training model is optimized by adopting the background loss function and the motion loss function, so that the training model can learn both background information and motion information through comparative learning, avoiding the situation in which the training model in the prior art can only learn background information but does not pay attention to the background deviation caused by motion information, improving the effect of video self-supervised learning, and solving the technical problem in the prior art that video self-supervised learning is easily affected by the image background.
[0186] Example 5
[0187] According to an embodiment of the present invention, a device for implementing the above data processing method is also provided. Figure 7 A data processing device according to embodiment 5 of the present application, such as Figure 7 As shown, the apparatus 700 includes:
[0188] An acquisition module 71 is used to obtain first feature information obtained by extracting features from a sample video clip by the model to be trained; a first determination module 72 is used to determine a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information; a second determination module 73 is used to determine a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; an optimization module 74 is used to optimize the model to be trained according to the background loss function and the motion loss function.
[0189] It should be noted that the acquisition module 71, the first determination module 72, the second determination module 73, and the optimization module 74 correspond to steps S301 to S302 in Example 3. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 3. It should be noted that the above modules, as part of the apparatus, can be run in the computing device 10 provided in Example 1.
[0190] As an optional embodiment, the above-mentioned acquisition module includes: an extraction sub-module, used to obtain the target video and randomly extract video clips from the target video to obtain sample video clips; a first training sub-module, used to input the sample video clips into the model to be trained to obtain the first feature information output by the model to be trained, wherein the model to be trained is a three-dimensional feature extraction model.
[0191] As an optional embodiment, the above-mentioned device also includes: a second feature information acquisition module, used to obtain the second feature information, wherein the second feature information acquisition module includes: a compressed data acquisition submodule, used to obtain compressed data corresponding to the target video; a key frame acquisition submodule, used to extract key frames corresponding to the sample video clips from the compressed data; a second training submodule, used to extract feature information of the key frames through a background feature extraction model to obtain the first and second feature information, wherein the background feature extraction model is a two-dimensional feature extraction model.
[0192] As an optional embodiment, the above-mentioned key frame acquisition submodule includes: a key frame determination submodule, used to determine the key frame of the target video clip from the compressed data, wherein the starting frame of the target video clip is earlier than the starting frame of the sample video clip by a first preset number of frames, and the ending frame of the target video clip is later than the ending frame of the sample video clip by a second preset number of frames; a key frame extraction submodule, used to extract any one frame from the key frames of the target video as the key frame corresponding to the sample video clip.
[0193] As an optional embodiment, the above-mentioned first determination module includes: a first pooling processing submodule, which is used to perform global average pooling processing on the three-dimensional first feature information and the two-dimensional second feature information, respectively, to obtain one-dimensional first feature information and one-dimensional second feature information; a first mapping submodule, which is used to perform a single mapping on the one-dimensional first feature information and the one-dimensional second feature information, to obtain the mapped one-dimensional first feature information and one-dimensional second feature information; a background loss function determination submodule, which is used to determine the first noise contrast estimation loss function based on the mapped one-dimensional first feature information and the one-dimensional second feature information, and determine the first noise contrast estimation loss function as the background loss function.
[0194] As an optional embodiment, the above-mentioned device also includes: a third feature information acquisition module, used to obtain third feature information, wherein the third feature information acquisition module includes: a motion vector extraction submodule, used to extract the motion vector of multiple consecutive frames of images after the sample video clip; a first motion feature information determination submodule, used to determine the first motion feature information according to the motion vector based on the motion feature extraction model; and a third feature information determination submodule, used to determine the first motion feature information as the third feature information.
[0195] As an optional embodiment, the above-mentioned second determination module includes: a second motion feature information prediction submodule, used to predict future second motion feature information based on the first feature information of the sample video clip; and a motion loss function determination submodule, used to determine the motion loss function based on the second motion feature information and the third feature information.
[0196] As an optional embodiment, the motion loss function determination submodule includes: a second mapping submodule, used to perform a single mapping on the second motion feature information and the third feature information to obtain the mapped second motion feature information and the third feature information; a motion loss function acquisition submodule, used to determine the second noise contrast estimation loss function based on the mapped second motion feature information and the third feature information, and determine the second noise contrast estimation loss function as the motion loss function.
[0197] As an optional embodiment, the above-mentioned optimization module includes: a construction submodule, which is used to construct a target loss function through a background loss function and a motion loss function based on preset hyperparameters; and a solution submodule, which is used to solve the target loss function to optimize the model to be trained.
[0198] It should be noted that the optional or preferred implementations of this embodiment can be found in the relevant descriptions in Examples 1, 2 and 3, and will not be repeated here.
[0199] Example 6
[0200] According to an embodiment of the present invention, a device for implementing the above data processing method is also provided. Figure 8 A data processing device according to embodiment 6 of the present application, such as Figure 8 As shown, the apparatus 800 includes:
[0201] The extraction module 81 is used to extract sample video clips from the target video; the first acquisition module 82 is used to obtain the first feature information obtained by the model to be trained through feature extraction of the sample video clips; the second acquisition module 83 is used to obtain the background feature information of the image in the sample video clip; the third acquisition module 84 is used to obtain the first motion feature information of the image after the sample video clip; the training module 85 is used to train the model to be trained based on the first feature information, background feature information and first motion feature information.
[0202] It should be noted that the extraction module 81, the first acquisition module 82, the second acquisition module 83, the third acquisition module 84, and the training module 85 correspond to steps S601 to S605 in Example 4. The examples and application scenarios implemented by the five modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiments 1, 2, and 3. It should be noted that the above-mentioned modules, as part of the apparatus, can be run in the computing device 10 provided in Example 1.
[0203] It should be noted that the optional or preferred implementations of this embodiment can be found in the relevant descriptions in Examples 1, 2 and 3, and will not be repeated here.
[0204] Example 7
[0205] An embodiment of the present invention further provides a storage medium, which includes a stored program, wherein the program code controls the device where the storage medium is located to execute the video processing method when the program is running.
[0206] Optionally, in this embodiment, the above-mentioned storage medium may be located in any one of the computing devices in the computing device group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0207] Optionally, the storage medium is configured to store program code for executing the following steps: receiving a video to be processed; performing feature extraction on the video to be processed through a video representation model to obtain video representation information of the video to be processed, wherein the video representation model is obtained by training an initial model, the initial model is a model obtained by training sample data, and the training task is related to the background feature information and motion feature information of the sample data; outputting the video representation information of the video to be processed, wherein the video representation information includes the background feature information of the video to be processed and the motion feature information of the video to be processed.
[0208] Optionally, the storage medium is configured to store program code for executing the following steps: after outputting the video representation information of the video to be processed, the method further includes at least one of the following: performing video classification on the video to be processed based on the video representation information to obtain a video label of the video to be processed; performing object detection on the video to be processed based on the video representation information to obtain a target object in the video to be processed; performing object tracking on the video to be processed based on the video representation information to obtain the position of the target object in each frame image of the video to be processed.
[0209] Optionally, the storage medium is configured to store program code for executing the following steps: obtaining an initial model, the steps of obtaining the initial model including: obtaining a target loss function of the model to be trained, wherein the target loss function is composed of background feature information and motion feature information; optimizing the model to be trained by solving the target loss function to obtain the initial model.
[0210] Optionally, the storage medium is configured to store program code for executing the following steps: obtaining a target loss function of the model to be trained, including: obtaining first feature information obtained by extracting features from a sample video clip by the model to be trained; determining a background loss function based on the first feature information and second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information; determining a motion loss function based on the first feature information and third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; determining the sum of the background loss function and the motion loss function as the target loss function. Optimizing the model to be trained based on the background loss function and the motion loss function to obtain an initial model.
[0211] It should be noted that the optional or preferred implementations of this embodiment can be found in the relevant descriptions in Examples 1, 2 and 3, and will not be repeated here.
[0212] Example 8
[0213] According to an embodiment of the present application, an embodiment of a computer terminal is also provided. The computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0214] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.
[0215] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the video processing method of the application: receiving the video to be processed; performing feature extraction on the video to be processed through a video representation model to obtain video representation information of the video to be processed, wherein the video representation model is obtained by training an initial model, and the initial model is a model obtained by training sample data, and the supervision task is related to the background feature information and motion feature information of the sample data; outputting the video representation information of the video to be processed, wherein the video representation information includes the background feature information of the video to be processed and the motion feature information of the video to be processed.
[0216] Optionally, Figure 9 is a structural block diagram of a computer terminal according to Example 8 of the present application, such as Figure 9 As shown, the computer terminal 900 may include: one or more (only one is shown in the figure) processors 902 , a memory 904 , and a peripheral interface 906 .
[0217] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the video processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned video processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 900 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0218] The processor is used to run the program. When the program is running, any of the above data processing methods is executed. The processor can call the information and application stored in the memory through the transmission device to perform the following steps:
[0219] Receive a video to be processed; extract features from the video to be processed through a video representation model to obtain video representation information of the video to be processed, wherein the video representation model is obtained by training an initial model, the initial model is a model obtained by training sample data, and the supervision task is related to the background feature information and motion feature information of the sample data; output the video representation information of the video to be processed, wherein the video representation information includes the background feature information of the video to be processed and the motion feature information of the video to be processed.
[0220] It can be understood by those skilled in the art that Figure 9 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 9 It does not limit the structure of the above electronic device. For example, the computer terminal 900 may also include Figure 9 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 9 Different configurations shown.
[0221] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0222] Example 9
[0223] According to an embodiment of the present application, a data training system is also provided, which includes: a processor; and a memory connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining first feature information obtained by extracting features from a sample video clip by a model to be trained; determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information; determining a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and optimizing the model to be trained according to the background loss function and the motion loss function.
[0224] It should be noted that the optional or preferred implementation of this embodiment can be found in the relevant descriptions in Examples 1 to 2, and will not be repeated here.
[0225] Example 10
[0226] According to an embodiment of the present invention, a device for implementing the above-mentioned video processing method is also provided. Figure 12 A video processing device according to embodiment 10 of the present application, such as Figure 12 As shown, the apparatus 1200 includes:
[0227] The first receiving module 1210 is used to receive the video to be processed; the feature extraction module 1220 is used to extract features from the video to be processed through a video representation model to obtain video representation information of the video to be processed, wherein the video representation model is obtained by training an initial model, and the initial model is a model obtained by training sample data, and the training task is related to the background feature information and motion feature information of the sample data; the output module 1230 is used to output the video representation information of the video to be processed, wherein the video representation information includes the background feature information of the video to be processed and the motion feature information of the video to be processed.
[0228] It should be noted that the first receiving module 1210, feature extraction module 1220, and output module 1230 described above correspond to steps S201 to S203 in Example 1. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Examples 1, 2, and 3. It should be noted that the above modules, as part of the apparatus, can be run in the computing device 10 provided in Example 1.
[0229] As an optional embodiment, the above-mentioned device also includes at least one of the following: a classification module, used to perform video classification on the video to be processed based on the video representation information to obtain the video label of the video to be processed; a detection module, used to perform object detection on the video to be processed based on the video representation information to obtain the target object in the video to be processed; a tracking module, used to perform object tracking on the video to be processed based on the video representation information to obtain the position of the target object in each frame image of the video to be processed.
[0230] As an optional embodiment, the above-mentioned device also includes: an initial model acquisition module, used to obtain the initial model, the initial model acquisition module includes: a target loss function acquisition sub-module, used to obtain the target loss function of the model to be trained, wherein the target loss function is composed of background feature information and motion feature information; an optimization sub-module, used to optimize the model to be trained by solving the target loss function to obtain the initial model.
[0231] As an optional embodiment, the target loss function acquisition submodule includes: a first feature acquisition submodule, used to obtain the first feature information obtained by extracting features from the sample video clip by the model to be trained; a background loss function determination submodule, used to determine the background loss function based on the first feature information and the second feature information, wherein the second feature information includes the background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information; a motion loss function determination submodule, used to determine the motion loss function based on the first feature information and the third feature information, wherein the third feature information includes the first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; a determination submodule, used to determine the sum of the background loss function and the motion loss function as the target loss function.
[0232] It should be noted that the optional or preferred implementations of this embodiment can be found in the relevant descriptions in Examples 1, 2 and 3, and will not be repeated here.
[0233] Example 11
[0234] According to an embodiment of the present invention, a device for implementing the above-mentioned video processing method is also provided. Figure 13 A video processing device according to embodiment 11 of the present application, such as Figure 13 As shown, the apparatus 1300 includes:
[0235] The second receiving module 1310 is used to receive the video to be processed; the first processing module 1320 is used to process the video to be processed through a video processing model to obtain a video label of the video to be processed, wherein the video processing model is obtained by training an initial model, and the initial model is a model obtained by training sample data, and the supervision task is related to the background feature information and motion feature information of the sample data; the first display module 1330 is used to display the video label of the video to be processed.
[0236] It should be noted that the second receiving module 1310, processing module 1320, and display module 1330 described above correspond to steps S1001 to S1003 in Example 2. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Examples 1, 2, and 3. It should be noted that the above modules, as part of the apparatus, can be run in the computing device 10 provided in Example 1.
[0237] As an optional embodiment, the above-mentioned device also includes at least one of the following: a recommendation module, used to recommend the video to be processed based on the label of the video to be processed; a modification module, used to display the video label of the video to be processed, receive proofreading information of the video label, and modify the video label based on the proofreading information.
[0238] It should be noted that the optional or preferred implementations of this embodiment can be found in the relevant descriptions in Examples 1, 2 and 3, and will not be repeated here.
[0239] Example 12
[0240] According to an embodiment of the present invention, a data processing method embodiment is also provided. Figure 14 : is a flow chart of a data training method according to embodiment 12 of the present invention, Figure 14 As shown, the method includes the following steps:
[0241] Step S1401: Receive live video.
[0242] The live video above is a video that needs to be represented for downstream tasks. The downstream tasks may be detection and recognition based on the video representation, as well as video labeling.
[0243] Live videos can be videos or video clips from various live streaming platforms. For example, a live video could be an e-commerce shopping video, and its corresponding downstream tasks could be object detection and object tracking to identify products in the live video and make accurate recommendations. Another example is an entertainment live video, and its corresponding downstream task could be video tagging to determine the video type (sports, beauty, baby girls, movies, etc.) through video tagging, and then make accurate recommendations for entertainment videos.
[0244] Step S1402: Process the live video through a video processing model to obtain a video tag of the live video, wherein the video tag is used to represent the product type of the target object in the live video. The video processing model is obtained by training an initial model, and the initial model is a model obtained by training sample data. The training task is related to the background feature information and motion feature information of the sample data.
[0245] The above-mentioned video representation model is a feature extraction model for extracting features from live videos. The video representation model can be a neural network model with a convolutional layer structure. The video label is used to represent the product type of the target object in the live video. The product type of the target object can be the product recommended or demonstrated in the live video. For example, a live shopping video of a product can use the name of the product being recommended as the video label. The video representation model can be obtained by further training through supervised learning on the basis of the above-mentioned initial model. Specifically, the above-mentioned initial model can be further trained by using sample data with preset topic labels, so that the video representation model can extract features from the preset topic content and use the preset topic as a video label.
[0246] The above-mentioned initial model can be a feature extraction model obtained after learning and training through sample data. The sample data used by the initial model is a video clip. During the training process, the supervision information includes background feature information and motion feature information, that is, the supervision task is related to both the background feature information and the motion feature information, so that the initial model learns both the background feature information and the motion feature information, avoiding the background bias problem that occurs in the learning of sample data in the prior art. In an optional embodiment, a live video clip can be obtained as sample data, and the foreground and background of a frame image containing the background in the live video clip are separated to obtain background information and motion information, and the features in the background information are extracted to obtain background feature information, and the features in the motion information are extracted to obtain motion feature information. The background feature information and the motion feature information are used to train the initial model respectively, so that the initial model can learn both background information and motion information. The above-mentioned initial model can also be trained using the method in Example 3 of the present application.
[0247] In an optional embodiment, the live video is a shopping live video on an e-commerce platform, the shopping live video includes a live room background, a goods shelf in the background, and a product being explained by the host. The preset topics can be the explanation of the product D by the host and the explanation of the product E by the host. The video feature extraction by the video representation model can identify the video of the explanation of the product D by the host and the video of the explanation of the product E by the host, and obtain the video labels “product D” and “product E” respectively. It should be noted that, since the video representation model can pay more attention to the motion information features in the live video, the video label determined according to the features extracted and output by the video representation model can accurately represent the content of the live video. In the above embodiment of the shopping live video, since the video representation model pays more attention to the motion feature information, the video representation model accurately extracts the features of the product being explained by the host in the foreground as the video label, avoids the influence of the products on the goods shelf in the background, and can more accurately represent the product type of the target object in the live video.
[0248] It should be noted that, since the video processing model can extract the background feature information and the motion feature information of the live video, the video label can be obtained based on the content in the motion feature information, and the video processing model can more accurately represent the live video, for example, in the above embodiment of the shopping live video on the e-commerce platform, the video processing model can extract the background feature information representing the live room background and the goods shelf, and the motion feature information representing the product D explained by the host, and then obtain “product D” as the label of the video, avoiding the video processing model from extracting the features of the product X in the goods shelf in the background and taking the product X as the label of the video.
[0249] In the same live video, multiple video labels can be included, for example, in the above embodiment of the shopping live video on the e-commerce platform, when the host explains the product D and the product E at the same time, “product D” and “product E” can be taken as the labels of the live video at the same time.
[0250] In step S1403, the video label of the live video is displayed.
[0251] After obtaining the video label of the live video, the video label can be displayed at the corresponding position of the target object in the live video, for example, in the above embodiment of the shopping live video on the e-commerce platform, the video label “product D” can be displayed on the product D explained by the host.
[0252] In this embodiment, live video is received and processed by a video processing model to obtain a video tag of the live video, wherein the video tag is used to represent the product type of the target object in the live video. The video processing model is obtained by training an initial model, and the initial model is a model obtained by training sample data. The training task is related to the background feature information and motion feature information of the sample data, so that the video representation model can extract the background feature information and motion feature information in the video to be processed respectively, and the video representation model pays more attention to the motion feature information during feature extraction, so that the video representation information can contain more motion feature information, thereby avoiding the problem that the video representation model is easily affected by the background in the live video during feature extraction, improving the accuracy of the video representation model in representing the video content, and solving the problem of inaccurate representation of the video content by the video representation model in the prior art.
[0253] Example 13
[0254] According to an embodiment of the present invention, a device for implementing the above-mentioned video processing method is also provided. Figure 15 A video processing device according to embodiment 13 of the present application, such as Figure 15 As shown, the apparatus 1500 includes:
[0255] The third receiving module 1510 is used to receive live video; the second processing module 1520 is used to process the live video through a video processing model to obtain a video tag of the live video, wherein the video tag is used to represent the product type of the target object in the live video, and the video processing model is obtained by training an initial model, and the initial model is a model obtained by training sample data, and the training task is related to the background feature information and motion feature information of the sample data; the second display module 1530 is used to display the video tag of the live video.
[0256] It should be noted that the third receiving module 1510, the second processing module 1520, and the second display module 1530 described above correspond to steps S1401 to S1403 in Example 12. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Examples 1, 2, 3, and 12. It should be noted that the above modules, as part of the apparatus, can be run in the computing device 10 provided in Example 1.
[0257] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0258] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0259] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0260] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0261] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0262] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
[0263] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A video processing method, characterized in that: include: Receive videos to be processed; Extracting features from the video to be processed using a video representation model to obtain video representation information of the video to be processed, wherein the video representation model is obtained by training an initial model, the initial model is a target loss function based on the model to be trained, and the model to be trained is optimized, the target loss function being composed of background feature information of a sample video clip and motion feature information of the sample video clip; Outputting video representation information of the video to be processed, wherein the video representation information includes background feature information of the video to be processed and motion feature information of the video to be processed; The method further includes: obtaining first feature information obtained by extracting features from the sample video clip by the model to be trained; determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information of the image in the sample video clip; determining a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and determining the sum of the background loss function and the motion loss function as the target loss function.
2. The method according to claim 1, characterized in that After outputting the video representation information of the video to be processed, the method further includes at least one of the following: Classifying the video to be processed based on the video representation information to obtain a video label of the video to be processed; Performing object detection on the video to be processed based on the video representation information to obtain a target object in the video to be processed; Object tracking is performed on the video to be processed based on the video representation information to obtain the position of the target object in each frame image of the video to be processed.
3. The method according to claim 1, characterized in that The method further includes: obtaining the initial model, and the step of obtaining the initial model includes: Obtaining the target loss function of the model to be trained; The model to be trained is optimized by solving the target loss function to obtain the initial model.
4. A video processing method, characterized in that: include: Receive videos to be processed; Processing the video to be processed using a video processing model to obtain a video label for the video to be processed, wherein the video processing model is obtained by training an initial model, the initial model is obtained by optimizing the model to be trained based on a target loss function of the model to be trained, the target loss function being composed of background feature information of a sample video clip and motion feature information of the sample video clip; Display the video tag of the video to be processed; The method further includes: obtaining first feature information obtained by extracting features from the sample video clip by the model to be trained; determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information of the image in the sample video clip; determining a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and determining the sum of the background loss function and the motion loss function as the target loss function.
5. The method according to claim 4, characterized in that After displaying the video tag of the video to be processed, the method further includes at least one of the following: Recommending the video to be processed based on the tag of the video to be processed; The video tag of the video to be processed is displayed, proofreading information of the video tag is received, and the video tag is modified based on the proofreading information.
6. A video processing method, characterized in that: include: Receive live video; Processing the live video using a video processing model to obtain a video tag for the live video, wherein the video tag is used to represent the product type of the target object in the live video, the video processing model is obtained by training an initial model, the initial model is obtained by optimizing the model to be trained based on a target loss function of the model to be trained, the target loss function being composed of background feature information of a sample video clip and motion feature information of the sample video clip; Display a video tag of the live video; The method further includes: obtaining first feature information obtained by extracting features from the sample video clip by the model to be trained; determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information of the image in the sample video clip; determining a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and determining the sum of the background loss function and the motion loss function as the target loss function.
7. A data training method, characterized in that: include: Obtaining first feature information obtained by extracting features from a sample video clip by a to-be-trained model; Determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information; determining a motion loss function based on the first feature information and third feature information, wherein the third feature information includes first motion feature information of an image following the sample video clip, and the motion loss function is used to characterize a degree of difference between second motion feature information predicted based on the first feature information and the first motion feature information; The model to be trained is optimized according to the background loss function and the motion loss function.
8. The method according to claim 7, characterized in that Acquiring first feature information obtained by extracting features from a sample video clip by the to-be-trained model, including: Obtaining a target video, and randomly extracting video clips from the target video to obtain the sample video clips; The sample video clip is input into the model to be trained to obtain the first feature information output by the model to be trained, wherein the model to be trained is a three-dimensional feature extraction model.
9. The method according to claim 8, characterized in that Before determining the background loss function based on the first feature information and the second feature information, the method further includes: acquiring the second feature information, wherein the step of acquiring the second feature information includes: Obtaining compressed data corresponding to the target video; Extracting key frames corresponding to the sample video clip from the compressed data; The feature information of the key frame is extracted by a background feature extraction model to obtain the second feature information, wherein the background feature extraction model is a two-dimensional feature extraction model.
10. The method according to claim 9, characterized in that Extracting key frames corresponding to the sample video clip from the compressed data includes: Determining a key frame of a target video segment from the compressed data, wherein a start frame of the target video segment is earlier than a start frame of the sample video segment by a first preset number of frames, and an end frame of the target video segment is later than an end frame of the sample video segment by a second preset number of frames; Any frame is extracted from the key frames of the target video as the key frame corresponding to the sample video segment.
11. The method according to claim 9, characterized in that Determining a background loss function based on the first feature information and the second feature information includes: Performing global average pooling processing on the three-dimensional first feature information and the two-dimensional second feature information respectively to obtain one-dimensional first feature information and one-dimensional second feature information; Mapping the one-dimensional first feature information and the one-dimensional second feature information into a single output to obtain mapped one-dimensional first feature information and one-dimensional second feature information; A first noise contrast estimation loss function is determined based on the mapped one-dimensional first feature information and the one-dimensional second feature information, and the first noise contrast estimation loss function is determined to be the background loss function.
12. The method according to claim 7, characterized in that Before determining the motion loss function based on the first feature information and the third feature information, the method further includes: acquiring the third feature information, wherein the step of acquiring the third feature information includes: Extracting motion vectors of a plurality of consecutive frames of images following the sample video clip; determining the first motion feature information according to the motion vector based on a motion feature extraction model; The first motion feature information is determined to be the third feature information.
13. The method according to claim 12, characterized in that Determining a motion loss function based on the first feature information and the third feature information includes: predicting future second motion feature information based on the first feature information of the sample video clip; The motion loss function is determined based on the second motion feature information and the third feature information.
14. The method according to claim 13, characterized in that Determining the motion loss function based on the second motion feature information and the third feature information includes: Mapping the second motion feature information and the third feature information into a single output to obtain mapped second motion feature information and third feature information; A second noise contrast estimation loss function is determined based on the mapped second motion feature information and the third feature information, and the second noise contrast estimation loss function is determined to be the motion loss function.
15. A video processing device, characterized in that: include: A first receiving module, configured to receive a video to be processed; a feature extraction module, configured to extract features from the video to be processed using a video representation model to obtain video representation information of the video to be processed, wherein the video representation model is obtained by training an initial model, the initial model being obtained by optimizing the model to be trained based on a target loss function, the target loss function being composed of background feature information of a sample video segment and motion feature information of the sample video segment; An output module, configured to output video representation information of the video to be processed, wherein the video representation information includes background feature information of the video to be processed and motion feature information of the video to be processed; The device is further used to: obtain first feature information obtained by extracting features from a sample video clip by the model to be trained; determine a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information of the image in the sample video clip; determine a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and determine the sum of the background loss function and the motion loss function as the target loss function.
16. A video processing device, characterized in that: include: A second receiving module, configured to receive the video to be processed; a first processing module, configured to process the video to be processed using a video processing model to obtain a video label for the video to be processed, wherein the video processing model is obtained by training an initial model, the initial model being obtained by optimizing the model to be trained based on a target loss function of the model to be trained, the target loss function being composed of background feature information of a sample video clip and motion feature information of the sample video clip; A first display module, configured to display a video tag of the video to be processed; The device is further used to: obtain first feature information obtained by extracting features from a sample video clip by the model to be trained; determine a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information of the image in the sample video clip; determine a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and determine the sum of the background loss function and the motion loss function as the target loss function.
17. A video processing device, characterized in that: include: A third receiving module is used to receive live video; a second processing module, configured to process the live video using a video processing model to obtain a video tag for the live video, wherein the video tag is used to represent the product type of the target object in the live video, the video processing model being obtained by training an initial model, the initial model being obtained by optimizing the model to be trained based on a target loss function of the model to be trained, the target loss function being composed of background feature information of a sample video clip and motion feature information of the sample video clip; A second display module is used to display the video tag of the live video; The device is further used to: obtain first feature information obtained by extracting features from a sample video clip by the model to be trained; determine a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information of the image in the sample video clip; determine a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between the second motion feature information predicted based on the first feature information and the first motion feature information; and determine the sum of the background loss function and the motion loss function as the target loss function.
18. A data training device, characterized in that: include: An acquisition module, configured to acquire first feature information obtained by extracting features from a sample video clip by a to-be-trained model; a first determining module, configured to determine a background loss function based on the first feature information and second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize a degree of difference between the first feature information and the background feature information; a second determining module, configured to determine a motion loss function based on the first feature information and third feature information, wherein the third feature information includes first motion feature information of an image following the sample video clip, and the motion loss function is configured to characterize a degree of difference between second motion feature information predicted based on the first feature information and the first motion feature information; An optimization module is used to optimize the model to be trained according to the background loss function and the motion loss function.
19. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the method according to any one of claims 1 to 14.
20. A processor, characterized in that: The processor is configured to run a program, wherein the program executes the method according to any one of claims 1 to 14 when running.
21. A data training system comprising: processor; as well as A memory is connected to the processor and is used to provide the processor with instructions for processing the following processing steps: obtaining first feature information obtained by extracting features from a sample video clip by the model to be trained; determining a background loss function based on the first feature information and the second feature information, wherein the second feature information includes background feature information of the image in the sample video clip, and the background loss function is used to characterize the degree of difference between the first feature information and the background feature information; determining a motion loss function based on the first feature information and the third feature information, wherein the third feature information includes first motion feature information of the image after the sample video clip, and the motion loss function is used to characterize the degree of difference between second motion feature information predicted based on the first feature information and the first motion feature information; and optimizing the model to be trained according to the background loss function and the motion loss function.
Citation Information
Patent Citations
Image foreground and background segmentation method, image foreground and background segmentation network model training method, and image processing method and device
CN107341805A
Method and device for image processing, computer readable storage medium, and electronic device
US20190377944A1