Video processing method and device
By inputting video data into pre-trained video understanding model and using multiple sets of sample video data for multimodal information training, the problem that the existing technology cannot fully capture video semantics is solved, and the accuracy and generalization ability of video content recognition are significantly improved.
Patent Information
- Application Number
- CN202510228964.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-27
AI Technical Summary
Existing video content understanding technologies cannot fully capture the complex semantics of video content, and the accuracy and generalization ability of identifying video content are low.
By obtaining the target video data, including image data, subtitle data and barrage data, and inputting it into a pre-trained video understanding model. The video understanding model is obtained through training of multiple groups of sample video data. Each group of sample video data includes sample video, sample subtitle information corresponding to sample video, and sample barrage information corresponding to sample video.
This multimodal information-combining training method allows the video understanding model to capture the complex semantics of video content more comprehensively, thereby significantly improving the accuracy and generalization ability of video content recognition.
Smart Images

Figure CN120220016A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of multimedia technology, and in particular, to a video processing method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of Internet technology and the popularization of intelligent devices, video content has become an indispensable part of people's daily lives. Whether it is social media, online education, or entertainment platforms, the generation and consumption of video data are increasing rapidly, posing higher requirements for the processing and understanding of video content. However, existing video content understanding technologies cannot comprehensively capture the complex semantics of video content, and the accuracy and generalization ability of identifying video content are relatively low.
[0003] It should be noted that the above content is not necessarily prior art and does not limit the patent protection scope of the present application. Summary of the Invention
[0004] Embodiments of the present application provide a video processing method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above-mentioned technical problems.
[0005] One aspect of embodiments of the present application provides a video processing method, the method comprising: Obtain target video data, the target video data including image data, subtitle data, and barrage data; Input the target video data into a pre-trained video understanding model, and output an understanding result for the target video data through the video understanding model, the understanding result including video analysis, identification, and / or recognition; Wherein, the video understanding model is trained by multiple groups of sample video data; each group of the sample video data includes a sample video, sample subtitle information corresponding to the sample video, and sample barrage information corresponding to the sample video.
[0006] Optionally, the video understanding model is obtained through multiple rounds of training operations, and different groups of sample video data are used in each round of training operation; wherein, each round of training operation includes: Obtain target sample video data, the target sample data including target sample image data, target sample subtitle data, and target sample barrage data; Obtain multiple sample feature data according to the target sample image data, target sample subtitle data, and target sample barrage data; Input the multiple sample feature data as training data into the model to be trained, and train the model to be trained.
[0007] Optionally, obtaining target sample video data includes: Obtaining the timestamps of multiple sample subtitle data and the timestamps of multiple sample barrage data; Obtaining target video data, where the target video data includes multiple sample image data in units of frames; and Forming the target sample video data according to the sample subtitle data, sample barrage data, and sample image data with the same timestamp.
[0008] Optionally, obtaining multiple sample feature data according to the sample image data, sample subtitle data, and sample barrage data includes: Obtaining initial visual features through the sample image data; Obtaining first initial text features through the sample subtitle data; Obtaining second initial text features through the sample barrage data; Adjusting the feature dimensions of the initial visual features, the first initial text features, and the first initial text features to obtain initial visual features, first initial text features, and second initial text features with unified feature dimensions; Performing multi-modal alignment and fusion on the visual features, first initial text features, and second initial text features with unified feature dimensions to obtain visual features, first text features, and second text features in the same spatial dimension one by one.
[0009] Optionally, performing multi-modal alignment and fusion on the visual features, first initial text features, and second initial text features with unified feature dimensions to obtain visual features, first text features, and second text features in the same spatial dimension one by one, including: Inputting the visual features, first initial text features, and second initial text features with unified feature dimensions into a contrastive learning model to output the aligned visual features, first initial text features, and second initial text features through the contrastive learning model; Processing the aligned visual features, first initial text features, and second initial text features based on a cross-attention mechanism to obtain visual features, first text features, and second text features fused with multi-modal information one by one.
[0010] Optionally, characterized in that inputting the multiple sample feature data as training data into a model to be trained and training the model to be trained includes: Inputting the multiple sample feature data as training data into the model to be trained and training the model to be trained based on a joint optimization loss; Wherein, the joint optimization loss includes at least two of a contrastive loss, an alignment loss, and a task loss.
[0011] Optionally, the video processing method further includes: generating one or more sets of enhanced video data for model training based on the target sample video data; wherein the enhanced video data is obtained by adjusting the sample image data, sample caption data, and / or sample bullet screen data in the target sample video data; the adjustment includes image flipping, cropping, text synonym replacement, and / or translation.
[0012] Another aspect of the embodiments of the present application provides a video processing apparatus, the apparatus includes: an acquisition module, configured to acquire target video data, where the target video data includes image data, caption data, and bullet screen data; an input module, configured to input the target video data into a pre-trained video understanding model, and output an understanding result for the target video data through the video understanding model, where the understanding result includes video analysis, identification, and / or recognition; wherein the video understanding model is trained by multiple sets of sample video data; each set of the sample video data includes a sample video, sample caption information corresponding to the sample video, and sample bullet screen information corresponding to the sample video.
[0013] Another aspect of the embodiments of the present application provides a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0014] Another aspect of the embodiments of the present application provides a computer-readable storage medium, where computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method as described above is implemented.
[0015] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.
[0016] The embodiments of the present application adopting the above technical solutions may include the following advantages: By inputting target video data into a pre-trained video understanding model. The video understanding model can output an understanding result for the target video data. Among them, the video understanding model is trained with multiple groups of sample video data, and each group of sample video data includes a sample video, sample caption information corresponding to the sample video, and sample barrage information corresponding to the sample video. This training method combining multi-modal information enables the video understanding model to capture the complex semantics of video content more comprehensively, thereby significantly improving the accuracy and generalization ability of video content recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments. The shown embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0018] Figure 1 Schematically shows an operating environment diagram of the video processing method according to Embodiment 1 of the present application; Figure 2 Schematically shows a flowchart of the video processing method according to Embodiment 1 of the present application; Figure 3 Schematically shows a flowchart of each training step of the video understanding model; Figure 4 Schematically shows Figure 3 a sub-step flowchart of step S300 in; Figure 5 Schematically shows Figure 3 a sub-step flowchart of step S302 in; Figure 6 Schematically shows Figure 5 a sub-step flowchart of step S508 in; Figure 7 Schematically shows an exemplary application flowchart; Figure 8 Schematically shows a block diagram of the video processing apparatus according to Embodiment 2 of the present application; and Figure 9 Schematically shows a hardware architecture diagram of a computer device according to Embodiment 3 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0020] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0021] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.
[0022] First, the following provides the term explanations involved in the present application: Video understanding: It refers to the process of analyzing and extracting semantic information from the visual content, audio content, and related text information (such as subtitles and bullet comments) presented in the video through algorithms or models to achieve video content classification, annotation, or recognition.
[0023] Bullet comment information: It refers to the scrolling or static comments sent by users in real time during video playback. These comments usually contain the audience's feedback on the video content, emotional expressions, or supplementary information.
[0024] Subtitle information: It refers to the text description or translation provided for the audience in the video, usually including dialogues, sound effect descriptions, or background information, aiming to improve the understandability of the video.
[0025] Multi-modal information fusion: It refers to the integration of multiple information modalities (such as video frames, subtitle texts, bullet comment texts, etc.) in the video to improve the model's understanding ability of video content.
[0026] Secondly, to facilitate the understanding of the technical solutions provided by the embodiments of the present application by those skilled in the art, the related technologies are described below: The applicant has learned that video content understanding technology usually relies on image classification models pre-trained on large-scale general datasets (such as ImageNet). This method depends on the richness and diversity of training data. It not only fails to comprehensively capture the complex semantics of video content, but also due to the insufficient data volume or uneven distribution of niche content (such as second-generation videos, specific characters in movies and TV shows, etc.), the recognition accuracy of the model is low and it is difficult to effectively generalize.
[0027] Therefore, the embodiments of the present application provide a video processing technical solution to overcome the above problems. See the following for details.
[0028] Finally, for the convenience of understanding, an exemplary operating environment is provided below.
[0029] As Figure 1 shown, the operating environment diagram includes: a service platform 2, and clients (4A, 4B,..., 4N).
[0030] The service platform 2 can be connected to the clients (4A, 4B,..., 4N) through a network.
[0031] The service platform 2 can be a single server, a server cluster, or a cloud computing service center.
[0032] The service platform 2 can provide services such as storage and reading to the clients.
[0033] The service platform 2 can be located in a data center such as a single location, or distributed in different geographical locations (for example, in multiple locations). The service platform 2 can provide services via a network. The network includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and combinations thereof, or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.
[0034] Clients (4A, 4B, …, 4N) can be configured to access the content and services of service platform 2. Clients (4A, 4B, …, 4N) can include electronic devices with or external to a display panel, such as mobile devices, tablet devices, laptop computers, workstations, virtual reality devices, gaming devices, digital streaming devices, vehicle terminals, smart TVs, set-top boxes, etc., and can also include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing device can load a virtual machine based on a virtual image and / or other data defining specific software (e.g., operating system, dedicated application, server) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices.
[0035] Clients (4A, 4B, …, 4N) can be associated with one or more users. A single user can also use one or more of clients (4A, 4B, …, 4N) to access service platform 2. Clients (4A, 4B, …, 4N) can travel to various locations and use different networks to access service platform 2.
[0036] Clients (4A, 4B, …, 4N) can include an interface. The interface can include a touchpad, a touch screen, a mouse, a keyboard, or other sensing elements. For example, the input element can be configured to receive user instructions, and the user instructions can cause clients (4A, 4B, …, 4N) to perform various operations, such as uploading a video, etc.
[0037] It should be noted that the above devices are exemplary, and in different scenarios or according to different requirements, the quantity and types of devices are adjustable.
[0038] The following takes clients or service platforms as the execution entities and introduces the technical solutions of this application through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be construed as being limited only to the embodiments described herein.
[0039] Embodiment 1 Figure 2 A flowchart of a video processing method according to Embodiment 1 of the present application is schematically shown.
[0040] As Figure 2 shown, the video processing method can include steps S200 to S202, where: Step S200, obtaining target video data, where the target video data includes image data, subtitle data, and bullet screen data; Step S202: Input the target video data into a pre-trained video understanding model, and output an understanding result for the target video data through the video understanding model. The understanding result includes video analysis, identification, and / or recognition. Among them, the video understanding model is trained with multiple groups of sample video data; each group of the sample video data includes a sample video, the sample caption information corresponding to the sample video, and the sample bullet screen information corresponding to the sample video.
[0041] The video processing method provided in this embodiment inputs the target video data into a pre-trained video understanding model. The video understanding model can output an understanding result for the target video data. Among them, the video understanding model is trained with multiple groups of sample video data, and each group of sample video data includes a sample video, the sample caption information corresponding to the sample video, and the sample bullet screen information corresponding to the sample video. This training method that combines multi-modal information enables the video understanding model to capture the complex semantics of video content more comprehensively, thereby significantly improving the accuracy and generalization ability of video content recognition.
[0042] The following combines Figure 2 , and elaborates in detail on each step in steps S200 - S202 and other optional steps.
[0043] Step S200 , obtain target video data, where the target video data includes image data, caption data, and bullet screen data.
[0044] The target video data includes multi-modal information such as image data, bullet screen data, caption data, and voice data. These multi-modal information together constitute the complete information source of the target video, providing rich inputs for subsequent video understanding. This data can be obtained through various methods, such as video streaming platforms, local storage devices, etc.
[0045] Step S202 , input the target video data into a pre-trained video understanding model, and output an understanding result for the target video data through the video understanding model. The understanding result includes video analysis, identification, and / or recognition. Among them, the video understanding model is trained with multiple groups of sample video data; each group of the sample video data includes a sample video, the sample caption information corresponding to the sample video, and the sample bullet screen information corresponding to the sample video.
[0046] The video understanding model can be a deep learning model based on, for example, a convolutional neural network (CNN), a recurrent neural network (RNN), or a Transformer architecture. This video understanding model can simultaneously process image data, subtitle data, and bullet screen data, and output the understanding result of the target video data. Specifically, the understanding result can include, but is not limited to: (1) Video analysis: The overall analysis of video content, such as the theme, sentiment tendency, scene classification, etc. of the video.
[0047] (2) Identification: The identification of specific objects in the video, such as face recognition, object detection, scene recognition, etc.
[0048] (3) Recognition: The recognition of specific events in the video, such as action recognition, event detection, etc.
[0049] Subtitles usually only contain dialogue information. If only the subtitle information in the video is extracted and used as auxiliary features to be input into the model, and natural language processing technology is combined for the extension of video semantics, it is difficult to comprehensively reflect the scene of the video content and the user's understanding. Therefore, each set of sample video data used to train the video understanding model in the embodiments of this application can include, but is not limited to: (1) Sample video: The original video data for training, containing an image sequence.
[0050] (2) Sample subtitle information: The subtitle data synchronized with the sample video, used to provide text context information.
[0051] (3) Sample bullet screen information: The bullet screen data synchronized with the sample video, which is the real-time feedback of users on the video content and contains rich emotional, opinion, and interaction information.
[0052] In some embodiments, through multimodal learning technology, the sample video, sample subtitle information, and sample bullet screen information can be fused to learn the deep features of the video content.
[0053] The video understanding model trained in this embodiment can be applied to various scenarios, such as: (1) Video content recommendation: According to the user's historical viewing records and the video understanding result, recommend video content that the user may be interested in.
[0054] (2) Video summary generation: Automatically generate a video summary through the video understanding result to help users quickly understand the core content of the video.
[0055] (3) Video content review: By identifying sensitive content in the video, automatically conduct content review to ensure that the video content complies with platform specifications as much as possible.
[0056] (4) Video search: Provide a more accurate video search function based on the video understanding results.
[0057] In this embodiment, by obtaining target video data including images, captions, and bullet screens and inputting it into a pre-trained video understanding model, video analysis, identification, or recognition results are output, thereby improving the accuracy and richness of video content understanding, supporting multi-modal information analysis, and enhancing the model's adaptability to complex scenarios.
[0058] Regarding the specific training process of the video understanding model: In an alternative embodiment, as Figure 3 shown, the video understanding model is obtained through multiple rounds of training operations, and different groups of sample video data are used in each round of training operation; wherein, each round of training operation includes: Step S300, obtain target sample video data, where the target sample data includes target sample image data, target sample caption data, and target sample bullet screen data; Step S302, obtain multiple sample feature data according to the target sample image data, target sample caption data, and target sample bullet screen data; Step S304, input the multiple sample feature data as training data into the model to be trained, and train the model to be trained.
[0059] The target sample image data refers to each frame of the video. These frames can be a sequence of static images or dynamic video frames. The target sample image data can include information such as RGB pixel values, resolution, and frame rate. The target sample caption data refers to the text information embedded in the video, which is used to provide dialogue, commentary, or other written explanations. The target sample caption data can include text content, font and style, and timestamp. Among them, the timestamp includes the start and end times of each caption, which is used to synchronize with the video frames. The target sample bullet screen data refers to the comments or messages sent by users in real time while watching the video. The target sample bullet screen data can include text content, user information, display style, timestamp, etc. Among them, the timestamp includes the appearance time of each bullet screen, which is used to synchronize with the video frames.
[0060] In this embodiment, through multiple rounds of training operations, diverse sample video data is used to extract multi-modal feature data to train the video understanding model, thereby improving the video understanding model's comprehensive understanding and generalization ability of video content.
[0061] In an alternative embodiment, as Figure 4 shown, step S300 may include: Step S400, obtain the timestamps of multiple sample caption data and the timestamps of multiple sample bullet screen data; Step S402, obtain target video data, where the target video data includes multiple sample image data in units of frames; and Step S404, form the target sample video data according to the sample caption data, sample barrage data, and sample image data with the same time stamp.
[0062] Regarding obtaining sample caption data, the sample caption data can be extracted from a video platform or a caption file, including caption text and its corresponding time stamp (start time and end time), etc. In some embodiments, the caption data can be preprocessed, such as removing redundant information, unifying the time format, etc.
[0063] Regarding obtaining sample barrage data, the sample barrage data can be extracted from a barrage file or a real-time barrage stream, including barrage text and its corresponding time stamp, etc. In some embodiments, the barrage data can be cleaned and filtered to remove irrelevant content (such as advertisements, duplicate barrages, etc.).
[0064] Align the time stamps of the sample image data, sample caption data, and sample barrage data so that the three are on the same time axis. In some embodiments, for the time stamps of missing data, interpolation or default values are used for filling. For example, when the start time of a barrage is n and the duration is 3 seconds, if the barrage data is missing at times n, n + 1, and n + 2, the barrage data is filled so that the barrage matches the corresponding video frame throughout its display period. In some embodiments, for the time stamps of overlapping data, it can be processed by setting priorities, merging data, displaying separately, or filtering data, etc.
[0065] In some embodiments, the target sample video data can be described in the form of sample data of video frame-caption-barrage pairs, that is, each sample contains a video frame, corresponding caption information, and barrage information.
[0066] In this embodiment, by aligning the video frames, captions, and barrages (i.e., sample image data, sample caption data, and sample barrage data) based on the same time stamp to generate the target sample video data, the data consistency can be improved, the multi-modal information fusion effect can be enhanced, and the content understanding accuracy can be improved.
[0067] The multiple sample feature data can include visual features, text features, etc. The multiple sample feature data can be obtained in the following manner: In an alternative embodiment, as Figure 5 shown, step S302 may include: Step S500, obtain initial visual features through the sample image data; Step S502, obtain first initial text features through the sample caption data; Step S504, obtaining second initial text features from the sample barrage data; Step S506, performing feature dimension adjustment on the initial visual features, the first initial text features, and the first initial text features to obtain initial visual features, first initial text features, and second initial text features with unified feature dimensions; Step S508, performing multimodal alignment and fusion on the visual features, first initial text features, and second initial text features with unified feature dimensions to obtain visual features, first text features, and second text features in the same spatial dimension one by one.
[0068] For obtaining the initial visual features, the initial visual features of the sample image data can be extracted through a pre-trained convolutional neural network (CNN) (such as ResNet, VGG, EfficientNet, etc.). For example, various visual features such as edges, textures, shapes, contours of objects, and layouts of scenes. The extracted initial visual features are usually high-dimensional vectors. Thus, subsequently, in order to reduce the computational complexity and improve the training efficiency of the video content understanding model, feature dimensionality reduction is performed on the initial visual features.
[0069] For obtaining the first initial text features, the initial text features (i.e., the first initial text features) of the target sample caption data can be extracted through natural language processing (NLP) techniques. NLP techniques can include word embedding (WordEmbedding), recurrent neural network (RNN), or Transformer architectures (such as BERT, GPT), etc. In some embodiments, the target sample caption data is cleaned and tokenized to remove irrelevant information such as stop words and punctuation marks.
[0070] For obtaining the second initial text features, the initial text features (i.e., the second initial text features) of the target sample barrage data can be extracted through NLP techniques. In some embodiments, if the sample barrage data is usually short and contains a large amount of user-generated content (UGC), then techniques such as sentiment analysis and keyword extraction can be used to further enrich the second initial text features.
[0071] In this embodiment, through multimodal feature extraction, dimension adjustment, and alignment and fusion, the visual, caption, and barrage features are unified into the same semantic space, thereby comprehensively capturing the multi-dimensional information of the video content, enriching the data used to train the video understanding model, so as to significantly improve the accuracy and generalization ability of the video understanding model subsequently.
[0072] In an alternative embodiment, as Figure 6 shown, step S508 may include: Step S600: Unify the visual features, the first initial text feature, and the second initial text feature in the contrastive learning model of the feature dimension, so as to output the aligned visual features, the first initial text feature, and the second initial text feature through the contrastive learning model. Step 602: Process the aligned visual features, the first initial text feature, and the second initial text feature based on the cross-attention mechanism to obtain the visual features, the first text feature, and the second text feature that fuse multi-modal information one by one.
[0073] The contrastive learning model can be trained by a contrastive loss function (Contrastive Loss) or a triplet loss function (Triplet Loss). The contrastive learning model is used to align the visual features, the first initial text feature (initial caption feature), and the second initial text feature (initial bullet screen feature) in the same semantic space. By processing multiple sample feature data through contrastive learning, the subsequent trained video understanding model can learn the similarities and differences between different modal features, thereby enhancing the alignment effect of the features.
[0074] Specific implementation of the contrastive learning model: (1) Positive sample pairs: Use the initial visual feature, the initial caption feature, and the initial bullet screen feature at the same timestamp as positive sample pairs.
[0075] (2) Negative sample pairs: Use the initial visual feature, the initial caption feature, and the initial bullet screen feature at different timestamps as negative sample pairs.
[0076] (3) Loss function: Use a contrastive loss function (such as InfoNCE Loss) to optimize the model, so that the features of the positive sample pairs are as close as possible in the semantic space, and the features of the negative sample pairs are as far away as possible.
[0077] (4) Output aligned features: After passing through the contrastive learning model, the visual features, the first initial text feature, and the second initial text feature are aligned in the semantic space, and the output is the aligned visual features, the first initial text feature, and the second initial text feature.
[0078] The cross-attention mechanism is used to further fuse the aligned visual features, the first initial text feature, and the second initial text feature to generate the visual features, the first text feature, and the second text feature that contain multi-modal information. By processing multiple sample feature data through the cross-attention mechanism, it is convenient for the subsequent trained video understanding model to dynamically capture the interaction relationships between different modal features, thereby enhancing the semantic expression ability of the features.
[0079] The specific implementation of the cross-attention mechanism can be as follows: (1)Construction of Query, Key, and Value: Use visual features as the query, and the first initial text feature and the second initial text feature as the key and value. Alternatively, use the first initial text feature as the query, and the visual feature and the second initial text feature as the key and value.
[0080] (2)Calculation of Attention Weights: Calculate the attention weights between the query and the key through dot-product attention or additive attention.
[0081] (3)Feature Fusion: Perform weighted summation on the value according to the attention weights to obtain the fused feature representation.
[0082] (4)Output the Fused Features: After being processed by the cross-attention mechanism, generate visual features, the first text feature, and the second text feature containing multimodal information. These feature representations are in the same spatial dimension and correspond one by one, capable of comprehensively describing the visual information, caption information, and barrage information of the video content.
[0083] In this embodiment, multi-modal feature alignment is achieved through contrastive learning, and the cross-attention mechanism is combined to deeply fuse multi-modal features, generating visual features, caption features (the first text feature), and barrage features (the second text feature) in a unified semantic space, thereby comprehensively capturing the multi-dimensional information of the video content. These multi-dimensional information are used as training data for training the video understanding model, which can enhance the model's understanding ability of complex video content and significantly improve the efficiency and accuracy of video analysis.
[0084] In an alternative embodiment, step S304 may include: Input the multiple sample feature data as training data into the model to be trained, and train the model to be trained based on the joint optimization loss; wherein, the joint optimization loss includes at least two of contrastive loss, alignment loss, and task loss.
[0085] Specifically, the visual features, the first text features (caption features), and the second text features (bullet screen features) after multi-modal alignment and fusion are used as training data and input into the model to be trained. These feature data are in the same semantic space and can comprehensively describe the visual, text, and user feedback information of the video content. The model to be trained can be a multi-modal deep learning model, such as a multi-modal fusion model based on Transformer, a multi-task learning model, etc. The input of the model is multi-modal feature data, and the output is the result of the video understanding task (such as video classification, sentiment analysis, content summary, etc.).
[0086] The joint optimization loss includes at least two of the contrastive loss, the alignment loss, and the task loss, and is used to comprehensively optimize the performance of the model. The specific implementation method of each loss function is as follows: Contrastive Loss: The visual features, caption features (i.e., the first text features), and bullet screen features (i.e., the second text features) at the same timestamp are used as positive sample pairs. The visual features, caption features, and bullet screen features at different timestamps are used as negative sample pairs. The contrastive loss function (such as InfoNCE Loss) is used to optimize the model, so that the features of the positive sample pairs are close in the semantic space, and the features of the negative sample pairs are far away. The alignment effect between different modal features is enhanced through the contrastive loss, ensuring that the visual features, caption features, and bullet screen features are as close as possible in the semantic space.
[0087] Alignment Loss: The distance between the visual features and the text features (the first text features and the second text features) can be calculated by the mean squared error (MSE) or the cosine similarity (CosineSimilarity). Minimize the distance between the visual features and the text features to enhance their alignment effect. The alignment effect of the multi-modal features is further optimized through the alignment loss, ensuring the consistency of the visual features and the text features in the semantic space.
[0088] Task Loss: Select the corresponding loss function according to the specific task. For example: for the classification task, use the cross-entropy loss (Cross-Entropy Loss). For the regression task, use the mean squared error loss (Mean SquaredError Loss). For the generation task, use the adversarial loss (Adversarial Loss) or the reconstruction loss (Reconstruction Loss). The performance of the video understanding model on specific tasks, such as video classification, sentiment analysis, content summary, etc., is optimized through the task loss.
[0089] The joint optimization loss is the weighted sum of the contrastive loss, the alignment loss, and the task loss, and the formula is as follows: Ljoint = α * L contrastive + β * L alignment + λ * L task Wherein, L joint represents the joint optimization loss, L contrastive represents the contrast loss, L alignmen represents the alignment loss, L task represents the task loss; α, β, and λ are the weight coefficients of each loss function, used to balance the importance of different loss functions.
[0090] In some embodiments, the joint optimization loss can be minimized by the Gradient Descent method or other optimization algorithms (such as Adam, RMSprop, etc.). In some embodiments, the model parameters can be updated through the backpropagation algorithm to gradually optimize the performance of the model.
[0091] In this embodiment, the video understanding model is trained through the joint optimization loss (including the contrast loss, alignment loss, and task loss) to achieve the alignment of multi-modal features and the optimization of task performance, thereby enhancing the comprehensive understanding ability of the video understanding model for video content and significantly improving the accuracy and efficiency of video analysis tasks.
[0092] In alternative embodiments, the video processing method may further include: generating one or more sets of enhanced video data for model training based on the target sample video data; wherein, the enhanced video data is obtained by adjusting the sample image data, sample caption data, and / or sample bullet screen data in the target sample video data; the adjustment includes image flipping, cropping, text synonym replacement, and / or translation.
[0093] Diverse image data can be generated by adjusting the video frames (sample image data) to enhance the robustness of the video understanding model. The adjustment of the sample image data can include: Image flipping: For example, horizontally or vertically flipping the video frames.
[0094] Image rotation: For example, rotating the video frames by a certain angle (such as 90°, 180°).
[0095] Image cropping: For example, cropping a part of the video frames to retain the key content.
[0096] Color adjustment: For example, adjusting the brightness, contrast, or saturation of the video frames.
[0097] Adding noise: For example, adding random noise to the video frames to simulate low-quality images.
[0098] Diverse subtitle data can be generated by adjusting subtitle text (sample subtitle data), which can enhance the model's text understanding ability. Adjustments to the sample subtitle data can include: Synonym replacement: Replace some words in the subtitle with synonyms (e.g., replace "happy" with "glad").
[0099] Text expansion: Add additional descriptive text to the subtitle (e.g., expand "This is a cat" to "This is a cute cat").
[0100] Text translation: Translate the subtitle into another language (e.g., from Chinese to English) for training multilingual models.
[0101] Text perturbation: Randomly delete or replace some words in the subtitle to simulate text noise.
[0102] Diverse barrage data can be generated by adjusting barrage text (sample barrage data), which can enhance the model's understanding ability of interactive content. Adjustments to the sample barrage data can include: Synonym replacement: Replace some words in the barrage with synonyms (e.g., replace "Hahaha" with "Laughing to death").
[0103] Emotion enhancement: Adjust the emotional tendency of the barrage (e.g., change a neutral barrage to a positive or negative one).
[0104] Text perturbation: Randomly delete or replace some words in the barrage to simulate barrage noise.
[0105] Barrage density adjustment: Increase or decrease the number of barrages to simulate scenarios with different barrage densities.
[0106] In this embodiment, the target sample video data is adjusted by methods such as image flipping, cropping, text synonym replacement, and translation to generate enhanced video data. This method can improve the model's robustness, enhance diversity, support complex scenarios, and improve the training efficiency of the video understanding model.
[0107] To make this application easier to understand, the following is combined with Figure 7 to provide an exemplary application.
[0108] It should be noted that the video understanding model of this application embodiment is obtained through multiple rounds of training operations, and different groups of data are used in each round of training operation; among them, each round of training operation includes: Step S11, collect data and preprocess the data. The data includes video frames, subtitles, and barrages.
[0109] (1) For collecting video frames: Extract key frames through a video processing program (such as FFmpeg) for video frame collection.
[0110] (2)For extracting bullet screens: Parse external subtitles or extract embedded subtitles through OCR technology.
[0111] (3)For crawling bullet screens: Crawl bullet screens and clean invalid information.
[0112] Step S12, obtain the timestamps of multiple subtitles and the timestamps of multiple bullet screens.
[0113] Step S13, according to the timestamps of subtitles and bullet screens (i.e., the same timestamps), match the bullet screens and subtitles (i.e., texts) with the corresponding video frames. It should be noted that if there is a situation where a certain timestamp has missing bullet screens or subtitles within a time period (e.g., 3 seconds), the bullet screens or subtitles can be supplemented or cropped.
[0114] Step S14, perform deep feature extraction on video frames, subtitles, and bullet screens to obtain visual features, subtitle features, and bullet screen features.
[0115] (1)For extracting visual features (video frame features): Use ResNet to extract visual features.
[0116] (2)For extracting subtitle / bullet screen features: Use BERT to extract text features. The text features include subtitle features and bullet screen features.
[0117] Step S15, perform feature dimension alignment on visual features, subtitle features, and bullet screen features to obtain visual features, subtitle features, and bullet screen features with the same feature dimensions.
[0118] (1)Use dimensionality reduction or feature mapping methods to unify visual features and text features.
[0119] (2)Map visual features and text features to the same feature space to ensure that the feature dimensions are as consistent as possible.
[0120] Step S16, perform multi-modal alignment and fusion on visual features, subtitle features, and bullet screen features with unified feature dimensions to obtain visual features, subtitle features, and bullet screen features in the same spatial dimension.
[0121] (1)Use contrastive learning to align visual features and text features.
[0122] (2)Use alignment loss (mean + covariance) to optimize the consistency of features (visual features and text features).
[0123] (3)Implement multi-modal interaction through cross-attention mechanism. Among them, implementing multi-modal interaction means generating visual features, subtitle features, and bullet screen features that fuse multi-modal information.
[0124] Step S17: Use the visual features, caption features, and bullet screen features that integrate multi-modal information and have the same spatial dimension as training data to input into the model to be trained, and perform model optimization and training.
[0125] Specifically, train the model to be trained based on the joint optimization loss. The joint optimization loss includes contrast loss, alignment loss, and task loss.
[0126] It should be noted that one or more sets of enhanced video data for model training can be generated based on the data (i.e., video frames, bullet screens, captions); among them, the enhanced video data is obtained through video frame enhancement (flipping, cropping), text enhancement (synonym replacement, data back-translation), etc.
[0127] This application proposes a multi-modal video understanding model training method by integrating video frames, caption information, and bullet screen information, which can make up for the shortage of niche video data and the limitations of single-modal models, and make full use of the user feedback value in the bullet screen, significantly improve the accuracy and generalization ability of niche video content recognition, and promote the development of video content understanding technology in the niche field.
[0128] Embodiment 2 Figure 8 Schematically shows a block diagram of a video processing device according to Embodiment 2 of the present application. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 8 As shown, the device 1000 may include: an acquisition module 1100, an input module 1200, where: The acquisition module 1100 is configured to acquire target video data, where the target video data includes image data, caption data, and bullet screen data; The input module 1200 is configured to input the target video data into a pre-trained video understanding model, and output an understanding result for the target video data through the video understanding model. The understanding result includes video analysis, identification, and / or recognition; Among them, the video understanding model is trained by multiple sets of sample video data; each set of the sample video data includes a sample video, sample caption information corresponding to the sample video, and sample bullet screen information corresponding to the sample video.
[0129] In an alternative embodiment, the video understanding model is obtained through multiple rounds of training operations, and different sets of sample video data are used in each round of training operation; among them, each round of training operation includes: Obtain target sample video data, where the target sample data includes target sample image data, target sample subtitle data, and target sample bullet screen data; Obtain multiple sample feature data according to the target sample image data, target sample subtitle data, and target sample bullet screen data; Use the multiple sample feature data as training data and input it into the model to be trained to train the model to be trained.
[0130] In an alternative embodiment, obtaining the target sample video data includes: Obtain the timestamps of multiple sample subtitle data and the timestamps of multiple sample bullet screen data; Obtain target video data, where the target video data includes multiple sample image data in units of frames; and Form the target sample video data according to the sample subtitle data, sample bullet screen data, and sample image data with the same timestamp.
[0131] In an alternative embodiment, obtaining multiple sample feature data according to the sample image data, sample subtitle data, and sample bullet screen data includes: Obtain initial visual features through the sample image data; Obtain first initial text features through the sample subtitle data; Obtain second initial text features through the sample bullet screen data; Adjust the feature dimensions of the initial visual features, the first initial text features, and the first initial text features to obtain initial visual features, first initial text features, and second initial text features with unified feature dimensions; Perform multi-modal alignment and fusion on the visual features, first initial text features, and second initial text features with unified feature dimensions to obtain visual features, first text features, and second text features in the same spatial dimension one by one.
[0132] In an alternative embodiment, performing multi-modal alignment and fusion on the visual features, first initial text features, and second initial text features with unified feature dimensions to obtain visual features, first text features, and second text features in the same spatial dimension one by one includes: Input the visual features, first initial text features, and second initial text features with unified feature dimensions into a contrastive learning model to output the aligned visual features, first initial text features, and second initial text features through the contrastive learning model; Process the aligned visual features, first initial text features, and second initial text features based on a cross-attention mechanism to obtain visual features, first text features, and second text features that integrate multi-modal information one by one.
[0133] In an alternative embodiment, inputting the multiple sample feature data as training data into a model to be trained, and training the model to be trained includes: Inputting the multiple sample feature data as training data into the model to be trained, and training the model to be trained based on a jointly optimized loss; wherein, the jointly optimized loss includes at least two of a contrast loss, an alignment loss, and a task loss.
[0134] In an alternative embodiment, the video processing apparatus further includes a generation module configured to: generate one or more sets of enhanced video data for model training based on the target sample video data; wherein, the enhanced video data is obtained by adjusting the sample image data, sample caption data, and / or sample barrage data in the target sample video data; the adjustment includes image flipping, cropping, text synonym replacement, and / or translation.
[0135] Embodiment III Figure 9 FIG. schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing the video processing method according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack-mounted server, a blade server, a tower server, or a cabinet server (including a stand-alone server or a server cluster composed of multiple servers), etc. As Figure 9 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the video processing method. In addition, the memory 10010 can also be used to temporarily store various data that have been output or will be output.
[0136] In some embodiments, the processor 10020 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other chip. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0137] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal via a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0138] It should be noted that Figure 9 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0139] In this embodiment, the video processing method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.
[0140] Embodiment 4 The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the video processing method in the embodiment are implemented.
[0141] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system and various application software installed on the computer device, such as the program code of the video processing method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.
[0142] Embodiment Five The embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.
[0143] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device. Thus, they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0144] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A video processing method, characterized in that: The method comprises: Acquire target video data, wherein the target video data includes image data, subtitle data, and bullet screen data; Inputting the target video data into a pre-trained video understanding model, and outputting an understanding result for the target video data through the video understanding model, wherein the understanding result includes video analysis, identification and / or recognition; The video understanding model is trained by multiple groups of sample video data; each group of sample video data includes a sample video, sample subtitle information corresponding to the sample video, and sample barrage information corresponding to the sample video.
2. The method according to claim 1, characterized in that The video understanding model is obtained through multiple rounds of training operations, each round of training operations using a different set of sample video data; wherein each round of training operations includes: Acquire target sample video data, wherein the target sample data includes target sample image data, target sample subtitle data, and target sample bullet screen data; Acquire a plurality of sample feature data according to the target sample image data, the target sample subtitle data and the target sample bullet screen data; The plurality of sample feature data are input as training data into the model to be trained, and the model to be trained is trained.
3. The method according to claim 2, characterized in that Get the target sample video data, including: Obtaining timestamps of multiple sample subtitle data and timestamps of multiple sample bullet comment data; Acquire target video data, wherein the target video data includes a plurality of sample image data in frames; and The target sample video data is formed according to the sample subtitle data, sample bullet screen data and sample image data with the same time stamp.
4. The method according to claim 2, characterized in that: According to the sample image data, the sample subtitle data and the sample bullet screen data, a plurality of sample feature data are obtained, including: Acquire initial visual features through the sample image data; Acquire a first initial text feature through the sample subtitle data; Acquire a second initial text feature through the sample bullet comment data; Adjusting the feature dimensions of the initial visual feature, the first initial text feature, and the second initial text feature to obtain the initial visual feature, the first initial text feature, and the second initial text feature with unified feature dimensions; Multimodal alignment and fusion are performed on the visual features, the first initial text features and the second initial text features with unified feature dimensions to obtain the visual features, the first text features and the second text features of the same spatial dimension in one-to-one correspondence.
5. The method according to claim 4, characterized in that Perform multimodal alignment and fusion on the visual features, the first initial text features, and the second initial text features with unified feature dimensions to obtain visual features, the first text features, and the second text features of the same spatial dimension in one-to-one correspondence, including: Comparing the visual features, the first initial text features, and the second initial text features with unified feature dimensions into a comparative learning model, so as to output the aligned visual features, the first initial text features, and the second initial text features through the comparative learning model; The aligned visual features, the first initial text features, and the second initial text features are processed based on a cross-attention mechanism to obtain visual features, first text features, and second text features that are integrated with multimodal information in a one-to-one correspondence.
6. The method according to any one of claims 2 to 5, characterized in that: Inputting the plurality of sample feature data as training data into a model to be trained, and training the model to be trained, comprising: Inputting the plurality of sample feature data as training data into a model to be trained, and training the model to be trained based on a joint optimization loss; The joint optimization loss includes at least two of contrast loss, alignment loss and task loss.
7. The method according to any one of claims 2 to 5, characterized in that: The method further comprises: Generating one or more sets of enhanced video data for model training based on the target sample video data; The enhanced video data is obtained by adjusting sample image data, sample subtitle data and / or sample bullet screen data in the target sample video data; the adjustment includes image flipping, cropping, text synonym replacement and / or translation.
8. A video processing device, characterized in that: The device comprises: An acquisition module, used to acquire target video data, wherein the target video data includes image data, subtitle data and bullet screen data; An input module, used for inputting the target video data into a pre-trained video understanding model, and outputting an understanding result for the target video data through the video understanding model, wherein the understanding result includes video analysis, identification and / or recognition; The video understanding model is trained by multiple groups of sample video data; each group of sample video data includes a sample video, sample subtitle information corresponding to the sample video, and sample barrage information corresponding to the sample video.
9. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 7 are implemented.
Citation Information
Cited By
Video bullet screen young user emotion feature analysis method fused with multi-modal information
CN121686432A