Video Processing Method, Apparatus, and Computer-Readable Storage Medium

By using attention modules to enhance the video features in video processing, the inaccurate audio-visual event detection problem caused by coarse-grained feature fusion in the prior art is solved, and more accurate audio-visual event detection is achieved.

CN114140708BActive Publication Date: 2025-07-01ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110937670.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-16
Publication Date
2025-07-01
Estimated Expiration
2041-08-16

AI Technical Summary

Technical Problem

The coarse-grained feature fusion method in existing video detection technology leads to inaccurate detection of audio-visual events in video.

Method used

By receiving the video to be processed, the initial video features and audio features are extracted, and the attention module uses the attention module to enhance the video features based on the weight parameters of multiple dimensions to obtain the enhanced video features and predict the audio-visual event based on this.

Benefits of technology

Through fine-grained modal fusion, background noise interference is reduced, and the capture performance of sound source locations in video is improved, thereby improving the accuracy of audio-visual event detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140708B_ABST
    Figure CN114140708B_ABST
Patent Text Reader

Abstract

The present invention discloses a video processing method, apparatus, and computer-readable storage medium. Among them, the method includes: receiving a video to be processed, and performing feature extraction on the video to be processed to obtain an initial video feature and an initial audio feature of the video to be processed; determining weight parameters in multiple dimensions through the initial audio feature, and enhancing the initial video feature by using the weight parameters in multiple dimensions based on a first attention module to obtain an enhanced video feature; predicting an audiovisual event in the video to be processed based on the enhanced video feature. The present invention solves the technical problem in the related art that the coarse-grained video detection method leads to inaccurate detection of audiovisual events in the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and in particular, to a video processing method, apparatus, and computer-readable storage medium. Background Art

[0002] The human perception system can fuse visual and auditory information to achieve an understanding of audiovisual events in the real world. Traditional video detection techniques are limited to visual methods and ignore other perception methods, making it impossible to accurately detect audiovisual events. In related technologies, by using a multi-modal event detection algorithm to fuse audio and video features, the detection of audiovisual events in a video can be achieved. However, existing multi-modal event detection algorithms adopt a coarse-grained feature fusion method. For example, audio features only participate in guiding video features in a single dimension, resulting in inaccurate detection of audiovisual events in the video.

[0003] Regarding the problem of inaccurate detection of audiovisual events in the video caused by the coarse-grained video detection method in the above-mentioned related technologies, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of the present invention provide a video processing method, apparatus, and computer-readable storage medium to at least solve the technical problem of inaccurate detection of audiovisual events in the video caused by the coarse-grained video detection method in related technologies.

[0005] According to one aspect of the embodiments of the present invention, there is provided a video processing method, including: receiving a video to be processed, and performing feature extraction on the video to be processed to obtain initial video features and initial audio features of the video to be processed; determining weight parameters in multiple dimensions through the initial audio features, and enhancing the initial video features based on the first attention module using the weight parameters in multiple dimensions to obtain enhanced video features; predicting audiovisual events in the video to be processed based on the enhanced video features.

[0006] According to one aspect of the embodiments of the present invention, there is provided a video processing method, including: obtaining a live video to be processed collected during a live broadcast; classifying and detecting the live video using a target detection model to obtain a prediction result of audiovisual events in the live video; adding label information to the live video based on the prediction result, where the target detection model is used to perform feature extraction on the live video to obtain initial video features and initial audio features of the live video; determining weight parameters in multiple dimensions through the initial audio features, and enhancing the initial video features based on the first attention module using the weight parameters in multiple dimensions to obtain enhanced video features; predicting audiovisual events based on the enhanced video features.

[0007] According to another aspect of the embodiments of the present invention, there is also provided a video processing apparatus, including: a receiving module, configured to receive a video to be processed, extract features from the video to be processed, and obtain initial video features and initial audio features of the video to be processed; an enhancement module, configured to determine weight parameters in multiple dimensions based on the initial audio features, and enhance the initial video features based on the first attention module by using the weight parameters in multiple dimensions to obtain enhanced video features; a prediction module, configured to predict audiovisual events in the video to be processed based on the enhanced video features.

[0008] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the method in any one of the above.

[0009] According to another aspect of the embodiments of the present invention, there is also provided a computer program. When the computer program runs, it executes the method in any one of the above.

[0010] According to another aspect of the embodiments of the present invention, there is also provided a video processing system, including: a processor; and a memory, connected to the processor, configured to provide instructions for the processor to perform the following processing steps: receive a video to be processed, extract features from the video to be processed to obtain initial video features and initial audio features of the video to be processed; determine weight parameters in multiple dimensions based on the initial audio features, and enhance the initial video features based on the first attention module by using the weight parameters in multiple dimensions to obtain enhanced video features; predict audiovisual events in the video to be processed based on the enhanced video features.

[0011] In the embodiments of the present invention, by receiving a video to be processed, extracting features from the video to be processed to obtain initial video features and initial audio features, enhancing the initial video features based on the first attention module by using the weight parameters in multiple dimensions to obtain enhanced video features, and predicting audiovisual events in the video to be processed based on the enhanced video features, through fine-grained modality fusion of audio and video features in multiple dimensions, the interference caused by background noise to audiovisual event detection is reduced, the position of the sound source in the video can be captured more accurately, thereby improving the accuracy of audiovisual event detection, and further solving the technical problem that the coarse-grained video detection method in the related art leads to inaccurate detection of audiovisual events in the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0013] Figure 1 It is a hardware structure block diagram of a computing device for implementing a data training method;

[0014] Figure 2 It is a flowchart of a video processing method according to an embodiment of the present invention;

[0015] Figure 3a It is a schematic diagram of an optional triple attention network structure according to an embodiment of the present invention;

[0016] Figure 3b It is a schematic diagram of an optional MFB module according to an embodiment of the present invention;

[0017] Figure 4a It is a schematic diagram of an optional dense cross-modal attention module structure according to an embodiment of the present invention;

[0018] Figure 4b It is a schematic diagram of an optional dense correlation weight calculation according to an embodiment of the present invention;

[0019] Figure 4c It is a schematic diagram of an optional grouped weighted average according to an embodiment of the present invention;

[0020] Figure 5 It is a schematic diagram of an optional video processing method according to an embodiment of the present invention;

[0021] Figure 6 It is a schematic diagram of an optional video processing method according to an embodiment of the present invention;

[0022] Figure 7 It is a schematic diagram of the influence of different balance hyperparameters on the detection result;

[0023] Figure 8 It is a flowchart of a video processing method according to an embodiment of the present invention;

[0024] Figure 9 It is a schematic diagram of a video processing device according to an embodiment of the present invention;

[0025] Figure 10 It is a structure block diagram of a computer terminal according to an embodiment of the present application. Detailed implementation manners

[0026] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] Embodiment 1

[0029] According to an embodiment of the present invention, an embodiment of a video processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0030] The method embodiment provided in the first embodiment of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 is a hardware structure block diagram of a computer terminal of a video processing method according to an embodiment of the present invention. As Figure 1 shown, the computing device 10 may include one or more (only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computing device 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0031] The memory 104 can be used to store software programs and modules of application software, such as program instructions / modules corresponding to the video processing method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the vulnerability detection method of the above-mentioned application program. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computing device 10 through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0032] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the computing device 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0033] Under the above operating environment, the present application provides a Figure 2 video processing method as shown. Figure 2 is a flowchart of the video processing method according to Embodiment 1 of the present invention. As Figure 2 shown, the method includes:

[0034] Step S201, receiving a video to be processed, and performing feature extraction on the video to be processed to obtain initial video features and initial audio features of the video to be processed.

[0035] The above video to be processed is a video for which an audiovisual event needs to be detected, and the audiovisual event is an event containing images and audio. For example, the audiovisual event can be a segment of video in the video to be processed that contains a voice conversation and images.

[0036] The video to be processed can be a video of any theme or application scenario, including but not limited to live videos obtained on live streaming platforms, traffic videos in traffic scenarios, teaching videos in the education field, medical examination videos in the medical field, etc.

[0037] The above initial video features and initial audio features can be extracted through a trained feature extraction model. The initial video features are used to represent the image features in the video to be processed, and the initial audio features are used to represent the sound features in the video to be processed.

[0038] Step S202: Determine weight parameters in multiple dimensions based on the initial audio features, and enhance the initial video features by using the weight parameters in multiple dimensions based on the first attention module to obtain enhanced video features.

[0039] By calculating the attention weight parameters in multiple dimensions in a fine-grained fusion manner, the initial video features and the initial audio features are fused, and enhanced video features are obtained. Compared with the initial video features, the enhanced video features highlight the event-related regions (the event-related regions are video segments with audiovisual events in the video to be processed), reduce the interference of background noise in the audiovisual event detection process, and significantly improve the performance of capturing the sound source position in the video.

[0040] In an optional embodiment, the above first attention module may be a triple attention module, and the above multiple dimensions may include a channel dimension, a spatial dimension, and a time dimension. The triple attention module obtains weight parameters in three dimensions of the channel dimension, the spatial dimension, and the time dimension based on the initial audio features, and then enhances the initial video features in a fine-grained manner in the three dimensions of channel, space, and time.

[0041] Figure 3a is a schematic diagram of an optional triple attention network structure according to an embodiment of the present invention, as Figure 3a shown. The triple attention network structure includes a channel attention module, a spatial attention module, and a time attention module. The spatial attention module may adopt a multi-modal factorized bilinear pooling module (MFB module), and input the initial audio feature a t (and ) and the initial video feature v t (and ) into the triple attention network model to implement the enhanced processing of the initial video feature by the initial audio feature in a fine-grained manner in the three dimensions of channel, space, and time, and obtain enhanced video features

[0042] Step S203: Predict the audiovisual events in the video to be processed based on the enhanced video features.

[0043] After obtaining the enhanced video features, the enhanced video features are fused with the audio features to obtain the fused features of audio and video, and the fused features can be used to predict the audiovisual events in the video to be processed.

[0044] In an alternative implementation, after predicting the audiovisual event in the video to be processed based on the enhanced video features, the above method further includes: outputting the prediction result of the audiovisual event, where the prediction result includes any one or more of whether the audiovisual event exists in the video to be processed, the video segment where the audiovisual event is located, and the category of the audiovisual event.

[0045] Specifically, the prediction result of the audiovisual event may include the audiovisual event-related segment and the category of the audiovisual event. The prediction result of the audiovisual event-related segment may include whether the audiovisual event exists in the video to be processed, and when the audiovisual event exists in the video to be processed, the video segment where the audiovisual event exists in the video to be processed. For example, the audiovisual event to be predicted may be the audiovisual event of an airplane taking off. The video to be processed obtained can be based on the above method to obtain enhanced video features, and the enhanced video features are input into the trained detection model to obtain a prediction result. The prediction result may include whether the video to be processed contains the audiovisual event of an airplane taking off, the video segment where the audiovisual event of an airplane taking off exists in the video to be processed, and the category of the audiovisual event. Based on the category of the audiovisual event, a label can be added to the detected audiovisual event. For example, "airplane taking off" can be used as the category label of the audiovisual event. In this embodiment, the audiovisual event is predicted based on the enhanced video features, which enhances the detection performance for distinguishing similar sound categories. For example, it can more accurately distinguish between noise and the audio features in the audiovisual event.

[0046] The video processing method in this embodiment can be used for the detection of audiovisual events in the video in various application scenarios such as video recommendation scenarios, video content review, video content understanding scenarios, and audio-video separation scenarios.

[0047] In this embodiment, the video to be processed is received, and feature extraction is performed on the video to be processed to obtain the initial video features and initial audio features of the video to be processed. The initial video features are enhanced based on the first attention module using weight parameters in multiple dimensions to obtain enhanced video features. The audiovisual event in the video to be processed is predicted based on the enhanced video features. By performing fine-grained modality fusion on the audio and video features in multiple dimensions, the interference caused by background noise to the detection of audiovisual events is reduced, the position of the sound source in the video can be captured more accurately, and thus the accuracy of the detection of audiovisual events is improved, solving the technical problem that the coarse-grained video detection method in the related art results in inaccurate detection of audiovisual events in the video.

[0048] As an alternative embodiment, performing feature extraction on the video to be processed to obtain the initial video features of the video to be processed includes: obtaining the image sequence of the video to be processed; extracting the feature map from the image sequence based on the image feature extraction model; and performing global average pooling on the feature map to obtain the initial video features.

[0049] The above image sequence can be images with a specified number of frames extracted from the video to be processed. The specified number of frames can be determined according to the image feature extraction model, which is not limited here. For example, 16 RGB images are extracted from the video to be processed as the above image sequence.

[0050] The above can be a convolutional neural network model, such as the VGG-19 network model. The image feature extraction model can be pre-trained based on an image dataset (such as the ImageNet dataset) for the VGG-19 network model.

[0051] The above feature map can be the feature map of a video segment with a specified time length. For example, to obtain the initial video feature, a sequence of 16 RGB images can be sampled from the video to be processed and input into the pre-trained VGG-19 network model to extract the pool5 feature map of a 1-second video segment. Global average pooling is used to obtain the initial video feature v at the segment level t , t ∈ [1, 10].

[0052] As an alternative embodiment, feature extraction is performed on the video to be processed to obtain the initial audio feature of the video to be processed, including: obtaining the audio segment in the video to be processed; converting the audio segment into a spectrogram; extracting a feature vector from the spectrogram based on the audio feature extraction model; and determining the feature vector as the initial audio feature.

[0053] The above audio segment can be audio with a specified time length extracted from the video to be processed. The specified time length can be determined according to the audio feature extraction model.

[0054] The above audio feature extraction model can be a pre-trained convolutional neural network model, such as the VGGish network model. Specifically, the audio feature extraction model can be pre-trained based on an audio dataset (such as the AudioSet dataset) for the VGGish network model.

[0055] For example, to obtain the initial audio feature, the audio segment of each 1 second in the video to be processed can be converted into a log-mel spectrogram, and a 128D feature vector is extracted based on the pre-trained VGGish network model as the initial audio feature a at the segment level t , t ∈ [1, 10].

[0056] As an alternative embodiment, the weight parameters in multiple dimensions include the first-dimensional attention weight parameter, the second-dimensional attention weight parameter, and the third-dimensional attention weight parameter. In step S202, the initial video feature is enhanced based on the first attention module using the weight parameters in multiple dimensions, including the following steps:

[0057] In step S2021, the initial video features are enhanced using the first - dimension attention weight parameters to obtain the first - dimension video features.

[0058] The above - mentioned first attention module can be a triple - attention module. The above - mentioned first dimension can be the channel dimension, the second dimension can be the spatial dimension, and the third dimension can be the temporal dimension. The triple - attention module enhances the initial video features in a fine - grained manner in the three dimensions of channel, spatial, and temporal based on the initial audio features.

[0059] In an alternative embodiment, weight parameters in multiple dimensions are determined from the initial audio features, including: performing non - linear transformation and activation processing on the initial audio features with respect to the initial video features to obtain the first - dimension attention weight parameters.

[0060] The first - dimension attention weight parameters can be channel attention weights. After obtaining the initial audio features and the initial video features the initial audio features and the initial video features can be projected and aligned to the same dimension through two non - linear transformations, and the channel attention weights are obtained through a squeeze - and - excitation module Specifically, the calculation process of the channel attention weights is as follows:

[0061]

[0062] where and are fully - connected layers using ReLU activation, AVP represents global average pooling in the spatial dimension, and represent two linear transformations respectively, δ represents the activation operation of ReLU, and σ represents the activation operation of sigmoid respectively.

[0063] The first - dimension attention weight parameters can be channel attention weights Using the channel attention weights to enhance the initial video features to obtain the video features of channel attention (i.e., the first - dimension video features), the specific process is as follows:

[0064]

[0065] where ⊙ represents element - wise multiplication.

[0066] Step S2022: Based on the second-dimensional attention weight parameter and the third-dimensional attention weight parameter, obtain the second-dimensional attention feature mapping weight. Among them, the second-dimensional attention weight parameter is obtained by fusing the initial audio feature and the first-dimensional video feature in the second dimension, and the third-dimensional attention weight parameter is obtained by fusing the initial audio feature and the first-dimensional video feature in the third dimension.

[0067] Specifically, the second-dimensional attention weight parameter is the spatial attention weight The third-dimensional attention weight parameter is the temporal attention weight Based on the spatial attention weight and the temporal attention weight, calculate and obtain the spatial attention feature mapping weight

[0068]

[0069] Among them, W3 represents a linear transformation.

[0070] Step S2023: Use the second-dimensional attention feature mapping weight to update the first-dimensional video feature to obtain an enhanced video feature

[0071]

[0072] Among them, is the spatial attention feature mapping weight. By using the spatial attention feature mapping weight to update the video feature of the channel attention An enhanced video feature of the audio in the three dimensions of channel, space, and time can be obtained

[0073] In an optional embodiment, weight parameters in multiple dimensions are determined through the initial audio feature, including: respectively expanding the dimensions of the initial audio feature and the first-dimensional video feature based on an activation function to obtain an expanded audio feature and an expanded video feature; determining the video feature units of the expanded video feature in the second dimension; based on a multimodal bilinear matrix factorization pooling module, fusing the video feature units in the second dimension and the expanded audio feature to obtain the second-dimensional attention weight parameter.

[0074] The above activation function can be the ReLU activation function. Use a fully connected layer activated by ReLU to expand the initial audio feature a t and the video feature of the channel attention to the same dimension kdo to obtain an expanded audio feature and an expanded video feature. The video feature units in the second dimension are the video features at each spatial position, and the second-dimensional attention weight parameter can be the spatial attention weight Perform fusion on the initial audio feature a in the spatial dimension t and the video feature of channel attention to obtain the spatial attention weight:

[0075] Specifically, the calculation process of the spatial attention weight is as follows:

[0076]

[0077]

[0078]

[0079] Among them, and are two learnable matrix parameters, which are respectively used to expand the initial audio feature a t and the video feature of channel attention to the same dimension kdo. SP(f, k) represents the sum pooling (i.e., sumpooling) operation with both the kernel and the stride being k, and D(·) represents the dropout layer, which is used to prevent overfitting.

[0080] By adopting the multi-modal bilinear matrix factorization pooling module (i.e., the MFB module), the video feature and the corresponding audio feature at each spatial position are finely fused using the shared MFB module. Based on the MFB module, the correlation between the audio and video features is calculated, replacing the simple element-wise multiplication correlation calculation method in the related technology, which can significantly improve the performance of capturing the sound source position in the video. Figure 3b is a schematic diagram of an optional MFB module according to an embodiment of the present invention. As Figure 3b shown, since the number of parameters contained in the weight matrix W in the spatial attention weight is too large, the MFB module can be used for decomposition to reduce the number of parameters. In addition, the MFB module introduces square normalization and normalization L2 to achieve stable training of the model.

[0081] In another optional embodiment, the third-dimensional attention weight parameter is the temporal attention weight which can be obtained by using a bi-directional LSTM (Long-Short Term Memory) network model in the temporal dimension to perform fine-grained audio-visual fusion modeling for each video feature spatial block. The specific steps are as follows:

[0082] Project the initial audio feature a t and the video feature of channel attention onto the same dimension do:

[0083] Among them, and are fully connected layers using ReLU activation.

[0084] The video features of each space and audio features are represented as:

[0085]

[0086] Input into a bi-directional LSTM network to perform temporal modeling on each space and obtain temporal attention weights

[0087]

[0088] As an alternative embodiment, step S203, predicting audiovisual events in the video to be processed based on the enhanced video features, includes the following steps:

[0089] Step S2031, input the initial audio features and the enhanced video features into the self-attention module respectively to obtain self-attention audio features and self-attention video features.

[0090] The above self-attention module consists of a multi-head attention module, a residual connection, and a layer normalization layer. The self-attention module pays more attention to the correlation within the features and uses the features themselves as weight parameters. For example, when inputting feature m into the self-attention module, self-attention feature m self = self(m) The calculation process is as follows:

[0091] Using m as the weight parameter, obtain the query (i.e., Q), key (i.e., K), and value (i.e., V) of self-attention:

[0092]

[0093]

[0094] M = Concat(m1, m2,..., m h )W O ;

[0095] M r = LayerNorm(M + m);

[0096] Self(m) = LayerNorm(δ(M r W2)W3 + W r )。

[0097] Specifically, the initial audio feature a t and the enhanced video feature are respectively input into two self-attention modules. Based on the above calculation process of the self-attention feature, we get: the self-attention audio feature a self = self(a) and the self-attention video feature v self = self(v).

[0098] Step S2032: Input the initial audio feature and the self-attention video feature into the second attention module to obtain the cross-attention audio feature, and input the enhanced video feature and the self-attention audio feature into the second attention module to obtain the cross-attention video feature. Then, fuse the cross-attention audio feature and the cross-attention video feature to obtain the fused feature.

[0099] The above second attention module can be a dense cross-modal attention module. The dense cross-modal attention module is a module that effectively fuses two-modal information by using the dense relationship within and between modalities. Figure 4a is a schematic diagram of an optional dense cross-modal attention module structure according to an embodiment of the present invention. As Figure 4a shown, the second attention module adopts a multi-head dense cross-modal attention module (DCMA module). The DCMA module uses the dense correlation weight (DCWC) calculation method to replace the sparse matrix calculation method in the related art. Specifically, input the features (x, y) of the two modalities into the dense cross-modal attention module. x is used as the query of the dense cross-modal attention module, and concat(x, y) is used as the key and value items. In the DCMA module, the correlation between the feature x and concat(x, y) is decomposed into Nx,x and Nx,y. Nx,x calculates the correlation within the modality through the classical matrix multiplication method, and Nx,y calculates the fine-grained cross-modal correlation through the dense correlation weight (DCWC) method. Figure 4b is a schematic diagram of an optional dense correlation weight calculation according to an embodiment of the present invention. As Figure 4b shown, in the dense correlation weight (DCWC) calculation method, the operation between elements adopts the grouped weighted average (GWA) calculation method to replace the traditional inner product. The calculation method of the dense correlation weight (DCWC) is as follows:

[0100] N x,y = DCWC(x, y),

[0101] GWA(x i , y j ) = sum((x ix × y j ) ⊙ W);

[0102] where GWA is the weighted average of the outer product of feature x i and feature y i The symbol × represents the outer product operation, and W is the weight matrix. Figure 4c is a schematic diagram of an optional grouped weighted average according to an embodiment of the present invention, as Figure 4c shown, x i is Figure 4c ai in Figure 4c , and yi is i bi in The elements in the matrix (x × yi) are divided into two groups: diagonal elements (corresponding to the original inner product operation) and other elements. The diagonal elements of the weight matrix W are α, and the weights corresponding to other elements are:

[0103] where

[0104] In an optional embodiment, the initial audio feature and the self-attention video feature are input into the second attention module to obtain the cross-attention audio feature, and the enhanced video feature and the self-attention audio feature are input into the second attention module to obtain the cross-attention video feature, including: based on the second attention module, performing grouped weighted average processing on the initial audio feature and the self-attention video feature to obtain the cross-attention audio feature; based on the second attention module, performing grouped weighted average processing on the enhanced video feature and the self-attention audio feature to obtain the cross-attention video feature.

[0105] In an optional embodiment, the initial audio feature a t and the self-attention video feature v self , and the enhanced video feature and the self-attention audio feature a self are respectively input into two multi-head dense cross-modal attention modules to calculate the cross-attention audio feature a cross and the cross-attention video feature v cross :

[0106] a cross = DCMA(a, v self );

[0107] v cross = DCMA(v, a self );

[0108] Figure 5 is a schematic diagram of an optional video processing method according to an embodiment of the present invention. As Figure 5 shown, based on Figure 5 the audiovisual fusion module shown, the input feature 51 can be the audio feature of cross-attention, and the input feature 52 can be the video feature of cross-attention. The process of fusing the audio feature of cross-attention and the video feature of cross-attention to obtain the fusion feature 53 is as follows:

[0109] f av = a cross ⊙ v cross , m a,v = Concat(a, v);

[0110]

[0111] q1 = f av w q , k 1,2 = m a,v w k ; v 1,2 = m a,v w v ;

[0112] Among them, O av is the finally obtained fusion feature, q1 is the query, k 1,2 is the key, and v 1,2 is the value item.

[0113] By fusing the audio feature of cross-attention and the video feature of cross-attention, high-semantic features of audio and video fusion can be obtained.

[0114] Step S2033, predicting audiovisual events based on the fusion feature.

[0115] Specifically, the fusion feature can be input into a preset detection model to obtain the detection result of the audiovisual event. The above detection result can include the event-related segment of the predicted audiovisual event (that is, whether the video to be processed contains an audiovisual event and the position where the audiovisual event is located) and the audiovisual event category, etc.

[0116] For example, the video to be processed can be a video containing a person's conversation and an airplane takeoff event. The fusion feature is obtained from the above video to be processed based on the above method, and the fusion feature is input into a preset detection model. The detection results of the audiovisual events of the person's conversation and the airplane takeoff in the above video to be processed can be obtained, as well as the category of each audiovisual event. Labels can be added to the detected audiovisual events based on the category.

[0117] Since the fused features are obtained through the above-mentioned fine-grained cross-modal fusion, using the fused features to detect audiovisual events in the video to be processed can improve the accuracy of audiovisual event detection. For example, when detecting the audiovisual event of an airplane taking off, the sound of people talking can be accurately distinguished as noise, reducing the interference of noise on the detection of audiovisual events.

[0118] As an optional embodiment, the above method further includes: obtaining a model to be trained, where the model to be trained is used to predict audiovisual events based on the fused features; determining a first classification loss function based on the fused features; determining a second classification loss function based on the self-attention video features; and optimizing the model to be trained according to the first classification loss function and the second classification loss function.

[0119] The above model to be trained is a detection model for detecting audiovisual events based on the fused features. The detection model can output a detection result according to the obtained fused features, where the detection result may include whether there is an audiovisual event in the video to be processed and the category of the audiovisual event. The first classification loss function is determined based on the fused features and can be a cross-modal constraint loss function, focusing on the classification ability of the fused features. The second classification loss function is determined based on the self-attention video features and can be a single-modal constraint loss function, focusing on the classification ability of the single-modal features.

[0120] In an optional implementation, to improve the accuracy of the model to be trained in detecting the category of audiovisual events at the video level, a first classification loss function is determined based on the fused features, and a second classification loss function (i.e., a single-modal constraint loss function) is determined based on the self-attention video features in the intermediate stage. This not only uses the fused features O av to calculate the cross-entropy loss, but also uses the self-attention video features v self (i.e., single-modal features) to calculate the cross-entropy loss, realizing the use of the single-modal constraint loss function to strengthen the classification ability of the single-modal features, combining the single-modal constraint loss function with the audiovisual event classification loss based on the fused features to further improve the ability to identify event categories using single-modal features, and thus enhancing the discrimination performance for similar audiovisual event classifications.

[0121] Specifically, first use the fused features O av to calculate the cross-entropy loss

[0122] Use the self-attention video features v self to calculate the cross-entropy loss:

[0123] Obtain the first classification loss function:

[0124] and a second classification loss function:

[0125] where K represents the number of audiovisual event categories, represents the indicator function: Optimizing the above-mentioned model to be trained by combining the first classification loss function and the second classification loss function can enhance the discrimination performance of the model to be trained for classifying similar audiovisual events.

[0126] In an alternative embodiment, the first classification loss function is the classification loss of audiovisual events with multi-label soft margin loss, and the second classification loss function can be the single-modal event classification constraint loss. Based on the first classification loss function and the second classification loss function, a weakly supervised loss function can be obtained

[0127]

[0128] where λ is a balancing hyperparameter, is the first classification loss function, is the second classification loss function.

[0129] As an alternative embodiment, a prediction loss function is determined based on the fusion features; the model to be trained is optimized according to the prediction loss function, the first classification loss function, and the second classification loss function.

[0130] The detection result of the audiovisual event based on the above-mentioned model to be trained also includes whether there is an audiovisual event in the video to be processed, that is, the detection result of the segment related to the audiovisual event. The above prediction loss function is used to optimize the accuracy of the detection result of the segment related to the audiovisual event by the model to be trained.

[0131] Specifically, the prediction loss function can be determined based on the binary cross-entropy loss function. First, the binary cross-entropy loss s can be calculated using the fusion feature O av s = Sigmoid(FC(O av ))

[0132] and then the prediction loss function is obtained:

[0133] where N represents the number of training samples, and FC represents the classifier.

[0134] After obtaining the prediction loss function, the first classification loss function, and the second classification loss function, the above-mentioned model to be trained can be optimized using the three loss functions respectively, or a final loss function can be constructed based on the three loss functions, and the final loss function is used to train the model to be trained.

[0135] In an alternative embodiment, optimizing the feature extraction model according to the prediction loss function, the first classification loss function, and the second classification loss function includes: constructing a fully supervised loss function based on the prediction loss function, the first classification loss function, and the second classification loss function with preset hyperparameters; solving the fully supervised loss function to optimize the model to be trained.

[0136] Specifically, based on the prediction loss function the first classification loss function and the second classification loss function to obtain the fully supervised loss function

[0137]

[0138] where λ is a balancing hyperparameter.

[0139] Using the fully supervised loss function to optimize the model to be trained can improve the accuracy of the detection results of the audiovisual events in the video to be processed.

[0140] After completing the optimization of the model to be trained, the final detection result is determined by the cross-entropy loss av calculated based on the fused feature O and the binary cross-entropy loss s together. A reasonable comparison threshold can be set to determine whether the detection result contains audiovisual events. For example, the comparison threshold can be set to 0.5. If s ≥ 0.5, it is determined that the video to be processed contains an audiovisual event, and the audiovisual event is the category of the audiovisual event; if s < 0.5, it is determined that the segment of the video to be processed is a background video segment and does not contain the above-mentioned audiovisual events.

[0141] In an alternative embodiment, Figure 6 is a schematic diagram of an alternative video processing method according to an embodiment of the present invention. As Figure 6 shown, a video clip 601 with a preset number of frames is sampled from the video to be processed and input into the VGG-19 network to extract the initial video feature Vt. The audio clip 602 in the video to be processed is converted into a log-mel spectrogram 603, and the log-mel spectrogram 603 is input into the VGGish network to extract the initial audio feature a t The initial video feature Vt and the initial audio feature a t are input into the audio-guided triple attention module 606 to enhance the initial video feature in a fine-grained manner in the three dimensions of channel, space, and time by the initial audio feature, and obtain the enhanced video feature

[0142] The enhanced video features are input into the intra-modal attention module 607 (i.e., the self-attention module) to obtain the self-attention video features, and the initial video features a t are input into the intra-modal attention module 608 (i.e., the self-attention module) to obtain the self-attention audio features. The above second attention module respectively includes a dense cross-modal attention module 609 and a dense cross-modal attention module 610. The enhanced video features and the self-attention audio features are input into the dense cross-modal attention module 610 to obtain... The self-attention video features and the initial audio features are input into the dense cross-modal attention module 609. By inputting into the audio-video fusion module 605, the final fusion features can be obtained. After the fusion features are processed by a classification model (i.e., a fully connected layer FC), the detection results of the audio-visual event-related segments and the audio-visual event types can be obtained.

[0143] In addition, a single-modal constraint loss function 604 can be constructed based on the self-attention video features output by the intra-modal attention module 607, and a classification loss function can be constructed based on the fusion features output by the audio-video fusion module 611. The single-modal constraint loss function 604 is used to enhance the classification ability of the single-modal features. By combining the single-modal constraint loss function 604 with the classification loss function, the classification model is trained to further improve the ability of the classification model to identify event categories using single-modal features, thereby enhancing the discrimination performance for similar audio-visual event classifications.

[0144] Based on the video processing method in this embodiment, under the condition of weak supervision, the accuracy of detecting audio-visual events can reach 74.3%, and under the condition of full supervision, the accuracy of detecting audio-visual events can reach 79.6%. Compared with the existing detection network, the accuracy of detecting audio-visual events is improved.

[0145] Figure 7 is a schematic diagram of the influence of different balance hyperparameters on the detection results. As Figure 7 shown, the abscissa is the value of the balance hyperparameter, and the ordinate is the detection result accuracy. Curve 71 is the accuracy curve of the detection results after optimizing the above-mentioned model to be trained based on the weak supervision loss function Curve 72 is the accuracy curve of the detection results after optimizing the above-mentioned model to be trained based on the full supervision loss function. According to the influence of different balance hyperparameters on the detection result accuracy, the appropriate balance hyperparameter is determined, which can improve the accuracy of the detection results of audio-visual events.

[0146] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0147] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0148] Embodiment 2

[0149] According to an embodiment of the present invention, an embodiment of a video processing method is also provided. Figure 8 is a flowchart of a video processing method according to an embodiment of the present invention, as Figure 8 shown, the method includes:

[0150] Step S801, obtain the live video to be processed collected during the live broadcast.

[0151] Step S802, use the target detection model to classify and detect the live video to obtain the prediction result of the audiovisual event in the live video.

[0152] The above-mentioned live video to be processed is the video in the live broadcast platform that needs to detect audiovisual events, and the prediction result is obtained by detecting the live video based on the target detection model.

[0153] Step S803, add label information to the live video based on the prediction result. Among them, the target detection model is used to extract features from the live video to obtain the initial video features and initial audio features of the live video; determine the weight parameters in multiple dimensions through the initial audio features, and use the weight parameters in multiple dimensions to enhance the initial video features based on the first attention module to obtain enhanced video features; predict the audiovisual event based on the enhanced video features.

[0154] The target detection model may include a feature extraction model. The above-mentioned initial video features and initial audio features can be extracted by the trained feature extraction model. The initial video features are used to represent the image features in the video to be processed, and the initial audio features are used to represent the sound features in the video to be processed.

[0155] Specifically, the prediction result of the audiovisual event may include the segment related to the audiovisual event and the category of the audiovisual event. The prediction result of the segment related to the audiovisual event may include whether there is an audiovisual event in the video to be processed, and when there is an audiovisual event in the video to be processed, the video segment where the audiovisual event exists in the video to be processed.

[0156] For example, the audiovisual event to be predicted may be that the anchor is singing. The obtained live video can be enhanced with video features based on the above method. The enhanced video features are input into the trained target detection model, and the prediction result can be obtained. The prediction result may include whether the video to be processed contains the audiovisual event of the anchor singing, the video segment where the audiovisual event exists, and the category of the audiovisual event. Based on the category of the audiovisual event, a label can be added to the detected audiovisual event. For example, "singing" is used as the label information of the audiovisual event. In this embodiment, the audiovisual event is predicted based on the enhanced video features, which enhances the detection performance of differentiating similar sound categories. For example, it can more accurately distinguish between noise and the audio features in the audiovisual event.

[0157] The above label information can be used to recommend live videos to users. For example, the live videos corresponding to the audiovisual events with the "singing" label are recommended to interested users.

[0158] In the live video review scenario, the live video to be processed can be the live video being broadcast on the video live platform. The above collection process can be to collect the live video before it is distributed to the user side. By classifying and detecting the audiovisual events in the collected live video, the content of the live video is reviewed to determine whether the live video being broadcast involves illegal content categories, and then corresponding preprocessing measures are taken to prevent the spread of live videos containing illegal content on the network platform.

[0159] In this embodiment, the weight parameters of the attention are calculated in a fine-grained fusion manner in multiple dimensions to fuse the initial video features and the initial audio features, and enhanced video features are obtained. Compared with the initial video features, the enhanced video features highlight the event-related regions (the event-related regions are the video segments where there are audiovisual events in the video to be processed), reduce the interference of background noise in the audiovisual event detection process, and significantly improve the performance of capturing the sound source position in the video.

[0160] Embodiment 3

[0161] According to an embodiment of the present invention, there is also provided an apparatus for implementing the above video processing method. Figure 9 It is a schematic diagram of a video processing apparatus according to an embodiment of the present invention, as Figure 9 shown. The apparatus includes:

[0162] A receiving module 91, configured to receive a video to be processed and extract features from the video to be processed, so as to obtain an initial video feature and an initial audio feature of the video to be processed; an enhancement module 92, configured to determine weight parameters in multiple dimensions based on the initial audio feature, and perform enhancement processing on the initial video feature by using the weight parameters in multiple dimensions based on a first attention module, so as to obtain an enhanced video feature; a prediction module 93, configured to predict an audiovisual event in the video to be processed based on the enhanced video feature.

[0163] It should be noted here that the above receiving module 91, enhancement module 92, and prediction module 93 correspond to steps S201 to S203 in Embodiment 1. The functions realized by the three modules and the corresponding steps are the same in terms of examples and application scenarios, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the apparatus, can run in the computing device 10 provided in Embodiment 1.

[0164] As an optional embodiment, the above prediction module is further configured to: after predicting the audiovisual event in the video to be processed based on the enhanced video feature, output a prediction result of the audiovisual event, where the prediction result includes any one or more of whether the audiovisual event exists in the video to be processed, the video segment where the audiovisual event is located, and the category of the audiovisual event.

[0165] As an optional embodiment, the above receiving module is further configured to: obtain an image sequence of the video to be processed; extract a feature map from the image sequence based on an image feature extraction model; perform global average pooling on the feature map to obtain an initial video feature.

[0166] As an optional embodiment, the above receiving module is further configured to: obtain an audio segment in the video to be processed; a conversion sub-module, configured to convert the audio segment into a spectrogram; extract a feature vector from the spectrogram based on an audio feature extraction model; determine the feature vector as the initial audio feature.

[0167] As an alternative embodiment, the weight parameters in multiple dimensions include the first-dimension attention weight parameter, the second-dimension attention weight parameter, and the third-dimension attention weight parameter. The enhancement module is further configured to: enhance the initial video features using the first-dimension attention weight parameter to obtain the first-dimension video features; based on the second-dimension attention weight parameter and the third-dimension attention weight parameter, obtain the second-dimension attention feature mapping weight, wherein the second-dimension attention weight parameter is obtained by fusing the initial audio features and the first-dimension video features in the second dimension, and the third-dimension attention weight parameter is obtained by fusing the initial audio features and the first-dimension video features in the third dimension; use the second-dimension attention feature mapping weight to update the first-dimension video features to obtain enhanced video features.

[0168] As an alternative embodiment, the enhancement module is further configured to: perform non-linear transformation and activation processing on the initial audio features with respect to the initial video features to obtain the first-dimension attention weight parameter.

[0169] As an alternative embodiment, the enhancement module is further configured to: respectively expand the dimensions of the initial audio features and the first-dimension video features based on an activation function to obtain the expanded audio features and the expanded video features; determine the video feature units of the expanded video features in the second dimension; based on the multi-modal bilinear matrix factorization pooling module, fuse the video feature units in the second dimension and the expanded audio features to obtain the second-dimension attention weight parameter.

[0170] As an alternative embodiment, the prediction module is further configured to: respectively input the initial audio features and the enhanced video features into the self-attention module to obtain the self-attention audio features and the self-attention video features; input the initial audio features and the self-attention video features into the second attention module to obtain the cross-attention audio features, and input the enhanced video features and the self-attention audio features into the second attention module to obtain the cross-attention video features, fuse the cross-attention audio features and the cross-attention video features to obtain the fused features; predict the audiovisual event based on the fused features.

[0171] As an alternative embodiment, the prediction module is further configured to: perform grouped weighted average processing on the initial audio features and the self-attention video features based on the second attention module to obtain the cross-attention audio features; perform grouped weighted average processing on the enhanced video features and the self-attention audio features based on the second attention module to obtain the cross-attention video features.

[0172] As an alternative embodiment, the above-mentioned device further includes: an acquisition module, configured to acquire a model to be trained, where the model to be trained is used to predict an audiovisual event based on fused features; a first determination module, configured to determine a first classification loss function based on the fused features; a second determination module, configured to determine a second classification loss function based on self-attention video features; and an optimization module, configured to optimize the model to be trained according to the first classification loss function and the second classification loss function.

[0173] As an alternative embodiment, the above-mentioned device further includes: a third determination module, configured to determine a prediction loss function based on the fused features; and the above-mentioned optimization module is further configured to optimize the model to be trained according to the prediction loss function, the first classification loss function, and the second classification loss function.

[0174] As an alternative embodiment, the above-mentioned optimization module is further configured to construct a fully supervised loss function based on preset hyperparameters through the prediction loss function, the first classification loss function, and the second classification loss function; and solve the fully supervised loss function to optimize the model to be trained.

[0175] It should be noted that the optional or preferred implementation manners of this embodiment can refer to the relevant descriptions in Embodiment 1, which will not be elaborated here.

[0176] Embodiment 4

[0177] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a stored program, where, when the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned search method for the target object.

[0178] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be located in any one of the computing devices in a computing device cluster in a computer network, or in any one of the mobile terminals in a mobile terminal cluster.

[0179] Optionally, in this embodiment, the computer-readable storage medium is set to store program codes for performing the following steps: receiving a video to be processed, and performing feature extraction on the video to be processed to obtain initial video features and initial audio features of the video to be processed; determining weight parameters in multiple dimensions through the initial audio features, and enhancing the initial video features based on the first attention module using the weight parameters in multiple dimensions to obtain enhanced video features; and predicting an audiovisual event in the video to be processed based on the enhanced video features.

[0180] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: performing feature extraction on the video to be processed to obtain the initial video features of the video to be processed, including: obtaining an image sequence of the video to be processed; extracting a feature map from the image sequence based on an image feature extraction model; performing global average pooling on the feature map to obtain the initial video features.

[0181] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: performing feature extraction on the video to be processed to obtain the initial audio features of the video to be processed, including: obtaining an audio segment in the video to be processed; converting the audio segment into a spectrogram; extracting a feature vector from the spectrogram based on an audio feature extraction model; determining the feature vector as the initial audio features.

[0182] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: the weight parameters in multiple dimensions include the first-dimensional attention weight parameter, the second-dimensional attention weight parameter, and the third-dimensional attention weight parameter. Based on the first attention module, enhancing the initial video features by using the weight parameters in multiple dimensions, including: enhancing the initial video features by using the first-dimensional attention weight parameter to obtain the first-dimensional video features; obtaining the second-dimensional attention feature map weight based on the second-dimensional attention weight parameter and the third-dimensional attention weight parameter, where the second-dimensional attention weight parameter is obtained by fusing the initial audio features and the first-dimensional video features in the second dimension, and the third-dimensional attention weight parameter is obtained by fusing the initial audio features and the first-dimensional video features in the third dimension; updating the first-dimensional video features by using the second-dimensional attention feature map weight to obtain the enhanced video features.

[0183] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining the weight parameters in multiple dimensions through the initial audio features, including: performing non-linear transformation and activation processing on the initial audio features with respect to the initial video features to obtain the first-dimensional attention weight parameter.

[0184] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining the weight parameters in multiple dimensions through the initial audio features, including: respectively performing dimensional expansion on the initial audio features and the first-dimensional video features based on an activation function to obtain the expanded audio features and the expanded video features; determining the video feature units of the expanded video features in the second dimension; fusing the video feature units in the second dimension and the expanded audio features based on a multi-modal bilinear matrix factorization pooling module to obtain the second-dimensional attention weight parameter.

[0185] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: predicting audiovisual events in a video to be processed based on enhanced video features, including: inputting initial audio features and enhanced video features into a self-attention module respectively to obtain self-attention audio features and self-attention video features; inputting the initial audio features and the self-attention video features into a second attention module to obtain cross-attention audio features, and inputting the enhanced video features and the self-attention audio features into the second attention module to obtain cross-attention video features, fusing the cross-attention audio features and the cross-attention video features to obtain fused features; predicting audiovisual events based on the fused features.

[0186] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: inputting the initial audio features and the self-attention video features into a second attention module to obtain cross-attention audio features, and inputting the enhanced video features and the self-attention audio features into the second attention module to obtain cross-attention video features, including: performing grouped weighted average processing on the initial audio features and the self-attention video features based on the second attention module to obtain cross-attention audio features; performing grouped weighted average processing on the enhanced video features and the self-attention audio features based on the second attention module to obtain cross-attention video features.

[0187] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining a model to be trained, where the model to be trained is used to predict audiovisual events based on fused features; determining a first classification loss function based on the fused features; determining a second classification loss function based on the self-attention video features; optimizing the model to be trained according to the first classification loss function and the second classification loss function.

[0188] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining a prediction loss function based on the fused features; optimizing the model to be trained according to the prediction loss function, the first classification loss function, and the second classification loss function.

[0189] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: optimizing a feature extraction model according to the prediction loss function, the first classification loss function, and the second classification loss function, including: constructing a fully supervised loss function based on the prediction loss function, the first classification loss function, and the second classification loss function through preset hyperparameters; solving the fully supervised loss function to optimize the model to be trained.

[0190] Example 5

[0191] According to an embodiment of the present application, an embodiment of a computer terminal is further provided. The computer terminal may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal may also be replaced with a terminal device such as a mobile terminal.

[0192] Optionally, in this embodiment, the above computer terminal may be located in at least one of multiple network devices in a computer network.

[0193] In this embodiment, the above computer terminal may execute the program code of the following steps in the video processing method of the application program: receiving a video to be processed, and performing feature extraction on the video to be processed to obtain the initial video features and initial audio features of the video to be processed; determining weight parameters in multiple dimensions through the initial audio features, and enhancing the initial video features based on the first attention module using the weight parameters in multiple dimensions to obtain enhanced video features; predicting the audiovisual events in the video to be processed based on the enhanced video features.

[0194] Optionally, Figure 10 is a structural block diagram of a computer terminal according to Embodiment 5 of the present application. As Figure 10 shown, the computer terminal 1000 may include: one or more (only one is shown in the figure) processors 1002, a memory 1004, and a peripheral interface 1006.

[0195] Among them, the memory may be used to store software programs and modules, such as the program instructions / modules corresponding to the video processing method and device in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above video processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided relative to the processor, and these remote memories may be connected to the computer terminal 1000 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0196] The processor is used to run programs and can call the information and application programs stored in the memory through a transmission device to perform the following steps: receiving a video to be processed, extracting features from the video to be processed to obtain the initial video features and initial audio features of the video to be processed; determining weight parameters in multiple dimensions through the initial audio features, and enhancing the initial video features based on the weight parameters in multiple dimensions using a first attention module to obtain enhanced video features; predicting audiovisual events in the video to be processed based on the enhanced video features.

[0197] Those of ordinary skill in the art can understand that Figure 10 the structure shown is only illustrative, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 10 It does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 1000 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 10 in the figure, or have a different configuration from that shown Figure 10 in the figure.

[0198] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0199] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0200] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0201] In the several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the units or modules can be in an electrical or other form.

[0202] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0203] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0204] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.

[0205] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A video processing method, characterized in that, Including: Receiving a video to be processed, and performing feature extraction on the video to be processed to obtain initial video features and initial audio features of the video to be processed; Determining weight parameters in multiple dimensions through the initial audio features, and enhancing the initial video features by using the weight parameters in multiple dimensions based on a first attention module to obtain enhanced video features, where the multiple dimensions at least include a channel dimension, a spatial dimension, and a temporal dimension; Predicting audiovisual events in the video to be processed based on the enhanced video features; Among them, the weight parameters in the multiple dimensions include a first dimension attention weight parameter, a second dimension attention weight parameter, and a third dimension attention weight parameter. The first dimension is the channel dimension, the second dimension is the spatial dimension, and the third dimension is the temporal dimension; enhancing the initial video features by using the weight parameters in multiple dimensions based on a first attention module includes: enhancing the initial video features by using the first dimension attention weight parameter to obtain first dimension video features; obtaining the enhanced video features based on the second dimension attention weight parameter, the third dimension attention weight parameter, and the first dimension video features. The second dimension attention weight parameter is obtained by fusing the initial audio features and the first dimension video features in the second dimension, and the third dimension attention weight parameter is obtained by fusing the initial audio features and the first dimension video features in the third dimension.

2. The video processing method according to claim 1, wherein After predicting the audiovisual events in the video to be processed based on the enhanced video features, the method further includes: Outputting a prediction result of the audiovisual events, where the prediction result includes any one or more of whether the audiovisual events exist in the video to be processed, the video segments where the audiovisual events are located, and the categories of the audiovisual events.

3. The video processing method according to claim 1, wherein Obtaining the enhanced video features based on the second dimension attention weight parameter, the third dimension attention weight parameter, and the first dimension video features includes: Obtaining a second dimension attention feature mapping weight based on the second dimension attention weight parameter and the third dimension attention weight parameter; Updating the first dimension video features by using the second dimension attention feature mapping weight to obtain the enhanced video features.

4. The video processing method according to claim 1, wherein Predicting the audiovisual events in the video to be processed based on the enhanced video features includes: Respectively inputting the initial audio features and the enhanced video features into a self-attention module to obtain self-attention audio features and self-attention video features; Inputting the initial audio features and the self-attention video features into a second attention module to obtain cross-attention audio features, and inputting the enhanced video features and the self-attention audio features into the second attention module to obtain cross-attention video features; Fusing the cross-attention audio features and the cross-attention video features to obtain a fused feature; Predicting the audiovisual events based on the fused feature.

5. The video processing method according to claim 4, characterized in that, Input the initial audio features and the self-attention video features into a second attention module to obtain cross-attention audio features, and input the enhanced video features and the self-attention audio features into the second attention module to obtain cross-attention video features, including: Based on the second attention module, perform grouped weighted average processing on the initial audio features and the self-attention video features to obtain the cross-attention audio features; Based on the second attention module, perform grouped weighted average processing on the enhanced video features and the self-attention audio features to obtain the cross-attention video features.

6. The video processing method according to claim 4, wherein The method further includes: Obtain a model to be trained, where the model to be trained is used to predict the audiovisual event based on the fusion features; Determine a first classification loss function based on the fusion features; Determine a second classification loss function based on the self-attention video features; Optimize the model to be trained according to the first classification loss function and the second classification loss function.

7. The video processing method according to claim 6, wherein The method further includes: Determine a prediction loss function based on the fusion features; Optimize the model to be trained according to the prediction loss function, the first classification loss function, and the second classification loss function.

8. A video processing method, characterized in that, Include: Obtain the live video to be processed collected during the live broadcast; Use an object detection model to classify and detect the live video to obtain the prediction result of the audiovisual event in the live video; Add label information to the live video based on the prediction result; Wherein, the object detection model is used to extract features from the live video to obtain the initial video features and initial audio features of the live video; determine weight parameters in multiple dimensions through the initial audio features, and based on a first attention module, use the weight parameters in multiple dimensions to enhance the initial video features to obtain enhanced video features; predict the audiovisual event based on the enhanced video features, and the multiple dimensions at least include a channel dimension, a spatial dimension, and a time dimension; Wherein, the weight parameters in multiple dimensions include a first-dimension attention weight parameter, a second-dimension attention weight parameter, and a third-dimension attention weight parameter, the first dimension is the channel dimension, the second dimension is the spatial dimension, and the third dimension is the time dimension; enhancing the initial video features using the weight parameters in multiple dimensions based on a first attention module includes: enhancing the initial video features using the first-dimension attention weight parameter to obtain first-dimension video features; based on the second-dimension attention weight parameter, the third-dimension attention weight parameter, and the first-dimension video features, obtain the enhanced video features, where the second-dimension attention weight parameter is obtained by fusing the initial audio features and the first-dimension video features in the second dimension, and the third-dimension attention weight parameter is obtained by fusing the initial audio features and the first-dimension video features in the third dimension.

9. A video processing device, characterized in that, Include: A receiving module, configured to receive a video to be processed, and perform feature extraction on the video to be processed to obtain initial video features and initial audio features of the video to be processed; An enhancement module, configured to determine weight parameters in multiple dimensions based on the initial audio features, and perform enhancement processing on the initial video features by using the weight parameters in multiple dimensions based on a first attention module to obtain enhanced video features, where the multiple dimensions at least include a channel dimension, a spatial dimension, and a temporal dimension; A prediction module, configured to predict audiovisual events in the video to be processed based on the enhanced video features; Wherein, the weight parameters in the multiple dimensions include a first-dimension attention weight parameter, a second-dimension attention weight parameter, and a third-dimension attention weight parameter, the first dimension is the channel dimension, the second dimension is the spatial dimension, and the third dimension is the temporal dimension; the enhancement module is further configured to use the first-dimension attention weight parameter to enhance the initial video features to obtain first-dimension video features; and obtain the enhanced video features based on the second-dimension attention weight parameter, the third-dimension attention weight parameter, and the first-dimension video features, where the second-dimension attention weight parameter is obtained by fusing the initial audio features and the first-dimension video features in the second dimension, and the third-dimension attention weight parameter is obtained by fusing the initial audio features and the first-dimension video features in the third dimension.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 8.

11. A computer program product, characterized in that, The computer program product, when running, executes the method according to any one of claims 1 to 8.

12. A video processing system, characterized in that, Comprising: A processor; And A memory, connected to the processor, for providing instructions for the processor to perform the following processing steps: receiving a video to be processed, and performing feature extraction on the video to be processed to obtain initial video features and initial audio features of the video to be processed; Determining weight parameters in multiple dimensions based on the initial audio features, and performing enhancement processing on the initial video features by using the weight parameters in multiple dimensions based on a first attention module to obtain enhanced video features, where the multiple dimensions at least include a channel dimension, a spatial dimension, and a temporal dimension; Predicting audiovisual events in the video to be processed based on the enhanced video features; Among them, the weight parameters on the multiple dimensions include the first-dimension attention weight parameter, the second-dimension attention weight parameter, and the third-dimension attention weight parameter. The first dimension is the channel dimension, the second dimension is the spatial dimension, and the third dimension is the time dimension. Enhancing the initial video features by using the weight parameters on multiple dimensions based on the first attention module includes: enhancing the initial video features by using the first-dimension attention weight parameter to obtain the first-dimension video features; obtaining the enhanced video features based on the second-dimension attention weight parameter, the third-dimension attention weight parameter, and the first-dimension video features. The second-dimension attention weight parameter is obtained by fusing the initial audio features and the first-dimension video features in the second dimension, and the third-dimension attention weight parameter is obtained by fusing the initial audio features and the first-dimension video features in the third dimension.

Citation Information

Patent Citations

  • Event sentiment classification method and device, electronic equipment and storage medium

    CN112598067A

  • Video detection method and device, electronic equipment and storage medium

    CN112651319A

  • Audio-visual event positioning method and device based on cross-modal attention mechanism

    CN112989977A