Video processing method, device, storage medium and computer equipment
Through modal fusion and timing network generation, high timing resolution features are generated, combined with curtain segmentation point prediction and integrity evaluation network, the target screen is generated, which solves the problems of low efficiency and low accuracy of existing curtain segmentation technology, and achieves efficient and accurate curtain segmentation.
Patent Information
- Application Number
- CN202210653764.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-09
AI Technical Summary
The existing screen segmentation technology has low computational efficiency and low results accuracy, and relies on lens segmentation processing, resulting in high computational cost and low efficiency.
Modal fusion is performed by obtaining video data and associated text data, multimodal features are generated, and high timing resolution features are generated based on the timing network. Then, the screen segmentation point is used to predict the probability of the screen segmentation point, and the integrity of the nomination interval is evaluated through the screen integrity evaluation network, and the target screen is generated by the two.
It improves the accuracy of the screen segmentation result, reduces the calculation cost, improves the efficiency of screen segmentation, and avoids the problem of oversegment.
Smart Images

Figure CN115115975B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and more particularly, to a video processing method, apparatus, storage medium, and computer device. Background Art
[0002] With the rapid progress of storage technology and communication technology, the main carrier of information has gradually shifted from text and images to videos. Compared with text and images, videos can carry more information and are closer to the world perceived by humans. Video data not only contains both time dimension and space dimension, but also carries audio information and text information, and its application scenarios are very rich. Therefore, related technologies for video understanding have received extensive attention from people.
[0003] As one of the related technologies for video understanding, scene segmentation can divide a complete video into multiple scenes according to different video presentation forms and narrative techniques, so as to carry out subsequent video creative work such as video montage or derivation. Currently, scene segmentation technology usually adopts a scheme of first performing shot segmentation, then performing shot aggregation, and then finding scene segmentation points. This scheme not only has low computational efficiency but also has low accuracy of scene segmentation results. Summary of the Invention
[0004] Embodiments of the present application provide a video processing method, apparatus, storage medium, and computer device, aiming to improve the accuracy of scene segmentation results and the computational efficiency of scene segmentation.
[0005] On the one hand, embodiments of the present application provide a video processing method, which includes: obtaining video data and text data associated with the video data for modality fusion to obtain multi-modal features; generating high temporal resolution features corresponding to the multi-modal features based on a temporal network; predicting scene segmentation points for the high temporal resolution features according to a scene segmentation point prediction network to obtain the probability that each temporal position is a scene segmentation point; performing a pooling operation on the high temporal resolution features to obtain low temporal resolution features; evaluating the scene integrity of the low temporal resolution features according to a scene integrity evaluation network to obtain the scene integrity evaluation scores of each nomination interval; and combining the probability that each temporal position is a scene segmentation point with the scene integrity evaluation scores of each nomination interval to generate multiple target scenes corresponding to the video data.
[0006] On the other hand, an embodiment of the present application further provides a video processing device, which includes: a modality fusion module for obtaining video data and text data associated with the video data for modality fusion to obtain multi-modal features; a feature generation module for generating high temporal resolution features corresponding to the multi-modal features based on a temporal network; a segmentation point prediction module for predicting the scene segmentation points of the high temporal resolution features according to a scene segmentation point prediction network to obtain the probability of each temporal position being a scene segmentation point; a feature pooling module for performing a pooling operation on the high temporal resolution features to obtain low temporal resolution features; an integrity evaluation module for evaluating the scene integrity of the low temporal resolution features according to a scene integrity evaluation network to obtain the scene integrity evaluation scores of each nomination interval; and a target scene generation module for generating multiple target scenes corresponding to the video data by combining the probability of each temporal position being a scene segmentation point and the scene integrity evaluation scores of each nomination interval.
[0007] On the other hand, an embodiment of the present application further provides a computer-readable storage medium storing program code, wherein when the program code is run by a processor, the above video processing method is executed.
[0008] On the other hand, an embodiment of the present application further provides a computer device, which includes a processor and a memory, and the memory stores computer program instructions. When the computer program instructions are called by the processor, the above video processing method is executed.
[0009] On the other hand, an embodiment of the present application further provides a computer program product or a computer program, which includes computer instructions stored in a storage medium. The processor of the computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions to enable the computer device to execute the steps in the above video processing method.
[0010] A video processing method provided by this application can obtain video data and text data associated with the video data for modality fusion to obtain multi-modal features, generate high temporal resolution features corresponding to the multi-modal features based on a temporal network, and then predict the scene segmentation points for the high temporal resolution features according to the scene segmentation point prediction network to obtain the probability that each temporal position is a scene segmentation point, and perform a pooling operation on the high temporal resolution features to obtain low temporal resolution features. Then, according to the scene integrity evaluation network, the low temporal resolution features are evaluated for scene integrity to obtain the scene integrity evaluation scores for each nomination interval. Furthermore, by combining the probability that each temporal position is a scene segmentation point with the scene integrity evaluation scores for each nomination interval, multiple target scenes corresponding to the video data are generated. In this way, while accurately locating the scene segmentation points in the scene segmentation point prediction branch, the scene integrity evaluation branch is used to evaluate the scene integrity of the nomination regions containing the scene segmentation points, suppressing the over-segmentation problem, thereby greatly improving the accuracy of the scene segmentation result. And it does not rely on the segmented processing of shot segmentation, thus reducing the computational cost of scene segmentation and improving the efficiency of scene segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0012] Figure 1 FIG. shows a schematic diagram of a video including multiple scenes provided by this application.
[0013] Figure 2 FIG. shows a schematic diagram of a system architecture provided by an embodiment of this application.
[0014] Figure 3 FIG. shows a schematic flowchart of a video processing method provided by an embodiment of this application.
[0015] Figure 4 FIG. shows a network structure diagram of a temporal network provided by an embodiment of this application.
[0016] Figure 5 FIG. shows a network structure diagram of a scene segmentation point prediction network provided by an embodiment of this application.
[0017] Figure 6 FIG. shows a network structure diagram of a scene integrity evaluation network provided by an embodiment of this application.
[0018] Figure 7 FIG. shows a schematic flowchart of another video processing method provided by an embodiment of this application.
[0019] Figure 8 Shows a schematic diagram of an application scenario provided by an embodiment of the present application.
[0020] Figure 9 Shows a schematic flowchart of modal fusion provided by an embodiment of the present application.
[0021] Figure 10 Shows a schematic flowchart of obtaining nomination features provided by an embodiment of the present application.
[0022] Figure 11 Shows a flowchart of a scene segmentation task provided by an embodiment of the present application.
[0023] Figure 12 Is a block diagram of a video processing device provided by an embodiment of the present application.
[0024] Figure 13 Is a block diagram of a computer device provided by an embodiment of the present application.
[0025] Figure 14 Is a block diagram of a computer-readable storage medium provided by an embodiment of the present application. Detailed implementation manners
[0026] The following details the implementation manners of the present application. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The implementation manners described below by referring to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.
[0027] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.
[0028] The scene segmentation task specifically divides the semantically coherent scenes in the video according to the narrative techniques and presentation forms of the video. Compared with concepts such as actions and shots, a scene is a higher-level semantic concept. Just like the relationship between an object and a scene, the combination of interrelated objects forms the concept of a scene, and a scene is a super-shot composed of a group of semantically coherent shots. The scene segmentation task is to segment the video according to the differences in the high-level semantics of the video segments.
[0029] Please refer to Figure 1 ,Figure 1 A video clip containing six scenes is shown. The duration of the video is 34 seconds. Taking the first scene as an example, the first scene is the video segment from the 0th second to the 13th second of the video, and the first scene contains four shots. Exemplarily, the first scene can be the process of a customer service representative communicating with a user by phone. Shot A and Shot E are scenes where the customer service representative is making a call at the workstation, and Shot B and Shot F are scenes where the user is making a call on the street. Although the customer service representative making a call at the workstation and the user making a call on the street are in different scenes, these four shots together express the semantics of making a call, so they form the first scene.
[0030] In recent years, with the continuous development of deep learning, scene segmentation has shown high application value in actual application scenarios. For example, in the face of a large number of advertising videos, it is time-consuming and laborious to perform temporal analysis on advertising videos manually. The scene segmentation technology can use deep learning algorithms to divide a complete advertising video into multiple scenes according to the different presentation forms and narrative techniques of the advertising video. Among them, each scene can be used as advertising material composed of a group of semantically coherent shots, and these advertising materials can be used for subsequent advertising video mixing, derivation and other advertising creative work, which is the basis of advertising creative work.
[0031] Currently, scene segmentation methods usually adopt a bottom-up solution idea. First, a shot segmentation algorithm is used to divide the video into independent shots, then the spatio-temporal features of each shot video segment are extracted, and finally a clustering algorithm is used to splice the shots that are temporally continuous and similar in features to form scenes. The boundary accuracy of this solution completely depends on the shot detection algorithm, and a two-stage training method must be adopted, so it is impossible to achieve end-to-end in the process of network training and inference, resulting in high computational costs and low efficiency. In addition, shots of different lengths have different importance levels, but they are all summarized as spatio-temporal features of equal length for clustering, ultimately resulting in low accuracy of the scene segmentation results.
[0032] To solve the above problems, through research, the inventors proposed the video processing method provided in the embodiments of this application. This method can obtain video data and text data associated with the video data for modality fusion to obtain multi-modal features, and generate high temporal resolution features corresponding to the multi-modal features based on a temporal network. Then, according to the scene segmentation point prediction network, the scene segmentation points of the high temporal resolution features are predicted to obtain the probability that each temporal position is a scene segmentation point, and according to the scene integrity evaluation network, the scene integrity of the high temporal resolution features is evaluated to obtain the scene integrity evaluation score of each nomination interval. Furthermore, by combining the probability that each temporal position is a scene segmentation point with the scene integrity evaluation score of each nomination interval, multiple target scenes corresponding to the video data are generated.
[0033] Thus, while the screen segmentation point prediction branch accurately locates the screen segmentation point, the screen integrity evaluation branch is used to evaluate the screen integrity of the nomination region containing the screen segmentation point, suppressing the over-segmentation problem, thereby greatly improving the accuracy of the screen segmentation result. And it does not rely on the segmented processing of shot segmentation, thereby reducing the computational cost of screen segmentation and improving the efficiency of screen segmentation.
[0034] First, the architecture of the system of the video processing method involved in this application will be introduced below.
[0035] As Figure 2 shown, the video processing method provided in the embodiment of this application can be applied in the system 300. The data acquisition device 320 is used to acquire training data. For the video processing method of the embodiment of this application, the training data may include video data, text data, and training labels used for training. Among them, the labels used for training may be manually pre-annotated or calculated labels. After the training data is acquired, the data acquisition device 320 can store these training data in the database 340, and the training device 360 trains the target model 301 based on the training data maintained in the database 340.
[0036] The training device 360 trains a preset neural network based on the input video data and text data until the preset neural network meets the preset conditions, and obtains the trained target model 301. Among them, the preset conditions may be: the total loss value of the target loss function is less than the preset value, the total loss value of the target loss function no longer changes, or the number of training times reaches the preset number, etc.
[0037] The above target model 301 can be used to implement the video processing method of the embodiment of this application. The target model 301 in the embodiment of this application can specifically be a deep neural network model, for example, a convolutional neural network. It should be noted that in actual applications, the training data maintained in the database 340 does not necessarily all come from the acquisition of the data acquisition device 320, and it may also be received from other devices. Additionally, it should be noted that the training device 360 does not necessarily train the target model 301 completely based on the training data maintained in the database 340, and it may also obtain training data from the cloud or other places for model training. The above description should not be used as a limitation to the embodiment of this application.
[0038] The target model 301 trained according to the training device 360 can be applied to different systems or devices, such as applied to Figure 2The execution device 310 shown, and the execution device 310 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an Augmented Reality (AR) / Virtual Reality (VR), etc., or can also be a server or the cloud, etc.
[0039] In Figure 2 this, the execution device 310 can be used for data interaction with external devices. For example, a user can use the client device 330 to input data to the execution device 310 through a network. The input data in the embodiments of the present application can include: the video to be processed input by the client device. During the preprocessing of the input data by the execution device 310 or during the relevant processing such as calculation by the calculation module 311 of the execution device 310, the execution device 310 can call data, code, etc. in the data storage system 350 for corresponding calculation processing, or can also store the data, instructions, etc. obtained from the corresponding calculation processing into the data storage system 350.
[0040] Finally, the execution device 310 returns the processing result, for example, the multiple target screens generated by the target model 301, to the client device 330 through the network, so as to provide it to the user. It should be noted that the training device 360 can generate the corresponding target model 301 based on different training data for different targets or different tasks, and the corresponding target model 301 can be used to achieve the above targets or complete the above tasks, so as to provide the required results for the user.
[0041] Optionally, Figure 2 the system shown can be a Client-Server (C / S) system architecture. The execution device 310 can be a server (such as a cloud server), and the client device 330 can be a client (such as a laptop computer). The user can use the screen segmentation software in the laptop computer to upload the video to be segmented through the network to the cloud server. When the cloud server receives the video to be segmented, it uses the target model 301 to perform screen segmentation to generate multiple target screens, and returns the multiple target screens to the laptop computer. Then the user can obtain the multiple target screens on the screen segmentation software.
[0042] It should be noted that Figure 2 is only a schematic diagram of a system architecture provided by the embodiments of the present application. The architecture and application scenarios of the system described in the embodiments of the present invention are for more clearly explaining the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. For example, Figure 2The data storage system 350 therein is an external memory relative to the execution device 310. In other cases, the data storage system 350 can also be placed in the execution device 310. The execution device 310 can also directly obtain the video to be segmented as a client and perform segmentation. At this time, the execution device 310 is the client device. As is known to those of ordinary skill in the art, with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present invention are equally applicable to similar technical problems.
[0043] Please refer to Figure 3 , Figure 3 which shows a schematic flowchart of a video processing method provided by an embodiment of the present application. In a specific embodiment, the video processing method is applied to a video processing device 500 as shown in Figure 12 and a computer device 600 configured with the video output device 500 ( Figure 13 ).
[0044] The following will take the computer device as an example to illustrate the specific process of this embodiment. It can be understood that the computer device applied in this embodiment can be a server or a terminal, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, blockchain, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The video processing method can specifically include the following steps:
[0045] Step S110: Obtain video data and text data associated with the video data for modal fusion to obtain multi-modal features.
[0046] Considering that the scene segmentation task is more complex than the video understanding problem related to actions, and the text in the video often summarizes the video content and is very helpful for the scene segmentation task. Therefore, this application takes video data and text data as inputs to help the model understand higher-level semantic information to improve the accuracy of the scene segmentation result.
[0047] Among them, the video data is the video to be segmented into scenes. For example, an advertising video promoting a certain product. By segmenting the advertising video into scenes, downstream tasks such as video recommendation and video creative derivation can be performed based on the multiple scenes generated by the scene segmentation. Optionally, the subtitles in the video data can be obtained by using Optical Character Recognition (OCR) technology, that is, the text data associated with the video data. The existing lines or copywriting of the video can also be directly obtained as the text data, which is not limited here.
[0048] In some embodiments, after obtaining the video data and the text data associated with the video data, the video features corresponding to the video data and the text features corresponding to the text data can be extracted separately using a feature extractor. Further, a Cross-Attention network is used to perform modality fusion on the features of the video features and the text features to obtain a multi-modal feature representation. Among them, the Cross-Attention network can be a deep neural network introducing an attention mechanism, which is used to achieve a linear space complexity to complete a specified-order explicit feature combination. In the embodiments of the present application, the Cross-Attention network can select the most appropriate text feature at each time series position according to the correlation between the video features and the text features, and then fuse the most appropriate text feature with the video features to obtain a multi-modal feature.
[0049] For example, a game advertising video uploaded by a user can be obtained, and the corresponding OCR text can be obtained based on the game advertising video. Then, the game advertising video is input into a video feature extractor to obtain video features, such as a vector representation (Embedding) containing video frame image information, and the OCR text is input into a text feature extractor to obtain text features, such as a vector representation containing text information.
[0050] Further, the video features and the text features can be input into a Cross-Attention network for modality fusion to obtain the multi-modal features corresponding to the game advertising video. It should be noted that since neither the video features nor the text features contain time series position information, the Cross-Attention network will add time series position information to the video features and the text features during the process of modality fusion.
[0051] Since the OCR text information contains semantic information helpful for the scene segmentation task, the multi-modal features generated by the modality fusion can play a good auxiliary role in the subsequent scene segmentation, thereby improving the accuracy of generating multiple target scenes in the scene segmentation task.
[0052] Step S120: Generate high-time-series-resolution features corresponding to the multi-modal features based on a time series network.
[0053] A scene originally stems from the concept of drama and refers to a combination of a series of related events occurring in the same scene. The scene segmentation problem refers to dividing the semantically coherent scenes in a video according to the narrative method and presentation form of the video, and outputting the start and end times of each scene in the video. Because the appearance features of video frames may change greatly in the same scene. For example, during the switching of multiple shots, there may be a situation where the overall semantics of multiple shots are coherent, but there are significant changes between the shots. That is, when looking at the shots (low-level features) alone, they are not coherent, but the semantics are coherent. This leads to a very easy situation of mis-segmentation during scene segmentation. Therefore, the model used for scene segmentation must have the ability of long-term semantic modeling.
[0054] In the embodiment of this application, the temporal network is a deep neural network with the ability of long-term semantic modeling. Please refer to Figure 4 , Figure 4 which shows the network structure diagram of a temporal network. This temporal network has a total of three temporal scale levels, and there is a 2-fold temporal scale relationship between each layer. Among them, the upward arrow represents the 2-fold upsampling operation (up sample) in the temporal dimension, which is implemented by the nearest neighbor interpolation operation, and the downward arrow represents the 2-fold downsampling operation (down sample) in the temporal dimension, which is implemented by a 1D (Dimension) convolution with a kernel size of 3 and a stride of 2. The horizontal arrow represents the convolution unit, including a 1D convolution with a kernel size of 3 and a stride of 1 and a batch normalization layer.
[0055] As an implementation, the first level of this temporal network is a high temporal resolution sub-network, and the second and third levels are low temporal resolution sub-networks. The first level, the second level, and the third level are connected in parallel. After obtaining the multi-modal features, the multi-modal features can be input into the first level of the temporal network, and then multi-scale repeated fusion is performed by repeatedly exchanging information (upsampling and downsampling) on multiple parallel temporal resolution sub-networks. Finally, high temporal resolution features corresponding to the multi-modal features are output through the first level.
[0056] This temporal network maintains high temporal resolution features throughout the calculation process. Therefore, it avoids the information loss caused by restoring high temporal resolution features from low temporal resolution features. At the same time, the features of multiple temporal resolutions are continuously fused, so that the finally obtained features can capture the long-range dependencies in the video and generate high temporal resolution features that can improve the understanding ability of the scene segmentation task for temporal semantics, thereby improving the accuracy of the scene segmentation task in generating multiple target scenes.
[0057] Step S130: Predict the scene segmentation points for the high temporal resolution features according to the scene segmentation point prediction network, and obtain the probability that each temporal position is a scene segmentation point.
[0058] Considering that the boundary accuracy of existing scene segmentation methods completely depends on the shot detection algorithm and must adopt a two-stage training method, which cannot achieve end-to-end. Moreover, shots of different lengths have different importance levels, but they are all summarized as equi-length spatio-temporal features for combining into scenes, and thus the accuracy of the obtained scenes is not high.
[0059] Therefore, this application proposes a scene segmentation point prediction network and a scene integrity evaluation network to perform end-to-end scene segmentation on videos from two dimensions: the whole and the local. Generally, a video is composed of many video frames arranged in a time series. The temporal position can be understood as the specific position of a certain video frame in the time series dimension. The input of the scene segmentation point prediction network adopts local features with high temporal resolution, and based on this, the specific time of semantic change in the video can be accurately found, that is, the probability that each temporal position is a scene segmentation point can be found.
[0060] In some embodiments, the step of performing scene segmentation point prediction on the high-temporal-resolution features according to the scene segmentation point prediction network to obtain the probability that each temporal position is a scene segmentation point may include:
[0061] (1) Generating a target prediction feature map corresponding to the high-temporal-resolution features based on at least four convolutional blocks.
[0062] Among them, the scene segmentation point prediction network includes at least four convolutional blocks. Each convolutional block includes a convolutional layer, a batch normalization layer, and a non-linear layer. The convolutional kernels in each convolutional block are the same and the convolutional strides are the same. Optionally, the number of convolutional blocks can also be set according to the actual application requirements in combination with experiments, and no limitation is made here for examples. Please refer to Figure 5 , Figure 5 which shows a network structure diagram of a scene segmentation point prediction network.
[0063] The scene segmentation point prediction network includes a first convolutional block, a second convolutional block, a third convolutional block, and a fourth convolutional block. Each convolutional block contains a 1D convolutional layer, a batch normalization layer, and a Relu (Randomized Leaky) non-linear layer. The convolutional kernel size of the 1D convolutional layer is 3, and the convolutional stride is 1 to keep the temporal dimension of the output of each layer the same as the input temporal dimension.
[0064] As an implementation, the high temporal resolution features output by the temporal network can be input into the first convolutional block for the first convolutional processing to obtain the first predicted feature map, and the first predicted feature map can be input into the second convolutional block to obtain the second predicted feature map. Then, the second predicted feature map can be input into the third convolutional block to obtain the third predicted feature map, and the third predicted feature map can be input into the fourth convolutional block to obtain the target predicted feature map. By performing feature fusion on the high temporal resolution features through consecutive convolutional blocks, the predicted scene segmentation points are mapped onto the target predicted feature map with finer temporal and spatial dimensions.
[0065] (2) Based on the target predicted feature map, use the first activation function to calculate the probability that each temporal position is a scene segmentation point.
[0066] As an implementation, the first activation function can be the Sigmoid function. The Sigmoid function can be used as the output layer of the scene segmentation point prediction network. Thus, after obtaining the target predicted feature map, the target predicted feature map can be used as the independent variable and input into the Sigmoid function. Then, the Sigmoid function can calculate the probability that each temporal position is a scene segmentation point. Optionally, the target predicted feature map can be used as the independent variable and input into the Tanh activation function to calculate the boundary correction offset for each temporal position.
[0067] Step S140: Perform a pooling operation on the high temporal resolution features to obtain low temporal resolution features.
[0068] Since the main task of scene integrity evaluation is to understand the overall semantic information within the nominated region, a too high temporal resolution is not required. Reducing the temporal resolution can greatly reduce the GPU video memory occupancy and improve the throughput without affecting the model performance, thereby improving the computational efficiency of scene segmentation.
[0069] As an implementation, the time resolution of the high temporal resolution features can be reduced by performing a pooling operation (pooling) on the high temporal resolution features through a maximum pooling layer, and then the low temporal resolution features corresponding to the high temporal resolution features can be obtained.
[0070] Step S150: Perform scene integrity evaluation on the low temporal resolution features according to the scene integrity evaluation network to obtain the scene integrity evaluation scores for each nominated interval.
[0071] Considering that the input of the scene segmentation point prediction network is local features with high temporal resolution, that is, only having a local receptive field, and lacking complete context information, the scene segmentation points output by it may be the shot cut points with large feature changes in the scene. Therefore, this application proposes a scene integrity evaluation network to solve this problem.
[0072] Among them, the nomination interval refers to a temporal interval (from the start boundary to the end boundary) that may contain action segments. The scene integrity evaluation network is responsible for evaluating the scene integrity of the nomination region based on the overall features, mainly used to suppress the over-segmentation problem that may occur in scene segmentation point detection, and can output the scene integrity evaluation score of each nomination interval. The higher the scene integrity evaluation score, the better the scene integrity of the corresponding nomination interval; conversely, the worse the scene integrity. Please refer to Figure 6 , Figure 6 which shows the network structure diagram of a scene integrity evaluation network.
[0073] In some embodiments, the step of evaluating the scene integrity of the high temporal resolution features according to the scene integrity evaluation network to obtain the scene integrity evaluation score of each nomination interval may include:
[0074] (1) Determine the nomination feature map based on the low temporal resolution features and the sampling weight matrix.
[0075] The scene integrity evaluation network samples the features within each nomination region based on the low temporal resolution features, and then judges the scene integrity of this nomination region. Among them, the sampling weight matrix is a matrix composed of sampling weight masks corresponding to N sampling points on each nomination interval, and then the nomination feature map is determined according to the sampling weight matrix.
[0076] As an implementation manner, the step of determining the nomination feature map based on the low temporal resolution features and the sampling weight matrix may include:
[0077] (1.1) Obtain multiple nomination intervals.
[0078] (1.2) Generate the sampling weight matrix based on the nomination intervals.
[0079] (1.3) Determine the nomination feature maps corresponding to multiple nomination intervals based on the dot product of the low temporal resolution features and the sampling weight matrix.
[0080] Specifically, a nomination interval can be extended to obtain an extended nomination interval after extension. For example, the left and right boundaries of the nomination interval are each extended by half of the interval length, and then sampling operations are performed in the extended nomination interval to obtain sampling weight masks corresponding to multiple sampling points, and the sampling weight matrix is determined based on the multiple sampling weight masks. It should be noted that for video segments of a fixed length, the sampling weight matrix of each video is the same, so it only needs to be generated once in advance.
[0081] Further, perform a dot product on the low temporal resolution features and the sampling weight matrix corresponding to the nomination interval to calculate the feature representation corresponding to the nomination interval. Further, by enumerating all possible nomination intervals, the nomination feature map of all possible final nomination intervals can be obtained.
[0082] (2) Perform feature fusion on the nomination feature map to obtain an intermediate evaluation feature map.
[0083] As an implementation, similar to the network structure of the scene segmentation point prediction network, 4 two-dimensional convolutional blocks (Conv2D×4) can be used to perform feature fusion on the nomination feature map. Each convolutional block contains a two-dimensional convolutional layer, a batch normalization layer, and a Relu non-linear layer. Among them, the size of the two-dimensional convolutional kernel is 3, and the convolutional stride is 1 to keep the feature dimension of each layer's output the same as the input feature dimension.
[0084] (3) Upsample the intermediate evaluation feature map to obtain a target evaluation feature map.
[0085] Since, during the calculation of the nomination feature map, it is specifically represented as low temporal resolution features. To have the same temporal resolution as the features output by the scene segmentation point prediction network, the intermediate evaluation feature map needs to be restored to the original high temporal resolution. Optionally, use the bilinear interpolation algorithm (Bilinear Interpolation) to upsample the intermediate evaluation feature map to obtain the target evaluation feature map.
[0086] (4) Based on the target evaluation feature map, use the second activation function to calculate the scene integrity evaluation score for each nomination interval.
[0087] As an implementation, the second activation function can be the Sigmoid function. After obtaining the target evaluation feature map, the target evaluation feature map can be used as the independent variable and output to the Sigmoid function. Furthermore, the Sigmoid function calculates the scene integrity evaluation score for each nomination interval.
[0088] Step S160: Combine the probability of each temporal position being a scene segmentation point and the scene integrity evaluation score of each nomination interval to generate multiple target scenes for the video data.
[0089] Among them, when the probability of each temporal position being a scene segmentation point and the scene integrity evaluation score of each nomination interval are obtained, the two can be fused to calculate the prediction score of each nomination interval. Furthermore, based on the prediction scores, a prediction set of all possible nomination intervals is obtained, and the positions of the scene segmentation results in the prediction set are fine-tuned to obtain multiple target scenes.
[0090] In some embodiments, the step of generating multiple target scenes corresponding to the video data by combining the probability of each temporal position being a scene segmentation point and the scene integrity evaluation score of each nomination interval may include:
[0091] (1) Obtain the attenuation coefficient corresponding to each nomination interval.
[0092] Among them, the attenuation coefficient is inversely proportional to the maximum value of the probability of the scene segmentation point within the nomination interval (excluding the interval endpoints). That is, when the nomination interval already contains a scene segmentation point with a relatively high possibility, the prediction score of this nomination interval is reduced. Specifically, the attenuation coefficient of each nomination interval can be calculated based on the maximum value of the probability of the scene segmentation point within the nomination interval (excluding the interval endpoints).
[0093] (2) Based on the probability of each temporal position being a scene segmentation point, determine the probability of the interval endpoints of each nomination interval being scene segmentation points.
[0094] (3) Based on the attenuation coefficient, the probability of the interval endpoints being scene segmentation points, and the scene integrity evaluation score, determine the prediction score of each nomination interval.
[0095] (4) Determine multiple preselected scenes according to the prediction scores of each nomination interval.
[0096] As an implementation manner, the probability of the interval endpoints of each nomination interval being scene segmentation points can be obtained from the probability of each temporal position being a scene segmentation point. Further, based on the attenuation coefficient of each nomination interval, the probability of the interval endpoints of each nomination interval being scene segmentation points, and the scene integrity evaluation score of each nomination interval, a product operation is performed to obtain the prediction score of this nomination interval. According to the prediction scores of each nomination interval, a prediction set of all possible nomination intervals can be obtained.
[0097] Further, a non-maximum suppression algorithm (Non-Maximum Suppression, NMS) with an overlap threshold of 0 can be used for the prediction set composed of the multiple preselected scenes to obtain non-overlapping scene segmentation results, that is, multiple preselected scenes.
[0098] (5) Perform a fine-tuning operation on the multiple preselected scenes to obtain multiple target scenes corresponding to the multiple preselected scenes.
[0099] The goal of scene segmentation point prediction is to achieve high-precision scene segmentation point detection. To achieve this goal, the input of the scene segmentation point prediction network uses local features with high temporal resolution to find the specific time of semantic change. At the same time, the boundary correction offset can also be used to fine-tune the scene segmentation time point based on the time anchor, and then obtain multiple corresponding and more accurate target scenes.
[0100] As an implementation, the step of performing fine-tuning operations on multiple preselected scenes to obtain multiple target scenes corresponding to the multiple preselected scenes may include:
[0101] (5.1) Obtain the boundary correction offset at each time sequence position.
[0102] (5.2) Perform fine-tuning operations on the multiple preselected scenes according to the boundary correction offset to obtain multiple target scenes corresponding to the multiple preselected scenes.
[0103] During the actual scene segmentation process, affected by the maximum resolution, the segmented video segments will have the phenomenon of frame flickering, which brings a bad viewing experience to users. Therefore, an offset of 0.5 seconds before and after can be predicted for the time sequence point, that is, the boundary correction offset, so that the result of scene segmentation is more accurate.
[0104] Specifically, after the scene segmentation point prediction network generates the target prediction feature map corresponding to the high time sequence resolution feature, the boundary correction offset at each time sequence position can be calculated using an activation function (such as, Tanh activation function). Further, based on the boundary correction offset, the scene segmentation boundary points of each preselected scene are fine-tuned to obtain the corresponding multiple target scenes.
[0105] In the embodiment of the present application, video data and text data associated with the video data can be obtained for modality fusion to obtain multi-modal features, and high time sequence resolution features corresponding to the multi-modal features are generated based on the time sequence network. Further, the high time sequence resolution features are input into the scene segmentation point prediction network, and it is judged whether it is a scene segmentation point according to the features near the current moment. In order to further improve the algorithm performance under the high-precision scene segmentation requirement, the position of the scene segmentation point is fine-tuned within a small range using the boundary correction offset, so that the scene segmentation task has the ability of fine segmentation point positioning and improves the accuracy of the scene segmentation result.
[0106] Further, the low time sequence resolution features obtained by the pooling operation are input into the scene integrity evaluation network. The scene integrity evaluation network samples the features within the entire nomination region, and then judges the scene integrity of this nomination region. The scene integrity evaluation branch enables the model to have the ability of long-term semantic modeling. Finally, combining the probability that each time sequence position is a scene segmentation point and the scene integrity evaluation score of each nomination interval, multiple target scenes corresponding to the video data are generated as the scene segmentation result. The scene segmentation method proposed in the present application does not depend on the lens detection result, so end-to-end network calculation can be realized, thus greatly improving the calculation efficiency of scene segmentation.
[0107] Combined with the method described in the above embodiments, the following will give further detailed examples.
[0108] The video processing method of this application involves Artificial Intelligence (AI) technology. Artificial Intelligence technology uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. In other words, Artificial Intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can react in a way similar to human intelligence. Artificial Intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0109] Artificial Intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of Artificial Intelligence generally include technologies such as sensors, dedicated Artificial Intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of Artificial Intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0110] Computer Vision (CV) technology, as a branch of Artificial Intelligence technology, is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition, detection, and measurement, and further performing graphics processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build an Artificial Intelligence system that can obtain information from images or multi-dimensional data.
[0111] Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0112] The video processing method provided in this embodiment specifically involves technologies such as computer vision in Artificial Intelligence. Below, an example will be given with the video processing device specifically integrated in a computer device, and it will be elaborated in detail for the Figure 7 shown process in combination with Figure 8 the shown application scenario. This computer device can be a server or a terminal device, etc. Please refer to Figure 7 , Figure 7 which shows another video processing method provided in the embodiment of this application. In a specific embodiment, this video processing method can be applied to, for example,Figure 8 in the shown advertising creative scenario.
[0113] An advertising creative service provider provides a server side, which includes a cloud training server 410 and a cloud execution server 430. The cloud training server 410 is used to train a screen segmentation model for advertising creativity, and the cloud execution server 430 is used to deploy the screen segmentation model used for advertising creative operations and perform screen segmentation on the video sent by the client. Among them, the client is the advertising creative software 420 opened on a laptop when the user uses the advertising creative service. The video processing method may specifically include the following steps:
[0114] Step 201: The computer device obtains a training data set.
[0115] The video processing method provided in the embodiments of the present application includes the training of a preset attention network, a preset temporal network, a preset segmentation network, and a preset evaluation network. It should be noted that the training of the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network can be pre-performed according to the obtained training sample data set. Subsequently, when the video needs to be segmented by screen each time, the trained cross-attention network, temporal network, screen segmentation point prediction network, and screen integrity evaluation network can be directly calculated, without the need to perform network training again each time the screen is segmented.
[0116] In the embodiments of the present application, the training data set includes video training features, text training features, segmentation point detection labels, boundary correction offset labels, and screen integrity evaluation labels. Among them, the video training features and text training features can be generated in advance by a feature extractor. For example, using Swin-Transformer as the video feature extractor to obtain the video training feature F′ video , using the pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT) as the text feature extractor to obtain the text training feature F′ text .
[0117] Optionally, define the screen segmentation point annotation of the video as where S i is the normalized time of the i-th screen segmentation point, N g is the number of screen segmentation points in the video, and the real screen interval For the temporal position t on the feature map, define its represented time neighborhood as [t - 0.5, t + 0.5]. The segmentation point detection label can be obtained according to the screen segmentation point annotation Boundary correction offset label For each position (s, e) on the video integrity assessment score map, the video integrity assessment label can be calculated according to the following formula:
[0118]
[0119] where is the maximum value of the IoU between the nomination interval (s, e) and all true video intervals.
[0120] Step 202: The computer device obtains a preset attention network, a preset temporal network, a preset segmentation network, and a preset evaluation network.
[0121] Exemplarily, the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network are all pre-set neural networks. Among them, the preset attention network can be constructed by multi-head attention in a Transformer based on the Attention structure, and the cross-attention network obtained after training by the preset attention network performs modality fusion. The network structures of the preset temporal network, the preset segmentation network, and the preset evaluation network can refer to the descriptions of the temporal network, the video segmentation point prediction network, and the video integrity assessment network in the embodiments of the present application, and will not be described here.
[0122] Step 203: The computer device performs end-to-end network joint training on the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network through a training data set until the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network meet the preset conditions, and obtains the trained cross-attention network, temporal network, video segmentation point prediction network, and video integrity assessment network.
[0123] Considering that the existing video segmentation technology adopts a two-stage training method of shot segmentation and shot splicing, which cannot achieve end-to-end training, resulting in low efficiency of network training, and the computational efficiency of the network in the inference (application) stage is also relatively poor. Therefore, the embodiments of the present application propose to perform end-to-end network joint training on the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network to improve the efficiency of network training.
[0124] As an implementation manner, the video training feature and the text training feature are input into the preset attention network to obtain a multi-modal training feature after modality fusion, and the multi-modal training feature is input into the preset temporal network to obtain a high temporal resolution training feature.
[0125] Further, the high temporal resolution training feature is input into the preset segmentation network to obtain the segmentation point probability Boundary offset The high temporal resolution is passed through a pooling layer to obtain low temporal resolution training features, and then the low temporal resolution training features are input into a preset evaluation network to obtain an integrity evaluation score
[0126] Further, based on the segmentation point probability and the segmentation point detection label determine the segmentation point prediction loss function. For example, use the cross-entropy loss function, and the calculation formula is as follows:
[0127]
[0128] Further, based on the boundary offset and the boundary correction offset label determine the boundary offset loss function. For example, use the smooth L1 loss function, and the calculation formula is as follows:
[0129]
[0130] Further, based on the curtain integrity evaluation label and the integrity evaluation score determine the evaluation loss function. For example, use the smooth L1 loss function, and the calculation formula is as follows:
[0131]
[0132] Further, according to the segmentation point prediction loss function L s , the boundary offset loss function L0, and the evaluation loss function L c determine that the overall target loss function L of the network is:
[0133] L = αL s + βL0 + γL c
[0134] where α, β, and γ are the loss function proportionality coefficients, which are determined according to the actual situation of the experiment. Furthermore, the computer device performs end-to-end network joint training on the preset attention network, preset temporal network, preset segmentation network, and preset evaluation network according to the target loss function until the preset attention network, preset temporal network, preset segmentation network, and preset evaluation network meet the preset conditions.
[0135] It should be noted that the preset conditions can be: the total loss value of the target loss function is less than the preset value, the total loss value of the target loss function no longer changes, or the number of training times reaches the preset number, etc. Optionally, an optimizer can be used to optimize the target loss function, and the learning rate, batch size during training, and number of epochs for training can be set based on experimental experience.
[0136] Exemplarily, in Figure 8 the advertising creative scenario shown, the cloud training server 410 of the server can obtain a training data set, and obtain a preset attention network, a preset temporal network, a preset segmentation network, and a preset evaluation network. Then, through the training data set, end-to-end network joint training is performed on the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network until the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network meet the preset conditions, and the trained cross-attention network, temporal network, screen segmentation point prediction network, and screen integrity evaluation network are obtained.
[0137] Furthermore, the cross-attention network, the temporal network, the screen segmentation point prediction network, and the screen integrity evaluation network are deployed on the cloud execution server 430, and the cloud execution server 430 can perform screen segmentation on the video sent by the client.
[0138] Step 204: The computer device obtains the video data and the text data associated with the video data.
[0139] Exemplarily, in Figure 8 the advertising creative scenario shown, the user can click and upload a game advertising video through the advertising creative software 420 on the laptop 440, and at the same time, the advertising creative software 420 can obtain the corresponding OCR text based on the game advertising video. Optionally, when receiving the game advertising video, the cloud execution server 430 can also obtain the corresponding OCR text based on the game advertising video. The acquisition method of the text data can be set according to the actual product development needs.
[0140] Step 205: The computer device extracts the video features corresponding to the video data based on the video feature extractor, and extracts the text features corresponding to the text data based on the text feature extractor.
[0141] Exemplarily, in Figure 8 the advertising creative scenario shown, the cloud execution server 430 can use Swin-Transformer to extract the video feature F corresponding to the game advertising video v , and can use Bert to extract the text feature F corresponding to the OCR text t .
[0142] One thing to note is that there may be no text information in some time periods of the video. Therefore, zero padding is required according to the OCR timestamps to ensure that the video features and text features are aligned in the temporal dimension.
[0143] Step 206: The computer device performs modality fusion based on the video features and text features to obtain multi-modal features.
[0144] Considering that modality fusion is sensitive to time sequence, but the original video features and text features do not contain time sequence position information. Therefore, it is necessary to add time sequence position information to the fused features. In the embodiments of this application, a cross-attention network is used for modality fusion. The cross-attention network can select the most appropriate text features at each time sequence position according to the correlation between the video features and text features. Please refer to Figure 9 , Figure 9 shows a schematic flow diagram of a modality fusion. The following will elaborate on the modality fusion process in conjunction with Figure 9 for a detailed elaboration.
[0145] As an implementation manner, the steps for the computer device to perform modality fusion based on the video features and text features to obtain multi-modal features may include:
[0146] (1) The computer device obtains position encoding.
[0147] (2) Based on the position encoding, the video intermediate features of the video features and the text intermediate features of the text features are respectively calculated.
[0148] Optionally, the position encoding can be obtained through learning or directly obtained as fixed. For example, the position encoding PE (Positional Encoding) is calculated based on sine and cosine functions, which is not limited herein. After the computer device obtains the position encoding, it can splice the position encoding in the channel dimension after the video features and text features respectively to obtain the video composite feature F v +PE and the text composite feature F t +PE.
[0149] (3) The computer device performs a linear transformation on the video composite feature and the text composite feature to obtain the query vector corresponding to the video intermediate features, and the key vector and value vector corresponding to the text intermediate features.
[0150] Exemplarily, the computer device can perform a linear transformation on the video composite feature and the text composite feature by using the projection parameter matrices W Q , W K and W V to obtain the query vector Q corresponding to the video composite feature v, and the key vector K corresponding to the text composite feature t and the value vector V t , and the specific calculation formula is as follows:
[0151] Q v = W Q (F v + PE)
[0152] K t = W K (F t + PE)
[0153] V t = W V (F t + PE)
[0154] (4) The computer device inputs the query vector, the key vector, and the value vector into the cross-attention network to obtain the intermediate text feature.
[0155] (5) The computer device generates the multi-modal feature according to the intermediate text feature and the video feature.
[0156] Exemplarily, the computer device can input the query vector Q v , the key vector K t and the value vector V t into the cross-attention network, and calculate the intermediate text feature F o based on the network structure of multi-head attention (Multi-Head Attention, MHA). Further, the intermediate text feature F o is concatenated with the video feature F v to obtain the multi-modal feature F fution .
[0157] Step 207: The computer device generates the high temporal resolution feature corresponding to the multi-modal feature based on the temporal network.
[0158] Exemplarily, after obtaining the multi-modal feature, the computer device can input the multi-modal feature into the first layer of the temporal network, and then perform multi-scale repeated fusion by repeatedly exchanging information (upsampling and downsampling) on multiple parallel temporal resolution sub-networks, and finally output the high temporal resolution feature corresponding to the multi-modal feature through the first layer.
[0159] Step 208: The computer device predicts the scene segmentation point for the high temporal resolution feature according to the scene segmentation point prediction network to obtain the probability that each temporal position is the scene segmentation point.
[0160] In some embodiments, the step of the computer device predicting the curtain segmentation points for the high temporal resolution features according to the curtain segmentation point prediction network to obtain the probability of each temporal position being a curtain segmentation point may include:
[0161] (1) The computer device generates a target prediction feature map corresponding to the high temporal resolution features based on at least four convolutional blocks.
[0162] (2) The computer device calculates the probability of each temporal position being a curtain segmentation point based on the target prediction feature map by using a first activation function.
[0163] Exemplarily, the curtain segmentation point prediction network may include four convolutional blocks, and each convolutional block contains a 1D convolutional layer, a batch normalization layer, and a Relu non-linear layer. The computer device inputs the high temporal resolution features into the curtain segmentation point prediction network to obtain the target prediction feature map F SSD .
[0164] Further, based on the target prediction feature map F SSD , the probability of each temporal position being a curtain segmentation point is calculated by using the Sigmoid activation function The specific calculation formula is as follows:
[0165]
[0166] Step 209: The computer device performs a pooling operation on the high temporal resolution features to obtain low temporal resolution features.
[0167] The computer device inputs the high temporal resolution features into a pooling layer to obtain low temporal resolution features. Exemplarily, a maximum pooling layer may be used to reduce the high temporal resolution features to low temporal resolution features F p , and the specific calculation formula is as follows:
[0168]
[0169] Step 210: The computer device performs a curtain integrity evaluation on the low temporal resolution features according to the curtain integrity evaluation network to obtain the curtain integrity evaluation scores of each nomination interval.
[0170] In some embodiments, the step of the computer device performing a curtain integrity evaluation on the low temporal resolution features according to the curtain integrity evaluation network to obtain the curtain integrity evaluation scores of each nomination interval may include:
[0171] (1) The computer device determines a nomination feature map based on the low temporal resolution features and a sampling weight matrix.
[0172] Please refer to Figure 10 ,Figure 10 A flowchart showing a process for obtaining nomination features is presented. The following will elaborate on the process of determining the nomination feature map in combination with Figure 10 As an implementation, the steps for the computer device to determine the nomination feature map based on the low temporal resolution features and the sampling weight matrix may include:
[0173] (1.1) Obtain multiple nomination intervals.
[0174] (1.2) Generate a sampling weight matrix based on the nomination intervals.
[0175] (1.3) Determine the nomination feature map corresponding to the multiple nomination intervals based on the dot product of the low temporal resolution features and the sampling weight matrix.
[0176] Exemplarily, as Figure 10 shown, a nomination interval r = (t s , t e ) is given. To obtain the context information near the nomination interval, both the left and right boundaries are extended by half of the interval length. The extended nomination interval where d = t e - t s , representing the length of the nomination interval.
[0177] Furthermore, N points are uniformly sampled within the extended nomination interval r extend . The features at these N points are concatenated as the feature representation of the nomination interval r. However, the temporal positions of the sampling points may not be integers and need to be interpolated using the two features at nearby positions. This operation can be achieved using a sampling weight mask. For the nth sampling point t extend within the extended nomination interval r n , its sampling weight mask is defined as:
[0178]
[0179] Thus, where dec is the function to take the fractional part and floor is the function to take the integer part. Then the nomination feature at the sampling point t n can be calculated by the following formula:
[0180]
[0181] Furthermore, writing the sampling weight masks of the N sampling points in matrix form, the sampling weight matrix of the nomination interval r = (t s , t e ) can be obtained Thus, the nomination interval feature representation f s,e is obtained:
[0182]
[0183] Furthermore, by enumerating all possible nomination intervals, a nomination feature map composed of the nomination features of all nomination intervals can be obtained. It should be noted that for video segments of a fixed length, the sampling weight matrix of each video is the same, so it only needs to be pre-generated once.
[0184] (2) The computer device performs feature fusion on the nomination feature map to obtain an intermediate evaluation feature map.
[0185] (3) The computer device performs upsampling on the intermediate evaluation feature map to obtain a target evaluation feature map.
[0186] (4) The computer device calculates the shot integrity evaluation score of each nomination interval based on the target evaluation feature map using the second activation function.
[0187] As an implementation, similar to the network structure of the shot segmentation point prediction network, 4 two-dimensional convolutional blocks (Conv2D×4) can be used to perform feature fusion on the nomination feature map. Each convolutional block contains a two-dimensional convolutional layer, a batch normalization layer, and a Relu non-linear layer. Among them, the size of the two-dimensional convolutional kernel is 3, and the convolutional stride is 1 to keep the feature dimension of each layer output the same as the input feature dimension.
[0188] Furthermore, perform feature fusion on the nomination feature map to obtain an intermediate evaluation feature map, perform upsampling on the intermediate evaluation feature map to obtain a target evaluation feature map, and then calculate the shot integrity evaluation score of each nomination interval using the second activation function.
[0189] p c = Sigmoid(UpSample(Conv(F m )))
[0190] Step 211: The computer device combines the probability of each time sequence position being a shot segmentation point with the shot integrity evaluation score of each nomination interval to generate multiple target shots corresponding to the video data.
[0191] As an implementation, the step of the computer device combining the probability of each time sequence position being a shot segmentation point with the shot integrity evaluation score of each nomination interval to generate multiple target shots corresponding to the video data may include:
[0192] (1) Obtain the attenuation coefficient corresponding to each nomination interval.
[0193] Among them, the attenuation coefficient is inversely proportional to the maximum value of the probability of the power segmentation point within the nomination interval (excluding the interval endpoints). That is, when the nomination interval already contains a power segmentation point with a relatively high probability, the prediction score of the nomination interval is reduced. Specifically, the attenuation coefficient of each nomination interval can be calculated based on the maximum value of the probability of the power segmentation point within the nomination interval (excluding the interval endpoints).
[0194] (2) Based on the probability of each time series position being a power segmentation point, determine the probability of the interval endpoints of each nomination interval being power segmentation points.
[0195] (3) Based on the attenuation coefficient, the probability of the interval endpoints being power segmentation points, and the power integrity evaluation score, determine the prediction score of each nomination interval.
[0196] Exemplarily, given a nomination interval (s, e), the fused prediction score can be calculated according to the following formula:
[0197]
[0198]
[0199] Among them, is the attenuation coefficient.
[0200] (4) The computer device determines multiple preselected segments according to the prediction scores of each nomination interval.
[0201] (5) The computer device performs a fine-tuning operation on the multiple preselected segments to obtain multiple target segments corresponding to the multiple preselected segments.
[0202] Exemplarily, the computer device determines multiple preselected segments according to the prediction scores of each nomination interval, denoted as the prediction set of all possible nomination intervals Further, after the computer device generates the target prediction feature map corresponding to the high time series resolution feature in the power segmentation point prediction network, it can use an activation function (such as, Tanh activation function) to calculate the boundary correction offset P of each time series position o = 0.5 × Tanh(F SSD ), P o ∈ (-0.5, 0.5).
[0203] Further, by using the non-maximum suppression algorithm with an overlapping threshold of 0 for the prediction set Φ, a non-overlapping power segmentation result can be obtained. Finally, the boundary correction offset P o is used to fine-tune the power segmentation boundary points to obtain the final power segmentation prediction result
[0204] Exemplarily, please refer to Figure 11 , Figure 11A flowchart of a scene segmentation task is shown. In a specific embodiment, the scene segmentation task can be applied to Figure 8 Specifically, after acquiring the advertising video sent by the advertising creative software 420, the cloud execution server 430 can perform feature extraction and modality fusion based on the video clip and OCR text of the advertising video to obtain multimodal features, and calculate the high temporal resolution features corresponding to the multimodal features based on the temporal network.
[0205] Furthermore, based on the high temporal resolution feature, the scene segmentation point prediction branch and the scene integrity evaluation branch are calculated respectively to obtain the scene segmentation point probability, boundary correction offset and integrity evaluation score, and then score fusion is performed based on the scene segmentation point probability, boundary correction offset and integrity evaluation score to obtain a prediction set composed of multiple pre-selected scenes, and the prediction set is fine-tuned to obtain the final scene segmentation result. Then, the cloud execution server 430 sends the scene segmentation result to the advertising creative software 420, and the user obtains the multiple scenes segmented from the advertising video. Through these multiple scenes, that is, the advertising materials, subsequent advertising creative work such as advertising video mixing can be carried out.
[0206] In the embodiment of the present application, a training data set can be obtained, and a preset timing network, a preset segmentation network and a preset evaluation network can be obtained, and then the preset timing network, the preset segmentation network and the preset evaluation network are end-to-end network joint training is performed through the training data set until the preset timing network, the preset segmentation network and the preset evaluation network meet the preset conditions, and the trained timing network, the scene segmentation point prediction network and the scene integrity evaluation network are obtained. Thus, a top-down end-to-end scene segmentation framework is realized, and efficient scene segmentation can be achieved without relying on the shot segmentation algorithm.
[0207] Furthermore, the video data and the text data associated with the video data can be obtained for modal fusion to obtain multimodal features, and high temporal resolution features corresponding to the multimodal features are generated based on the temporal network. Then, the high temporal resolution features are input into the scene segmentation point prediction network, and whether it is a scene segmentation point is determined based on the features near the current moment. Furthermore, the low temporal resolution features obtained through the pooling operation are input into the scene integrity assessment network, which samples the features in the entire nominated area and then determines the scene integrity of this nominated area. The scene integrity assessment branch enables scene segmentation to have the ability of long-term semantic modeling and avoid the problem of over-segmentation.
[0208] Further, by combining the probability of each temporal position being a scene segmentation point with the scene integrity evaluation score of each nomination interval, multiple target scenes corresponding to the video data are generated as the scene segmentation result. To further improve the algorithm performance under the requirement of high-precision scene segmentation, the position of the scene segmentation point is slightly adjusted within a small range by using the boundary correction offset, so that the scene segmentation task has the ability to accurately locate the segmentation point and improves the accuracy of the scene segmentation result.
[0209] Please refer to Figure 12 , which shows a structural block diagram of a video processing device 500 provided by an embodiment of the present application. The video processing device 500 includes: a modality fusion module 510, configured to obtain video data and text data associated with the video data for modality fusion to obtain multi-modal features; a feature generation module 520, configured to generate high temporal resolution features corresponding to the multi-modal features based on a temporal network; a segmentation point prediction module 530, configured to perform scene segmentation point prediction on the high temporal resolution features according to a scene segmentation point prediction network to obtain the probability of each temporal position being a scene segmentation point; a feature pooling module 540, configured to perform a pooling operation on the high temporal resolution features to obtain low temporal resolution features; an integrity evaluation module 550, configured to perform scene integrity evaluation on the low temporal resolution features according to a scene integrity evaluation network to obtain the scene integrity evaluation score of each nomination interval; and a target scene generation module 560, configured to combine the probability of each temporal position being a scene segmentation point with the scene integrity evaluation score of each nomination interval to generate multiple target scenes corresponding to the video data.
[0210] In some embodiments, the target scene generation module 560 may include: a coefficient acquisition unit, configured to acquire an attenuation coefficient corresponding to each nomination interval; a probability determination unit, configured to determine the probability of the interval endpoint position of each nomination interval being a scene segmentation point based on the probability of each temporal position being a scene segmentation point; a score determination unit, configured to determine the prediction score of each nomination interval based on the attenuation coefficient, the probability of the interval endpoint position being a scene segmentation point, and the scene integrity evaluation score; a preselected scene determination unit, configured to determine multiple preselected scenes according to the prediction score of each nomination interval; and a target scene generation unit, configured to perform a fine-tuning operation on the multiple preselected scenes to obtain multiple target scenes corresponding to the multiple preselected scenes.
[0211] In some embodiments, the target scene generation unit may be specifically configured to: acquire a boundary correction offset of each temporal position; and perform a fine-tuning operation on the multiple preselected scenes according to the boundary correction offset to obtain multiple target scenes corresponding to the multiple preselected scenes.
[0212] In some embodiments, the modality fusion is performed by a cross-attention network. The video processing device 500 may further include: a training data acquisition module that acquires a training data set, where the training data set includes video training features, text training features, segmentation point detection labels, boundary correction offset labels, and scene integrity evaluation labels; a preset network acquisition module that is configured to acquire a preset attention network, a preset temporal network, a preset segmentation network, and a preset evaluation network; and a preset network training module that is configured to perform end-to-end network joint training on the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network using the training data set until the preset temporal network, the preset segmentation network, and the preset evaluation network meet the preset conditions, thereby obtaining a trained cross-attention network, a temporal network, a scene segmentation point prediction network, and a scene integrity evaluation network.
[0213] In some embodiments, the scene segmentation point prediction network includes at least four convolutional blocks. The segmentation point prediction module 530 may include: a predicted feature map generation unit that is configured to generate a target predicted feature map corresponding to high-temporal-resolution features based on at least four convolutional blocks, where each convolutional block includes a convolutional layer, a batch normalization layer, and a non-linear layer, and the convolutional kernels and convolutional strides in each convolutional block are the same; and a segmentation point probability calculation unit that is configured to calculate the probability that each temporal position is a scene segmentation point and the boundary correction offset of each temporal position based on the target predicted feature map using a first activation function.
[0214] In some embodiments, the predicted feature map generation unit may specifically be configured to: input the high-temporal-resolution features into a first convolutional block for first convolutional processing to obtain a first predicted feature map; input the first predicted feature map into a second convolutional block to obtain a second predicted feature map; input the second predicted feature map into a third convolutional block to obtain a third predicted feature map; and input the third predicted feature map into a fourth convolutional block to obtain the target predicted feature map.
[0215] In some embodiments, the integrity evaluation module 550 may include: a nomination feature map determination unit that is configured to determine a nomination feature map based on low-temporal-resolution features and a sampling weight matrix; a feature fusion unit that is configured to perform feature fusion on the nomination feature map to obtain an intermediate evaluation feature map; an upsampling unit that is configured to perform upsampling on the intermediate evaluation feature map to obtain a target evaluation feature map; and an integrity evaluation unit that is configured to calculate the scene integrity evaluation score of each nomination interval based on the target evaluation feature map using a second activation function.
[0216] In some embodiments, the nomination feature map determination unit may include: an interval acquisition subunit that is configured to acquire a plurality of nomination intervals; a matrix generation subunit that is configured to generate a sampling weight matrix based on the nomination intervals; and a feature map determination subunit that is configured to determine the nomination feature map based on the dot product of the low-temporal-resolution features and the sampling weight matrix.
[0217] In some embodiments, the matrix generation subunit may be specifically configured to: perform interval expansion on each nomination interval to obtain an expanded nomination interval; perform a sampling operation in the expanded nomination interval to obtain sampling weight masks corresponding to a plurality of sampling points; and determine a sampling weight matrix based on the plurality of sampling weight masks.
[0218] In some embodiments, the modality fusion module 510 may include: a data acquisition unit configured to acquire video data and text data associated with the video data; a video feature extraction unit configured to extract video features corresponding to the video data based on a video feature extractor; a text feature extraction unit configured to extract text features corresponding to the text data based on a text feature extractor; and a modality fusion unit configured to perform modality fusion based on the video features and the text features to obtain multi-modal features.
[0219] In some embodiments, the modality fusion unit may be specifically configured to: obtain a position encoding; generate a query vector corresponding to the video features based on the position encoding; generate a key vector and a value vector corresponding to the text data based on the position encoding; input the query vector, the key vector, and the value vector into a cross-attention network to obtain intermediate text features; and generate multi-modal features based on the intermediate text features and the video features.
[0220] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules may refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0221] In several embodiments provided in the present application, the coupling between modules may be electrical, mechanical, or other forms of coupling.
[0222] In addition, in each of the embodiments of the present application, the various functional modules may be integrated in one processing module, or each module may exist physically alone, or two or more modules may be integrated in one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0223] The solution provided by this application can obtain video data and text data associated with the video data for modality fusion to obtain multi-modal features, generate high temporal resolution features corresponding to the multi-modal features based on a temporal network, and then predict the scene segmentation points for the high temporal resolution features according to the scene segmentation point prediction network to obtain the probability of each temporal position being a scene segmentation point, and perform a pooling operation on the high temporal resolution features to obtain low temporal resolution features. Then, according to the scene integrity evaluation network, the low temporal resolution features are evaluated for scene integrity to obtain the scene integrity evaluation scores of each nomination interval. Furthermore, by combining the probability of each temporal position being a scene segmentation point with the scene integrity evaluation scores of each nomination interval, multiple target scenes corresponding to the video data are generated.
[0224] In this way, while the scene segmentation point prediction branch accurately locates the scene segmentation points, the scene integrity evaluation branch is used to evaluate the scene integrity of the nomination regions containing the scene segmentation points, suppressing the over-segmentation problem, thereby greatly improving the accuracy of the scene segmentation results. And it does not rely on the segmented processing of shot cuts, thus reducing the computational cost of scene segmentation and improving the efficiency of scene segmentation.
[0225] As Figure 13 shown, an embodiment of this application also provides a computer device 600. The computer device 600 includes a processor 610, a memory 620, a power supply 630, and an input unit 640. The memory 620 stores computer program instructions. When the computer program instructions are called by the processor 610, the various method steps provided by the above embodiments can be executed. Those skilled in the art can understand that the structure of the computer device shown in the figure does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Among them:
[0226] The processor 610 may include one or more processing cores. The processor 610 utilizes various interfaces and circuits to connect various parts within the entire battery management system. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 620, invoking data stored in the memory 620, performing various functions of the battery management system and processing data, as well as performing various functions of the computer device and processing data, the overall control of the computer device is achieved. Optionally, the processor 610 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 610 may integrate a combination of one or several of a central processing unit 610 (CPU), a graphics processing unit 610 (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 610 and can be implemented separately through a communication chip.
[0227] The memory 620 may include a random access memory 620 (RAM), and may also include a read-only memory 620 (ROM). The memory 620 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 620 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the computer device (such as a phone book and audio-video data), etc. Correspondingly, the memory 620 may also include a memory controller to provide the processor 610 with access to the memory 620.
[0228] The power supply 630 may be logically connected to the processor 610 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 630 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0229] An input unit 640, which can be used to receive input numerical or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0230] Although not shown, the computer device 600 may also include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 610 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 620 according to the following instructions, and the processor 610 will run the application programs stored in the memory 620, so as to implement the various method steps provided in the foregoing embodiments.
[0231] As Figure 14 shown, an embodiment of the present application also provides a computer-readable storage medium 700, in which computer program instructions 610 are stored, and the computer program instructions 710 can be called by a processor to execute the method described in the foregoing embodiments.
[0232] The computer-readable storage medium may be an electronic memory such as flash memory, EEPROM (electrically erasable programmable read-only memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium 700 has a storage space for program codes for executing any method steps in the above methods. These program codes can be read out from or written into one or more computer program products. The program codes can be compressed in an appropriate form, for example.
[0233] According to one aspect of the present application, there is provided a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementation manners provided in the foregoing embodiments.
[0234] The above are only the preferred embodiments of the present application, and there is no limitation to the present application in any form. Although the present application has been disclosed above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to equivalent embodiments with equivalent changes within the scope of the technical solution of the present application. However, as long as it does not depart from the content of the technical solution of the present application, any brief modifications, equivalent changes and modifications made to the above embodiments according to the technical essence of the present application still fall within the scope of the technical solution of the present application.
Claims
1. A video processing method, characterized in that The method includes: Obtaining video data and text data associated with the video data for modality fusion to obtain multi-modal features; Generating high temporal resolution features corresponding to the multi-modal features based on a temporal network; Predicting scene segmentation points for the high temporal resolution features according to a scene segmentation point prediction network to obtain the probability that each temporal position is a scene segmentation point; Performing a pooling operation on the high temporal resolution features to obtain low temporal resolution features; Evaluating the scene integrity of the low temporal resolution features according to a scene integrity evaluation network to obtain the scene integrity evaluation scores for each nomination interval; Combining the probability that each temporal position is a scene segmentation point with the scene integrity evaluation scores for each nomination interval to generate multiple target scenes corresponding to the video data, including: obtaining the attenuation coefficient corresponding to each nomination interval, where the attenuation coefficient corresponding to one nomination interval is inversely proportional to the maximum value of the probability that the temporal positions other than the interval endpoints in the nomination interval are scene segmentation points; determining the probability that the interval endpoints of each nomination interval are scene segmentation points based on the probability that each temporal position is a scene segmentation point; determining the prediction score for each nomination interval based on the attenuation coefficient, the probability that the interval endpoints are scene segmentation points, and the scene integrity evaluation score; determining multiple preselected scenes according to the prediction scores for each nomination interval; and performing a fine-tuning operation on the multiple preselected scenes to obtain multiple target scenes corresponding to the multiple preselected scenes.
2. The method according to claim 1, wherein The performing a fine-tuning operation on the multiple preselected scenes to obtain multiple target scenes corresponding to the multiple preselected scenes includes: Obtaining the boundary correction offset for each temporal position; Performing a fine-tuning operation on the multiple preselected scenes according to the boundary correction offset to obtain multiple target scenes corresponding to the multiple preselected scenes.
3. The method according to any one of claims 1 to 2, characterized in that The modality fusion is performed by a cross-attention network, and the cross-attention network, the temporal network, the scene segmentation point prediction network, and the scene integrity evaluation network are obtained through training by the following steps: Obtaining a training data set, where the training data set includes video training features, text training features, segmentation point detection labels, boundary correction offset labels, and scene integrity evaluation labels; Obtaining a preset attention network, a preset temporal network, a preset segmentation network, and a preset evaluation network; Performing end-to-end network joint training on the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network through the training data set until the entire network composed of the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network meets a preset condition, to obtain the trained cross-attention network, temporal network, scene segmentation point prediction network, and scene integrity evaluation network.
4. The method according to claim 1, characterized in that The scene segmentation point prediction network includes at least four convolutional blocks, and the predicting scene segmentation points for the high temporal resolution features according to the scene segmentation point prediction network to obtain the probability that each temporal position is a scene segmentation point includes: Generate a target prediction feature map corresponding to the high temporal resolution feature based on the at least four convolutional blocks, where each convolutional block includes a convolutional layer, a batch normalization layer, and a non-linear layer, and the convolutional kernels and convolutional strides in each convolutional block are the same; Based on the target prediction feature map, calculate the probability that each temporal position is a segmentation point and the boundary correction offset of each temporal position by using a first activation function.
5. The method according to claim 4, wherein The generating the target prediction feature map corresponding to the high temporal resolution feature based on the at least four convolutional blocks includes: Input the high temporal resolution feature into a first convolutional block for first convolutional processing to obtain a first prediction feature map; Input the first prediction feature map into a second convolutional block to obtain a second prediction feature map; Input the second prediction feature map into a third convolutional block to obtain a third prediction feature map; Input the third prediction feature map into a fourth convolutional block to obtain a target prediction feature map.
6. The method according to claim 1, characterized in that, The performing scene integrity evaluation on the low temporal resolution feature by a scene integrity evaluation network to obtain the scene integrity evaluation scores of each nomination interval includes: Determine a nomination feature map based on the low temporal resolution feature and a sampling weight matrix; Perform feature fusion on the nomination feature map to obtain an intermediate evaluation feature map; Perform upsampling on the intermediate evaluation feature map to obtain a target evaluation feature map; Based on the target evaluation feature map, calculate the scene integrity evaluation scores of each nomination interval by using a second activation function.
7. The method according to claim 6, characterized in that, The determining the nomination feature map based on the low temporal resolution feature and the sampling weight matrix includes: Obtain a plurality of nomination intervals; Generate a sampling weight matrix based on the nomination intervals; Determine the nomination feature maps corresponding to the plurality of nomination intervals based on the dot product of the low temporal resolution feature and the sampling weight matrix.
8. The method according to claim 7, wherein The generating the sampling weight matrix based on the nomination intervals includes: Perform interval expansion on each nomination interval to obtain an expanded nomination interval; Perform a sampling operation in the expanded nomination interval to obtain sampling weight masks corresponding to a plurality of sampling points; Determine a sampling weight matrix based on the plurality of sampling weight masks.
9. The method according to claim 1, wherein The obtaining video data and text data associated with the video data for modality fusion to obtain multimodal features includes: Obtain the video data and the text data associated with the video data; Extract video features corresponding to the video data based on a video feature extractor; Extract text features corresponding to the text data based on a text feature extractor; Perform modality fusion based on the video features and the text features to obtain multimodal features.
10. The method according to claim 9, characterized in that, The performing modality fusion based on the video features and the text features to obtain multimodal features includes: Obtain a position encoding; Calculate a video intermediate feature of the video features and a text intermediate feature of the text features respectively based on the position encoding; Perform a linear transformation on the video intermediate feature and the text intermediate feature to obtain a query vector corresponding to the video intermediate feature, and a key vector and a value vector corresponding to the text intermediate feature; Input the query vector, the key vector, and the value vector into a cross-attention network to obtain intermediate text features; Generate multimodal features based on the intermediate text features and the video features.
11. A video processing apparatus, characterized in that, The device includes: A modality fusion module for obtaining video data and text data associated with the video data for modality fusion to obtain multimodal features; A feature generation module for generating high temporal resolution features corresponding to the multimodal features based on a temporal network; A segmentation point prediction module for predicting the scene segmentation points of the high temporal resolution features according to a scene segmentation point prediction network to obtain the probability that each temporal position is a scene segmentation point; A feature pooling module for performing a pooling operation on the high temporal resolution features to obtain low temporal resolution features; An integrity evaluation module for evaluating the scene integrity of the low temporal resolution features according to a scene integrity evaluation network to obtain the scene integrity evaluation scores for each nomination interval; A target scene generation module for combining the probability that each temporal position is a scene segmentation point with the scene integrity evaluation scores for each nomination interval to generate multiple target scenes corresponding to the video data; The target scene generation module includes: a coefficient acquisition unit for acquiring the attenuation coefficient corresponding to each nomination interval, where the attenuation coefficient corresponding to one nomination interval is inversely proportional to the maximum value of the probability that the temporal positions other than the interval endpoints in the nomination interval are scene segmentation points; a probability determination unit for determining the probability that the interval endpoints of each nomination interval are scene segmentation points based on the probability that each temporal position is a scene segmentation point; a score determination unit for determining the prediction score of each nomination interval based on the attenuation coefficient, the probability that the interval endpoints are scene segmentation points, and the scene integrity evaluation score; a preselected scene determination unit for determining multiple preselected scenes according to the prediction scores of each nomination interval; a target scene generation unit for performing a fine-tuning operation on the multiple preselected scenes to obtain multiple target scenes corresponding to the multiple preselected scenes.
12. The device according to claim 11, characterized in that, The target scene generation unit is configured to: obtain the boundary correction offset of each temporal position; Perform a fine-tuning operation on the multiple preselected scenes according to the boundary correction offset to obtain multiple target scenes corresponding to the multiple preselected scenes.
13. The device according to claim 11 or 12, characterized in that, The modality fusion is performed by a cross-attention network, and the video processing device further includes: a training data acquisition module for acquiring a training data set, where the training data set includes video training features, text training features, segmentation point detection labels, boundary correction offset labels, and scene integrity evaluation labels; A preset network acquisition module for acquiring a preset attention network, a preset temporal network, a preset segmentation network, and a preset evaluation network; A preset network training module, which is used to perform end-to-end network joint training on the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network through the training data set until the entire network composed of the preset attention network, the preset temporal network, the preset segmentation network, and the preset evaluation network meets the preset conditions, and obtain the trained cross-attention network, temporal network, scene segmentation point prediction network, and scene integrity evaluation network.
14. The device according to claim 11, wherein The scene segmentation point prediction network includes at least four convolutional blocks, and the segmentation point prediction module includes: A predicted feature map generation unit, which is used to generate a target predicted feature map corresponding to the high temporal resolution feature based on the at least four convolutional blocks, wherein each convolutional block includes a convolutional layer, a batch normalization layer, and a non-linear layer, and the convolutional kernels and convolutional strides in each convolutional block are the same; A segmentation point probability calculation unit, which is used to calculate the probability that each temporal position is a scene segmentation point and the boundary correction offset of each temporal position based on the target predicted feature map by using a first activation function.
15. The device according to claim 14, characterized in that, The predicted feature map generation unit is configured to: input the high temporal resolution feature into the first convolutional block for first convolutional processing to obtain a first predicted feature map; Input the first predicted feature map into the second convolutional block to obtain a second predicted feature map; Input the second predicted feature map into the third convolutional block to obtain a third predicted feature map; Input the third predicted feature map into the fourth convolutional block to obtain a target predicted feature map.
16. The device according to claim 11, wherein The integrity evaluation module includes: A nomination feature map determination unit, which is used to determine a nomination feature map based on the low temporal resolution feature and the sampling weight matrix; A feature fusion unit, which is used to perform feature fusion on the nomination feature map to obtain an intermediate evaluation feature map; An upsampling unit, which is used to perform upsampling on the intermediate evaluation feature map to obtain a target evaluation feature map; An integrity evaluation unit, which is used to calculate the scene integrity evaluation score of each nomination interval based on the target evaluation feature map by using a second activation function.
17. The device according to claim 16, characterized in that, The nomination feature map determination unit includes: an interval acquisition subunit, which is used to acquire a plurality of nomination intervals; A matrix generation subunit, which is used to generate a sampling weight matrix based on the nomination intervals; A feature map determination subunit, which is used to determine the nomination feature maps corresponding to the plurality of nomination intervals based on the dot product of the low temporal resolution feature and the sampling weight matrix.
18. The device according to claim 17, characterized in that, The matrix generation subunit is used to: expand each nomination interval to obtain an expanded nomination interval after expansion; Perform a sampling operation in the expanded nomination interval to obtain sampling weight masks corresponding to a plurality of sampling points; Determine a sampling weight matrix based on the plurality of sampling weight masks.
19. The device according to claim 11, characterized in that, The modality fusion module includes: A data acquisition unit, which is used to acquire video data and text data associated with the video data; A video feature extraction unit, which is used to extract video features corresponding to the video data based on a video feature extractor; A text feature extraction unit, which is used to extract text features corresponding to the text data based on a text feature extractor; A modality fusion unit, configured to perform modality fusion based on the video feature and the text feature to obtain a multi-modal feature.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program codes, and the program codes can be called by a processor to execute the method according to any one of claims 1 to 10.
21. A computer device, characterized in that, Comprising: A memory; One or more processors, coupled to the memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the method according to any one of claims 1 to 10.
22. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a storage medium, a processor of a computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions, so that the computer device executes the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Image segmentation method and device, computer equipment and storage medium
CN112818955A
Video processing method and device, computer equipment and storage medium
CN114363695A