Video processing method and device

By using the target model based on the Transformer structure in the training of multimodal large models to determine the encoding parameters of the video and encode it, the problem of high bandwidth cost in the video preprocessing link is solved, and the video code rate and bandwidth cost are reduced are achieved.

CN119967187APending Publication Date: 2025-05-09ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510121976.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

During the training of multimodal large models, multiple video upload and download operations are involved in the video preprocessing link, resulting in high bandwidth costs.

Method used

By acquiring the video and determining suitable encoding parameters using a target model based on the Transformer structure, the video is then encoded using the target encoder, reducing the bit rate of the video, thereby reducing bandwidth costs.

Benefits of technology

It effectively reduces the bandwidth cost of video during transmission and provides a lower bandwidth video preprocessing link for training multimodal large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967187A_ABST
    Figure CN119967187A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video processing method and device. The method comprises the following steps: acquiring a first video; based on the first video, a first coding parameter is determined through a target model, the target model is obtained through training based on a sample video and a sample coding parameter corresponding to the sample video, the sample coding parameter is related to a target encoder, and the target model is a model based on a Transform structure; and encoding the first video by using the target encoder configured with the first encoding parameter to obtain a first encoded video. According to the method and the device, the appropriate low-bit-rate first coded video is obtained based on the first coding parameter which is more adaptive to the first video, and a basis is provided for reducing the bandwidth cost in a video preprocessing link in the training process of a multi-modal large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a video processing method and device. Background Art

[0002] At present, multimodal research plays an important role in various fields, so the training of multimodal large models is particularly important. In some scenarios, the training of multimodal large models depends on massive videos. Videos require higher bandwidth than images and text data. In particular, the preprocessing link of multimodal data such as videos involves multiple large-scale video upload and download operations, which often introduces high bandwidth costs. Therefore, how to provide an improved video processing method to reduce the bandwidth cost in the video preprocessing link in the above scenario has become an urgent problem to be solved. Summary of the invention

[0003] One or more embodiments of the present specification provide a video processing method and device to achieve better reduction of the video bit rate, and further provide a basis for reducing the bandwidth cost in the video preprocessing link during the training of a large multimodal model.

[0004] According to a first aspect, a video processing method is provided, comprising:

[0005] Get the first video;

[0006] Based on the first video, determining a first encoding parameter through a target model, wherein the target model is trained based on a sample video and its corresponding sample encoding parameters, the sample encoding parameters are related to a target encoder, and the target model is a model based on a Transformer structure;

[0007] The first video is encoded using the target encoder configured with the first encoding parameter to obtain a first encoded video.

[0008] According to a second aspect, a video processing device is provided, comprising:

[0009] A first acquisition module, configured to acquire a first video;

[0010] The first determination module is configured to determine a first encoding parameter based on the first video through a target model, wherein:

[0011] The target model is obtained by training based on a sample video and its corresponding sample encoding parameters, the sample encoding parameters are related to the target encoder, and the target model is a model based on a Transformer structure;

[0012] The first encoding module is configured to encode the first video using the target encoder configured with the first encoding parameter to obtain a first encoded video.

[0013] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in the first aspect.

[0014] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in the first aspect is implemented.

[0015] According to the video processing method and device provided in the embodiments of this specification, a first video is obtained; based on the first video, a first encoding parameter is determined through a target model, wherein the target model is trained based on a sample video and its corresponding sample encoding parameter, the sample encoding parameter is related to a target encoder, and the target model is a model based on a Transformer structure; the first video is encoded using a target encoder configured with the first encoding parameter to obtain a first encoded video. In the above process, based on the first video, a first encoding parameter that is more suitable for the first video is determined through a trained target model based on a Transformer structure, and then the first video is encoded using a target encoder configured with the first encoding parameter to obtain a first encoded video, thereby encoding the first video to reduce the video bit rate, thereby reducing the bandwidth cost of the encoded video during transmission, such as uploading and downloading. In some examples, a basis is further provided for reducing the bandwidth cost in the video preprocessing link during the training of a large multimodal model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A schematic diagram of an implementation framework of an embodiment disclosed in this specification;

[0018] Figure 2 A schematic diagram of a process of constructing training data of a target model provided in an embodiment;

[0019] Figure 3 An example diagram of the training process of the target model provided in the embodiment;

[0020] Figure 4A schematic diagram of the structure of an initial encoding parameter generation model provided in an embodiment;

[0021] Figure 5 A schematic diagram of a flow chart of a video processing method provided in an embodiment;

[0022] Figure 6 A schematic diagram of another flow chart of the video processing method provided in the embodiment;

[0023] Figure 7 A schematic diagram of a video processing scenario provided by an embodiment;

[0024] Figure 8 A schematic block diagram of a video processing device provided in an embodiment. DETAILED DESCRIPTION

[0025] The technical solutions of the embodiments of this specification will be described in detail below with reference to the accompanying drawings.

[0026] The embodiments of this specification disclose a video processing method and device. The application scenarios and technical concepts of the method are first introduced as follows:

[0027] As mentioned above, in some scenarios, such as the training of multimodal large models, it is necessary to rely on massive videos, which require higher bandwidth than images and text data. In particular, the preprocessing link of multimodal data such as videos involves multiple large-scale video upload and download operations, which often introduces high bandwidth costs. Therefore, how to provide an improved video processing method to reduce the bandwidth cost in the video preprocessing link in the above scenario has become an urgent problem to be solved.

[0028] In view of this, the inventor proposes a video processing method. Figure 1 A schematic diagram of an implementation scenario according to an embodiment disclosed in this specification is shown. In this implementation scenario, a first video is obtained, and a first encoding parameter is determined based on the first video through a target model, wherein the target model is trained based on a sample video and its corresponding sample encoding parameter, the sample encoding parameter is related to a target encoder, and the target model is a model based on a Transformer structure; the first video is encoded using a target encoder configured with the first encoding parameter to obtain a first encoded video.

[0029] In the above process, based on the first video, a first encoding parameter that is more suitable for the first video is determined through a trained target model based on a Transformer structure, and then the first video is encoded using a target encoder configured with the first encoding parameter to obtain a first encoded video, thereby encoding the first video to reduce the video bit rate, thereby reducing the bandwidth cost of the encoded video during transmission, such as uploading and downloading. In some examples, a basis is provided for reducing the bandwidth cost in the video preprocessing link during the training of a multimodal large model.

[0030] The video processing method provided in this specification is described in detail below in conjunction with specific embodiments.

[0031] In the video processing flow provided in the embodiments of this specification, it is necessary to use a trained target model to determine encoding parameters related to a target encoder that are more suitable for the video to be processed, and then use the aforementioned determined encoding parameters to configure the target encoder, and use the target encoder configured with the aforementioned determined encoding parameters to encode the video to obtain a corresponding encoded video, so as to reduce the bit rate of the video and provide a basis for reducing the bandwidth cost during transmission processes such as video uploading and downloading.

[0032] For a clear layout, the following first introduces the training process of the target model and the construction process of the training data corresponding to the target model.

[0033] in, Figure 2 A flowchart of the process of constructing the training data corresponding to the target model is shown in FIG. Figure 2 As shown, the construction process may include the following steps S210-S240:

[0034] In step S210, a sample video is obtained.

[0035] In some exemplary scenarios, multiple videos of different video scenes and different resolutions can be obtained from different video sources to form a video set. Any video in the video set can be used as a sample video for training the model, and accordingly, any video can be obtained from the video set as a sample video. Exemplarily, the aforementioned videos of different video scenes may include: videos for animals, videos for plants, videos for people, movies, MVs, and other types of videos. The videos in the video set may include various short videos.

[0036] In some cases, the number of videos in the video set may exceed the specified number so as to better train the model and obtain a model with better output accuracy and stability.

[0037] It can be understood that the process of taking any video in the video set as a sample video and then determining its corresponding sample coding parameters is similar. The following takes any video j as a sample video as an example to introduce the process of determining its corresponding sample coding parameters. The process of determining the sample coding parameters corresponding to other videos in the video set can refer to the process of determining the sample coding parameters corresponding to video j, which will not be repeated here.

[0038] In a specific example, in the aforementioned step S210, any video j can be obtained from the aforementioned video set as a sample video. Next, in step S220, multiple groups of candidate encoding parameters for the target encoder are obtained.

[0039] The target encoder may be any encoder for encoding a video. In some specific examples, the target encoder may be an AV1 (AOMedia Video 1) encoder.

[0040] The multiple sets of candidate encoding parameters may be randomly set based on a feature program or manually set based on experience. A single set of candidate encoding parameters may include one or more parameters related to a target encoder, wherein the parameters may include, but are not limited to, at least one of the following: rate control mode, key frame interval, bit rate, filter and tuning related parameters, denoising and sharpening related parameters, and input and output formats, etc.

[0041] Thereafter, in step S230, the sample videos are encoded using the target encoders configured with the respective groups of candidate encoding parameters to obtain sample encoded videos corresponding to the respective groups of candidate encoding parameters.

[0042] In this step, for each group of candidate coding parameters (the i-th group of candidate coding parameters is used as an example for explanation, i is an integer between [1, R], and R is the total number of groups of candidate coding parameters obtained), the i-th group of candidate coding parameters is used to configure the target encoder, and then the sample video, i.e., the aforementioned video j, is encoded using the target encoder configured with the i-th group of candidate coding parameters to obtain a sample encoded video corresponding to the i-th group of candidate coding parameters. In this way, the sample encoded videos corresponding to each group of candidate coding parameters are obtained.

[0043] Next, in step S240, a sample coded video whose quality meets the preset quality condition and has the lowest bit rate is determined from the sample coded videos corresponding to each group of candidate coding parameters, and its corresponding candidate coding parameters are determined as the sample coding parameters corresponding to the aforementioned sample video.

[0044] In this step, a specific video evaluation index may be used to determine the index value corresponding to the sample coded video corresponding to each group of candidate coding parameters, and then based on the index value and the corresponding bit rate corresponding to the sample coded video corresponding to each group of candidate coding parameters, a sample coded video whose quality meets the preset quality condition and has the lowest bit rate is determined from the sample coded videos corresponding to each group of candidate coding parameters. The aforementioned index value may indicate the quality of the corresponding sample coded video. Then, the sample coded video whose quality meets the preset quality condition and has the lowest bit rate is determined as the sample coding parameter corresponding to the sample video.

[0045] Exemplarily, the specific video evaluation index may include but is not limited to at least one of the following indicators: SSIM (Structural Similarity Index), PSNR (Peak Signal-to-Noise Ratio) and VMAF (Video Multimethod Assessment Fusion), etc. It can be understood that the above is only an exemplary example of a specific video evaluation index, and any indicator in the related art that can evaluate the quality of the encoded video can be applied to the embodiments of this specification. For example, the aforementioned specific video evaluation indicators may also include LPIPS (Learned Perceptual Image Patch Similarity) and the like.

[0046] Exemplarily, taking the specific video evaluation index including SSIM as an example, the process of determining the index value corresponding to the sample coded video corresponding to each group of candidate coding parameters includes: for each group of candidate coding parameters, using SSIM, determining the structural similarity value between the sample coded video corresponding to the group of candidate coding parameters and the sample video as the index value corresponding to the sample coded video corresponding to the group of candidate coding parameters; and so on, obtaining the index value corresponding to the sample coded video corresponding to each group of candidate coding parameters. The larger the structural similarity value between the sample coded video and the sample video, that is, the larger the index value corresponding to the sample coded video corresponding to the group of candidate coding parameters, the better the quality of the sample coded video can be represented.

[0047] Continuing with the above example, the process of determining a sample coded video whose quality meets the preset quality conditions and has the lowest bit rate is introduced. In some possible examples, the process may include: determining a sample coded video with the lowest bit rate from multiple sample coded videos whose corresponding indicator values ​​are within a specified range, as the sample coded video with the quality meeting the preset quality conditions and the lowest bit rate.

[0048] Exemplarily, the lower limit value of the specified range may not be lower than the specified index value threshold value to ensure the quality of the encoded video. In the above manner, based on the principle of JND (Just Noticed Distortion), it is possible to determine from each group of candidate encoding parameters the encoding parameters that can make the corresponding encoded video bit rate low but do not affect the perceived quality of the video, and obtain sample encoding parameters that are more suitable for the sample video, that is, by configuring the target encoder of the more suitable sample encoding parameters, a sample encoded video with good video quality and the lowest bit rate can be obtained. Afterwards, the corresponding model is trained by such sample videos and their corresponding sample encoding parameters, so that the trained model can output encoding parameters that are more suitable for the input video, and then the encoded video with good video quality and low bit rate corresponding to the input video can be obtained, which provides a better basis for reducing the bandwidth cost in the transmission process such as video uploading and downloading.

[0049] After that, after obtaining a plurality of sample videos and their corresponding sample encoding parameters in the above manner, the plurality of sample videos and their corresponding sample encoding parameters can be used to train a subsequent initial encoding parameter generation model to obtain a target model. Figure 3 As shown, a schematic diagram of a process flow of a target model training process is shown, wherein the training process may include the following steps S310-S320;

[0050] In step S310, the sample video and its corresponding sample encoding parameters are used to train an initial encoding parameter generation model to obtain a first encoding parameter generation model that meets a preset convergence condition.

[0051] The encoding parameter generation model is a model based on a Transformer structure. In some possible examples, the encoding parameter generation model may be a 3D-Swin-Transformer model.

[0052] like Figure 4 As shown in FIG. , an exemplary structure of an initial coding parameter generation model is provided. Figure 4As shown, the initial encoding parameter generation model includes a 3D partitioning layer (as shown in the figure "3DPatch Partition") and multiple feature processing layers. Among them, the first feature processing layer may include a linear embedding sublayer (as shown in the figure "Linear Embedding") and several Swin Transformer blocks (as shown in the figure "Video Swin Transfoemer Block"); other feature processing layers may include a patch merging sublayer (as shown in the figure "Patch Merging") and several SwinTransformer blocks. Among them, each feature processing layer includes several Transformer blocks, which process their inputs based on the attention mechanism and Swin (Shifted Window) sliding window mechanism). The number of SwinTransformer blocks between different feature processing layers may be the same or different.

[0053] The aforementioned 3D blocking layer can block the input video, such as a video frame sequence of size T×H×W×3, according to a preset size a×b×c to obtain a corresponding patch sequence, where T represents the number of frames of the input video, i.e., the video frame sequence, and H and W represent the height and width of the video frame in the input video, respectively, where 3 represents the number of channels, i.e., RGB channels, respectively. a, b, and c are all positive integers, and exemplarily, a, b, and c can be equal, and exemplarily, b and c can be equal, and a can be taken as 1 / 2 of b or c.

[0054] Assume that a is 2, b and c are both 4, and the preset size is 2×4×4. Figure 4 As shown, a video frame sequence of size T×H×W×3 can obtain non-overlapping T / 2×H / 4×W / 4 patch blocks after 3D blocking layer blocking, where the feature dimension of each patch block is 2×4×4×3=96.

[0055] Afterwards, each patch block is processed by the linear embedding sublayer of the first feature processing layer to project each patch block onto an arbitrary dimension C to obtain the token corresponding to each patch block. Figure 4 As shown, a token sequence of size T / 2×H / 4×W / 4×C can be obtained, which is used as the input of the SwinTransformer block connected to the linear embedding sublayer of the first feature processing layer.

[0056] For example, the Swin Transformer blocks in each feature processing layer appear in pairs, that is, the number of Swin Transformer blocks in each feature processing layer can be expressed as 2k, where k represents a positive integer, and the value of k between different feature processing layers can be the same or different. Figure 4 As shown, the first feature processing layer 1 may include 2k1 Swin Transformer blocks, the second feature processing layer 2 may include 2k2 Swin Transformer blocks, the third feature processing layer 3 may include 2k3 Swin Transformer blocks, ... the Mth feature processing layer M may include 2km Swin Transformer blocks.

[0057] Among them, in the paired Swin Transformer blocks, the first block uses the W-MSA (Window Multi-Head Self-Attention) structure, and the second block uses the SW-MSA (Shifted Window Multi-Head Self-Attention) structure. The way the W-MSA structure and the SW-MSA structure process data can refer to the data processing method of the corresponding structure in the relevant technology, which will not be repeated here.

[0058] Exemplarily, in each Swin Transformer block, the W-MSA structure (or SW-MSA structure) may also include a first LayerNorm sublayer before it, which is used to normalize the input of the Swin Transformer block; after the W-MSA structure (or SW-MSA structure), a first residual connection sublayer is also provided, which is used to add the output of the W-MSA structure (or SW-MSA structure) to the input of the Swin Transformer block; after the first residual connection sublayer, a second LayerNorm sublayer is provided, which is used to normalize the output of the first residual connection sublayer; after the second LayerNorm sublayer, an MLP (Multi-Layer Perceptron) is provided, which is used to further process its input (the output of the second LayerNorm layer); then, after the MLP, a second residual connection sublayer is also connected, which is used to normalize the output of the MLP and the output of the first residual connection sublayer; then, after the second residual connection sublayer, the next Swin The Transformer block, or, is connected to the next feature processing layer, or its output is used as the model output.

[0059] Exemplarily, the patch merging sublayer in the feature processing layer can be used to downsample its input, such as Figure 4 As shown, the patch merging sublayer of the second feature processing layer downsamples the size of the output of the first feature processing layer from T / 2×H / 4×W / 4×C to T / 2×H / 8×W / 8×2C, for example, merging multiple tokens corresponding to multiple patch blocks within a specific size range (e.g., 1×2×2) in the output of the first feature processing layer (exemplarily, the merging includes: splicing multiple tokens in the feature dimension and then linearly processing the splicing result); the patch merging sublayer of the third feature processing layer downsamples the size of the output of the second feature processing layer from T / 2×H / 8×W / 8×2C to T / 2×H / 16×W / 16×4C.

[0060] In step S310, the sample video may be input into the coding parameter generation model, so that the sample video is processed by the coding parameter generation model to obtain the predicted coding parameters, and then based on the preset loss function, the predicted coding parameters are combined with the sample coding video corresponding to the sample video to determine the current loss, and the parameters of the coding parameter generation model are adjusted with the goal of minimizing the current loss. The preset loss function may be, for example, but not limited to, a cross entropy loss function, a mean square error loss function, etc.

[0061] It is understandable that the parameters of the encoding parameter generation model can be adjusted cyclically and iteratively using multiple sample videos and their corresponding multiple sample encoding parameters to obtain a first encoding parameter generation model that meets the preset convergence condition. Exemplarily, the preset convergence condition may include but is not limited to one of the following: the training time exceeds the specified time, the number of model iterations exceeds the specified number, and the current loss obtained is lower than the specified loss threshold.

[0062] In some possible examples, step S310 may include the following steps 11-12:

[0063] In step 11, a specified number of video frames are extracted from the sample video. In this step, a specified frame extraction method can be adopted to extract a specified number of video frames from the sample video to obtain a video frame sequence sorted according to the training sequence in the sample video. The specified frame extraction method can be a random frame extraction method, or a frame extraction method based on a specified interval, etc.

[0064] In step 12, the initial coding parameter generation model is trained using the sample coding parameters corresponding to the specified number of video frames and the sample video. In this step, the specified number of video frames, i.e., the aforementioned video frame sequence, is input into the initial coding parameter generation model, so that the video frame sequence is processed by the initial coding parameter generation model to obtain the corresponding generated coding parameters; then, the difference between the generated coding parameters and the sample coding parameters is combined to adjust the parameters of the initial coding parameter generation model, thereby training the initial coding parameter generation model.

[0065] In the above method, by extracting frames from the sample video, the amount of data processed by the model is reduced, which can improve the model processing speed, processing efficiency, and time performance to a certain extent.

[0066] Afterwards, in step S320, the first encoding parameter generation model is adjusted according to a preset optimization method to obtain a second encoding parameter generation model, wherein the optimization direction of the preset optimization method includes: reducing the parameter amount and / or calculation amount of the first encoding parameter generation model.

[0067] In some possible examples, the preset optimization method includes: reducing the number of attention heads in the first encoding parameter generation model, reducing the network layers of the first encoding parameter generation model, and / or reducing the resolution of the input video of the first encoding parameter generation model. Exemplarily, the Swin Transformer blocks in certain feature processing layers can be reduced to reduce the network layers of the first encoding parameter generation model, and again exemplary, certain feature processing layers can also be reduced to reduce the network layers of the first encoding parameter generation model.

[0068] It can be understood that while reducing the number of attention heads in the first encoding parameter generation model, the number of attention heads in each Swin Transformer block can be kept the same.

[0069] In some possible examples, before a video, such as a sample video, a subsequent test video, and a first video, is input into the encoding parameter generation model, the video may be preprocessed to adjust the size, i.e., the resolution, of the video to a specified size, i.e., a specified resolution. The above-mentioned reduction of the resolution of the input video of the first encoding parameter generation model may be to reduce the specified size, i.e., the specified resolution, so as to reduce the amount of data processing of the encoding parameter generation model, thereby improving the temporal performance of the model.

[0070] In some other possible examples, after obtaining the first coding parameter generation model, before executing step S320, a verification video and its corresponding verification coding parameters may be obtained to verify the first coding parameter generation model. After the verification result meets expectations, step S320 may be executed again to better ensure that the accuracy of the generation result of the first coding parameter generation model obtained is high. The principle of obtaining the verification video is similar to the principle of obtaining the aforementioned sample video, and its acquisition method may refer to the aforementioned method of obtaining the sample video; the principle of determining the verification coding parameters corresponding to the verification video is similar to the principle of determining the sample coding parameters corresponding to the aforementioned sample video, and its determination method may refer to the aforementioned method of determining the sample coding parameters corresponding to the sample video, which will not be elaborated here. The verification process may refer to the verification process of the model in the relevant technology, which will not be elaborated here.

[0071] After obtaining the second encoding parameter generation model, in step S330, the second encoding parameter generation model is tested using several test videos and their corresponding test encoding parameters to obtain test results, wherein the test encoding parameters are related to the target encoder and are encoding parameters involved in the target encoder.

[0072] It can be understood that there is a corresponding relationship between the test video and the test encoding parameters. The principle of obtaining the test video is similar to the principle of obtaining the sample video mentioned above, and its acquisition method can refer to the aforementioned method of obtaining the sample video; the principle of determining the test encoding parameters corresponding to the test video is similar to the principle of determining the sample encoding parameters corresponding to the sample video mentioned above, and its determination method can refer to the aforementioned method of determining the sample encoding parameters corresponding to the sample video, which will not be repeated here.

[0073] In this step, for any test video X among several test videos, the test video X can be input into the second encoding parameter generation model, and the test video X can be processed by the second encoding parameter generation model to obtain output encoding parameters corresponding to the test video X; then, it is determined whether the output encoding parameters corresponding to the test video X are the same as the test encoding parameters corresponding to the test video X, and the judgment result corresponding to the test video X is obtained; through the aforementioned method, several judgment results corresponding to the several test videos are obtained; then, based on the several judgment results corresponding to the several test videos, the test result is obtained.

[0074] Exemplarily, the aforementioned process of obtaining the test result based on the several judgment results corresponding to the several test videos may be: determining the number of judgment results indicating that the output coding parameters corresponding to the test video are the same as the test coding parameters from the several judgment results as the first number, calculating the ratio between the first number and the total number of the several judgment results, and taking the ratio as the test result. The ratio may indicate the accuracy of the coding parameter generation result of the second coding parameter generation model. The larger the ratio, the higher the accuracy of the coding parameter generation result of the second coding parameter generation model.

[0075] Afterwards, in step S340, if the test result indicates that the accuracy of the generation result of the second encoding parameter generation model is within a specified allowable range, a target model is determined based on the second encoding parameter generation model.

[0076] In this step, if the test results indicate that the accuracy of the generation result of the second encoding parameter generation model is within the specified allowable range, it can be considered that the adjusted second encoding parameter generation model is reasonable and the accuracy of its generation result meets expectations. Accordingly, the target model is determined based on the second encoding parameter generation model.

[0077] In some possible examples, when the test result indicates that the accuracy of the generation result of the second encoding parameter generation model is within the specified allowable range, step S320 can be continued to be performed according to the preset optimization method, that is, in the optimization direction of reducing the parameter amount and / or calculation amount of the first encoding parameter generation model, to adjust the second encoding parameter generation model to obtain the adjusted second encoding parameter generation model, and then step S330 is performed for the adjusted second encoding parameter generation model to obtain the test result corresponding to the adjusted second encoding parameter generation model; if the test result corresponding to the adjusted second encoding parameter generation model indicates that the accuracy of the generation result of the adjusted second encoding parameter generation model is not within the specified allowable range, then the target model is determined based on the second encoding parameter generation model; if the test result corresponding to the adjusted second encoding parameter generation model indicates that the accuracy of the generation result of the adjusted second encoding parameter generation model is still within the specified allowable range, then steps S320-S330 can be continued to be performed for the adjusted second encoding parameter generation model until the test result corresponding to the repeatedly adjusted second encoding parameter generation model indicates that the accuracy of the generation result of the repeatedly adjusted second encoding parameter generation model is not within the specified allowable range, and the target model is determined based on the encoding parameter generation model obtained by the penultimate adjustment.

[0078] In some possible examples, the process of generating a model based on the second encoding parameter and determining a target model in step S340 may include steps 21-23:

[0079] In step 21, the format of the second encoding parameter generation model is converted into a specified format to obtain a third encoding parameter generation model. In this step, any model format conversion method in the relevant technology (for example, based on the conversion method provided by pytorch mentioned later) can be used to convert the format of the second encoding parameter generation model into a specified format to obtain a third encoding parameter generation model. In some examples, the format of the second encoding parameter generation model can be a related format of pytorch, and the specified format can be ONNX (Open Neural Network Exchange). Among them, pytorch is an open source machine learning library that can implement model definition and training. ONNX is a cross-platform model exchange format for sharing and using models between different deep learning frameworks. It allows users to train models in a deep learning framework and convert them to ONNX format, and then the model can be imported into another framework that supports ONNX for reasoning.

[0080] It can be understood that the format of the second encoding parameter generation model can also be other model formats that can be used for model training, and the specified format can also be other formats that can achieve cross-platform model exchange.

[0081] Then, in step 22, the third encoding parameter generation model is processed using a specified model optimization tool to reduce redundant parameters and calculations of the third encoding parameter generation model to obtain a fourth encoding parameter generation model. In some possible examples, when the aforementioned specified format is ONNX (Open Neural Network Exchange), the specified model optimization tool may be an ONNX SIN optimization tool. The ONNX SIN optimization tool is a tool specifically used to simplify and optimize models in the ONNX format. It eliminates redundant parameters and calculations in the model by performing a series of graph transformations and constant folding operations, thereby reducing the complexity and size of the model and improving the reasoning efficiency of the model.

[0082] In step 22, a designated model optimization tool may be used to load the third encoding parameter generation model and optimize the third encoding parameter generation model to reduce redundant parameters and calculations of the third encoding parameter generation model, thereby obtaining a fourth encoding parameter generation model.

[0083] Afterwards, in step 23, a target model is determined based on the fourth encoding parameter generation model. In this step, in one case, the fourth encoding parameter generation model can be directly determined as the target model. In another case, the precision of the parameters of the fourth encoding parameter generation model can be further lowered to obtain the target model. For example, the parameters of the fourth encoding parameter generation model are half-precisioned to obtain the target model; illustratively, assuming that the precision of the weight parameters of the fourth encoding parameter generation model is a 32-bit floating point value, the precision of the weight parameters of the fourth encoding parameter generation model can be lowered to a 16-bit floating point value to obtain the target model.

[0084] In the above process, based on the aforementioned preset optimization method, based on the processing of the specified model optimization tool and the application of half-precision reasoning, the first encoding parameter generation model that has been trained is optimized to obtain the target model. Under the premise of ensuring the accuracy of its generation process, the structural complexity of the target model obtained and the amount of calculation in the reasoning process are better reduced, thereby improving the processing efficiency of the subsequent video processing process of the target model, namely the reasoning process, improving the model's time performance, and reducing the comprehensive cost of video processing, including reducing the occupancy rate of processor resources during the reasoning process.

[0085] In some possible examples, the target model may include 4 feature processing layers, wherein the first feature processing layer, the second feature processing layer, and the fourth feature processing layer may respectively include 2 Swin Transformer blocks (i.e., a pair of Swin Transformer blocks), and the third feature processing layer may include 6 Swin Transformer blocks (i.e., three pairs of Swin Transformer blocks arranged in series).

[0086] After obtaining the target model, the target model can be used in conjunction with the target encoder to encode the video. Specifically, Figure 5 A flow chart of a video processing method in one embodiment of the present specification is shown. The method is executed by an electronic device, and the electronic device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities.

[0087] In the video processing, Figure 5 As shown, the method includes the following steps S510-S530:

[0088] In step S510, a first video is obtained. In this step, a video of any resolution and any video scene can be obtained from any video source as the first video. In some possible examples, the first video can be any video in a dataset used to train a multimodal large model.

[0089] Next, in step S520, based on the first video, a first encoding parameter is determined through a target model, wherein the target model is trained based on a sample video and its corresponding sample encoding parameter, the sample encoding parameter is related to a target encoder, and the target model is a model based on a Transformer structure. In this step, the first video can be processed based on the aforementioned specified preprocessing to obtain a preprocessed first video, and then the preprocessed first video can be input into the target model, and the preprocessed first video can be processed through the target model to obtain a first encoding parameter that is more suitable for the first video. As another example, a specified number of first video frames can be extracted from the first video based on the aforementioned specified frame extraction method to form a first video frame sequence, and then the first encoding parameter can be determined based on the first video frame sequence through the target model.

[0090] Then, in step S530, the first video is encoded using the target encoder configured with the first encoding parameters to obtain a first encoded video. In this step, the target encoder is first configured based on the first encoding parameters, and then the first video is encoded using the target encoder configured with the first encoding parameters to obtain a first encoded video with good quality and low bit rate.

[0091] In the above process, based on the first video, a first encoding parameter that is more suitable for the first video is determined through a trained target model based on a Transformer structure, and then the first video is encoded using a target encoder configured with the first encoding parameter to obtain a first encoded video, thereby encoding the first video to reduce the video bit rate, thereby reducing the bandwidth cost of the encoded video during transmission, such as uploading and downloading. In some examples, a basis is provided for reducing the bandwidth cost in the video preprocessing link during the training of a large multimodal model.

[0092] In some possible examples, Figure 5 Based on the process shown, Figure 6 As shown, the video processing flow may further include the following steps S540-S550:

[0093] In step S540, based on the specified video evaluation index, it is determined whether the quality of the first encoded video meets the standard. Exemplarily, the specified video evaluation index may include but is not limited to at least one of the following indicators: SSIM (Structural Similarity Index), PSNR (Peak Signal-to-Noise Ratio), and VMAF (Video Multimethod Assessment Fusion). The specified video evaluation index may be the same as or different from the aforementioned specific video evaluation index.

[0094] Exemplarily, based on the specified video evaluation index, the index value corresponding to the first encoded video can be determined, wherein the index value corresponding to the first encoded video can indicate the quality of the first encoded video; then the index value corresponding to the first encoded video is compared with the specified index threshold; if the index value corresponding to the first encoded video is not less than the specified threshold, it can be considered that the quality of the first encoded video meets the standard; conversely, if the index value corresponding to the first encoded video is less than the specified threshold, it can be considered that the quality of the first encoded video does not meet the standard.

[0095] In some possible examples, when the aforementioned video evaluation index includes the structural similarity index SSIM, in order to better improve the time performance of the model in processing the video and improve the efficiency of the target model in processing the video; step S540 may include the following steps:

[0096] Based on the structural similarity index SSIM, an adaptive frame skipping method is used to determine whether the quality of the first encoded video meets the standard.

[0097] In the above implementation, based on SSIM, an adaptive frame skipping method is adopted to combine the first encoded video and several target video frame pairs that have a corresponding relationship between the first video to determine the index value corresponding to the first encoded video, and then based on the index value corresponding to the first encoded video and the aforementioned specified threshold, it is judged whether the quality of the first encoded video meets the standard. Exemplarily, each pair of target video frame pairs may include a video frame in a frame of the encoded video and a video frame in a frame of the first video that have a corresponding relationship, and the several target video frame pairs are determined from the first encoded video and the first video respectively based on the adaptive frame skipping method. In this implementation, the time performance of the model in the process of processing videos can be better improved by the adaptive frame skipping method, and the efficiency of the target model in processing videos can be improved.

[0098] Then, in step S550, if the judgment result is that the quality of the first encoded video meets the standard, the first encoded video is stored. In this step, if the judgment result is that the quality of the first encoded video meets the standard, the first encoded video can be stored in a designated storage space for use in subsequent task processes.

[0099] In some other possible examples, such as Figure 6 As shown, the video processing flow may further include the following steps S560-S580:

[0100] In step S560, if the judgment result is that the quality of the first encoded video does not meet the standard, the first encoding parameter is adjusted to obtain the second encoding parameter with the goal of improving the quality of the encoded video. In this step, if the judgment result is that the quality of the first encoded video does not meet the standard, the first encoding parameter can be adjusted to obtain the second encoding parameter that can make the quality of the corresponding encoded video meet the standard on the basis of the first encoding parameter. Exemplarily, when the first encoding parameter includes the bit rate, the bit rate in the first encoding parameter can be appropriately increased to obtain the second encoding parameter; and again exemplarily, when the first encoding parameter includes other types of encoding parameters, such as key frame intervals and / or filter and tuning related parameters, the value of the corresponding type of encoding parameter in the first encoding parameter can be adjusted based on experience to obtain the second encoding parameter, which makes the quality of the corresponding encoded video meet the standard. Among them, the aforementioned making the quality of the corresponding encoded video meet the standard can refer to the indicator value corresponding to the corresponding encoded video being not less than the aforementioned specified threshold.

[0101] In some possible examples, the aforementioned process of adjusting the first encoding parameter may include: adjusting the first encoding parameter in response to a user operation, wherein the user operation may carry a specific encoding parameter to be adjusted and a corresponding adjustment target, and the adjustment target may include an adjustment value corresponding to the specific encoding parameter, or include an adjustment direction and an adjustment step corresponding to the specific encoding parameter.

[0102] After obtaining the second encoding parameters, in step S570, the first video is encoded using the target encoder configured with the second encoding parameters to obtain a second encoded video. And in step S580, the second encoded video is stored.

[0103] In the above steps, the electronic device configures the target encoder using the second encoding parameters, and then uses the target encoder configured with the second encoding parameters to encode the first video to obtain the second encoded video. Afterwards, the second encoded video is stored in the designated storage space for use in subsequent task processes. In the above example, the first encoding parameter corresponding to the first encoded video whose objective index value is less than the threshold is re-encoded by adaptive encoding parameters through re-encoding, so as to obtain an encoded video that meets the conditions, i.e., the quality meets the standard, based on the newly obtained new encoding parameters.

[0104] In some possible examples, in order to better ensure that the quality of the second encoded video obtained meets the standards, after obtaining the second encoded video, the indicator value corresponding to the second encoded video can be determined based on the aforementioned specified video evaluation indicators; if it is determined that the indicator value corresponding to the second encoded video is not less than the aforementioned specified threshold, that is, it is determined that the quality of the second encoded video meets the standards, then the second encoded video can be stored in the specified storage space; if it is determined that the indicator value corresponding to the second encoded video is less than the aforementioned specified threshold, it is determined that the quality of the second encoded video does not meet the standards, and the second encoding parameters are adjusted with the goal of improving the quality of the encoded video to obtain the adjusted second encoding parameters, and then the first video is encoded using the target encoder configured with the adjusted second encoding parameters to obtain the third encoded video; and so on, until an encoded video that meets the quality standards is obtained, and then the corresponding encoded video that meets the quality standards is stored in the specified storage space for use in subsequent task processes.

[0105] It can be understood that, in some possible examples, the aforementioned first video can be any video in the data set used to train the multimodal large model. Figure 7 As shown, the data set used to train the multimodal large model is stored in a database based on OSS (Object Storage Service), and the first video can be any video obtained from the database. It can be understood that in some exemplary scenarios, before training the multimodal large model, various video preprocessing is required for each video in the data set. The video preprocessing process involves multiple video downloads and uploads. If the original video, i.e., the video before encoding, is downloaded and uploaded, a large bandwidth will be occupied.

[0106] In order to reduce the bandwidth occupancy of the video preprocessing process during the training of the multimodal large model, the video in the above-mentioned database can be processed in sequence based on the video processing flow provided in the embodiments of this specification to obtain the encoded video corresponding to each video, and then the encoded video corresponding to each video can be stored back in the database.

[0107] Subsequently, in the aforementioned video preprocessing process, each encoded video is downloaded from the database, and the corresponding operators of various video preprocessing methods (such as Figure 7 Operator 1, operator 2 ... operator N) are shown to perform corresponding video preprocessing on each encoded video, and then use each encoded video after video preprocessing to train a multimodal large model to better reduce bandwidth costs and improve the training efficiency of the multimodal large model. The aforementioned video preprocessing may include but is not limited to: watermark removal, storyboard processing, video classification processing, and video cropping processing.

[0108] In the video processing flow provided in the embodiments of this specification, a more suitable encoding parameter can be determined for each video through a trained target model. After that, the video can be encoded based on the corresponding encoding parameters determined by the target model for the corresponding video to obtain a suitable encoded video and store it in a database. This type of encoded video can ensure the model performance of the subsequent multimodal large model and can better reduce the bandwidth occupied by the above-mentioned video preprocessing process. In some cases, the bit rate of the encoded video obtained by the video processing flow provided in the embodiments of this specification can be reduced by about 77% compared to the original video, i.e., the video before encoding, and in the above-mentioned video preprocessing process, the encoded video is processed and then the multimodal large model is trained, and the comprehensive cost can be reduced by about 51%.

[0109] As another example, the aforementioned first video may also be other videos that need to reduce the bit rate, so as to reduce the bandwidth consumption during video transmission and / or processing. The first video is processed through the above process to obtain encoding parameters related to the target encoder that are more suitable for the first video, and the first video is encoded based on the encoding parameters to obtain an encoded video with a lower bit rate and suitable quality, so as to reduce the bandwidth consumption of the video during transmission and processing, and reduce the bandwidth cost.

[0110] The foregoing describes certain embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in an order different from that in the embodiments, and the desired results may still be achieved. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0111] Corresponding to the above method embodiment, the present specification embodiment provides a video processing device 800, whose schematic block diagram is as follows: Figure 8 As shown, including:

[0112] A first acquisition module 810 is configured to acquire a first video;

[0113] A first determination module 820 is configured to determine a first encoding parameter based on the first video through a target model, wherein the target model is trained based on a sample video and its corresponding sample encoding parameters, the sample encoding parameters are related to a target encoder, and the target model is a model based on a Transformer structure;

[0114] The first encoding module 830 is configured to encode the first video using the target encoder configured with the first encoding parameters to obtain a first encoded video.

[0115] Some possible examples include:

[0116] A second acquisition module (not shown in the figure), configured to acquire a sample video;

[0117] A third acquisition module (not shown in the figure), configured to acquire multiple groups of candidate encoding parameters for the target encoder;

[0118] A second encoding module (not shown in the figure) is configured to encode the sample video using a target encoder configured with each group of candidate encoding parameters to obtain a sample encoded video corresponding to each group of candidate encoding parameters;

[0119] The second determination module (not shown in the figure) is configured to determine the sample encoded video whose quality meets the preset quality conditions and has the lowest bit rate from the sample encoded videos corresponding to each group of alternative encoding parameters, and determine the corresponding alternative encoding parameters as the sample encoding parameters corresponding to the sample video.

[0120] Some possible examples include:

[0121] A first training module (not shown in the figure) is configured to use the sample video and its corresponding sample encoding parameters to train an initial encoding parameter generation model to obtain a first encoding parameter generation model that meets a preset convergence condition;

[0122] A first adjustment module (not shown in the figure) is configured to adjust the first encoding parameter generation model according to a preset optimization method to obtain a second encoding parameter generation model, wherein the optimization direction of the preset optimization method includes: reducing the parameter amount and / or calculation amount of the first encoding parameter generation model;

[0123] A first test module (not shown in the figure) is configured to test the second encoding parameter generation model using a plurality of test videos and their corresponding test encoding parameters to obtain a test result, wherein the test encoding parameters are related to the target encoder;

[0124] The third determination module (not shown in the figure) is configured to determine the target model based on the second encoding parameter generation model if the test result indicates that the accuracy of the generation result of the second encoding parameter generation model is within a specified allowable range.

[0125] In some possible examples, the preset optimization method includes: reducing the number of attention heads in the first encoding parameter generation model, reducing the network layers of the first encoding parameter generation model, and / or reducing the resolution of the input video of the first encoding parameter generation model.

[0126] In some possible examples, the third determination module is specifically configured to convert the format of the second encoding parameter generation model into a specified format to obtain a third encoding parameter generation model;

[0127] Using a designated model optimization tool, processing the third encoding parameter generation model to reduce redundant parameters and calculations of the third encoding parameter generation model, thereby obtaining a fourth encoding parameter generation model;

[0128] A model is generated based on the fourth encoding parameter, and the target model is determined.

[0129] In some possible examples, the first training module is specifically configured to extract a specified number of video frames from the sample video;

[0130] An initial encoding parameter generation model is trained using the specified number of video frames and the sample encoding parameters corresponding to the sample video.

[0131] Some possible examples include:

[0132] A judgment module (not shown in the figure), configured to judge whether the quality of the first encoded video meets the standard based on a specified video evaluation index;

[0133] The first storage module (not shown in the figure) is configured to store the first encoded video if the judgment result is up to standard.

[0134] Some possible examples include:

[0135] A second adjustment module (not shown in the figure) is configured to adjust the first encoding parameter to obtain a second encoding parameter with the goal of improving the quality of the encoded video if the judgment result is that the quality of the first encoded video does not meet the standard;

[0136] A third encoding module (not shown in the figure), configured to encode the first video using the target encoder configured with the second encoding parameter to obtain a second encoded video;

[0137] The second storage module (not shown in the figure) is configured to store the second encoded video.

[0138] In some possible examples, the video evaluation index includes a structural similarity index SSIM;

[0139] The judgment module is specifically configured to judge whether the quality of the first encoded video meets the standard by using an adaptive frame skipping method based on the structural similarity index SSIM.

[0140] The above device embodiments correspond to the method embodiments. For specific descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, please refer to the corresponding method embodiments.

[0141] The embodiments of the present specification also provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the video processing method provided in the present specification.

[0142] An embodiment of the present specification further provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the video processing method provided in the present specification is implemented.

[0143] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the storage medium and computing device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0144] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of the present invention may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0145] The specific implementation methods described above further describe the purpose, technical solutions and beneficial effects of the embodiments of the present invention in detail. It should be understood that the above description is only a specific implementation method of the embodiments of the present invention and is not intended to limit the scope of protection of the present invention. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solution of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A video processing method, comprising: Get the first video; Based on the first video, determining a first encoding parameter through a target model, wherein the target model is trained based on a sample video and its corresponding sample encoding parameters, the sample encoding parameters are related to a target encoder, and the target model is a model based on a Transformer structure; The first video is encoded using the target encoder configured with the first encoding parameter to obtain a first encoded video.

2. The method of claim 1, further comprising: Get sample video; Acquire multiple sets of candidate encoding parameters for the target encoder; Encoding the sample videos using target encoders configured with each set of candidate encoding parameters to obtain sample encoded videos corresponding to each set of candidate encoding parameters; From the sample coded videos corresponding to each group of candidate coding parameters, a sample coded video whose quality meets the preset quality condition and has the lowest bit rate is determined, and its corresponding candidate coding parameters are determined as the sample coding parameters corresponding to the sample video.

3. The method of claim 1, further comprising: Using the sample video and its corresponding sample encoding parameters, an initial encoding parameter generation model is trained to obtain a first encoding parameter generation model that meets a preset convergence condition; According to a preset optimization method, the first encoding parameter generation model is adjusted to obtain a second encoding parameter generation model, wherein the optimization direction of the preset optimization method includes: reducing the parameter amount and / or calculation amount of the first encoding parameter generation model; Using a plurality of test videos and their corresponding test encoding parameters, testing the second encoding parameter generation model to obtain a test result, wherein the test encoding parameters are related to the target encoder; If the test result indicates that the accuracy of the generation result of the second encoding parameter generation model is within a specified allowable range, the target model is determined based on the second encoding parameter generation model.

4. The method of claim 3, wherein: The preset optimization method includes: reducing the number of attention heads in the first encoding parameter generation model, reducing the network layers of the first encoding parameter generation model, and / or reducing the resolution of the input video of the first encoding parameter generation model.

5. The method of claim 3, wherein: The generating a model based on the second encoding parameter and determining the target model comprises: Converting the format of the second encoding parameter generation model into a specified format to obtain a third encoding parameter generation model; Using a designated model optimization tool, processing the third encoding parameter generation model to reduce redundant parameters and calculations of the third encoding parameter generation model, thereby obtaining a fourth encoding parameter generation model; A model is generated based on the fourth encoding parameter, and the target model is determined.

6. The method of claim 3, wherein: The using the sample video and its corresponding sample encoding parameters to train an initial encoding parameter generation model includes: Extracting a specified number of video frames from the sample video; An initial encoding parameter generation model is trained using the specified number of video frames and the sample encoding parameters corresponding to the sample video.

7. The method according to any one of claims 1 to 6, further comprising: Based on a specified video evaluation index, determining whether the quality of the first encoded video meets the standard; If the judgment result is that the standard is met, the first encoded video is stored.

8. The method of claim 7, further comprising: If the judgment result is that the quality of the first encoded video does not meet the standard, adjusting the first encoding parameter to obtain a second encoding parameter with the goal of improving the quality of the encoded video; Encoding the first video using the target encoder configured with the second encoding parameter to obtain a second encoded video; The second encoded video is stored.

9. The method of claim 7, wherein: The video evaluation index includes a structural similarity index SSIM; The determining, based on a preset video quality indicator, whether the quality of the first encoded video meets the quality standard includes: Based on the structural similarity index SSIM, an adaptive frame skipping method is used to determine whether the quality of the first encoded video meets the standard.

10. A video processing device, comprising: A first acquisition module, configured to acquire a first video; The first determination module is configured to determine a first encoding parameter based on the first video through a target model, wherein: The target model is obtained by training based on a sample video and its corresponding sample encoding parameters, the sample encoding parameters are related to the target encoder, and the target model is a model based on a Transformer structure; The first encoding module is configured to encode the first video using the target encoder configured with the first encoding parameters to obtain a first encoded video.

11. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 9 is implemented.