Video processing method and device, electronic equipment and storage medium

By segmenting the video and encoding the target video frames, combined with the content description text, the problem of video file size reduction in existing technologies is solved, achieving low bitrate encoding and efficient storage.

CN121056684APending Publication Date: 2025-12-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410683642.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to effectively reduce video file size while maintaining video quality, resulting in high bandwidth consumption and storage costs.

Method used

By segmenting the original video, target video frames whose image quality information meets the preset screening conditions are identified and encoded. Combined with the content description text, the video frames and text are encoded using different compression rates to generate low bitrate video encoded files.

Benefits of technology

It achieves low bitrate video encoding, reducing bandwidth usage and storage costs, while preserving the details of video segments, facilitating high-quality decoding and restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121056684A_ABST
    Figure CN121056684A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video processing method and device, electronic equipment and a storage medium. The method comprises the following steps: segmenting an original video to obtain at least two video segments; determining a content description text of each video segment according to the content information of each video segment; obtaining a target video frame in each video segment; the original video, the content description text of each video segment and the target video frame are coded to obtain a coded video, a coded video frame and a coded text, and the code rate of the coded video is smaller than that of the coded video frame; the video coding file is determined according to the coding video, the coding video frame and the coding text, so that the bandwidth occupation is reduced, and the storage cost is reduced. Besides, the coded video frame has a higher code rate compared with the coded video, and the content description text can represent the content of the video segment, so that the high-quality target video can be conveniently restored based on the decoded video frame corresponding to the target video frame and the content description text in the decoding stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to data processing technology, and more particularly to a video processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Video compression can reduce the size of video files as much as possible while maintaining video quality, thus facilitating video transmission, storage, and playback.

[0003] Currently, video codecs are typically based on encoding and decoding standard algorithms, such as the h.26x series or MPEG (Moving Pictures Experts Group) series, to encode raw video to achieve video compression. Video compression can reduce bandwidth consumption and lower storage costs in scenarios such as transmitting video over a network, watching video on mobile devices, or storing video in the cloud. Therefore, how to generate compressed video with a lower bitrate has become a pressing issue. Summary of the Invention

[0004] This disclosure provides a video processing method, apparatus, electronic device, and storage medium that can reduce video bitrate.

[0005] In a first aspect, embodiments of this disclosure provide a video processing method, including:

[0006] The original video is segmented to obtain at least two video segments;

[0007] Determine the content description text for each video segment based on the content information of each video segment;

[0008] Obtain target video frames from each video segment, wherein the target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions;

[0009] The original video, the content description text of each video segment, and the target video frame are encoded to obtain an encoded video, an encoded video frame, and encoded text. The bitrate of the encoded video is less than the bitrate of the encoded video frame.

[0010] The video encoding file is determined based on the encoded video, encoded video frames, and encoded text.

[0011] Secondly, embodiments of this disclosure also provide a video processing method, including:

[0012] Obtain a video encoding file, wherein the video encoding file represents a compressed file of the original video, including encoded video, encoded video frames and encoded text, the bitrate of the encoded video is less than the bitrate of the encoded video frames, the encoded video frames represent the encoding information of target video frames in each video segment, the target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions, and the encoded text represents the encoding information of the content description text of each video segment.

[0013] The video encoded file is decoded to obtain decoded video, decoded video frames, and content description text;

[0014] The decoded video is repaired based on the decoded video frames and content description text to obtain the target video corresponding to the video encoding file.

[0015] Thirdly, embodiments of this disclosure also provide a video processing apparatus, the apparatus comprising:

[0016] The segmentation module is used to segment the original video to obtain at least two video segments;

[0017] The text determination module is used to determine the content description text of each video segment based on the content information of each video segment.

[0018] The acquisition module is used to acquire target video frames in each video segment, wherein the target video frame represents a video frame in each video segment whose image quality information meets preset filtering conditions.

[0019] The encoding module is used to encode the original video, the content description text of each video segment, and the target video frame to obtain encoded video, encoded video frame, and encoded text, wherein the bit rate of the encoded video is less than the bit rate of the encoded video frame.

[0020] The file determination module is used to determine the video encoding file based on the encoded video, encoded video frames, and encoded text.

[0021] Fourthly, embodiments of this disclosure also provide a video processing apparatus, the apparatus comprising:

[0022] The acquisition module is used to acquire video encoded files, wherein the video encoded file represents a compressed file of the original video, including encoded video, encoded video frames and encoded text, the bitrate of the encoded video is less than the bitrate of the encoded video frames, the encoded video frames represent the encoding information of target video frames in each video segment, the target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions, and the encoded text represents the encoding information of the content description text of each video segment;

[0023] The decoding module is used to decode the video encoded file to obtain decoded video, decoded video frames, and content description text;

[0024] The repair module is used to repair the decoded video based on the decoded video frames and content description text to obtain the target video corresponding to the video encoding file.

[0025] Fifthly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0026] One or more processors;

[0027] Storage device for storing one or more programs.

[0028] When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method as described in any embodiment of this disclosure.

[0029] Sixthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video processing method as described in any embodiment of this disclosure.

[0030] This disclosure provides a video processing method that involves segmenting an original video to obtain multiple video segments, determining content description text for each video segment based on its content information, and acquiring target video frames from each video segment. The target video frames of each video segment are then encoded to obtain encoded video frames, and the original video is encoded to obtain an encoded video. The bitrate of the encoded video is lower than that of the encoded video frames. Efficient encoding of the original video results in a lower bitrate, reducing bandwidth usage and storage costs. Furthermore, since the encoded video frames have a higher bitrate than the encoded video, they can better preserve the detailed content of the target video frames in each video segment, and the content description text can characterize the content of the video segment, facilitating the reconstruction of a high-quality target video during the decoding stage based on the decoded video frames corresponding to the target video frames and the content description text. Attached Figure Description

[0031] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0032] Figure 1 This is a schematic flowchart of a video processing method provided in an embodiment of the present disclosure;

[0033] Figure 2This is a schematic diagram of a video encoding method provided in an embodiment of the present disclosure;

[0034] Figure 3 This is a schematic flowchart of another video processing method provided in an embodiment of the present disclosure;

[0035] Figure 4 This is a schematic diagram of a video decoding method provided in an embodiment of the present disclosure;

[0036] Figure 5 This is a schematic diagram of an image quality restoration method provided in an embodiment of the present disclosure;

[0037] Figure 6 This is a schematic diagram of a video frame compression method provided in an embodiment of the present disclosure;

[0038] Figure 7 This is a schematic diagram of the structure of a video processing device provided in an embodiment of the present disclosure;

[0039] Figure 8 This is a schematic diagram of another video processing apparatus structure provided in an embodiment of the present disclosure;

[0040] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0041] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0042] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0043] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0044] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0045] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0046] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0047] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0048] Figure 1 This is a flowchart illustrating a video processing method provided in an embodiment of this disclosure. This embodiment is applicable to video encoding and decoding scenarios, such as encoding a video before sending it to the cloud via a client, or encoding a video before transmitting it over a network. This method can be executed by a video processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server.

[0049] like Figure 1 As shown, the method includes:

[0050] S110. Segment the original video to obtain at least two video segments.

[0051] The original video can represent the video to be encoded. For example, the original video may include video to be transmitted over a network. The original video may be video obtained through a network. Alternatively, the original video may include video on a mobile device to be uploaded to the cloud. The original video may be video stored locally on a mobile device. Alternatively, the original video may also include video on a mobile device to be played, etc. A video segment represents a fragment consisting of at least three video frames from the original video. A video segment includes the target video frame.

[0052] For example, transition video frames are obtained from the original video, wherein the transition video frames represent video frames used to connect different video segments; the original video is segmented according to the transition video frames to obtain at least two video segments.

[0053] In this embodiment, a transition video frame represents a video frame that connects two video segments with a transition effect, making the video smoother and more natural. Therefore, transition video frames in the original video can be identified, and the video segments before and after the transition video frame can be divided into different video segments. For example, using the transition video frame as a dividing node, the video frames between the current transition video frame and the previous transition video frame, along with the current transition video frame itself, can be considered as a single video segment.

[0054] Optionally, for any two adjacent video frames in the original video, the pixel difference between the two adjacent video frames is determined; the original video is segmented based on each adjacent video frame whose pixel difference exceeds a preset segmentation threshold to obtain at least two video segments.

[0055] The preset segmentation threshold is used to characterize the degree of difference between adjacent video frames, and the preset segmentation threshold can be set according to the actual situation.

[0056] In this embodiment of the disclosure, video frames in the original video are traversed, the pixel difference between adjacent video frames is calculated, and the pixel difference is compared with a preset segmentation threshold to obtain video frames whose pixel difference exceeds the preset segmentation threshold. Adjacent video frames whose pixel difference exceeds the preset segmentation threshold are divided into different video segments to obtain at least two video segments.

[0057] Optionally, if no transition video frame appears within the set time interval, and no adjacent video frames have a pixel difference exceeding a preset segmentation threshold, the original video is divided according to the set segmentation time. For example, if the original video does not have a transition video frame or adjacent video frames with a pixel difference exceeding the preset segmentation threshold within x consecutive seconds, the continuous x-second video sequence is segmented according to the set segmentation time.

[0058] Optionally, the video segments can be sorted according to their corresponding timestamps to obtain a video segment sequence. Optionally, if the order of the video segments is not considered, a video segment set can also be constructed based on the video segments.

[0059] S120. Determine the content description text for each video segment based on the content information of each video segment.

[0060] The content description text represents the content of the video segment. A pre-defined content understanding model interprets the video segment into a text description, i.e., the content description text. For example, a video segment can be input into the pre-defined content understanding model, which performs frame extraction on the video segment, identifies the content in the extracted video frames, and summarizes the identified content to output a text description.

[0061] For example, for each of at least two video segments, a set number of video frames are obtained from the video segment;

[0062] Based on the content information of the set number of video frames, determine the content description text of the video segment.

[0063] It should be noted that the number of video frames to be extracted can be determined based on the number of video frames contained in the video segment. After obtaining the set number of video frames from the video segment, image recognition is performed on the video frames to determine their content information. Combined with the content information corresponding to the set number of video frames, the content description text of the video segment is determined. For example, a video segment is input into a preset content understanding model. The model performs frame extraction on the input video segment, identifies the content information of the extracted video frames, and summarizes the identified content information to output the content description text of the input video segment.

[0064] S130. Obtain target video frames from each video segment, wherein the target video frame represents a video frame in each video segment whose image quality information meets preset filtering conditions.

[0065] The preset filtering conditions can include conditions related to the image quality of video frames. For example, preset filtering conditions may include sharpness conditions, resolution conditions, sharpness conditions, or color gamut conditions.

[0066] In some embodiments, the preset filtering conditions may include conditions related to the image quality of the video frames and conditions related to the frame interval time. By setting the frame interval time-related conditions, suitable target video frames are selected to balance the stability of the repaired content and the compression rate of the video. Where the time interval is too long, it will affect the stability of the repaired video frame content; where the time interval is too short, it will increase the data volume of the compressed video file, thereby affecting the compression rate of the original video.

[0067] For example, for each of at least two video segments, at least two video frames in the video segment whose time interval and image quality information meet preset filtering conditions are obtained as the target video frames of the video segment.

[0068] For example, the target video frame includes the first frame and the last frame. This disclosure does not specifically limit the meaning of the preset filtering conditions. For each video segment in a video segment sequence or video segment set, the first and last frames of the current video segment are obtained as the target video frame for that current video segment.

[0069] S140. Encode the original video, the content description text of each video segment, and the target video frame to obtain an encoded video, encoded video frames, and encoded text, wherein the bitrate of the encoded video is less than the bitrate of the encoded video frames.

[0070] Bitrate represents the amount of video data per unit of time. It determines the degree of compression the encoder applies to the original video, and also determines the image quality and file size of the encoded video file. The lower the bitrate, the higher the compression level, the lower the image quality, and the smaller the storage space occupied by the file.

[0071] The encoded video represents the encoding result of the original video. A first encoder can be used to encode the original video. The first encoder is used to encode the original video using a first compression rate, outputting the encoded video as a video bitstream. The encoded video frame represents the encoding result of the target video frame for each video segment in the original video. The first encoder can be used to encode the target video frame. If the target video frame includes the first and last frames of each video segment, the first encoder is used to encode the first and last frames using a second compression rate, outputting the encoded first and last frames as a video bitstream. The encoded text represents the encoding result of the content description text. A second encoder can be used to encode the content description text. The second encoder is used to encode the text, outputting the encoded text as a text bitstream.

[0072] For example, the original video is encoded based on a first compression ratio to obtain an encoded video; the first and last frames of each video segment are encoded based on a second compression ratio to obtain encoded first and last frames of each video segment, wherein the first compression ratio is higher than the second compression ratio, and the encoded video frame includes the encoded first and last frames of each video segment. Since the first encoder encodes the first and last frames of each video segment based on a lower compression ratio, more detailed information can be preserved. Using the first encoder to encode the original video based on a higher compression ratio allows for the preservation of the original video content with a lower bitrate. Although the first and last frames of each video segment use a higher bitrate to preserve detailed information, their proportion in the entire original video is small, and the high bitrate of the first and last frames does not affect the bitrate of the original video, thus achieving efficient compression of the original video.

[0073] Assume the encoder includes a first encoder and a second encoder, etc. The first encoder may include a video compression module and an image compression module, etc. The original video segments are input into the encoder. The image compression module compresses the first and last frames of each video segment based on encoding parameter 'a', resulting in encoded video frames with a quality level of 'x'. The video compression module compresses the original video segments based on encoding parameter 'b', resulting in encoded video with a quality level of 'y'. It should be noted that quality level 'y' is lower than quality level 'x'. Encoding parameters include compression ratio, etc., and the compression ratios corresponding to the first and last frames are lower than the compression ratio corresponding to the original video, to achieve efficient encoding of the original video and to retain more detail in the first and last frames of each video segment.

[0074] S150. Determine the video encoding file based on the encoded video, encoded video frames, and encoded text.

[0075] Among them, a video encoded file can represent a compressed file of the original video. For example, a video encoded file can be presented as a video compressed package, which is transmitted during video transmission.

[0076] For example, segmentation information of the original video is obtained, wherein the segmentation information represents the mapping relationship between video segments and video frames. Based on the encoded video, the segmentation information, and the encoded first frame, encoded last frame, and encoded text of each video segment, the corresponding video encoding file of the original video is determined. Appending the segmentation information to the video encoding file allows the segmentation method of the original video to be transmitted to the decoding device, enabling the decoding device to determine each video segment in the decoded video based on the segmentation information. Then, the image quality of each video segment in the decoded video can be optimized based on the decoded first frame and decoded last frame of the current video segment. Specifically, the pixel distribution of each video frame in the current video segment of the decoded video can be adjusted based on the pixel distribution of the decoded first frame and decoded last frame of the current video segment, combined with the content description text.

[0077] Optionally, the bitstreams corresponding to the encoded text and encoded video frames can be grouped based on the video segment to which the encoded text and encoded video frames belong. Then, a video encoded file is constructed based on the bitstream of the encoded video corresponding to the original video, the segmentation information, and the bitstreams corresponding to the encoded text and encoded video frames in each video segment.

[0078] Figure 2 This is a schematic diagram illustrating a video encoding method provided in an embodiment of this disclosure. Figure 2As shown, the original video 200 includes n video frames. The original video is segmented to obtain multiple video segments 210 without disrupting the continuity of the video segments, facilitating a thorough understanding of their content and thus summarizing the content description text of each video segment. Segmentation information is determined based on the mapping relationship between video segments 210 and the video frames of the original video 200. For example, segmentation information may include the first video segment comprising the first to fifth frames of the original video, the second video segment comprising the sixth to eighth frames of the original video, and so on. Video segment 210 may include at least three video frames. The content understanding module 220 performs frame extraction processing on each video segment 210 and summarizes the content description text of each video segment 210 based on the content of the extracted video frames. Then, the content description text is input into the second encoder 230, which outputs encoded text 240. The dual-frame extraction module 250 extracts the first and last frames 260 corresponding to each video segment 210. The first and last frames of each video segment 210 are input to the first encoder 270. The first encoder 270 is used to encode the first and last frames 260 using a low compression rate, outputting encoded first and encoded last frames 280, and to encode the original video using a high compression rate, outputting encoded video 290. A video encoded file is constructed based on the encoded video 290, segmentation information, encoded text 240 corresponding to each video segment 210, and the encoded first and encoded last frames 280.

[0079] The technical solution of this disclosure involves segmenting the original video to obtain multiple video segments, then determining the content description text for each video segment based on its content information, and obtaining the target video frames in each video segment. The target video frames of each video segment are then encoded to obtain encoded video frames, and the original video is further encoded to obtain an encoded video. The bitrate of the encoded video is lower than that of the encoded video frames. Efficient encoding of the original video results in a lower bitrate, reducing bandwidth usage and storage costs. Furthermore, since the encoded video frames have a higher bitrate than the encoded video, they can better preserve the detailed content of the target video frames in each video segment, and the content description text can characterize the content of the video segment, facilitating the reconstruction of a high-quality target video based on the decoded video frames corresponding to the target video frames and the content description text during the decoding stage.

[0080] Figure 3This is a flowchart illustrating another video processing method provided in an embodiment of this disclosure. This disclosure applies to video encoding and decoding scenarios, such as decoding a video encoded file sent by a client after receiving it from the cloud, or decoding a video before viewing it on a mobile device. This method can be executed by a video processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server.

[0081] like Figure 3 As shown, the method includes:

[0082] S310, Obtain the video encoding file.

[0083] The video encoding file represents a compressed file of the original video, including encoded video, encoded video frames, and encoded text. The bitrate of the encoded video is lower than that of the encoded video frames. The encoded video frames represent the encoding information of target video frames in each video segment. The target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions. The encoded text represents the encoding information of the content description text of each video segment.

[0084] For example, a video encoded file includes encoded video, segmentation information, encoded text corresponding to each video segment, encoded first frame, and encoded last frame. The bitrate of the encoded video is lower than that of the encoded first and last frames, resulting in a lower overall bitrate for the encoded video file, thus achieving the effect of occupying less bandwidth and storage space. Because the bitrate of the encoded first and last frames is higher, more content details of the first and last frames are preserved after compression. Therefore, during the decoding stage, the encoded first and last frames of each video segment can be decoded to obtain the decoded first and last frames of each video segment. These decoded first and last frames are then combined to reconstruct the high-quality target video.

[0085] For example, a video encoded file transmitted via a communication network is obtained. The video encoded file can be obtained by encoding the original video by other electronic devices, and the video encoded file can be transmitted to the current electronic device through media such as communication networks, where the current electronic device performs decoding and restoration of the video encoded file.

[0086] S320. Decode the video encoding file to obtain decoded video, decoded video frames, and content description text.

[0087] The quality of the decoded video is inferior to that of the decoded video frame. The decoded video frame includes the first decoded frame and the last decoded frame of each video segment. The first decoded frame represents the decoding result of the encoded information of the first frame of each video segment, and the last decoded frame represents the decoding result of the encoded information of the last frame of each video segment.

[0088] Decoded video represents the decoding result of encoded video. A first decoder can be used to decode the encoded video. Decoded video frames represent the decoding result of encoded video frames. Since encoded video frames include encoded first frames and encoded last frames, the first frame of each video segment in the original video is encoded to obtain the edited first frame, and the last frame of each video segment is encoded to obtain the encoded last frame. A first decoder can be used to decode the encoded video frames. For example, by decoding the encoded first frame and encoded last frame using the first decoder, the decoded first frame and decoded last frame are output. Decoded text represents the decoding result of encoded text. A second decoder can be used to encode the encoded text.

[0089] For example, the encoded video in the video encoding file is obtained, and the encoded video is decoded to obtain a decoded video; the encoded video frames in the video encoding file are obtained, and the encoded video frames are decoded to obtain decoded video frames, wherein the decoded video frames include the decoded first frame and decoded last frame of each video segment, the decoded first frame represents the decoding result corresponding to the first frame of each video segment, and the decoded last frame represents the decoding result corresponding to the last frame of each video segment; the encoded text of each video segment in the video encoding file is decoded to obtain the content description text of each video segment.

[0090] S330. Repair the decoded video based on the decoded video frames and content description text to obtain the target video corresponding to the video encoding file.

[0091] For example, video segments in the decoded video are determined based on segmentation information in the video encoding file, wherein the segmentation information represents the mapping relationship between video segments and video frames; prompt information corresponding to each video segment is generated based on the decoded first frame, decoded last frame, and content description text of each video segment; and the decoded video is subjected to image quality restoration based on the prompt information corresponding to each video segment to obtain the target video.

[0092] The prompt information can be used to provide constraints for repairing decoded videos using a preset diffusion model. The preset diffusion model can be a model that fits the pixel distribution.

[0093] In this embodiment, the preset diffusion model includes a compression module, a noise prediction module, and a decompression module. The compression module compresses the decoded first and last frames into low-dimensional first and last frame features to reduce data processing and improve the fitting efficiency of distribution information. The compression module also compresses the decoded video into a sequence of low-dimensional video frame features. The noise prediction module, under the constraint of the prompt information, repairs the decoded video based on the distribution information of the first and last frame features to obtain a video with restored image quality. The noise prediction module may include a Unet model. The decompression module restores the video with restored image quality output by the noise prediction model to its original size to obtain the target video.

[0094] The pre-diffusion model maps video frames to a noise map containing only noise by adding noise t times to the video frames in the decoded video. This continuous addition of noise to the video frames is called forward diffusion, or the noise addition process. The pre-diffusion model also learns how to transform the noisy data into noise-free data; that is, it restores the noise map to a high-quality video frame through t noise reduction operations. For example, before each noise reduction operation, the input to the pre-diffusion model is the data z at the current time step t. t , for z t Denoising is performed, and the output data z t-1 The mean and variance of the noise distribution (e.g., the noise follows a Gaussian distribution δ∈N(μ)). t-1 , σ t-1 Among them, μ t-1 Representing data z t-1 The noise distribution at time step t-1 has a mean σ t-1 Representing data z t-1 The variance of the noise distribution at time step t-1. Then, for the data z t-1 The mean and variance of the noise distribution are sampled to obtain the noise δ. t-1 Then, from the data z t Subtract noise δ from the noise t-1 Obtain data z t-1 The same method is used for the data z at time step t-1. t-1 Denoising is performed to obtain z t-2 This will not be elaborated upon here. Continue in this manner until z0 is obtained.

[0095] The forward noise addition process of the pre-defined diffusion model yields z it , z it Represented as:

[0096]

[0097] Among them, z itIt represents the noise feature at time step t. Since the three video frames are trained together, i represents the latent vector corresponding to the i-th frame, ∈ t This is noise extracted from the standard normal distribution, and ⊙ represents a multiplication operation. β t This represents a hyperparameter used to control the amount of noise added in each step.

[0098] The inverse noise reduction process of the pre-defined diffusion model is based on the diffused data z. it The original data z was inferred by gradually reducing noise. i0 .

[0099] It should be noted that the noise addition and reduction methods of the preset diffusion model in this disclosure are not limited to those listed above. If there are other methods to speed up sampling, they can also be reused in the diffusion model of this disclosure.

[0100] The target video represents the video sequence after the decoded video has undergone image quality restoration. The image quality of the target video depends on the compression rate of the first and last frames of each video segment.

[0101] Since encoding does not alter the mapping between video frames and video segments, the mapping between each video frame and video segment in the decoded video can be determined based on the segmentation information, thus identifying the video segments within the decoded video. Then, the image quality of the corresponding video segment in the decoded video can be improved based on the decoded first frame, decoded last frame, and content description text of each video segment.

[0102] Figure 4 This is a schematic diagram of a video decoding method provided in an embodiment of this disclosure. Figure 4 As shown, the encoded video is decoded by the first decoder 400 to obtain the decoded video 410. The first decoder 400 decodes the encoded first frame and encoded last frame of each video segment to obtain the decoded first frame and decoded last frame of each video segment 420. The second decoder 430 decodes the encoded text of each video segment to obtain the content description text 440 of each video segment. The text description text is mapped to text features by the text encoder 450. The decoded first frame and decoded last frame are input into a preset diffusion model 460 to obtain the first frame noise features and last frame noise features output by the preset diffusion model 460. The first frame noise features represent the noise distribution information of the decoded first frame at each time step. The last frame noise features represent the noise distribution information of the decoded last frame at each time step. The text features, first frame noise features, and last frame noise features of the same video segment are concatenated to obtain the prompt information of the corresponding video segment. The decoded video is segmented according to the segmentation information in the video encoding file to obtain the video segments of the decoded video. The image quality of each video segment in the decoded video is restored by using a preset diffusion model 460 based on the prompt information of each video segment, resulting in the target video 470.

[0103] In some embodiments, generating prompt information corresponding to each video segment based on the decoded first frame, decoded last frame, and content description text includes:

[0104] Text features are determined based on the content description text corresponding to each video segment. First-frame and last-frame features are determined based on the decoded first and last frames of each video segment. Forward diffusion processing is performed on the first-frame and last-frame features using a preset diffusion model to obtain the first-frame and last-frame noise features corresponding to each video segment. For each video segment in the decoded video, the text features, first-frame noise features, and last-frame noise features corresponding to the video segment are concatenated to obtain the corresponding prompt information for that video segment.

[0105] Figure 5 This is a schematic diagram of an image quality restoration method provided in an embodiment of this disclosure. Figure 5 As shown, a decoded video includes a first video segment, a second video segment, ..., the xth video segment, where the first video segment includes video frame 1, video frame 2, ..., video frame k, the second video segment includes video frame k+1, ..., and so on. Taking the restoration process of video frames in the first video segment as an example, the text encoder 510 converts the content description text corresponding to the first video segment into text embeddings, which serve as a constraint condition for the noise prediction module 520 of the preset diffusion model. The decoded first frame and decoded last frame of the first video segment are input into the compression module 530 of the preset diffusion model to obtain the first frame feature z_0 and the last frame feature z_k output by the compression module 530. Random noise is gradually added to the first frame feature z_0 and the last frame feature z_k by the noise prediction module 520, so that the first frame feature z_0 and the last frame feature z_k are mapped to the pure noise form of the first frame noise feature z_0t and the last frame noise feature z_kt, respectively, which serves as another constraint condition for the noise prediction module 520. Here, t represents the time step, indicating that t noises are added. By concatenating the first frame noise feature z_0t and the last frame noise feature z_kt of the text corresponding to the first video segment, the prompt information corresponding to the first video segment is obtained. The same method can be used to determine the prompt information corresponding to other video segments, which will not be elaborated here.

[0106] In some embodiments, the decoded video is subjected to image quality restoration based on the prompt information corresponding to each video segment to obtain the target video, including:

[0107] For each video segment in the decoded video, a video frame sequence is formed based on N video frames in each video segment, resulting in multiple video frame sequences. For each video frame sequence, optical flow information is obtained, and each video frame in the video frame sequence is compressed based on the optical flow information to obtain video frame features. The optical flow information is used to describe the motion information of pixels in the video frame sequence over time. The video frame features corresponding to each video frame sequence are forward diffused using a preset diffusion model to obtain target noise features corresponding to each video frame sequence. The target noise features are then reverse diffused using a preset diffusion model based on the prompt information corresponding to the video segment to which each video frame sequence belongs, resulting in repaired video frame sequences corresponding to each video frame sequence. The repaired video frame sequences corresponding to each video frame sequence are then concatenated to obtain the target video.

[0108] The video frame sequence represents the sequence of N video frames in the video segment of the decoded video, arranged in order of timestamps, so as to facilitate the calculation of optical flow information of the video frames in the video frame sequence.

[0109] Since adjacent video frames are not isolated, and considering the continuity of video frames, the compression module of the pre-defined diffusion model compresses video frames based on the optical flow information of each frame relative to its adjacent frames. Optical flow information describes the motion of pixels in a video frame sequence over time. For example, the speed and direction of an object's movement can be estimated by the change in grayscale values ​​of pixels between two adjacent video frames. Optical flow information can be used for motion compensation in video compression. By calculating the optical flow vectors between video frames, motion compensation can be performed during video compression, achieving efficient compression.

[0110] Figure 6 This is a schematic diagram of a video frame compression method provided in an embodiment of this disclosure. Figure 6 As shown, each three adjacent video frames in the decoded video are used as input to the compression module 600 in the preset diffusion model. The optical flow estimation submodule 610 in the compression module 600 obtains the optical flow information corresponding to the three video frames, and inputs the optical flow information into the residual network module 620 so that the residual network module 620 compresses the three video frames based on the optical flow information to obtain three low-dimensional video frame features.

[0111] Reference Figure 5Taking the first three video frames in the first video segment as an example, video frames 1, 2, and 3 are input into a preset diffusion model. The compression module 530 estimates the optical flow information between adjacent video frames and compresses the three video frames based on the optical flow information to obtain low-dimensional video frame features, denoted as z_0, z_1, and z_2. The noise prediction module 520 gradually adds random noise to z_0, z_1, and z_2 to map z_0, z_1, and z_2 into pure noise forms z_0t, z_1t, and z_2t, respectively. The prompt information corresponding to the first video segment and z_0t, z_1t, and z_2t are input into the noise prediction module 520. The noise prediction module 520 performs progressive noise reduction processing on z_0t, z_1t, and z_2t based on the prompt information to achieve image quality restoration and obtain the restored video frame features, denoted as z_0′, z_1′, and z_2′. Input z_0′, z_1′, and z_2′ into the decompression module 540. The decompression module 540 restores z_0′, z_1′, and z_2′ to obtain repaired video frame 1, repaired video frame 2, and repaired video frame 3. The sizes of repaired video frame 1, repaired video frame 2, and repaired video frame 3 are the same as those of video frame 1, video frame 2, and video frame 3.

[0112] Decoding encoded video frames of quality level x and edited video frames of quality level y yields the decoding results of the first and last encoded frames of each video segment, as well as the decoding result of the encoded video itself—the decoded first frame, decoded last frame, and decoded video. Combining the decoded first and last frames of each video segment with the content description text, the decoded video is repaired to bring its quality level closer to x, achieving the effect of repairing highly compressed decoded videos based on the first and last frames of video segments.

[0113] The technical solution of this disclosure involves decoding a video encoded file to obtain a decoded video, decoded video frames, and content description text. Based on the decoded video frames and content description text, prompt information is generated. The decoded video is then repaired based on the prompt information to obtain the target video corresponding to the video encoded file. Since the bitrate of the encoded video is lower than that of the encoded video frames, and a higher bitrate can retain more detail, the image quality of the decoded video frames obtained by decoding the encoded video frames is higher than that of the decoded video obtained by decoding the encoded video. Repairing the decoded video based on the decoded video frames and content description text allows the image quality of the repaired target video to approach that of the decoded video frames, improving the image quality of the decoded video and achieving high video compression while reducing video quality loss.

[0114] Figure 7This is a schematic diagram of a video processing device provided in an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware, and optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.

[0115] like Figure 7 As shown, the device includes: a segmentation module 710, a text determination module 720, an acquisition module 730, an encoding module 740, and a file determination module 750.

[0116] The segmentation module 710 is used to segment the original video to obtain at least two video segments;

[0117] The text determination module 720 is used to determine the content description text of each video segment based on the content information of each video segment.

[0118] The acquisition module 730 is used to acquire target video frames in each video segment, wherein the target video frame represents a video frame in each video segment whose image quality information meets preset filtering conditions.

[0119] The encoding module 740 is used to encode the original video, the content description text of each video segment, and the target video frame to obtain an encoded video, an encoded video frame, and encoded text, wherein the bit rate of the encoded video is less than the bit rate of the encoded video frame.

[0120] The file determination module 750 is used to determine the video encoding file based on the encoded video, encoded video frames, and encoded text.

[0121] Optionally, the segmentation module 710 is specifically used for:

[0122] Obtain transition video frames from the original video, wherein the transition video frames represent video frames used to connect different video segments;

[0123] The original video is segmented based on the transition video frames to obtain at least two video segments.

[0124] Optionally, the segmentation module 710 is specifically used for:

[0125] For any two adjacent video frames in the original video, determine the pixel difference between the two adjacent video frames;

[0126] The original video is segmented based on adjacent video frames whose pixel differences exceed a preset segmentation threshold, resulting in at least two video segments.

[0127] Optionally, the text determination module 720 is specifically used for:

[0128] For each of at least two video segments, a set number of video frames are obtained from the video segment;

[0129] Based on the content information of the set number of video frames, determine the content description text of the video segment.

[0130] Optionally, the acquisition module 730 is specifically used for:

[0131] For each of the at least two video segments, at least two video frames in the video segment whose time interval and image quality information meet the preset filtering conditions are obtained as the target video frames of the video segment.

[0132] Optionally, the target video frame includes a first frame and a last frame; the encoding module 740 is specifically used for:

[0133] The original video is encoded based on the first compression ratio to obtain an encoded video;

[0134] The first and last frames of each video segment are encoded based on the second compression ratio to obtain the encoded first and last frames of each video segment. The first compression ratio is higher than the second compression ratio, and the encoded video frame includes the encoded first and last frames of each video segment.

[0135] Optionally, the file determination module 750 is specifically used for:

[0136] Obtain the segmentation information of the original video, wherein the segmentation information represents the mapping relationship between video segments and video frames;

[0137] Based on the encoded video, segmentation information, and the encoded first frame, encoded last frame, and encoded text of each video segment, the video encoding file corresponding to the original video is determined.

[0138] The video processing apparatus provided in this disclosure can execute the video processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0139] Figure 8 This is a schematic diagram of another video processing device structure provided in an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware, and optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.

[0140] like Figure 8 As shown, the device includes: an acquisition module 810, a decoding module 820, and a repair module 830.

[0141] The acquisition module 810 is used to acquire a video encoding file, wherein the video encoding file represents a compressed file of the original video, including encoded video, encoded video frames and encoded text, the bitrate of the encoded video is less than the bitrate of the encoded video frames, the encoded video frames represent the encoding information of target video frames in each video segment, the target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions, and the encoded text represents the encoding information of the content description text of each video segment;

[0142] The decoding module 820 is used to decode the video encoded file to obtain decoded video, decoded video frames, and content description text;

[0143] Repair module 830 is used to repair the decoded video based on the decoded video frames and content description text to obtain the target video corresponding to the video encoding file.

[0144] Optionally, the decoding module 820 is specifically used for:

[0145] Obtain the encoded video from the video encoding file, and decode the encoded video to obtain the decoded video;

[0146] The encoded video frames in the video encoding file are obtained, and the encoded video frames are decoded to obtain decoded video frames. The decoded video frames include the first decoded frame and the last decoded frame of each video segment. The first decoded frame represents the decoding result corresponding to the first frame of each video segment, and the last decoded frame represents the decoding result corresponding to the last frame of each video segment.

[0147] The encoded text of each video segment in the video encoding file is decoded to obtain the content description text of each video segment.

[0148] Furthermore, the repair module 830 is specifically used for:

[0149] The video segments in the decoded video are determined based on the segmentation information in the video encoding file, wherein the segmentation information represents the mapping relationship between video segments and video frames;

[0150] Based on the first and last decoded frames and content description text of each video segment, generate prompt information corresponding to each video segment;

[0151] The decoded video is then repaired based on the prompts corresponding to each video segment to obtain the target video.

[0152] Furthermore, the step of generating corresponding prompt information for each video segment based on the decoded first frame, decoded last frame, and content description text includes:

[0153] Determine text features based on the content description text corresponding to each video segment;

[0154] Based on the decoded first frame and decoded last frame of each video segment, determine the first frame feature and last frame feature of each video segment; perform forward diffusion processing on the first frame feature and last frame feature using a preset diffusion model to obtain the first frame noise feature and last frame noise feature of each video segment.

[0155] For each video segment in the decoded video, the text features, first frame noise features, and last frame noise features corresponding to the video segment are concatenated to obtain the prompt information corresponding to the video segment.

[0156] Further, the step of performing image quality restoration on the decoded video based on the prompt information corresponding to each video segment to obtain the target video includes:

[0157] For each video segment in the decoded video, a video frame sequence is formed based on N video frames in each video segment, resulting in multiple video frame sequences;

[0158] For each video frame sequence, optical flow information of the video frame sequence in the video segment is obtained, and video frame features are obtained by compressing each video frame in the video frame sequence based on the optical flow information. The optical flow information is used to describe the motion information of pixels in the video frame sequence over time.

[0159] The video frame features corresponding to each video frame sequence are forward diffused using a preset diffusion model to obtain the target noise features corresponding to each video frame sequence.

[0160] Based on the prompt information corresponding to the video segment to which each video frame sequence belongs, a preset diffusion model is used to perform image quality restoration processing and reverse diffusion processing on the target noise features corresponding to the video segment to obtain the restored video frame sequence corresponding to each video frame sequence of the video segment.

[0161] The target video is obtained by splicing together the video frame sequences corresponding to the video frames of each video segment.

[0162] The video processing apparatus provided in this disclosure can execute the video processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0163] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0164] Figure 9This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 9 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 9 The diagram below shows the structure of the terminal device or server 900. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0165] like Figure 9 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An edit / output (I / O) interface 905 is also connected to the bus 904.

[0166] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0167] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.

[0168] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0169] The electronic device provided in this embodiment and the video processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0170] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video processing method provided in the above embodiments.

[0171] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0172] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0173] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0174] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:

[0175] The original video is segmented to obtain at least two video segments;

[0176] Determine the content description text for each video segment based on the content information of each video segment;

[0177] Obtain target video frames from each video segment, wherein the target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions;

[0178] The original video, the content description text of each video segment, and the target video frame are encoded to obtain an encoded video, an encoded video frame, and encoded text. The bitrate of the encoded video is less than the bitrate of the encoded video frame.

[0179] The video encoding file is determined based on the encoded video, encoded video frames, and encoded text.

[0180] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:

[0181] Obtain a video encoding file, wherein the video encoding file represents a compressed file of the original video, including encoded video, encoded video frames and encoded text, the bitrate of the encoded video is less than the bitrate of the encoded video frames, the encoded video frames represent the encoding information of target video frames in each video segment, the target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions, and the encoded text represents the encoding information of the content description text of each video segment.

[0182] The video encoded file is decoded to obtain decoded video, decoded video frames, and content description text;

[0183] The decoded video is repaired based on the decoded video frames and content description text to obtain the target video corresponding to the video encoding file.

[0184] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0185] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0186] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0187] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0188] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0189] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0190] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0191] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video processing method, characterized in that, include: The original video is segmented to obtain at least two video segments; Determine the content description text for each video segment based on the content information of each video segment; Obtain target video frames from each video segment, wherein the target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions; The original video, the content description text of each video segment, and the target video frame are encoded to obtain an encoded video, an encoded video frame, and encoded text. The bitrate of the encoded video is less than the bitrate of the encoded video frame. The video encoding file is determined based on the encoded video, encoded video frames, and encoded text.

2. The method according to claim 1, characterized in that, The process of segmenting the original video to obtain at least two video segments includes: Obtain transition video frames from the original video, wherein the transition video frames represent video frames used to connect different video segments; The original video is segmented based on the transition video frames to obtain at least two video segments.

3. The method according to claim 1, characterized in that, The process of segmenting the original video to obtain at least two video segments includes: For any two adjacent video frames in the original video, determine the pixel difference between the two adjacent video frames; The original video is segmented based on adjacent video frames whose pixel differences exceed a preset segmentation threshold, resulting in at least two video segments.

4. The method according to claim 1, characterized in that, The step of determining the content description text of each video segment based on the content information of each video segment includes: For each of at least two video segments, a set number of video frames are obtained from the video segment; Based on the content information of the set number of video frames, determine the content description text of the video segment.

5. The method according to claim 1, characterized in that, The step of obtaining the target video frame in each video segment includes: For each of the at least two video segments, at least two video frames in the video segment whose time interval and image quality information meet the preset filtering conditions are obtained as the target video frames of the video segment.

6. The method according to claim 1, characterized in that, The target video frame includes the first frame and the last frame; The original video and the target video frames of each video segment are encoded to obtain encoded video and encoded video frames, including: The original video is encoded based on the first compression ratio to obtain an encoded video; The first and last frames of each video segment are encoded based on the second compression ratio to obtain the encoded first and last frames of each video segment. The first compression ratio is higher than the second compression ratio, and the encoded video frame includes the encoded first and last frames of each video segment.

7. The method according to claim 6, characterized in that, The step of determining the video encoding file based on the encoded video, encoded video frames, and encoded text includes: Obtain the segmentation information of the original video, wherein the segmentation information represents the mapping relationship between video segments and video frames; Based on the encoded video, segmentation information, and the encoded first frame, encoded last frame, and encoded text of each video segment, the video encoding file corresponding to the original video is determined.

8. A video processing method, characterized in that, include: Obtain a video encoding file, wherein the video encoding file represents a compressed file of the original video, including encoded video, encoded video frames and encoded text, the bitrate of the encoded video is less than the bitrate of the encoded video frames, the encoded video frames represent the encoding information of target video frames in each video segment, the target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions, and the encoded text represents the encoding information of the content description text of each video segment. The video encoded file is decoded to obtain decoded video, decoded video frames, and content description text; The decoded video is repaired based on the decoded video frames and content description text to obtain the target video corresponding to the video encoding file.

9. The method according to claim 8, characterized in that, The process of decoding the video encoded file to obtain decoded video, decoded video frames, and content description text includes: Obtain the encoded video from the video encoding file, and decode the encoded video to obtain the decoded video; The encoded video frames in the video encoding file are obtained, and the encoded video frames are decoded to obtain decoded video frames. The decoded video frames include the first decoded frame and the last decoded frame of each video segment. The first decoded frame represents the decoding result corresponding to the first frame of each video segment, and the last decoded frame represents the decoding result corresponding to the last frame of each video segment. The encoded text of each video segment in the video encoding file is decoded to obtain the content description text of each video segment.

10. The method according to claim 9, characterized in that, The step of repairing the decoded video based on the decoded video frames and content description text to obtain the target video corresponding to the video encoding file includes: The video segments in the decoded video are determined based on the segmentation information in the video encoding file, wherein the segmentation information represents the mapping relationship between video segments and video frames; Based on the first frame, the last frame, and the content description text of each video segment, generate corresponding prompt information for each video segment; The decoded video is then repaired based on the prompts corresponding to each video segment to obtain the target video.

11. The method according to claim 10, characterized in that, The step of generating corresponding prompt information for each video segment based on the decoded first frame, decoded last frame, and content description text includes: Determine text features based on the content description text corresponding to each video segment; Based on the decoded first frame and decoded last frame of each video segment, determine the first frame features and last frame features of each video segment. The first frame features and the last frame features are forward diffused using a preset diffusion model to obtain the first frame noise features and the last frame noise features corresponding to each video segment. For each video segment in the decoded video, the text features, first frame noise features, and last frame noise features corresponding to the video segment are concatenated to obtain the prompt information corresponding to the video segment.

12. The method according to claim 10, characterized in that, The step of performing image quality restoration on the decoded video based on the prompt information corresponding to each video segment to obtain the target video includes: For each video segment in the decoded video, a video frame sequence is formed based on N video frames in each video segment, resulting in multiple video frame sequences; For each video frame sequence, the optical flow information of the video frame sequence is obtained, and the video frame features are obtained by compressing each video frame in the video frame sequence based on the optical flow information. The optical flow information is used to describe the motion information of pixels in the video frame sequence over time. By performing forward diffusion processing on the video frame features corresponding to each video frame sequence using a preset diffusion model, the target noise features corresponding to each video frame sequence are obtained. By using a preset diffusion model based on the prompt information corresponding to the video segment to which each video frame sequence belongs, the target noise features are reverse diffused to obtain the repaired video frame sequence corresponding to each video frame sequence. The target video is obtained by splicing together the repaired video frame sequences corresponding to each video frame sequence.

13. A video processing apparatus, characterized in that, include: The segmentation module is used to segment the original video to obtain at least two video segments; The text determination module is used to determine the content description text of each video segment based on the content information of each video segment. The acquisition module is used to acquire target video frames in each video segment, wherein the target video frame represents a video frame in each video segment whose image quality information meets preset filtering conditions. The encoding module is used to encode the original video, the content description text of each video segment, and the target video frame to obtain encoded video, encoded video frame, and encoded text, wherein the bit rate of the encoded video is less than the bit rate of the encoded video frame. The file determination module is used to determine the video encoding file based on the encoded video, encoded video frames, and encoded text.

14. A video processing apparatus, characterized in that, include: The acquisition module is used to acquire video encoded files, wherein the video encoded file represents a compressed file of the original video, including encoded video, encoded video frames and encoded text, the bitrate of the encoded video is less than the bitrate of the encoded video frames, the encoded video frames represent the encoding information of target video frames in each video segment, the target video frames represent video frames in each video segment whose image quality information meets preset filtering conditions, and the encoded text represents the encoding information of the content description text of each video segment; The decoding module is used to decode the video encoded file to obtain decoded video, decoded video frames, and content description text; The repair module is used to repair the decoded video based on the decoded video frames and content description text to obtain the target video corresponding to the video encoding file.

15. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method as described in any one of claims 1-12.

16. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the video processing method as described in any one of claims 1-12.