Video encoding and decoding methods

The video encoding and decoding methods improve accuracy by employing pixel point grouping and entropy decoding techniques, leveraging reference information and machine learning models to adaptively encode and decode video frames.

WO2026001914A1PCT designated stage Publication Date: 2026-01-02ZHEJIANG DAHUA TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102874
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-18
Filing Date
2025-06-23
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing video frame encoding and decoding processes are rigid and lack adaptability, leading to low encoding and decoding accuracy.

Method used

A video encoding method that utilizes pixel point grouping and reference information to enhance encoding, and a decoding method that employs entropy decoding and inverse transformation to improve decoding accuracy.

Benefits of technology

Enhances the accuracy and efficiency of video frame encoding and decoding by utilizing reference information and advanced machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102874_02012026_PF_FP_ABST
    Figure CN2025102874_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a video encoding and decoding method. The video decoding method comprises: obtaining encoded data and first reference information from, the first reference information including feature information of at least one decoded video frame before a target video frame; and obtaining a decoded target video frame based on the encoded data and the first reference information.
Need to check novelty before this filing date? Find Prior Art

Description

VIDEO ENCODING AND DECODING METHODSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to Chinese Application No. 202410831847.8, filed on June 25, 2024, and Chinese Application No. 202411639744.8, filed on November 18, 2024, the entire contents of each of which are incorporated herein by reference.TECHNICAL FIELD

[0002] The present disclosure relates to the field of image processing technology, and in particular to a video encoding and decoding method.BACKGROUND

[0003] Since an original video frame occupies a large amount of data, it usually needs to encode and then decode the original video frame during a transmission process of the video frame. Existing video frame encoding and decoding process generally encodes and decodes the entire video frame based on a universal encoding standard. Accordingly, the mode used in the encoding and decoding process is relatively single and rigid, which makes it difficult to perform encoding adaptively and decoding in a refined manner, resulting in low encoding and decoding accuracy.

[0004] Therefore, it is desirable to provide a video encoding and decoding method to improve the encoding and decoding accuracy of the video frame.SUMMARY

[0005] One or more embodiments of the present disclosure provide a video decoding method, implemented by a decoding terminal, comprising: obtaining encoded data and first reference information from an encoding terminal, the first reference information including feature information of at least one decoded video frame before a target video frame; and obtaining a decoded target video frame based on the encoded data and the first reference information.

[0006] One or more embodiments of the present disclosure provide an image decoding method, implemented by a decoding terminal, comprising: obtaining pixel point grouping information of an original video frame from an encoding terminal, the pixel point grouping information being obtained by performing pixel point grouping on the original video frame based on an image content of the original video frame; obtaining entropy decoding data by performing entropy decoding on encoded data based on the pixel point grouping information; obtaining inverse quantization data by performing inverse quantization on the entropy decoding data; and obtaining a decoded target video frame corresponding to the original video frame by performing an inverse transformation on the inverse quantization data.

[0007] One or more embodiments of the present disclosure provide a video encoding method, implemented by an encoding terminal, comprising: obtaining an original video frame and second reference information, the second reference information including feature information of at least one decoded video frame before the original video frame; obtaining encoded data based on the original video frame and the second reference information; and sending the encoded data to a decoding terminal.

[0008] One or more embodiments of the present disclosure provide an image encoding method, implemented by an encoding terminal, comprising: obtaining pixel point grouping information of an original video frame by performing pixel point grouping on the original video frame based on an image content of the original video frame; obtaining second transformation data by transforming the original video frame; obtaining second quantization data by quantizing the second transformation data; obtaining encoded data by performing entropy encoding on the second quantization data based on the pixel point grouping information; and sending the encoded data to a decoding terminal.

[0009] One or more embodiments of the present disclosure provide a system, comprising: at least one storage device including a set of instructions; and at least one processor in communication with the at least one storage device, wherein when executing the set of instructions, the at least one processor is directed to perform the method described in any embodiment of the present disclosure.

[0010] One or more embodiments of the present disclosure provide a non-transitory computer-readable storage medium, comprising at least one set of instructions, wherein when executed by one or more processors of a computing device, the at least one set of instructions causes the computing device to perform the method described in any embodiment of the present disclosure.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The present disclosure will be further illustrated by way of exemplary embodiments, which will be described in detail by means of the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbering indicates the same structure, wherein:

[0012] FIG. 1 is a schematic diagram illustrating an application scenario of a video encoding and decoding process according to some embodiments of the present disclosure;

[0013] FIG. 2 is a schematic diagram illustrating software / hardware of a computing device according to some embodiments of the present disclosure;

[0014] FIG. 3 is a block diagram illustrating an exemplary video encoding and decoding system according to some embodiments of the present disclosure;

[0015] FIG. 4 is a flowchart illustrating a video encoding method according to some embodiments of the present disclosure;

[0016] FIG. 5 is a flowchart illustrating an exemplary video decoding method according to some embodiments of the present disclosure;

[0017] FIG. 6 is a schematic diagram illustrating video encoding and decoding according to some embodiments of the present disclosure;

[0018] FIG. 7 a flowchart illustrating an exemplary video decoding method according to some embodiments of the present disclosure;

[0019] FIG. 8 is a flowchart illustrating an exemplary process of pixel point grouping according to some embodiments of the present disclosure;

[0020] FIG. 9A is a schematic diagram illustrating a preset continuous path according to some embodiments of the present disclosure;

[0021] FIG. 9B is a schematic diagram illustrating a preset continuous path according to some embodiments of the present disclosure;

[0022] FIG. 10 is a flowchart illustrating an exemplary process of pixel point grouping according to some embodiments of the present disclosure;

[0023] FIG. 11 a flowchart illustrating an exemplary video decoding method according to some embodiments of the present disclosure;

[0024] FIG. 12 a flowchart illustrating an exemplary video encoding method according to some embodiments of the present disclosure;

[0025] FIG. 13 a flowchart illustrating an exemplary video decoding method according to some embodiments of the present disclosure; and

[0026] FIG. 14 is a schematic diagram illustrating video encoding and decoding according to some embodiments of the present disclosure.DETAILED DESCRIPTION

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required to be used in the description of the embodiments are briefly described below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present disclosure, and it is possible for a person of ordinary skill in the art to apply the present disclosure to other similar scenarios in accordance with these drawings without creative labor. Unless obviously obtained from the context or the context illustrates otherwise, the same numeral in the drawings refers to the same structure or operation.

[0028] It should be understood that the terms “system, ” “device, ” “unit” and / or “module” used herein are a way to distinguish between different components, elements, parts, sections, or assemblies at different levels. However, the terms may be replaced by other expressions if other words accomplish the same purpose.

[0029] As shown in the present disclosure and in the claims, unless the context clearly suggests an exception, the words “one, ” “a, ” “an, ” “one kind, ” and / or “the” do not refer specifically to the singular, but may also include the plural. Generally, the terms “including” and “comprising” suggest only the inclusion of clearly identified steps and elements, however, the steps and elements that do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0030] Flowcharts are used in the present disclosure to illustrate the operations performed by a system according to embodiments of the present disclosure, and the related descriptions are provided to aid in a better understanding of the magnetic resonance imaging method and / or system. It should be appreciated that the preceding or following operations are not necessarily performed in an exact sequence. Instead, steps can be processed in reverse order or simultaneously. Also, it is possible to add other operations to these processes or to remove a step or steps from these processes.

[0031] FIG. 1 is a schematic diagram illustrating an application scenario of a video encoding and decoding process according to some embodiments of the present disclosure. As illustrated in FIG. 1, an application scenario 100 may include a processing device 110, a storage device 120, a network 130, an image acquisition device 140, and a terminal device 150. In some embodiments, the processing device 110, the storage device 120, and the image acquisition device 140 may be connected to and / or communicate with each other via a wireless connection, a wired connection, or a combination thereof.

[0032] The processing device 110 may be configured to process data and / or information correlated with a video encoding and decoding method. For example, the processing device 110 may obtain encoded data based on an original video frame and second reference information. As another example, the processing device 110 may obtain a decoded target video frame corresponding to the original video frame based on the encoded data and first reference information. As another example, the processing device 110 may obtain pixel point grouping information of the original video frame by performing pixel point grouping on the original video frame based on an image content of the original video frame; obtain transformation data (hereinafter referred to as second transformation data) by transforming the original video frame; obtain quantization data (hereinafter referred to as second quantization data) by quantizing the transformation data; and obtain the encoded data by performing entropy encoding on the quantization data based on the pixel point grouping information. As another example, the processing device 110 may obtain entropy decoding data by performing entropy decoding on the encoded data based on the pixel point grouping information; obtain inverse quantization data by performing inverse quantization on the entropy decoding data; and obtain the decoded target video frame corresponding to the original video frame by performing inverse transformation on the inverse quantization data.

[0033] In some embodiments, the processing device 110 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processing device 110 may be local or remote.

[0034] The storage device 120 may be configured to store data, instructions, and / or any other information correlated with the video encoding and decoding method. In some embodiments, the storage device 120 may be configured to store data obtained from the processing device 110, the image acquisition device 140, and / or the terminal device 150. For example, the storage device 120 may be configured to store the original video frame, the second reference information, the encoded data, prompt information, second prompt data input by a user, and the pixel point grouping information of the original video frame. As another example, the storage device 120 may be configured to store the first reference information, the decoded target video frame, first prompt data input by the user, etc.

[0035] In some embodiments, the storage device 120 may store data and / or instructions that the processing device 110 may execute or use to perform exemplary methods described in the present disclosure.

[0036] In some embodiments, the storage device 120 may include a mass storage, removable storage, a volatile read-and-write memory, a read-only memory (ROM) , or the like, or any combination thereof. An exemplary mass storage may include a magnetic disk, an optical disk, a solid-state drive, etc. An exemplary removable storage may include a flash drive, a floppy disk, an optical disk, a memory card, a zip disk, a magnetic tape, etc. An exemplary volatile read-and-write memory may include a random access memory (RAM) . An exemplary RAM may include a dynamic RAM (DRAM) , a double date rate synchronous dynamic RAM (DDR SDRAM) , a static RAM (SRAM) , a thyristor RAM (T-RAM) , and a zero-capacitor RAM (Z-RAM) , etc. An exemplary ROM may include a mask ROM (MROM) , a programmable ROM (PROM) , an erasable programmable ROM (EPROM) , an electrically erasable programmable ROM (EEPROM) , a compact disk ROM (CD-ROM) , and a digital versatile disk ROM, etc. In some embodiments, the storage device 120 may be implemented on a cloud platform. In some embodiments, the storage device 120 may be integrated into the processing device 110 and / or the image acquisition device 140.

[0037] The network 130 may include any suitable network that may facilitate the exchange of information and / or data for the application scenario 100. In some embodiments, one or more components (e.g., the processing device 110, the storage device 120, the image acquisition device 140, and the terminal device 150) of the application scenario 100 may communicate information and / or data with one or more other components of the application scenario 100 via the network 130. For example, the processing device 110 may obtain the original video frame from the image acquisition device 140 via the network 130. As another example, the processing device 110 may obtain the second prompt data and / or the first prompt data input by the user from the terminal device 150 via the network 130.

[0038] In some embodiments, the network 130 may be and / or include a public network (e.g., the Internet) , a private network (e.g., a local area network (LAN) , a wide area network (WAN) ) ) , a wired network (e.g., an Ethernet network) , a wireless network (e.g., an 802.11 network, a Wi-Fi network) , a cellular network (e.g., a Long Term Evolution (LTE) network) , a frame relay network, a virtual private network (VPN) , a satellite network, a telephone network, routers, hubs, switches, server computers, and / or any combination thereof. In some embodiments, the network 130 may include one or more network access points. For example, the network 130 may include wired and / or wireless network access points such as base stations and / or internet exchange points through which the one or more components of the application scenario 100 may be connected to the network 130 to exchange data and / or information.

[0039] The image acquisition device 140 may be configured to acquire a video and / or an image. In some embodiments, the image acquisition device 140 may include an image acquisition device, an auxiliary device for mounting the image acquisition device, etc. For example, the image acquisition device may include a camera and a pan-tilt device, such as a monocular camera, a dual-light pan-tilt, etc. In some embodiments, the image acquisition device 140 may be configured to acquire video or image information of a target subject in a target region. In some embodiments, the image acquisition device 140 may send the acquired video to the storage device 120 for storage, or to the processing device 110 for related data processing (e.g., encoding, and decoding) via the network 130. In some embodiments, the image acquisition device 140 may further include a built-in storage device and a processor, which are respectively used to store and process the video acquired by the image acquisition device.

[0040] The terminal device 150 enables user interaction between the user and the application scenario 100. For example, the terminal device 150 may transmit instructions or data (e.g., the second prompt data, and the first prompt data) input by the user to the processing device 110, and / or display processing results (e.g., the encoded data, and decoded data) generated by the processing device 110 to the user. The terminal device 150 may include a mobile device 151, a tablet computer 152, a laptop computer 153, or the like, or any combination thereof. In some embodiments, the mobile device 151 may include a smart home device, a wearable device, a smart mobile device, a virtual reality device, an augmented reality device, or the like, or any combination thereof. In some embodiments, the terminal device 150 may be part of the processing device 110.

[0041] In some embodiments, the terminal device 150 may include any device with algorithm processing capability, which may be integrated with an artificial intelligence (AI) platform. The AI platform refers to a platform used to develop machine learning models. For example, applications (APPs) , plug-in components, and / or websites of the AI platform may be installed on the terminal device 150.

[0042] It should be noted that the above description is merely provided for the purposes of illustration, and not intended to limit the scope of the present disclosure. For those having ordinary skills in the art, a plurality of variations and modifications may be made under the teachings of the present disclosure. However, those variations and modifications do not depart from the scope of the present disclosure.

[0043] FIG. 2 a schematic diagram illustrating software / hardware of a computing device according to some embodiments of the present disclosure.

[0044] In some embodiments, a computing device 200 may include a server, a personal computer, a laptop computer, a smart phone, a tablet computer, a smart mobile phone, etc.

[0045] As shown in FIG. 2, the computing device 200 may include a processor 210, a storage device, an input / output 230, and a communication port 240. The storage device may include a non-volatile storage medium 225 and a memory 223. The processor 210, the storage device (e.g., the memory 223, and the non-volatile storage medium 225) , and the input / output 230 may be connected via a system bus 250. The communication port 240 may be connected to the system bus 250 via the input / output 230.

[0046] The processor 210 may execute computer instructions (e.g., a program code) and may perform functions of a processing device in accordance with the techniques described in the present disclosure (e.g., the video encoding method, and the video decoding method) . In some embodiments, the processor 210 may include one or more hardware processors, such as a microcontroller, microprocessor, etc. For illustrative purposes only, only one processor is described in the computing device 200. However, it is noted that the computing device 200 may also include a plurality of processors. Operations and / or methods described in the present disclosure that are performed by a single processor may also be performed by a plurality of processors together or separately. For example, if the processor of the computing device 200 described in the present disclosure performs an operation A and an operation B, it should be appreciated that the operation A and the operation B may also be performed by two or more different processors of the computer device 200 jointly or separately (e.g., a first processor executes the operation A and a second processor executes the operation B, or the first processor and the second processor jointly execute the operations A and B) .

[0047] The storage device may store data / information obtained from the image acquisition device 140, and / or any other component of the application scenario 100. For example, the non-volatile storage medium 225 may store an operating system, a computer program, and a database. The memory 2 23 may provide an environment for an operation of the operating system and the computer programs in the non-volatile storage medium 225. The database may be configured to store video encoded data (e.g., an original video frame, second reference information, encoded data, prompt information, second prompt data, pixel point grouping information of the original video frame, etc. ) , video decoded data (e.g., first reference information, a decoded target video frame, first prompt data, etc. ) . The processor 2 10 may execute the computer programs to implement the video encoding method and / or the video decoding method described herein.

[0048] The input / output 230 may be configured to exchange information between the processor 210 and an external device, such as the image acquisition device 140. In some embodiments, the input / output 230 may include an input device and an output device. The input device may include a keyboard, a mouse, a touch screen, a microphone, or the like, or any combination thereof. The output device may include a display device, a loudspeaker, a printer, a projector, or the like, or any combination thereof.

[0049] The communication port 240 may be configured to communicate with an external terminal (e.g., the image acquisition device 140) via a network connection. The connection may be a wired connection, a wireless connection, any connection that enables data transmission and / or reception, or the like. or any combination thereof.

[0050] It should be understood that the descriptions of FIG. 2 are only provided for the illustration of illustration and do not constitute a limitation to the present disclosure. For those skilled in the art, various changes and modifications can be made under the guidance of the present disclosure. Features, structures, manners, and other characteristics of embodiments of the present disclosure can be combined in various ways to obtain other and / or alternative embodiments. However, such changes and modifications do not exceed the scope of the present disclosure.

[0051] FIG. 3 is a block diagram illustrating an exemplary video encoding and decoding system according to some embodiments of the present disclosure.

[0052] As shown in FIG. 3, in some embodiments, a video encoding and decoding system 300 may include an encoding module 310 and a decoding module 320. In some embodiments, one or more modules of the video encoding and decoding system 300 may be connected with each other via a connection. The connection may be a wireless connection or a wired connection.

[0053] The encoding module 310 may be configured to encode an original video frame.

[0054] In some embodiments, the encoding module 310 may be configured to obtain encoded data based on the original video frame and second reference information. For example, the encoding module 310 may obtain the encoded data by processing the original video frame and the second reference information through a second generative machine learning model.

[0055] In some embodiments, the encoding module 310 may be configured to obtain pixel point grouping information of the original video frame by performing pixel point grouping on the original video frame based on an image content of the original video frame; obtain second transformation data by transforming the original video frame; obtain second quantization data by quantizing the second transformation data; and obtain the encoded data by performing entropy encoding on the second quantization data based on the pixel point grouping information.

[0056] In some embodiments, the encoding module 310 may further include an encoding acquisition unit 311, an encoding prompt unit 313, and an encoding unit 315.

[0057] The encoding acquisition unit 311 may be configured to acquire the original video frame and / or the second reference information.

[0058] The encoding prompt unit 313 may be configured to generate prompt information and / or second prompt data.

[0059] In some embodiments, the encoding prompt unit 313 may be configured to generate the prompt information based on the original video frame and the second reference information through the second generative machine learning model. In some embodiments, the encoding prompt unit 313 may be configured to generate the prompt information based on the original video frame through the second generative machine learning model. The prompt information reflects at least one of following information related to the original video frame: pixel value information, scaling information, timing information, segmented semantic information, and classified semantic information.

[0060] In some embodiments, the encoding prompt unit 313 may be configured to determine the second prompt data input by a user. The second prompt data may be configured to indicate to adjust the image content of the original video frame.

[0061] The encoding unit 315 may be configured to obtain the encoded data by processing the original video frame and the second reference information through the second generative machine learning model. In some embodiments, the encoding unit 315 may encode the prompt information, and obtain a bitstream that fuses feature information of the original video frame and prompt information by fusing encoded prompt information and the encoded data. In some embodiments, the encoding unit 315 may obtain the bitstream that fuses the feature information of the original video frame and prompt information by encoding the original video frame, the second reference information, and the prompt information. In some embodiments, the encoding unit 315 may adjust the bitstream that fuses the feature information of the original video frame and prompt information based on the second prompt data. In some embodiments, the encoding unit 315 may adjust the encoded data of the original video frame and prompt information based on the second prompt data, and obtain an adjusted bitstream by encoding adjusted encoded data and prompt information through the second generative machine learning model.

[0062] In some embodiments, the encoding unit 315 may be configured to obtain pixel point grouping information of the original video frame by performing pixel point grouping on the original video frame based on the image content of the original video frame; obtain second fusion data by fusing the original video frame and the second reference information; obtain first transformation data by transforming the second fusion data; obtain first quantization data by quantizing the first transformation data; and obtain the encoded data by performing entropy encoding on the first quantization data based on the pixel point grouping information.

[0063] The decoding module 320 may be configured for decoding. In some embodiments, the decoding module 320 may be configured to obtain a decoded target video frame corresponding to the original video frame based on the encoded data. In some embodiments, the decoding module 320 may be configured to obtain the decoded target video frame based on the encoded data and first reference information. For example, the decoding module 320 may obtain the decoded target video frame by processing the encoded data and the first reference information through a first generative machine learning model. As another example, the decoding module 320 may obtain the decoded target video frame by processing the encoded data, the first reference information, and the prompt information through the first generative machine learning model. As another example, the decoding module 320 may obtain the decoded target video frame by processing the encoded data, the first reference information, the prompt information, and the first prompt data through the first generative machine learning model. The first prompt data may be configured to indicate an adjustment of an image content of the decoded target video frame during a process of obtaining the decoded target video frame. As another example, the decoding module 320 may obtain the decoded target video frame by processing the encoded data, the first reference information, and the second prompt data and / or the first prompt data through the first generative machine learning model.

[0064] In some embodiments, the decoding module 320 may be configured to obtain entropy decoding data by performing entropy decoding on the encoded data based on the pixel point grouping information; obtain inverse quantization data by performing inverse quantization on the entropy decoding data; and obtain the decoded target video frame corresponding to the original video frame by performing inverse transformation on the inverse quantization data.

[0065] In some embodiments, the decoding module 320 may further include a decoding acquisition unit 321, a decoding prompt unit 323, and a decoding unit 325.

[0066] The decoding acquisition unit 321 may be configured to acquire the encoded data and / or the first reference information. The first reference information may include feature information of at least one decoded video frame before a target video frame.

[0067] In some embodiments, the decoding acquisition unit 321 may be configured to acquire the prompt information. In some embodiments, the decoding acquisition unit 321 may be configured to acquire the second prompt data. In some embodiments, the decoding acquisition unit 321 may be configured to acquire the pixel point grouping information of the original video frame.

[0068] The decoding prompt unit 323 may be configured to determine the first prompt data input by the user. In some embodiments, the decoding prompt unit 323 may directly obtain customized first prompt data input by the user. In some embodiments, the decoding prompt unit 323 may determine the first prompt data based on user input information.

[0069] The decoding unit 325 may be configured to obtain the decoded target video frame based on the encoded data and the first reference information. In some embodiments, the decoding unit 325 may be configured to obtain the decoded target video frame by processing the encoded data and the first reference information through the first generative machine learning model. In some embodiments, the decoding unit 325 may be configured to obtain the entropy decoding data by performing entropy decoding on the encoded data based on the pixel point grouping information; obtain inverse quantization data by performing inverse quantization on the entropy decoding data; obtain first fusion data by fusing the inverse quantization data with the first reference information; and obtain the decoded target video frame by performing inverse transformation on the first fusion data.

[0070] It can be understood that the video encoding and decoding system 300 can be configured to implement the method in any embodiment of the present disclosure. More descriptions may be found in the detailed descriptions of the embodiments of the method below, which are not repeated here.

[0071] It should be noted that the above description of the video encoding and decoding system 300 is only for example and explanation, and does not limit the scope of application of the present disclosure. For those skilled in the art, various modifications and changes can be made to the video encoding and decoding system 300 under the guidance of the present disclosure. However, these modifications and changes are still within the scope of the present disclosure. For example, the encoding module 310 may include a user prompt unit and an information prompt unit. The user prompt unit may be configured to determine the second prompt data, and the information prompt unit may be configured to determine the prompt information. As another example, the decoding module 320 may include a user prompt unit and an information prompt unit. The user prompt unit may be configured to determine the first prompt data, and the information prompt unit may be configured to determine the prompt information. As another example, the encoding module 310 and the decoding module 320 may share an information prompt module for determining the prompt information. For example, the encoding module 310 may include an encoding prompt unit configured to determine the second prompt data, and the decoding module 320 may include a decoding prompt unit configured to determine the first prompt data. The encoding module 310 and the decoding module 320 may be externally connected to the information prompt module. The information prompt module may be configured to determine the prompt information.

[0072] As mentioned above, a video encoding and decoding process aims to achieve efficient transmission of video data. A video encoding process can compress the video data to obtain data with a relatively low bit rate for data transmission. At the decoding terminal, the compressed data is expected to restore the original video image with high fidelity. To this end, people have been committed to continuously improving the accuracy of video data encoding and decoding. Accordingly, one of the embodiments of the present application provides a video encoding and decoding method and system. The video encoding method obtains encoded data based on the original video frame and feature information of at least one decoded video frame before the original video frame. The video decoding method obtains a decoded target video frame based on the encoded data and feature information of at least one decoded video frame before a target video frame. The above method combines the video encoding and decoding technology based on the video frame generation technology. When a video frame to be encoded (e.g., the original video frame) and a corresponding reference video frame (e.g., the at least one decoded video frame before the original video frame / the target video frame) are obtained, feature information corresponding to the reference video frame is obtained, such that the quality of the bitstream and the decoded target video frame corresponding to the original video frame is enhanced using the reference video frame during the process of encoding and decoding, thereby improving the accuracy and efficiency of video frame encoding and decoding.

[0073] The above video encoding and decoding method is described in detail below with reference to specific embodiments.

[0074] FIG. 4 is a flowchart illustrating a video encoding method according to some embodiments of the present disclosure.

[0075] The video encoding method shown in FIG. 4 may be implemented by an encoding terminal. The encoding terminal refers to a terminal that implements an encoding operation. The encoding terminal may process a video frame (e.g., an original video frame) through an encoder (e.g., the processing device 110, the computing device 200, or the encoding module 310) , and convert visual information of the video frame into compact feature representation information (e.g., a bitstream) to facilitate the transmission of video data via a network. The compact feature representation information is machine readable.

[0076] In some embodiments, the encoding terminal may include a business executor (e.g., a service provider for providing an encoding service) , or a software and hardware system (e.g., the processing device 110, the computing device 200, or the encoding module 310) that specifically performs the encoding operation. The software and hardware system may include a second generative machine learning model and has functions related to the encoding operation such as fusion, transformation, quantization, entropy encoding, auxiliary transformation, auxiliary quantization, auxiliary entropy encoding, etc.

[0077] In some embodiments, a process 400 may be performed by the processing device 110, the computing device 200 (e.g., the processor 210) , or the video encoding and decoding system 300 (e.g., the encoding module 310) . As shown in FIG. 4, in some embodiments, the process 400 may include the following operations.

[0078] In 410, an original video frame and second reference information may be obtained. The second reference information includes feature information of at least one decoded video frame before the original video frame. In some embodiments, the operation 410 may be performed by the encoding module 310 (e.g., the encoding acquisition unit 311) .

[0079] In this embodiment, the original video frame refers to a current video frame to be encoded, or an image to be encoded. In some embodiments, the original video frame may include any video frame to be encoded in a video. The original video frame may include a first frame in a shooting timing sequence of the video, or another frame after the first frame.

[0080] In some embodiments, the encoding operation may be performed sequentially (e.g., encoding the first frame of the video first) according to a shooting time (or the shooting timing sequence) of the video frame, or in a reverse order (e.g., firstly encoding the last frame of the video) .

[0081] In some embodiments, a video frame that is before the original video frame in the shooting timing sequence may be a video frame that completes the encoding operation or a video frame that does not complete the encoding operation. For example, a video captured by the image acquisition device 140 contains a total of 24 frames (e.g., frams 1-24) . If the encoding operation for frames 1-10 is completed, frames 11-24 are video frames to be encoded, and the original video frame may be any one of the frames 11-24. As another example, a video contains a total of 24 frames (e.g., frams 1-24) . If the encoding operation for frames 1-7 and frames 12-15 is completed, frames 8-11 and frames 16-24 are video frames to be encoded, and the original video frame may be any one of the frames 8-11 and frames 16-24.

[0082] The at least one decoded video frame refers to a decoded video frame obtained by the decoding operation based on the encoded data, which is also referred to as a reconstructed video frame. The reconstructed video frame may correspond to the original video frame in a one-to-one manner. The at least one decoded video frame before the original video frame refers to a video frame that completes the encoding operation and a decoding operation before the original video frame is encoded.

[0083] The encoding operation is performed based on the second reference information, for example, the original video frame is encoded with reference to related information of the at least one decoded video frame before the original video frame. Accordingly, the at least one decoded video frame before the original video frame is also referred to as a reference video frame of the original video frame. The reference video frame corresponding to each original video frame may be the same, partially the same, or different.

[0084] In some embodiments, the reference video frame of the original video frame may include a decoded video frame of a video frame that is obtained before the original video frame in the shooting timing sequence, and / or a decoded video frame of a video frame that is obtained after the original video frame in the shooting timing sequence. For example, a video contains 1-24 video frames that are obtained in order in the shooting timing sequence, an 11th frame is a current video frame to be encoded (e.g., the original video frame) , and the decoding operation for frames 1-7 and frames 12-15 is completed (e.g., encoded data of the frames 1-7 and frames 12-15 is sent to the decoding terminal, respectively, and the decoding terminal obtains, based on the encoded data, decoded video frames corresponding to the frames 1-7 and frames 12-15, respectively) , a reference video frame corresponding to the 11th frame may be one or more of the decoded video frames corresponding to the frames 1-7 and frames 12-15. As another example, a video contains 1-24 video frames that are obtained in order in the shooting timing sequence, an 11th frame is a current video frame to be encoded (e.g., the original video frame) , and the decoding operation for frames 1-10 is completed, a reference video frame corresponding to the 11th frame may be one or more of the decoded video frames corresponding to the frames 1-10. As another example, a video contains 1-24 video frames that are obtained in order in the shooting timing sequence, an 11th frame is a current video frame to be encoded (e.g., the original video frame) , and the decoding operation for frames 12-24 is completed, decoded video frames before the original video frame may be one or more of decoded video frames corresponding to the frames 12-24.

[0085] In some embodiments, the reference video frame of the original video frame may include a decoded video frame of a video frame that is obtained before the original video frame in the shooting timing sequence and adjacent to the original video frame. For example, referring to the above example, the 11th frame is the current video frame to be encoded, and a decoded video frame of a video frame that is obtained before the 11th frame in the shooting timing sequence and adjacent to the 11th frame may be a decoded video frame corresponding to a 10th frame. In some embodiments, the reference video frame of the original video frame may include a plurality of decoded video frames (e.g., 2, 3, 5, 10 decoded video frames, etc. ) that are obtained before the original video frame in the shooting timing sequence, and the plurality of decoded video frames may include adjacent video frames and / or decoded video frames of adjacent video frames. If the original video frame does not correspond to a previously decoded video frame, the reference video frame of the original video frame may be an empty set.

[0086] In some embodiments, the at least one decoded video frame before the original video frame may be obtained from the decoding terminal. For example, after the encoding module 310 obtains the at least one decoded video frame obtained by the decoding operation from the decoding terminal, the at least one decoded video frame may be sorted according to a decoding time to obtain one or more decoded video frames before the original video frame. In some embodiments, the at least one decoded video frame before the original video frame may be directly obtained from the encoding terminal. In this case, the encoding terminal may be provided with an encoding module and a decoding module simultaneously. The encoding module may be configured to encode the original video frame, and the decoding module may be configured to decode the encoded data or the bitstream. In some embodiments, the at least one decoded video frame before the original video frame may be obtained from a storage device. For example, after the decoding terminal obtains the at least one decoded video frame based on the encoded data / the bitstream, the at least one decoded video frame may be sent to the storage device 120 for storage, such that the storage device 120 stores a decoded video frame set containing a plurality of decoded video frames. The encoding module 310 may obtain the decoded video frame set from the storage device 120 to obtain one or more decoded video frames before the original video frame, or directly obtain the one or more decoded video frames before the original video frame from the storage device 120.

[0087] The feature information of the at least one decoded video frame may include at least one of following information: the at least one decoded video frame, feature data of the at least one decoded video frame, and some intermediate data of the at least one decoded video frame during the decoding process.

[0088] The feature data of the at least one decoded video frame may include a feature vector of the at least one decoded video frame. In some embodiments, the encoding module 310 may obtain the feature vector of the at least one decoded video frame by performing feature extraction on the at least one decoded video frame using a preset algorithm (e.g., deep neural networks (DNNs) , convolutional neural networks (CNNs) , recurrent neural networks (RNNs) , long short-term memory (LSTM) , and a generative pre-training Transformer model, etc. ) .

[0089] The intermediate data of the at least one decoded video frame during the decoding process may include a feature vector corresponding to data of any stage of the decoding operation. For example, the intermediate data of the at least one decoded video frame may include a feature vector corresponding to entropy decoding data, inverse quantization data, auxiliary entropy decoding information, auxiliary inverse quantization information, or auxiliary inverse transformation information of the at least one decoded video frame.

[0090] In 420, encoded data may be obtained based on the original video frame and the second reference information. In some embodiments, the operation 420 may be performed by the encoding module 310 (e.g., the encoding unit 315) .

[0091] The encoded data refers to an encoded data stream, which may include the original video frame, or related information of the original video frame and the reference video frame. For example, the encoded data may include a binary data sequence reflecting an image feature of the original video frame.

[0092] In some embodiments, the encoded data may be presented in a form of a bitstream. The bitstream refers to an encoded data stream, which is a sequence of binary data. By encoding the original video frame or the image into the bitstream, the storage and transmission of the video or the image can be facilitated.

[0093] In some embodiments, the encoding terminal may obtain the encoded data by encoding the original video frame and the second reference information. For example, the encoding module 310 may fuse the original video frame and the reference video frame corresponding to the original video frame (e.g., fusing by splicing or a channel superposition) , and obtain a fusion feature vector by performing feature extraction on the original video frame and the reference video frame after fusion, and obtain the encoded data by encoding the fusion feature vector. As another example, the encoding module 310 may perform feature extraction the original video frame and the reference video frame corresponding to the original video frame, respectively, obtain a fusion feature vector by fusing a feature vector of the original video frame and a feature vector of the reference video frame, and obtain the encoded data by encoding the fusion feature vector.

[0094] In some embodiments, the encoding terminal may obtain the encoded data by process the original video frame and the second reference information through a second generative machine learning model. For example, the second generative machine learning model may determine difference information or residual information between the original video frame and the reference video frame based on the original video frame and the second reference information, and obtain the encoded data of the original video frame by compressing and encoding the original video frame based on the difference information or the residual information.

[0095] A generative machine learning model is a model for data generation using probabilistic machine learning model. The generative machine learning model regards data generation as a process of extracting samples from a prior distribution and generate new data accordingly instead of simply classifying existing data.

[0096] The second generative machine learning model refers to a generative machine learning model provided at the encoding terminal. In some embodiments, the second generative machine learning model may include at least one of a multi-head self-attention model (e.g., a multi-head Transformer) , a vector quantization generative adversarial network (e.g., VQ-GAN) , and a diffusion model.

[0097] Taking the second generative machine learning model being a generative model based on a multi-head Transformer as an example, the second generative machine learning model is composed of a plurality of identical structural layers, and each of the plurality of identical structural layers includes a multi-head attention layer, a feed forward network layer, and a layer normalization layer connected in sequence. The multi-head attention layer allows the second generative machine learning model to consider a plurality of different features or information points (e.g., the feature information of the original video frame, the feature information of the reference video frame, etc. ) when processing input data (e.g., the original video frame and the second reference information) , thereby improving the understanding ability of the second generative machine learning model. The layer normalization layer can connect an output of the multi-head attention layer and the feed forward network layer with an input data residual and perform normalization, thereby stabilizing the training process and improving the performance of the second generative machine learning model.

[0098] In some embodiments, the second generative machine learning model may be obtained by training an initial generative model based on first sample data. The initial generative model may be obtained based on a language model. For example, the language model may include but is not limited to DNNs, CNNs, RNNs, LSTM, and a generative pre-training Transformer model, etc., which is not limited on the specific construction and deployment of the language model. The first sample data may include a sample original video frame and sample second reference information.

[0099] In some embodiments, the input data of the second generative machine learning model may include the original video frame and the second reference information. For example, the original video frame and the second reference information are input into the second generative machine learning model as the input data, and the second generative machine learning model performs feature extraction, compression, vector quantization and other processing on the original video frame and the second reference information to obtain the encoded data of the original video frame and output the encoded data of the original video frame.

[0100] In some embodiments, the input data of the second generative machine learning model may include fusion data of the original video frame and the second reference information. For example, the encoding module 310 may obtain second fusion data by performing operations such as splicing, addition, or multiplication on the original video frame and the reference video frame (in this case, the second reference information includes the at least one decoded video frame before the original video frame) , and input the second fusion data into the second generative machine learning model, and the second generative machine learning model may output the encoded data corresponding to the original video frame by analyzing and processing the second fusion data.

[0101] In some embodiments, after the first sample data is input into the initial generative model as the input data, parameters of each module in the initial generative model (e.g., a plurality of structural layers consisting of the multi-head attention layer, the feed forward network layer, and the layer normalization layer) may be adjusted such that a loss function converges or a value of the loss function is less than a first loss threshold, thereby obtaining a trained second generative machine learning model. For example, a first loss function D + λR is constructed, such that a value of the first loss function D + λR is less than the first loss threshold, and the trained second generative machine learning model is obtained, where D denotes a difference between the sample original video frame and at least one sample decoded video frame. The at least one sample decoded video frame refers to a reconstructed video frame obtained by decoding based on sample encoded data. A set of sample original video frames and sample decoded video frames may correspond to the same video frame. The sample encoded data is a bitstream obtained by the initial generative model based on the sample original video frame and the sample second reference information during a model training process. R denotes a data volume (the unit may be Byte) of the sample encoded data output during the model training process, also referred to as a code rate (in this case, the second generative machine learning model determines the encoded data based on the original video frame and the second reference information) . In some embodiments, R may denote a data volume of a sample bitstream corresponding to the sample encoded data and prompt information (in this case, the second generative machine learning model may determine a bitstream containing the prompt information corresponding to the original video frame based on the original video frame and the second reference information) .

[0102] In some embodiments, the second generative machine learning model and a first generative machine learning model provided at the decoding terminal may be jointly trained. More descriptions regarding joint training may be found in the present disclosure below.

[0103] In some embodiments, the encoding terminal may determine encoding prompt data matching the original video frame. The encoding prompt data may be sent to the decoding terminal together with the encoded data. The encoding prompt data may be configured to indicate a feature that needs to be noticed during the encoding process of the original video frame.

[0104] In some embodiments, the encoding prompt data may include fixed encoding prompt information and supplementary encoding prompt information. The fixed encoding prompt information refers to prompt information generated based on the original video frame and the second reference information, or based on the original video frame, which is also referred to as prompt information. The supplementary encoding prompt information refers to prompt data obtained based on input information of a user of the encoding terminal (e.g., a user of the encoding terminal) , which is also referred to as second prompt data.

[0105] The prompt information reflects intra-frame information (e.g., pixel value information, segmented semantic information, classified semantic information, etc. ) and / or inter-frame information (e.g., timing information, scaling information, etc. ) of the original video frame and / or the reference video frame.

[0106] In some embodiments, the prompt information reflects at least one of the following information related to the original video frame: pixel value information, scaling information, timing information, segmented semantic information, and classified semantic information.

[0107] The pixel value information refers to related information of pixel points contained in the video frame. For example, the pixel value information may include at least one of a texture feature, a brightness feature, a color feature, etc., of the original video frame, or at least one of a texture feature, a brightness feature, a color feature, etc., of the original video frame and the reference video frame corresponding to the original video frame.

[0108] The segmented semantic information refers to semantic information of different regions after the video frame is segmented into a plurality of regions (e.g., the original video frame is segmented into n*n sub-regions) . For example, the segmented semantic information may include the semantic information of each sub-region in the original video frame, or the semantic information of each sub-region in the original video frame and the reference video frame. The semantic information reflects the meaning and function of the image content, such as a type of an object, an object attribute (e.g., color, and size) , a scenario to which the object belongs, or a correlation of objects (e.g., a dog stands at the door, a cat sits on a chair, a person drives a car, etc. ) , and other information contained in the image.

[0109] The classified semantic information refers to semantic information of regions with similar contents in the video frame. For example, if the original video frame is divided into n*n sub-regions, and the n*n sub-regions are clustered into a sky region, a seawater region, a human region, and an animal region according to the contents contained therein, the classified semantic information may include semantic information of the sky region, semantic information of the seawater region, semantic information of the human region, and semantic information of the animal region. Similar to other information, the classified semantic information may include the classified semantic information of the original video frame, or the classified semantic information of the original video frame and the reference video frame.

[0110] The scaling information refers to a scaling ratio of the original video frame with respect to the reference video frame. For example, the scaling information is that the original video frame is reduced by 30%compared to the decoded video frame at a previous moment.

[0111] The timing information refers to a shooting timing sequence relationship between the original video frame and the corresponding reference video frame, such as in a video, the original video frame being obtained before or after the reference video frame according to the shooting timing sequence, and / or a count of frames (e.g., adjacent (e.g., the count of frames is 0) , 1 frame, 2 frames, etc. ) between the original video frame and the reference video frame.

[0112] In some embodiments, the encoding terminal may obtain a prompt template corresponding to an encoding prompt; generate an encoding prompt of the original video frame using the prompt template; and generate prompt information of the original video frame based on the encoding prompt and the original video frame. By generating the prompt information based on the prompt template and the original video frame, the features of the original video frame that need to be noticed during the encoding process are clarified, thereby improving the encoding accuracy.

[0113] In some embodiments, the encoding terminal may determine an inter-frame difference between the original video frame and the reference video frame based on the original video frame and the second reference information; generate the encoding prompt of the original video frame based on the inter-frame difference; and generate prompt information of the original video frame based on the encoding prompt and the inter-frame difference. By determining the prompt information based on the original video frame and the second reference information, the adaptability of the prompt information to the corresponding original video frame can be improved.

[0114] In some embodiments, the encoding terminal may generate the prompt information based on the original video frame and the second reference information through the second generative machine learning model. For example, the encoding terminal may generate a model encoding prompt for the original video frame using the encoding prompt unit 313, and obtain the prompt information of the original video frame by inputting the original video frame, the second reference information, and the model encoding prompt into the second generative machine learning model.

[0115] The encoding prompt unit may be configured to introduce an additional prior vector for the original video frame to guide the learning of a content feature related to the prior vector. In some embodiments, the encoding prompt unit 313 may be coupled with the encoding unit 315, and the encoding prompt unit 313 may generate the model encoding prompt for the original video frame after obtaining the original video frame and the second reference information corresponding to the original video frame, thereby introducing the encoding prompt information into the encoding process to achieve a higher accuracy of the encoding process.

[0116] An input of the encoding prompt unit 313 may include digital signal information of any mode such as a voice, a text, or a picture. In some embodiments, the input of the encoding prompt unit may include a feature vector corresponding to the digital signal information such as the voice, the text, or the picture, which is not limited in the present disclosure.

[0117] In some embodiments, the encoding prompt unit 313 may generate a fixed model encoding prompt for the original video frame based on a fixed prompt template, or may generate an adaptive model encoding prompt for the original video frame based on the feature information of the original video frame and the reference video frame.

[0118] By obtaining the prompt information of the original video frame based on the second reference information, the obtained prompt information can be comprehensive.

[0119] In some embodiments, the encoding terminal may generate the prompt information based on the original video frame through the second generative machine learning model. For example, the encoding module 310 may generate the model encoding prompt for the original video frame, at least input the original video frame and the model encoding prompt into the second generative machine learning model, perform feature extraction on the input original video frame based on the model encoding prompt using the second generative machine learning model, and fuse information contained in the model encoding prompt to obtain the prompt information of the original video frame.

[0120] In some embodiments, the prompt information of the original video frame may include intermediate output data of the second generative machine learning model. For example, referring to the above embodiment, if the second generative machine learning model is constructed based on a Transformer network, the encoding module 310 may use an output of one multi-head attention layer of the second generative machine learning model as the prompt information of the original video frame.

[0121] In some embodiments, the encoding terminal may generate the prompt information based on the original video frame and the second reference information through a feature extraction model. In some embodiments, the encoding terminal may generate the prompt information based on the original video frame through the feature extraction model. For example, feature extraction model may include but is not limited to DNNs, CNNs, RNNs, LSTM, a generative pre-training Transformer model, etc.

[0122] In some embodiments, the prompt information of the original video frame may include prompt information of the reference video frame. For example, the prompt information may further include pixel value information, scaling information, timing information, segmented semantic information, and classified semantic information of the reference video frame.

[0123] With the continuous development of AI technology, a machine learning model may adjust an image / a video based on a user instruction to obtain a processing result expected by a user. For example, the machine learning model may receive the user instruction (e.g., a prompt word text, voice) and data to be processed, and add an additional content to the data to be processed based on the user instruction, thereby outputting a target image / video with the additional content added to the data to be processed. However, during image / video processing, the user is required to transmit the data to be processed to a service platform or an online machine learning model that can provide an image processing service via the network, such that the service platform or the online machine learning model can process the data to be processed based on the user instruction.

[0124] Taking the image as an example, during the transmission of the data to be processed, it generally needs to first encode the image to be processed (e.g., the original video frame) to obtain a bitstream, and send the bitstream of the image to be processed to the service platform or the online machine learning model. The service platform or the online machine learning model needs to first decode the bitstream to obtain a decoded image (e.g., a decoded target video frame) corresponding to the image to be processed, and then adjust the decoded image based on the user instruction. An adjusted decoded image can be encoded and transmitted to a terminal device used by the user via the network again. The terminal device also decode firstly to obtain the adjusted image. In practical applications, the user may also transmit the adjusted image to other terminal devices. Accordingly, the entire image processing is relatively complicated as relating to encoding, decoding, and other operations, which becomes complicated especially when the data to be processed is a video containing a plurality of video frames.

[0125] The second prompt data may be configured to indicate an adjustment of an image content and / or an image size of the original video frame. For example, the second prompt data may include adding a subject (e.g., “adding a puppy to the original video frame” ) , a text (e.g., “adding a title to the original video frame” ) , etc., to an image (e.g., the original video frame) . As another example, the second prompt data may include removing a certain content (e.g., a text such as “removing sensitive information from the original video frame” ) from an image (e.g., the original video frame) . The sensitive information may include a sensitive word, a sensitive object, etc., which may be preset by the user. As another example, the second prompt data may include cropping (e.g., “reducing the size of the original video frame by half” , “reducing the size of the original video frame by one third” , “cut off the region outside the people in the original video frame” , etc. ) an image (e.g., the original video frame) to change a size (e.g., an aspect ratio) of the image or retain a target region of the image.

[0126] The second prompt data may include information understandable by human beings, including but not limited to a picture, a voice, a text, etc. The second prompt data may also include information understandable by machines, including but not limited to a feature vector.

[0127] In some embodiments, the encoding terminal may obtain the second prompt data input by the user. For example, the user may input customized second prompt data through the encoding prompt unit 313. In some embodiments, the encoding terminal may determine the second prompt data based on user input information. For example, the user may input guidance information through the terminal device 150, and the encoding prompt unit 313 may determine the second prompt data based on the guidance information input by the user.

[0128] In some embodiments, the encoding prompt unit 313 may obtain a data stream corresponding to the second prompt data using the second generative machine learning model. For example, the guidance information input by the user, or the customized second prompt data directly input by the user may be input into the second generative machine learning model together with the original video frame and the second reference information to obtain the encoded data and an embedding vector corresponding to the second prompt data.

[0129] In some embodiments, the encoding prompt unit 313 may obtain the data stream corresponding to the second prompt data using a first embedding layer network structure. The first embedding layer network structure may include but is not limited to DNNs, CNNs, and RNNs. For example, the encoding prompt unit 313 may be provided with the first embedding layer network structure, and the guidance information input by the user, or the customized second prompt data directly input by the user may be input into the first embedding layer network structure, and the first embedding layer network structure may obtain the corresponding embedding vector by processing the user input information (e.g., the guidance information, or the customized second prompt data) . The first embedding layer network structure may be a part of the second generative machine learning model, or independent of the second generative machine learning model.

[0130] In some embodiments, the first embedding layer network structure and the second generative machine learning model may be jointly trained. More descriptions regarding the joint training may be found in the present disclosure below.

[0131] By determining the second prompt data input by the user, it facilitates the user of the encoding terminal to guide information generation during the encoding process, thereby adjusting the video bitstream, improving the matching degree between the encoding and decoding process and the intention of the user, increasing the personalization and interest of video frame encoding and decoding, and improving the encoding and decoding efficiency.

[0132] In some embodiments, the prompt information and / or the second prompt data may be determined by an external module of the encoding terminal. For example, the encoding terminal may be externally provided with a user prompt module, the user prompt module may be configured to determine the prompt information and / or the second prompt data, and send the determined prompt information and / or second prompt data to the encoding terminal, and the encoding terminal may convert the prompt information and / or the second prompt data into a bitstream and send the bitstream to the decoding terminal. As another example, the user prompt module may be connected with the decoding terminal, and after the user prompt module determines the prompt information and / or the second prompt data, the user prompt module may convert the prompt information and / or the second prompt data into the bitstream and directly send the bitstream to the decoding terminal.

[0133] If the prompt information and / or the second prompt data are directly determined by the encoding terminal, joint training and parameter adjustment may be performed on a module for determining the prompt information, a module for determining the second prompt data, and a module for determining the encoded data. For example, joint training may be performed on the second generative machine learning model for determining the prompt information and the encoded data and the embedding network structure for determining the second prompt data.

[0134] In some embodiments, the encoding terminal may obtain pixel point grouping information of the original video frame by performing pixel point grouping on the original video frame based on an image content of the original video frame; obtain second fusion data by fusing the original video frame and the second reference information; obtain first transformation data by transforming the second fusion data; obtain first quantization data by quantizing the first transformation data; and obtain the encoded data by performing entropy encoding on the first quantization data based on the pixel point grouping information. The pixel point grouping information may be sent to the decoding terminal together with the encoded data. More descriptions regarding determining the encoded data based on the pixel point grouping information may be found in FIG. 7 and the related descriptions thereof.

[0135] In 430, the encoded data may be sent to a decoding terminal. In some embodiments, the operation 430 may be performed by the encoding module 310.

[0136] The decoding terminal refers to a terminal that implements a decoding operation. In some embodiments, the decoding terminal may include a business executor (e.g., a service provider for providing a decoding business) , or a software and hardware system (e.g., the processing device 110, the computing device 200, and the decoding module 320) that specifically performs the decoding operation. The software and hardware system may include a first generative machine learning model, and has functions related to the decoding operation such as fusion, inverse transformation, inverse quantization, entropy decoding, inverse auxiliary transformation, inverse auxiliary quantization, auxiliary entropy decoding, etc.

[0137] The encoding terminal and the decoding terminal are respectively provided with the generative machine learning model (e.g., the second generative machine learning model, and the first generative machine learning model) , such that both the encoding terminal and the decoding terminal can obtain a supplementary prompt (e.g., the second prompt data, and first prompt data) given by the user (e.g., the user of the encoding terminal or the decoding terminal) , thereby determining the guidance information input by the user and improving the matching degree between the encoding and decoding process and the intention of the user.

[0138] In some embodiments, after the encoding terminal completes encoding of the original video frame, the encoded data may be transmitted to the decoding terminal via the network.

[0139] In some embodiments, the prompt information may be sent to the decoding terminal together with the encoded data. In some embodiments, the fusion data of the prompt information and the encoded data may be sent to the decoding terminal. For example, the encoding module 310 may encode the prompt information, and fuse encoded prompt information with the encoded data to obtain a bitstream fused with the feature information of the original video frame and prompt information, and send the obtained bitstream to the decoding terminal. In some embodiments, the encoding terminal may encode the original video frame, the second reference information, and the prompt information to obtain an encoded bitstream fused with the feature information of the original video frame and the reference video frame, and prompt information, and send the encoded bitstream to the decoding terminal. For example, the encoding module 310 may obtain the encoded bitstream by inputting the original video frame, the second reference information, and the prompt information into the second generative machine learning model. In some embodiments, the encoding terminal may obtain the encoded bitstream by performing weighted fusion on the original video frame, the second reference information, and the prompt information, and inputting the weighted fusion of the original video frame, the second reference information, and the prompt information into the second generative machine learning model.

[0140] By encoding and fusing the feature information and the prompt information, the obtained bitstream can more accurately describe the feature information of the original video frame, such that the encoding accuracy is improved, and a more accurate decoded target video frame (e.g., a decoded target video frame that is consistent or substantially consistent with the original video frame) can be obtained during decoding.

[0141] In some embodiments, the second prompt data may be sent to the decoding terminal together with the encoded data. In some embodiments, the second prompt data may be sent to the decoding terminal together with the encoded data and the prompt information.

[0142] In some embodiments, the encoding terminal may adjust the bitstream fused with the feature information of the original video frame and the reference video frame, and prompt information based on the second prompt data, and send an adjusted bitstream to the decoding terminal. In some embodiments, the encoding terminal may obtain the adjusted bitstream by adjusting the encoded data of the original video frame and the prompt information based on the second prompt data through the second generative machine learning model. For example, after the encoding module 310 generates the bitstream fused with the feature information of the original video frame and the reference video frame and prompt information based on the original video frame, the second reference information, and the prompt information through the second generative machine learning model, the bitstream may be further adjusted based on the second prompt data to obtain the adjusted bitstream. In this case, in the process of training the second generative machine learning model, a second loss function D1-D2+ λR may be constructed. A value of the second loss function may be adjusted to be less than the second loss threshold by parameter adjustment, thereby obtaining a trained second generative machine learning model. Where D1 denotes a difference between a sample original video frame and a sample decoded target video frame. The sample decoded target video frame is a reconstructed video frame obtained by decoding based on sample encoded data during the model training process, and the sample encoded data is a bitstream obtained by encoding based on the sample original video frame, sample second reference information, the prompt information, and the second prompt data during the model training process. D2 denotes a difference between the sample decoded target video frame and a user indication content corresponding to the second prompt data, i.e., a difference between the sample decoded target video frame and a video frame expected by the user. R denotes a data volume of the sample decoded data output during the model training process, or a data volume of a sample bitstream corresponding to the sample encoded data and the prompt information.

[0143] For example, the encoding prompt unit 313 may include the first embedding layer network structure, the first embedding layer network structure may be connected to the second generative machine learning model, and the second prompt data may include a text “a puppy appears in the image” input by the user. The first embedding layer network structure may learn a potential feature of the second prompt data and convert the potential feature of the second prompt data into a corresponding embedding vector, and input the embedding vector into the second generative machine learning model. The second generative machine learning model may analyze the original video frame and the reference video frame. If targets that can be detected in the original video frame and the reference video frame do not include the puppy, the feature information of the original video frame may be adjusted such that adjusted feature information includes feature information corresponding to the puppy; and the prompt information may be adjusted such that adjusted prompt information includes semantic information related to the puppy. Similarly, if the targets that can be detected in the original video frame and the reference video frame include a plurality of puppies, the feature information of the original video frame may be adjusted such that the adjusted feature information only includes feature information corresponding to one puppy; and the prompt information may be adjusted such that the adjusted prompt information only includes semantic information related to one puppy. That is to say, the content in the bitstream corresponding to the original video frame may be added, deleted, or modified at the encoding terminal based on the second prompt data, thereby improving the matching degree between the encoding and decoding process and the encoding intention of the user, and achieving a higher degree of freedom during the encoding process.

[0144] By encoding the video frame based on the second prompt data related to the user input information, it facilitates the user of the encoding terminal to guide information generation during the encoding process, thereby adjusting the video bitstream, improving the matching degree between the encoding and decoding process and the intention of the user, increasing the personalization and interest of video frame encoding and decoding, and improving the encoding and decoding efficiency.

[0145] After the encoded data, the prompt information, and the second prompt data are sent to the decoding terminal, the decoding terminal may perform decoding to obtain the decoded target video frame. More descriptions regarding decoding may be found in FIG. 5 and the related description thereof.

[0146] FIG. 5 is a flowchart illustrating an exemplary video decoding method according to some embodiments of the present disclosure. The video decoding method shown in FIG. 5 is implemented by a decoding terminal.

[0147] In some embodiments, a process 500 may be performed by the processing device 110, the computing device 200 (e.g., the processor 210) , or the video encoding and decoding system 300 (e.g., the decoding module 320) . As shown in FIG. 5, in some embodiments, the process 500 may include the following operations.

[0148] In 510, encoded data and first reference information may be obtained from an encoding terminal. In some embodiments, the operation 510 may be performed by the decoding module 320 (e.g., the decoding acquisition unit 321) .

[0149] The first reference information may include feature information of at least one decoded video frame before a target video frame.

[0150] The target video frame refers to a video frame that needs to be decoded currently, which is correlated with an original video frame corresponding to the encoded data. For example, if the original video frame corresponding to the encoded data obtained from the encoding terminal is a 1st frame, the target video frame is the 1st frame of a video. As another example, if the original video frame corresponding to the encoded data obtained from the encoding terminal is a 14th frame, the target video frame is a 14th frame of the video.

[0151] Referring to the above, the feature information of the at least one decoded video frame may include at least one of following information: the at least one decoded video frame, feature data of the at least one decoded video frame, and some intermediate data of the at least one decoded video frame during the decoding process.

[0152] Since the target video frame corresponds to the original video frame, the first reference information essentially includes the feature information of the at least one decoded video frame before the original video frame. That is, the first reference information and the second reference information may contain the same content, and “first” and “second’ may be used to distinguish the information of the encoding terminal and the decoding terminal.

[0153] In some embodiments, the decoding terminal may use the second reference information obtained from the encoding terminal as the first reference information. In some embodiments, the decoding terminal may obtain the first reference information from a storage device (e.g., the storage device 120) . For example, the decoding acquisition unit 321 may obtain one or more decoded video frames before a current target video frame in terms of shooting timing sequence from the storage device 120, use the one or more decoded video frames as reference video frames, and determine the first reference information based on the reference video frames. For example, the decoding terminal may use the one or more decoded video frames as the first reference information, or extract the feature data of each of the one or more decoded video frames as the first reference information, or use the certain intermediate data of each of the reference video frames during the decoding process as the first reference information. More descriptions regarding the first reference information may be found in the related descriptions of the second reference information in FIG. 4, which are not repeated here.

[0154] In 520, a decoded target video frame may be obtained based on the encoded data and the first reference information. In some embodiments, the operation 520 may be performed by the decoding module 320 (e.g., the decoding unit 325) .

[0155] The decoding terminal may obtain decoding information by decoding the encoded data based on the first reference information, and obtain the decoded target video frame (i.e., a reconstructed video frame) based on the decoding information by reconstruction.

[0156] By obtaining the decoded target video frame based on the first reference information and introducing the information contained in the reference video frame in the encoding and decoding process, information with a higher correlation with the encoding and decoding process can be adaptively obtained, thereby improving encoding and decoding efficiency and accuracy.

[0157] In some embodiments, the decoding terminal may obtain the decoded target video frame by processing the encoded data and the first reference information through a first generative machine learning model. For example, the encoded data and the first reference information may be input into the first generative machine learning model, and the first generative machine learning model may decode the encoded data to obtain decoding feature information; the first reference information, the decoding feature information, and the prompt information may be fused to obtain fusion decoding information; and the fusion decoding information may be reconstructed to obtain the decoded target video frame. A fusion mode may include but is not limited to splicing and channel superposition.

[0158] The first generative machine learning model refers to a generative machine learning model provided at the decoding terminal.

[0159] The first generative machine learning model and the second generative machine learning model may have the same network structure. For example, both the first generative machine learning model and the second generative machine learning model may be a multi-head self-attention model, a VQ-GAN, or a diffusion model. Accordingly, a training process of the first generative machine learning model may be similar to the training process of the second generative machine learning model. For example, an initial generative model may be trained based on second sample data to obtain the second sample data including sample encoded data and sample first reference information, after the second sample data is input into the initial generative model as input data, parameter adjustment is performed on the initial generative model based on a third loss function such that the third loss function converges or a value of the third loss function is less than the third loss threshold, thereby obtaining a trained first generative machine learning model. More descriptions regarding the first generative machine learning model may be found in the related descriptions of the second generative machine learning model in FIG. 4, which are not be repeated here.

[0160] In some embodiments, the first generative machine learning model and the second generative machine learning model may be obtained by training separately or jointly. For example, taking a generative model based on a multi-head Transformer as an example, a joint initial generative model may include an input layer, an encoder layer (consisting of a plurality of multi-head attention layers, feed forward network layers, and layer normalization layers) , a decoder layer (consisting of a plurality of masked multi-head attention layers, feed forward network layers, and layer normalization layers) , and a reconstruction layer. The joint training of the first generative machine learning model and the second generative machine learning model may include inputting training data containing sample original video frame and sample reference information into the joint initial generative model, the sample original video frame and the sample reference information entering the encoder layer through the input layer, and the encoder layer outputting sample encoded data after processing; the decoder layer processing the sample encoded data and the sample reference information output by the encoder layer and inputting the sample encoded data and the sample reference information into the reconstruction layer, and the reconstruction layer outputting sample reconstructed video frame after processing. In some embodiments, during the joint training, joint parameter adjustment may be performed by constructing a first joint loss function such that the first joint loss function converges or a value of the first joint loss function is less than a first joint loss threshold, thereby obtaining a trained joint generative model. For example, the constructed first joint loss function is D0 + λR0, and parameters of the encoder layer, the decoder layer, and the reconstruction layer are jointly adjusted such that a value of the D0 +λR0 is less than the first joint loss threshold. Where D0 denotes a difference between the sample original video frame and the sample reconstructed video frame; R0 denotes a data volume of the sample encoded data. The trained joint generative model deployed in part (e.g., an encoding part including the input layer and the encoder layer) or in full at the encoding terminal for encoding may be the second generative machine learning model. Similarly, the trained joint generative model deployed in part (e.g., the input layer, the decoder layer, and the reconstruction layer, or a decoding part including the decoder layer and the reconstruction layer) or in full at the decoding terminal for decoding may be the first generative machine learning model.

[0161] Referring to the above, encoding prompt data (e.g., the prompt information, and second prompt data) may be sent to the decoding terminal together with the encoded data. That is, the decoding terminal may obtain the prompt information and / or the second prompt data from the encoding terminal in addition to the encoded data and the first reference information. In this case, the decoding terminal may obtain the decoded target video frame based on the encoded data, the first reference information, and the prompt information and / or the second prompt data.

[0162] In some embodiments, the decoding terminal may determine difference information between feature information of the current target video frame and the feature information of the reference video frame based on the encoded data, the first reference information, and the prompt information, and reconstruct the reference video frame and the difference information to obtain the decoded target video frame. In some embodiments, the decoding terminal may obtain the fusion decoding information based on the encoded data, the first reference information, and the prompt information, and reconstruct the fusion decoding information to obtain the decoded target video frame.

[0163] For example, if fusion data of the prompt information and the encoded data is sent to the decoding terminal, the decoding module 320 may decode the fusion data to obtain the decoding feature information corresponding to the encoded data and decoding prompt information corresponding to the prompt information, and fuse the first reference information, the decoding feature information, and the decoding prompt information in a preset fusion mode to obtain the fusion decoding information, such that the fusion decoding information includes reference information in more dimensions. Furthermore, the decoding module 320 may reconstruct the fusion decoding information to obtain the decoded target video frame, such that the reconstruction process of the decoded target video frame is jointly guided by the reference video frame and the prompt information, and the fusion decoding information fused with multiple information is reconstructed into the decoded target video frame, thereby improving the accuracy of the decoded target video frame.

[0164] In some embodiments, the decoding terminal may obtain the decoded target video frame by processing the encoded data, the first reference information, and the prompt information through the first generative machine learning model. For example, the decoding module 320 may input the encoded data, the first reference information, and the prompt information into the first generative machine learning model, and the first generative machine learning model may analyze and reconstruct the encoded data, the first reference information, and the prompt information to output the decoded target video frame.

[0165] In some embodiments, the decoding terminal may adjust the decoding feature information and the decoding prompt information based on the second prompt data to obtain adjusted decoding feature information and adjusted decoding prompt information adapted to the second prompt data, and perform decoding reconstruction to obtain the decoded target video frame.

[0166] In some embodiments, the decoding terminal may obtain the decoded target video frame by processing the encoded data, the first reference information, the prompt information, and the second prompt data through the first generative machine learning model.

[0167] In some embodiments, the decoding terminal may determine first prompt data input by the user. The first prompt data is correlated with guidance information input by the user through the decoding terminal.

[0168] The first prompt data is configured to indicate an adjustment of an image content and / or an image size of the decoded target video frame during a process of obtaining the decoded target video frame.

[0169] Similar to the second prompt data, the first prompt data may include information understandable by human beings, including but not limited to a picture, a voice, a text etc. The first prompt data may also include information understandable by machines, including but not limited to a feature vector. For example, the first prompt data may include adding a subject, a text, etc., to an image (e.g., the original video frame) ; or removing a certain content from an image (e.g., the original video frame) ; or cropping an image (e.g., the original video frame) to change a size of the image or retain a target region of the image.

[0170] In some embodiments, the decoding terminal may obtain the first prompt data (also referred to as decoding supplementary information) input by the user. For example, the user may input customized first prompt data through the decoding prompt unit 323. In some embodiments, the encoding terminal may determine the first prompt data based on the user input information. For example, the user may input the guidance information through the terminal device 150, and the decoding prompt unit 323 may determine the first prompt data based on the guidance information input by the user.

[0171] By obtaining the first prompt data, it facilitates the user of the decoding terminal to guide information generation during the decoding process, thereby adjusting the decoding information, and improving the matching degree between the decoding process and the intention of the user.

[0172] In some embodiments, the decoding prompt unit 323 may obtain information corresponding to the first prompt data using the first generative machine learning model. For example, the guidance information input by the user, or the customized first prompt data directly input by the user may be input into the first generative machine learning model to obtain an embedding vector corresponding to the first prompt data. A process of determining the first prompt data may be similar to the process of determining the second prompt data. More descriptions may be found in the related descriptions of the second prompt data in FIG. 4, which are not repeated here.

[0173] In some embodiments, the decoding terminal may adjust the decoding feature information and the decoding prompt information based on the first prompt data to obtain adjusted decoding feature information and adjusted decoding prompt information adapted to the first prompt data, and then perform decoding reconstruction to obtain the decoded target video frame.

[0174] In some embodiments, the decoding terminal may obtain the decoded target video frame by processing the encoded data, the first reference information, the prompt information, and the first prompt data through the first generative machine learning model. For example, the encoded data, the first reference information, the prompt information, and the first prompt data may be input into the first generative machine learning model, and the first generative machine learning model may decode the encoded data and the prompt information to obtain the decoding feature information and the decoding prompt information; decoding guidance information corresponding to the first prompt data may be determined by analyzing the first prompt data; the decoding feature information and the decoding prompt information may be adjusted based on the decoding guidance information; the fusion decoding information may be obtained by fusing the first reference information (e.g., the feature information of the reference video frame) , the adjusted decoding feature information and the adjusted decoding prompt information; and the decoded target video frame may be obtained reconstructing the fusion decoding information.

[0175] For example, the decoding prompt unit 323 may include a second embedding layer network structure. The second embedding layer network structure may be connected to the first generative machine learning model. The first prompt data may include a text “a dog with a tilted head” input by the user. The text “a dog with a tilted head” input by the user may learn a potential feature through the second embedding layer network structure and the potential feature may be converted into a corresponding embedding vector, and the embedding vector may be input into the first generative machine learning model. The first generative machine learning model may adjust the decoding feature information such that the adjusted decoding feature information includes feature information corresponding to the dog with a tilted head, and adjust the decoding prompt information such that the adjusted decoding prompt information includes classification information and description information related to the dog. Similarly, if the decoding feature information includes information related to the dog but a posture of the dog does not match (i.e., in this embodiment, the dog's head is not tilted) the potential feature corresponding to the first prompt data, the decoding feature information may be adjusted such that the adjusted decoding feature information includes the feature information corresponding to the dog with a tilted head, and the decoding prompt information may be adjusted such that the adjusted decoding prompt information includes description information related to the dog with a tilted head. That is to say, the decoding terminal can add, delete, or modify the content in a bitstream corresponding to the decoded target video frame through the first prompt data to achieve a higher degree of freedom during the decoding process.

[0176] Similar to the first embedding layer network structure of the encoding terminal, the second embedding layer network structure of the decoding terminal may also include but is not limited to DNNs, CNNs, and RNNs. The second embedding layer network structure may be a part of the first generative machine learning model, or independent of the first generative machine learning model. In some embodiments, the second embedding layer network structure and the first generative machine learning model may be jointly trained.

[0177] In some embodiments, the first embedding layer network structure of the encoding terminal and the second embedding layer network structure of the decoding terminal may be the same network structure. In some embodiments, the first embedding layer network structure, the second embedding layer network structure, the second generative machine learning model, and the first generative machine learning model may be jointly or separately trained.

[0178] For example, taking the joint training of the first embedding layer network structure, the second embedding layer network structure, the second generative machine learning model, and the first generative machine learning model as an example, a first user instruction of the encoding terminal may be input into the first embedding layer network structure, and the sample original video frame and the sample reference information may be input into a joint initial generative model. The first embedding layer network structure may send sample second prompt data obtained based on the first user instruction to the joint initial generative model, and an encoder part (e.g., the encoding part including the input layer and the encoder layer) of the joint initial generative model may obtain the sample encoded data based on the sample original video frame, the sample reference information, and the sample second prompt data, and the sample encoded data may be input into the decoder layer of the joint initial generative model. A second user instruction of the decoding terminal may be input into the second embedding layer network structure, and the second embedding layer network structure may send sample first prompt data obtained based on the second user instruction to a decoder part (e.g., the part including the decoder layer and the reconstruction layer) of the joint initial generative model. The decoder part may obtain the sample reconstructed video frame based on the sample encoded data and the sample first prompt data. During the training process, a second joint loss function Da-Db+λRa may be constructed. Joint parameter adjustment may be performed on the first embedding layer network structure, the second embedding layer network structure, and each module (e.g., the encoder part, and the decoder part) of the joint initial generative model based on the second joint loss function, such that a value of the second joint loss function is less than a second joint loss threshold or is a minimum value, thereby obtaining a trained joint generative model, a trained first embedding layer network structure, and a trained second embedding layer network structure. Where Da denotes a difference between the sample original video frame and the sample reconstructed video frame; Ra denotes a data volume of the sample encoded data; and Db denotes a difference between the sample decoded target video frame and a user indication content corresponding to the sample second prompt data, and / or a difference between the sample decoded target video frame and a user indication content corresponding to the sample first prompt data.

[0179] In some embodiments, the decoding terminal may obtain the decoded target video frame by processing the encoded data, the first reference information, the second prompt data, and the first prompt data through the first generative machine learning model.

[0180] It should be noted that the user of the encoding terminal is usually different from the user of the decoding terminal. The user of the encoding terminal may be a manufacturer, and the user of the decoding terminal may be a customer.

[0181] By introducing the reference video frame in both the encoding process and decoding process, the quality of the information is enhanced in the encoding process and the decoding process using the reference video frame; the information included in the reference video frame and the encoding prompt data are introduced in the encoding process and the decoding process to adaptively obtain information that has a higher correlation with each encoding and decoding process, thereby improving the encoding and decoding accuracy in the complete encoding and decoding process.

[0182] In some embodiments, the decoding terminal may obtain pixel point grouping information of the original video frame; obtain entropy decoding data by performing entropy decoding on the encoded data based on the pixel point grouping information; obtain inverse quantization data by performing inverse quantization on the entropy decoding data; obtaining first fusion data by fusing the inverse quantization data with the first reference information; and obtain the decoded target video frame by performing an inverse transformation on the first fusion data. In some embodiments, the decoding terminal may obtain first context information by processing the pixel point grouping information through a first context model; generate a first prediction probability based on the first context information through a first probability model; and obtain the entropy decoding data by performing arithmetical decoding on the encoded data based on the first prediction probability. The first context model may include a bidirectional Transformer network. More descriptions regarding obtaining the decoded target video frame based on the pixel point information may be found in FIG. 11 and the related description thereof.

[0183] It should be noted that the above description of the processes 400 and 500 is only for example and illustration, and does not limit the scope of application of the present disclosure. For those skilled in the art, various modifications and changes can be made to the processes 400 and 500 under the guidance of the present disclosure. However, these modifications and changes are still within the scope of the present disclosure.

[0184] FIG. 6 is a schematic diagram illustrating video encoding and decoding according to some embodiments of the present disclosure.

[0185] In FIG. 6, an encoding terminal may include a first user prompt module configured to determine second prompt data input by a user (e.g., a manufacturer) of the encoding terminal; a decoding terminal may include a second user prompt module configured to determine first prompt data input by a user (e.g., a customer) of the decoding terminal. Both the encoding terminal and the decoding terminal may be coupled with an information prompt module. The information prompt module may be configured to determine the prompt information.

[0186] Taking a current target video frame (i.e., an original video frame) being a third frame in a video frame sequence, and first and first generative machine learning models being Transformer-based generative models as an example, the encoding terminal may select a decoded video frame corresponding to a reconstructed second frame that is located before the third frame in shooting timing sequence from a decoded video frame set as a reference video frame. The encoding terminal may splice the original video frame of the third frame and the decoded video frame of the second frame to be input into the second generative machine learning model. The second generative machine learning model may perform analysis and feature extraction on a spliced video frame to obtain image feature information corresponding to the third frame, and perform entropy encoding on the image feature information corresponding to the third frame to obtain a bitstream corresponding to the third frame. Optionally, the encoding terminal may generate a model encoding prompt for the third frame using the information prompt module, and input the model encoding prompt into the second generative machine learning model. The second generative machine learning model may perform feature extraction on the spliced video frame with reference to the model encoding prompt, and fuse information contained in the model encoding prompt to obtain prompt information corresponding to the third frame. Optionally, the encoding terminal may analyze supplementary information input by the use using the first user prompt module to determine the second prompt data, and input the second prompt data into the second generative machine learning model. The second generative machine learning model may adjust the image feature information and the prompt information corresponding to the third frame based on user guidance information contained in the second prompt data to obtain adjusted image feature information and adjusted prompt information, and perform entropy encoding on the adjusted image feature information and the adjusted prompt information to obtain the bitstream corresponding to the third frame. The encoding terminal may transmit the bitstream corresponding to the third frame to the decoding terminal.

[0187] The decoding terminal may input the received bitstream corresponding to the third frame into the first generative machine learning model, and the first generative machine learning model may decode the bitstream to obtain decoding feature information corresponding to the third frame. The decoding terminal may select a decoded video frame corresponding to a reconstructed second frame that is located before the third frame in the shooting timing sequence from the decoded video frame set as the reference video frame, and input the reference video frame into the first generative machine learning model. The first generative machine learning model may fuse the feature information of the second decoded video frame and the decoding feature information corresponding to the third frame to obtain fusion decoding information; and reconstruct the fusion decoding information to obtain the decoded target video frame corresponding to the third frame and output the decoded target video frame corresponding to the third frame. Optionally, the decoding terminal may obtain the prompt information from the encoding terminal, input the prompt information into the first generative machine learning model, and the first generative machine learning model may decode the prompt information to obtain decoding prompt information corresponding to the third frame. Alternatively, a model encoding prompt may be generated for the third frame using the information prompt module, and the model encoding prompt may be input into the first generative machine learning model to obtain the decoding prompt information corresponding to the third frame. Optionally, the decoding terminal may analyze the supplementary information input by the user using the second user prompt module to determine first prompt data, and input the first prompt data into the first generative machine learning model. The first generative machine learning model may analyze the first prompt data to determine decoding guidance information corresponding to the first prompt data; adjust the decoding feature information and the decoding prompt information corresponding to the third frame based on the decoding guidance information; fuse the feature information of the decoded video frame of the second frame, the adjusted decoding feature information and the adjusted decoding prompt information corresponding to the third frame to obtain the fusion decoding information; and reconstruct the fusion decoding information to obtain the decoded target video frame corresponding to the third frame and output the decoded target video frame corresponding to the third frame. The decoded target video frame corresponding to the third frame may be placed in the decoded video frame set as a latest decoded target video frame.

[0188] It should be noted that the above description is only for example and explanation, and does not limit the scope of application of the present disclosure. For those skilled in the art, various modifications and changes can be made to FIG. 6 and the related description sunder the guidance of the present disclosure. However, these modifications and changes are still within the scope of the present disclosure. For example, the encoding terminal and the decoding terminal may each be coupled with an information prompt module (e.g., the information prompt module may be a feature extraction network constructed by a CNN, LSTM, or Transformer, etc., for determining the prompt information based on the original video frame, or the original video frame and the reference video frame) . As another example, the original video frame may be another frame in the video frame sequence / the video.

[0189] FIG. 7 a flowchart illustrating an exemplary video decoding method according to some embodiments of the present disclosure. A process 700 may be implemented by an encoding terminal. In some embodiments, the process 700 may be performed by the processing device 110, the computing device 200 (e.g., the processor 210) , or the video encoding and decoding system 300 (e.g., the encoding module 310) . As shown in FIG. 7, in some embodiments, the process 700 may include the following operations.

[0190] In 710, pixel point grouping information of an original video frame may be obtained by performing pixel point grouping on the original video frame based on an image content of the original video frame.

[0191] The image content refers to information related to pixel values of pixel points in a video frame / an image. For example, the image content of the original video frame may include the pixel value of each pixel point in the original video frame.

[0192] In some embodiments, the encoding terminal may group the pixel points in the original video frame based on a correlation (e.g., a similarity between the pixel values) of the pixel values of the pixel points in the original video frame. Furthermore, the encoding terminal may mark the pixel points in the original video frame according to the grouping of the pixel points, thereby obtaining the pixel point grouping information of the original video frame. In some embodiments, in the same pixel point group, a grouping mark corresponding to each pixel point may be the same, and grouping marks corresponding to pixel points in different pixel point groups may be different.

[0193] In some embodiments, the encoding terminal may perform pixel point grouping on the original video frame based on the image content of the original video frame based on a preset continuous path. Specifically, the encoding module 310 may obtain a plurality of initial pixel point groups by traversing each of the pixel points in the original video frame in turn based on the preset continuous path; and obtain a plurality of target pixel point groups of the original video frame by adjusting pixel points in at least a portion of the plurality of initial pixel point groups.

[0194] More descriptions regarding performing pixel point grouping based on the preset continuous path may be found in FIG. 8 the related descriptions thereof.

[0195] In some embodiments, the encoding terminal may obtain the plurality of target pixel point groups of the original video frame by processing the pixel values of the pixel points in the original video frame through a classification model.

[0196] More descriptions regarding performing pixel point grouping using the classification model may be found in FIG. 10 and the related descriptions thereof.

[0197] The pixel point grouping information may be correlated with the grouping mark corresponding to each pixel point in each of the target pixel point groups. The grouping mark may be represented by a numerical value, and each of the target pixel point groups may correspond to a different grouping mark. For example, the grouping mark may be 0, 1, 2, 3, 4, 5, etc. In some embodiments, the numerical value of the grouping mark may be positively correlated with a count of the pixel points contained in each of the target pixel point groups. For example, a target pixel point group with a grouping mark of 0 contains the least count of pixel points, a target pixel point group with a grouping mark of 1 contains more pixel points than the target pixel point group with the grouping mark of 0, a target pixel point group with a grouping mark of 2 contains more pixel points than the target pixel point group with the grouping mark of 1, etc.

[0198] In some embodiments, the encoding terminal may assign a corresponding grouping mark to each of the target pixel point groups, determine one or more mask images corresponding to the original video frame based on the grouping marks of the plurality of target pixel point groups, and obtain the pixel point grouping information of the original video frame based on the one or more mask images. For example, the pixel point grouping information may be the one or more mask images corresponding to the original video frame, or may be information obtained after processing (e.g., transformation, transformation and quantization, or transformation, quantization and entropy encoding, etc. ) the one or more mask images.

[0199] The mask refers to a matrix composed of a plurality of elements reflecting pixel point information in the original video frame, each of the elements corresponding to one of the pixel points in the original video frame. Accordingly, a mask image is a matrix image corresponding to the mask. In some embodiments, the pixel values of the pixel points (i.e., each element in the mask matrix) in the one or more mask images corresponding to the original video frame may represent grouping categories (i.e., the grouping marks) of the pixel points in the original video frame.

[0200] For example, the encoding terminal may determine a mask image of the same size as the original video frame based on the plurality of target pixel point groups, and the pixel values of the pixel points in the mask image may be the grouping marks of the corresponding pixel points in the original video frame. For example, as shown in FIG. 9A or FIG. 9B, the mask image may include a plurality of pixel points with values of 0, 1, 2, and 3, indicating that the pixel points of the corresponding original video frame are divided into four pixel point groups, and the grouping marks of the four pixel point groups are 0, 1, 2, and 3, respectively.

[0201] As another example, the encoding end may determine a plurality of mask images of the same size as the original video frame based on the plurality of target pixel point groups, and each of the mask images may correspond to one of the pixel point groups of the original video frame. Referring to FIG. 9A or FIG. 9B, the encoding terminal obtains four mask sub-images based on the grouping marks (e.g., the numbers in the figure) of the mask image shown in FIG. 9A or FIG. 9B, i.e., determines the four mask images corresponding to the original video frame. A first mask sub-image corresponds to a grouping mark 0, and the pixel values of the pixel points in the first mask sub-image corresponding to the pixel points with the grouping mask of 0 in the original video frame are 1, and the pixel values of the remaining pixel points are 0; a second mask sub-image corresponds to a grouping mark 1, and the pixel values of the pixel points in the second mask sub-image corresponding to the pixel points with the grouping mask of 1 in the original video frame are 1, and the pixel values of the remaining pixel points are 0; a third mask sub-image corresponds to a grouping mark 2, and the pixel values of the pixel points in the third mask sub-image corresponding to the pixel points with the grouping mask of 2 in the original video frame are 1, and the pixel values of the remaining pixel points are 0; a fourth mask sub-image corresponds to a grouping mark 3, and the pixel values of the pixel points in the fourth mask sub-image corresponding to the pixel points with the grouping mask of 3 in the original video frame are 1, and the pixel values of the remaining pixel points are 0.

[0202] In some embodiments, the encoding terminal may determine the one or more mask images of the original video frame based on the plurality of target pixel point groups; and obtain the pixel point grouping information of the original video frame by transforming and / or quantizing the one or more mask images.

[0203] Transforming refers to reducing a mask data volume or a bit rate of the one or more mask images by performing nonlinear downsampling on the one or more mask images. The nonlinear downsampling may include but is not limited to bilinear downsampling, transformation downsampling based on CNN, transformation downsampling based on MLP neural networks, etc. For example, the encoding terminal may perform a dimensionality reduction transformation on one of the mask images corresponding to the original video frame to obtain a feature vector of h*w*c, where h denotes length, w denotes width, and c denotes channel (c is 1 by default) , and h*w*c may be less than a dimension of the original video frame. For example, h*w*c=210*200*5, and the dimension of the original video frame is 256*256*5.

[0204] Quantizing refers to a process of mapping a continuous range of values to a finite count of discrete values. For example, values in a continuous interval [1, 5] are mapped to five discrete values 1, 2, 3, 4, and 5 according to a certain rule (e.g., an integer operation, a rounding operation, etc. ) . Quantizing does not change the dimension of the image.

[0205] For example, the encoding terminal may obtain transformation information of the one or more mask images by transforming the one or more mask images, and determine the transformation information of the one or more mask images as the pixel point grouping information of the original video frame. As another example, the encoding terminal may obtain the transformation information of the one or more mask images by transforming the one or more mask images; obtain quantization information of the one or more mask images by quantizing the transformation information of the one or more mask images; and determine the quantization information of the one or more mask images as the pixel point grouping information of the original video frame.

[0206] In some embodiments, the encoding terminal may obtain the transformation information of the one or more mask images by transforming the one or more mask images; obtain the quantization information of the one or more mask images by quantizing the transformation information of the one or more mask images; obtain entropy encoding information of the one or more mask images by performing entropy encoding on the quantization information of the one or more mask images; and determine the entropy encoding information of the one or more mask images as the pixel point grouping information of the original video frame.

[0207] In some embodiments, the encoding terminal may determine the one or more mask images as the pixel point grouping information of the original video frame.

[0208] In some embodiments, the numerical value of the grouping mark may be correlated with an order in which pixel information or feature information is processed in a second context model. For example, the smaller the numerical value of the grouping mark, the earlier the pixel information or the feature information of the pixel points in a corresponding target pixel point group is processed by the second context model. More descriptions regarding the second context model may be found in the present disclosure below.

[0209] In some embodiments, the pixel point grouping information and the encoded data may be sent to the decoding terminal together. In some embodiments, the encoding terminal may transmit the pixel point grouping information and the encoded data independently (i.e., the pixel point grouping information and the encoded data are not fused) to the decoding terminal, or transmit the pixel point grouping information and the encoded data to the decoding terminal after fusing the pixel point grouping information and the encoded data. For example, the encoding terminal may transmit a mask bitstream of the one or more mask images and the encoded data to the decoding terminal independently, or merge transmit the mask bitstream of the one or more mask images and the encoded data to the decoding terminal after fusing (e.g., embedding the mask bitstream into the encoded data) the mask bitstream of the one or more mask images and the encoded data.

[0210] By transmitting the pixel point information and the encoded data (independently transmitting or transmitting after fusion) to the decoding terminal together, mask information corresponding to the original video frame can be obtained at the decoding terminal, such that the decoding terminal can perform decoding and reconstruction based on the mask information, thereby improving the decoding accuracy and the image quality of the decoded target video frame.

[0211] In some embodiments, the pixel point grouping information may be transmitted to the decoding terminal in sequence according to the corresponding grouping masks. For example, if the pixel point grouping information is information related to a plurality of mask images corresponding to the original video frame, the encoding terminal may transmit the pixel point grouping information corresponding to each of the target pixel point group to the decoding terminal in sequence based on a size of each of the grouping masks.

[0212] In 720, second fusion data may be obtained by fusing the original video frame and second reference information.

[0213] In some embodiments, a fusion mode may include but is not limited to splicing, addition, multiplication, weighted summation, etc. For example, if the second reference information is at least one decoded video frame before the original video frame, the encoding module 310 may obtain the second fusion data by splicing the original video frame and the reference video frame. As another example, if the second reference information is the feature information of the at least one decoded video frame before the original video frame, the encoding module 310 may obtain the second fusion data by performing the weighted summation on the feature information of the original video frame and the second reference information.

[0214] In 730, first transformation data may be obtained by transforming the second fusion data.

[0215] Referring to the above, transforming refers to performing nonlinear downsampling (e.g., bilinear downsampling, transformation downsampling based on CNN, transformation downsampling based on MLP neural networks, etc. ) on data to be processed (e.g., the second fusion data) so as to represent a main feature of the data to be processed using a more compact expression, and reduce the dimension or the data volume of the data to be processed.

[0216] In some embodiments, the encoding terminal may perform the nonlinear downsampling on the second fusion data using the CNN so as to express a main feature of the second fusion data using the more compact expression (e.g., a bitstream) , and obtain the first transformation data.

[0217] In 740, first quantization data may be obtained by quantizing the first transformation data.

[0218] Referring to the above, quantizing refers to a process of mapping a continuous range of values to a finite count of discrete values, which is an integer operation on the data. For example, the first quantization data (e.g., presented as a feature vector) may be obtained by mapping pixel values corresponding to the first transformation data to a finite count of discrete values by quantization.

[0219] In 750, encoded data may be obtained by performing entropy encoding on the first quantization data based on the pixel point grouping information.

[0220] Since the second fusion data of the original video frame and the second reference information is transformed and quantization, in order to improve the encoding accuracy, in this embodiment, preferably, the quantization information obtained after the one or more mask images corresponding to the original video frame are transformed and quantization may be determined as the pixel point grouping information. It can be understood that in some embodiments, the pixel point grouping information may be the transformation information corresponding to the one or more mask images.

[0221] Entropy encoding is a data compression technique that assigns a code length based on a probability of occurrence of a symbol (e.g., a classification mark) . A symbol with a high probability of occurrence is assigned with a relatively short code, and a symbol with a low probability of occurrence is assigned with a relatively long code. Accordingly, an overall code length of the obtained encoded data is reduced, thereby achieving efficient data transmission and / or storage.

[0222] An entropy encoding mode may include but is not limited to variable length encoding and arithmetic encoding.

[0223] In some embodiments, the encoding terminal may obtain second context information by processing the pixel point grouping information through the second context model; generate a second prediction probability based on the second context information through a second probability model; and obtain the encoded data by performing arithmetical encoding on the first quantization data based on the second prediction probability.

[0224] The second context model refers to a model configured to provide context information by the encoding terminal. For example, the second context model may be configured to provide the second probability model with the second context information related to the first quantization data to guide the second probability model to perform probability prediction, such that the second probability model can more accurately predict the probability of each character.

[0225] The second context model may include a bidirectional Transformer network.

[0226] The second probability model may be configured to determine a prediction probability for each character in a character string to be encoded for arithmetic encoding. The second probability model may be a mapping table that records a correspondence between each character and the probability of the character. When the entropy encoding is performed, the encoding terminal may determine the prediction probability corresponding to each character by searching the mapping table to generate the second prediction probability. The second probability model may be an algorithm or a program code capable of calculating a probability value of each character in the character string to be encoded.

[0227] In some embodiments, the second probability model may be a dynamic probability model. Since the probability value of each character in the character string to be encoded changes dynamically based on context information, establishing the dynamic probability model can improve the accuracy of the entropy encoding.

[0228] For example, the encoding terminal may input the quantization information corresponding to the one or more mask images of the original video frame into the second context model, and the second context model may analyze and process the quantization information to output the second context information corresponding to the original video frame. The second probability model may determine the second prediction probability corresponding to the first quantization data based on the second context information. Furthermore, the encoding terminal may perform the arithmetic encoding on the first quantization data based on the second prediction probability, thereby embedding the quantization information corresponding to the one or more mask images into a main bitstream of the original video frame to obtain the encoded data.

[0229] In some embodiments, the second context model and the second probability model may be two modules of the same machine learning model. For example, the second generative machine learning model described in FIG. 4 may include a transformation module, a quantization module, a pixel point grouping information determination module, and an entropy encoding module. The transformation module may be configured to transform the second fusion data to obtain the first transformed data. The quantization module may be configured to quantize the first transformation data to obtain the first quantization data. The pixel point grouping information determination module may be configured to perform pixel point grouping on the original video frame based on the image content of the original video frame to obtain the pixel point grouping information of the original video frame. The entropy encoding module may be configured to perform the entropy encoding on the first quantization data or the second quantization data based on the pixel point grouping information to obtain the encoded data. The entropy encoding module may include a network structure composed of the second context model and the second probability model, and an arithmetic encoding unit. The pixel point grouping information of the original video frame may be input into the entropy encoding module, and the second context model of the entropy encoding module may analyze and process the pixel point grouping information to output the second context information and input the second context information into the second probability model. The second probability model may determine the second prediction probability corresponding to the first quantization data or the second quantization data based on the second context information. The arithmetic encoding unit may perform the arithmetic encoding on the first quantization data or the second quantization data based on the second prediction probability to obtain the encoded data (e.g., a bitstream fused with the mask information) . The transformation module, the quantization module, the pixel point grouping information determination module, and the entropy encoding module of the second generative machine learning model may be jointly trained. The joint training may be similar to the training process of the second generative machine learning model described above (e.g., the related descriptions in FIG. 4 or FIG. 5) , which is not repeated here. In some embodiments, the transformation module, the quantization module, the pixel point grouping information determination module, and the entropy encoding module of the second generative machine learning model may be a machine learning model for implementing the corresponding functions, or a partial structure of the machine learning model for implementing the corresponding fu nctions.

[0230] In some embodiments, the second context model may include a logical determination module and a generation module. The logical determination module may be configured to sort the plurality of pixel point groups based on the pixel point grouping information, and determine pixel value information corresponding to each of the pixel point groups in the first quantization data. The generation module may be configured to process pixel value information corresponding to the pixel points in each of the pixel point groups in the first quantization data to obtain the second context information corresponding to each of the pixel point groups.

[0231] In some embodiments, the encoding terminal may sort (e.g., sort the pixel points in an ascending order based on the grouping marks) the plurality of pixel point groups (e.g., the plurality of target pixel point groups) based on the pixel point grouping information; determine the pixel value information corresponding to each of the pixel point groups in the first quantization data; and determine the encoded data based on an order of the plurality of pixel point groups.

[0232] Since there is a correspondence between the pixel points in the original video frame and the pixel points in the mask images, accordingly, the first quantization data obtained after transforming and quantizing the original video frame (or the second fusion data of the original video frame) and the pixel point information obtained after transforming and quantizing the one or more mask images corresponding to the original video frame may also have a pixel point correspondence. For example, determining the pixel value information corresponding to each of the pixel point groups in the first quantization data may include determining a pixel value of each pixel point in the first quantization data obtained after transforming and quantizing the original video frame corresponding to each pixel point in the quantization information obtained after transforming and quantizing the one or more mask images.

[0233] Taking the sorting of the pixel points in an ascending order based on the grouping marks as an example, determining the encoded data based on an order of the plurality of pixel point groups may specifically include the following operations for an Nth pixel point group (N is a positive integer) of the plurality of pixel point groups.

[0234] In response to determining that N is 1, the encoding terminal may assign a fixed character probability to first quantization data corresponding to a first pixel point group (i.e., a pixel point group with the smallest grouping mark) , and perform the arithmetic encoding on the pixel value information corresponding to each of pixel points in the first pixel point group in the first quantization data in turn based on the fixed character probability to obtain encoded data (e.g., a bitstream) corresponding to the first pixel point group. Alternatively, the encoding terminal may input the pixel value information corresponding to the first pixel point group in the first quantization data into the second probability model to obtain a second prediction probability corresponding to the first pixel point group, and perform the arithmetic encoding on the pixel value information corresponding to each of the pixel points in the first pixel point group in the first quantization data in turn based on the second prediction probability corresponding to the first pixel point group to obtain the encoded data corresponding to the first pixel point group.

[0235] In response to determining that N is 2, the encoding terminal may obtain second context information corresponding to the first pixel point group by inputting, into a second context model, the pixel value information corresponding to the first pixel point group in the first quantization data; generate a second prediction probability corresponding to a second pixel point group based on the second context information corresponding to the first pixel point group through the second probability model; and obtain encoded data corresponding to the second pixel point group by performing the arithmetical encoding on pixel value information corresponding to each of the pixel points in the second pixel point group in the first quantization data based on the second prediction probability corresponding to the second pixel point group.

[0236] In response to determining that N is greater than 2, the encoding terminal may input pixel value information corresponding to an (N-1) th pixel point group in the first quantization data into the second context model, the second context model determining second context information corresponding to the (N-1) th pixel point group based on second context information corresponding from the first pixel point group to an (N-2) th pixel point group and pixel value information corresponding to the (N-1) th pixel point group; generate a second prediction probability corresponding to the Nth pixel point group based on the second context information corresponding to the (N-1) th pixel point group through the second probability model; and obtain encoded data corresponding to the Nth pixel point group by performing the arithmetical encoding on pixel value information corresponding to each of the pixel points in the Nth pixel point group in the first quantization data based on the second prediction probability corresponding to the Nth pixel point group.

[0237] As described above, at the encoding terminal, the second context model may generate the second context information of the pixel point group corresponding to each grouping mark by processing the first quantization data in sequence based on the grouping mark. However, at the decoding terminal, due to the lack of the first quantization data in the encoding process, the context model (also referred to as the first context model) at the decoding terminal can predict context information (also referred to as first context information) of a pixel point group that sorted later only relying on the decoded data corresponding to a pixel point group that sorted earlier. However, when the arithmetical decoding is performed on the encoded data corresponding to the first pixel point group, since there is no reference information of a previous pixel point group, the decoded data corresponding to the first pixel point group can only be decoded based on the fixed character probability. Accordingly, in order to be consistent with the decoding terminal, the encoding terminal uses the fixed character probability consistent with the decoding terminal when performing the arithmetical encoding on the pixel value information corresponding to the first pixel point group in the first quantization data. That is to say, when predicting the context information (e.g., the second context information and the first context information) , the second context model at the encoding terminal and the first context model at the decoding terminal are limited to the known information that can be acquired by the decoding terminal.

[0238] In some embodiments, in order to further improve the prediction accuracy of the context information (e.g., the second context information and the first context information) , auxiliary inverse transformation information in an auxiliary bitstream may be introduced. Since input data of an auxiliary transformation of auxiliary bitstream processing is the transformation data in the encoding process, the result of the auxiliary inverse transformation may be similar to the transformation data or the quantization data in the encoding process.

[0239] In some embodiments, the encoding terminal may obtain the auxiliary transformation information by performing the auxiliary transformation on the first transformation data; obtain auxiliary quantization information by performing auxiliary quantization on the auxiliary transformation information; and obtain the auxiliary bitstream by performing auxiliary entropy encoding on the auxiliary quantization information, the auxiliary bitstream and the encoded data being sent to the decoding terminal together.

[0240] Furthermore, in some embodiments, the encoding terminal may obtain second auxiliary entropy decoding information by performing second auxiliary entropy decoding on the auxiliary bitstream; obtain second auxiliary inverse quantization information by performing auxiliary inverse quantization on the second auxiliary entropy decoding information; and obtain second auxiliary inverse transformation information by performing the auxiliary inverse transformation on the second auxiliary inverse quantization information.

[0241] In some embodiments, the encoding terminal may obtain the auxiliary inverse transformation information from the decoding terminal. To distinguish from related operations of the encoding terminal, the auxiliary inverse transformation information obtained from the decoding terminal is also referred to as first auxiliary inverse transformation information. For example, after the auxiliary bitstream is sent to the decoding terminal, the decoding terminal may obtain first auxiliary entropy decoding information by performing first auxiliary entropy decoding on the auxiliary bitstream; obtain first auxiliary inverse quantization information by performing auxiliary inverse quantization on the first auxiliary entropy decoding information; and obtain the first auxiliary inverse transformation information by performing the auxiliary inverse transformation on the first auxiliary inverse quantization information.

[0242] Enhanced information is generated by obtaining the first auxiliary transformation information or the second auxiliary transformation information based on the auxiliary bitstream, such that the auxiliary transformation information is input into the second context model, thereby further improving the accuracy of the context information input into the second context model.

[0243] In some embodiments, the encoding terminal may sort a plurality of pixel point groups (e.g., a plurality target pixel point groups) based on the pixel point grouping information; determine pixel value information corresponding to each of the plurality of pixel point groups in the first quantization data; and determine the encoded data according to the first auxiliary transformation information or the second auxiliary transformation information obtained from the decoding terminal based on an order of the plurality of pixel point groups.

[0244] Taking sorting the pixel points in an ascending order based on the grouping marks as an example, determining the encoded data according to the first auxiliary transformation information based on the order of the plurality of pixel point groups may specifically include the following operations for the Nth pixel point group (N is a positive integer) of the plurality of pixel point groups.

[0245] In response to determining that N is 1, the encoding terminal may assign the fixed character probability to the first quantization data corresponding to the first pixel point group (i.e., the pixel point group with the smallest grouping mark) , and perform the arithmetic encoding on the pixel value information corresponding to each of the pixel points in the first pixel point group in the first quantization data in turn based on the fixed character probability to obtain the encoded data (e.g., the bitstream) corresponding to the first pixel point group. Alternatively, the encoding terminal may input the pixel value information corresponding to the first pixel point group in the first quantization data into the second probability model to obtain the second prediction probability corresponding to the first pixel point group, and perform the arithmetic encoding on the pixel value information corresponding to each of the pixel points in the first pixel point group in the first quantization data in turn based on the second prediction probability corresponding to the first pixel point group to obtain the encoded data corresponding to the first pixel point group. Alternatively, the encoding terminal may obtain initial second context information by inputting at least a portion of data input of the second auxiliary inverse transformation information (e.g., the pixel value information corresponding to the first pixel point group in the second auxiliary inverse transformation information, or the second auxiliary inverse transformation information) into the second context model; generate the second prediction probability corresponding to the first pixel point group based on the initial second context information through the second probability model; and obtain the encoded data corresponding to the first pixel point group by performing the arithmetic encoding on the pixel value information corresponding to each of the pixel points in the first pixel point group in the first quantization data in turn based on the second prediction probability corresponding to the first pixel point group.

[0246] In response to determining that N is 2, the encoding terminal may obtain the second context information corresponding to the first pixel point group by inputting, into the second context model, the pixel value information (i.e., the pixel value information of each of the pixel points corresponding to each of the pixel points in the first pixel point group in the first quantization data) corresponding to the first pixel point group in the first quantization data and at least the portion of the data input of the second auxiliary inverse transformation information; generate a second prediction probability corresponding to a second pixel point group based on the second context information corresponding to the first pixel point group through the second probability model; and obtain encoded data corresponding to the second pixel point group by performing the arithmetical encoding on pixel value information corresponding to each of the pixel points in the second pixel point group in the first quantization data based on the second prediction probability corresponding to the second pixel point group.

[0247] In response to determining that N is greater than 2, the encoding terminal may input pixel value information (i.e., pixel value information of each of the pixel points in the (N-1) th group corresponding to the pixel points in the first quantization data) corresponding to the (N-1) th pixel point group in the first quantization data and at least the portion of the data input of the second auxiliary inverse transformation information into the second context model, the second context model determining second context information corresponding to the (N-1) th pixel point group based on second context information corresponding from the first pixel point group to the (N-2) th pixel point group, the pixel value information corresponding to the (N-1) th pixel point group, and at least the portion data of the second auxiliary inverse transformation information; generate the second prediction probability corresponding to the Nth pixel point group based on the second context information corresponding to the (N-1) th pixel point group through the second probability model; and obtain the encoded data corresponding to the Nth pixel point group by performing the arithmetical encoding on the pixel value information corresponding to the Nth pixel point group in the first quantization data based on the second prediction probability corresponding to the Nth pixel point group.

[0248] At least the portion of the data input of the second auxiliary inverse transformation information may include pixel value information corresponding to the pixel point groups in the second auxiliary inverse transformation information. For example, for the arithmetical encoding corresponding to the second pixel point group, the encoding terminal may obtain the second context information corresponding to the first pixel point group by inputting the pixel value information of the pixel point corresponding to each of the pixel points in the first pixel point group in the first quantization data, and the pixel value information of the pixel point corresponding to each of the pixel points in the second pixel point group in the second auxiliary inverse transformation information into the second context model. As another example, for the arithmetical encoding corresponding to the (N-1) th pixel point group, the encoding terminal may input the pixel value information corresponding to each of the pixel points in the (N-2) th pixel point group in the first quantization data, and the pixel value information corresponding to each of the pixel points in the (N-1) th pixel point group in the second auxiliary inverse transformation information into the second context model. The second context model may determine the second context information corresponding to the (N-1) th pixel point group based on second context information corresponding from the first pixel point group to an (N-3) th pixel point group, the pixel value information corresponding to each of the pixel points in the (N-2) th pixel point group in the first quantization data, and the pixel value information corresponding to each of the pixel points in the (N-1) th pixel point group in the second auxiliary inverse transformation information.

[0249] In some embodiments, in each encoding operation, all the second auxiliary inverse transformation information may be input into the second context model. In this case, each of the pixel point groups may share the same second auxiliary inverse transformation information. For example, for the arithmetical encoding corresponding to the (N-1) th pixel point group, the encoding terminal may input the pixel value information corresponding to each of the pixel points in the (N-2) th pixel point group in the first quantization data and the second auxiliary inverse transformation information into the second context model.

[0250] By constructing the context model (e.g., the second context model) based on the generative model, and determining the prediction probability (e.g., the second prediction probability) of the character under entropy encoding using the context model to guide the probability model (e.g., the second probability model) based on the masks and / or the auxiliary bitstream information of the original video frame, the accuracy of entropy encoding can be improved, thereby improving the encoding efficiency and accuracy.

[0251] By introducing the second auxiliary inverse transformation information in the process of generating the second context information at the encoding terminal, more reference information (e.g., the first context model at the decoding terminal can predict the first context information based on more information) can be provided for the prediction of the context information, thereby improving the accuracy of context prediction.

[0252] It should be noted that the above description of the process 700 is only for example and explanation, and does not limit the scope of application of the present disclosure. For those skilled in the art, various modifications and changes can be made to the process 700 under the guidance of the present disclosure. However, these modifications and changes are still within the scope of the present disclosure. For example, the operation 720 in the process 700 may be omitted. In this case, in the operation 730, the encoding terminal may obtain the second transformation data by transforming the original video frame. Correspondingly, in the operation 740, the encoding terminal may obtain the second quantization data by quantizing the second transformation data; in the operation 750, the encoding terminal may obtain the encoded data by performing entropy encoding on the second quantization data based on the pixel point grouping information.

[0253] FIG. 8 is a flowchart illustrating an exemplary process of pixel point grouping according to some embodiments of the present disclosure.

[0254] A process 800 may be implemented by an encoding terminal. In some embodiments, the process 800 may be performed by the processing device 110, the computing device 200 (e.g., the processor 210) , or the video encoding and decoding system 300 (e.g., the encoding module 310) . As shown in FIG. 8, in some embodiments, the process 800 may include the following operations.

[0255] In 810, a plurality of initial pixel point groups may be obtained by traversing each of pixel points in an original video frame in turn based on a preset continuous path.

[0256] The original video may be a single frame in a video / aa video frame sequence, which is also referred to as an original image.

[0257] The preset continuous path refers to a preset path in which all pixel points of the original video frame can be traversed. In some embodiments, the preset continuous path may include a hollow square path (e.g., as shown in FIG. 9A) , a Z-shaped path (e.g., as shown in FIG. 9B) , etc.

[0258] For example, the preset continuous path may be a hollow square path in which each pixel point is traversed in sequence from an outermost periphery to a center of the original video frame (e.g., from the outside to the inside in a direction of an arrow shown in FIG. 9A) . Alternatively, the preset continuous path may be a hollow square path in which each pixel point is traversed in sequence from the center to the outermost periphery of the original video frame (e.g., from the inside to the outside in an opposite direction of the arrow shown in FIG. 9A) .

[0259] As another example, the preset continuous path may be that the original video frame is traversed from a first row from left to right in a top-to-bottom direction or a bottom-to-top direction, and after the first row ends, a second row is traversed from right to left, and after the second row ends, a third row is traversed from left to right again, ..., until the last pixel point in the last row (e.g., each pixel point is traversed in sequence from the first row to the last row in the direction indicated by the arrow shown in FIG. 9B) . Alternatively, the preset continuous path may be that the original video frame is traversed from a first column from top to bottom in a left-to-right direction or a right-to-left direction, and after the first column ends, a second column is traversed from bottom to top, until the last pixel point in the last column.

[0260] A count of the initial pixel point groups may be set as required, which is not specifically limited in the present disclosure. For example, 2, 3, 4, or 5 initial pixel point groups may be obtained. A count of pixel points contained in each of the initial pixel point groups may be set as required, which is not specifically limited in the present disclosure. For example, the count of the pixel points contained in each of the initial pixel point groups may be increased in sequence according to an order of division (e.g., according to a traversal order of the preset continuous path) . For example, the count of the pixel points contained in N initial pixel point groups may be {8, 20, 40, ..., Mn} , where Mn denotes the count of the pixel points contained in an Nth initial pixel point group.

[0261] FIG. 9A and FIG. 9B are schematic diagrams illustrating a preset continuous path according to some embodiments of the present disclosure. Images shown in FIG. 9A and FIG. 9B are mask images, where numbers in the images represent grouping marks, and an arrow direction represents a traversal direction of the preset continuous path.

[0262] For example, referring to FIG. 9A, the preset continuous path is a hollow square path from an outermost periphery to a center. An encoding terminal may traverse each of pixel points in an original video frame in sequence according to the preset continuous path, and divide the 1st to 9th pixel points into an initial pixel point group 1, which corresponds to a grouping mark 0; 10th to 40th pixel points into an initial pixel point group 2, which corresponds to a grouping mark 1; 41st-96th pixel points into an initial pixel point group 3, which corresponds to a grouping mark 2; and 97th to 150th pixel points into an initial pixel point group 4, which corresponds to a grouping mark 3, so as to obtain four initial pixel point groups.

[0263] As another example, referring to FIG. 9B, the preset continuous path may be a Z-shaped path from top to bottom. The encoding terminal may traverse each of the pixel points in the original video frame in sequence according to the preset continuous path, and divide 1st to 9th pixel points into an initial pixel point group 1, which corresponds to a grouping mark 0; 10th to 34th pixel points into an initial pixel point group 2, which corresponds to a grouping mark 1; 35th to 88th pixel points into an initial pixel point group 3, which corresponds to a grouping mark 2; and 89th to 150th pixel points into an initial pixel point group 4, which corresponds to a grouping mark 3, so as to obtain four initial pixel point groups.

[0264] In 820, a plurality of target pixel point groups of the original video frame may be obtained by adjusting pixel points in at least a portion of the plurality of initial pixel point groups.

[0265] In some embodiments, the encoding terminal may obtain the plurality of target pixel point groups of the original video frame by adjusting pixel points in two adjacent initial pixel point groups among the plurality of initial pixel point groups. Specifically, the encoding terminal may obtain adjusted target pixel point groups by migrating pixel points satisfying a preset condition in a second initial pixel point group of the two adjacent initial pixel point groups to a first initial pixel point group of the two adjacent initial pixel point groups, the first initial pixel point group (e.g., an initial pixel point group i, i being an integer) being located before the second initial pixel point group (e.g., an initial pixel point group i+1) according to the order of division.

[0266] The two adjacent initial pixel point groups refer to two initial pixel point groups that are adjacent to each other according to the order of division. For example, referring to the above embodiment, among the four initial pixel point groups, the initial pixel point group 1 and the initial pixel point group 2 are the two adjacent initial pixel point groups, the initial pixel point group 2 and the initial pixel point group 3 are the two adjacent initial pixel point groups, and the initial pixel point group 3 and the initial pixel point group 4 are the two adjacent initial pixel point groups.

[0267] The preset condition refers to a condition that needs to be satisfied for migrating / expanding pixel points. The preset condition may include that in the first initial pixel point group and the second initial pixel point group of the two adjacent initial pixel point groups, a mean square difference between values of pixel points in the first initial pixel point group adjacent to the second initial pixel point group and a value of a kth pixel point in the second initial pixel point group is less than a mean square difference threshold, where k is a positive integer. For example, the two adjacent initial pixel point groups are the initial pixel point group 3 and the initial pixel point group 4, the initial pixel point group 3 contains the 35th to 88th pixel points of the original video frame, and the initial pixel point group 4 contains the 89th to 150th pixel points of the original video frame; the pixel point in the initial pixel point group 3 adjacent to the initial pixel point group 4 is the 88th pixel point of the original video frame, i.e., the last pixel point of the initial pixel point group 3; the kth pixel point in the initial pixel point group 4 is an 89th + (k-1) th pixel point of the original video frame.

[0268] For example, for the two adjacent initial pixel point groups, the encoder terminal may adjust the pixel points use the following mathematical equation (1) : Mi′=Mi+k, if MSE (Pl, Pl+k) <Thr (1)

[0269] wherein, Mi′ denotes a count of pixel points in a current initial pixel point group i after an adjustment; Mi denotes (before the adjustment) a count of pixel points in the current initial pixel point group i (i.e., the first initial pixel point group) ; k denotes a count of pixel points to be expanded in the current initial pixel point group i, i.e., a count of pixel points to be migrated from a next initial pixel point group i+1 (i.e., the second initial pixel point group) adjacent to the current initial pixel point group; Pl denotes a pixel value of an adjacent pixel point in the current initial pixel point group i to the next initial pixel point group i+1; Pl+k denotes a pixel value of a kth pixel point in the next initial pixel point group i+1; MSE denotes mean square error calculation; and Thr denotes a mean square error threshold. A similarity between the adjacent pixel point and other pixel points may be determined by a difference fed back by the mean square error threshold.

[0270] The if statement denotes a condition for pixel point migration (i.e., the preset condition) . If the condition is satisfied, the k pixel points migrate from the initial pixel point group i+1 to the initial pixel point group i. In addition, k ∈ Q, where Q denotes a count threshold. The mean square error threshold and the count threshold may be set as required.

[0271] For the two adjacent initial pixel point groups i and i+1, multiple iterations may be performed based on the mathematical equation (1) until an MSE calculation result is not less than the mean square error threshold. In this case, the expansion of the count of the pixel points in the current initial pixel point group i is completed, and an adjusted pixel point group i (i.e., a target pixel point group i) and an adjusted pixel point group i+1 are obtained. In some embodiments, the adjusted pixel point group i+1 may also be used as the first initial pixel point group in a next group of two adjacent initial pixel point groups for pixel expansion (e.g., k pixel points migrate from a pixel point group i+2 to the adjusted pixel point group i+1) , and an expanded pixel point group i+1 may be a target pixel point group i+1. In some embodiments, when the expansion of the count of the pixel points in the pixel point group i+1 is not performed, the adjusted pixel point group i+1 may be used as the target pixel point group i+1.

[0272] For example, taking an original video frame with a dimension of length*width*channel=768*512*3 as an example, if the preset continuous path is a hollow square  path from the outermost periphery to the center, the original video frame is divided into four initial pixel point groups, and the count of the pixel points in each of the initial pixel point groups is [39321, 78643, 117964, 157288] , respectively, and the pixel points in each of the initial pixel point groups are adjusted using the mathematical equation (1) . Assuming that the mean square error threshold is 20, the value of K corresponding to each two adjacent initial pixel point groups is 200, 300, and 150, respectively (i.e., the count of the pixel points migrating from the initial pixel point group 2, the initial pixel point group 3, and the initial pixel point group 4 is 200, 300, and 150, respectively) , then the finally obtained count of the pixel points in the four target pixel point groups is [39321+200=39521, 78643-200+300=78743, 117964-300+150=117814, 157288-150=157138] , respectively, where the value of K is equal to a sum of the values of k corresponding to the multiple iteration calculations, which may be equal to or greater than k. For example, if the adjustment of the pixel points of the two adjacent initial pixel point groups is completed after three iteration calculations, K = k1+k2+k3, where k1, k2 and k3 are the values of k in the first, second and third calculations respectively.

[0273] In some embodiments, the encoding terminal may adjust the pixel points in each two adjacent initial pixel point groups among the plurality of initial pixel point groups based on the preset condition. For example, for N initial pixel point groups, the encoding terminal may obtain a target pixel point group 1, a target pixel point group 2, a target pixel point group 3, ..., a target pixel point group n by adjust the pixel points in an initial pixel point group 1, an initial pixel point group 2, an initial pixel point group 3, ..., an initial pixel point group n-1 in sequence based on the preset condition. Since the expanded pixel points in the initial pixel point group n-1 migrate from the initial pixel point group n, the count of the pixel points in the initial pixel point group n also changes, and the target pixel point group n is obtained.

[0274] By adjusting the pixel points in the two adjacent initial pixel point groups, the pixel points in the same pixel point group can have a higher similarity, such that the results obtained by subsequent processing based on the pixel point group information are more accurate.

[0275] In 830, pixel point grouping information of the original video frame may be obtained based on the plurality of target pixel point groups.

[0276] Referring to the above, the encoding terminal may determine one or more mask images of the original video frame based on the plurality of target pixel point groups; and obtain the pixel point grouping information of the original video frame by processing the one or more mask images. For example, the encoding terminal may determine one or more mask images of the same size as the original video frame based on the plurality of target pixel point groups, a pixel value of each of the pixel points in the mask images being a grouping mark of a corresponding pixel point in the original video frame. Furthermore, the encoding terminal may determine the one or more mask images as the pixel point grouping information of the original video frame; or determine transformation information of the one or more mask images as the pixel point grouping information of the original video frame; or determine quantization information of the one or more mask images as the pixel point grouping information of the original video frame; or determine entropy encoding information of the one or more mask images as the pixel point grouping information of the original video frame. More descriptions may be found in the related descriptions of FIG. 7, which are not repeated here.

[0277] For example, taking transformation downsampling based on CNN as an example, referring to the above embodiment, for the four target pixel point groups whose counts of pixel points are [39521, 78743, 117814, 157138] , respectively, the encoding terminal may normalize mask images obtained based on the four target pixel point groups, and input the mask images into four 3*3 CNNs in turn for transformation to obtain the transformation information, wherein convolution kernels of the four CNNs move with stride = 2 (i.e., downsampling by one time) . Furthermore, the transformation information is denormalized, cropped to a range of [0, 3] , and then quantization to obtain the quantization information, and the quantization information is determined as the pixel point grouping information.

[0278] FIG. 10 is a flowchart illustrating an exemplary process of pixel point grouping according to some embodiments of the present disclosure.

[0279] In a process 1000, an encoding terminal may perform pixel point grouping on pixel points in an original video frame using a classification model based on scenario information. In some embodiments, the process 1000 may be performed by the processing device 110, the computing device 200 (e.g., the processor 210) , or the video encoding and decoding system 300 (e.g., the encoding module 310) . As shown in FIG. 10, in some embodiments, the process 1000 may include the following operations.

[0280] In 1010, a plurality of target pixel point groups of an original video frame may be obtained by processing pixel values of pixel points in the original video frame through a classification model.

[0281] The classification model may be configured to perform scenario classification on the pixel points contained in the original video frame based on the pixel values of the pixel points.

[0282] The scenario classification refers to dividing the pixel points into different scenario groups based on the pixel values. For example, if the original video frame is a landscape image, the pixel points contained in the original video frame may be divided into four scenario groups: sky, sea, ground, and plants, and each scenario group may correspond to one target pixel point group.

[0283] In some embodiments, as shown in FIG. 10, the encoding terminal may input an original video frame 1011 into a classification model 1012 as input data, and the classification model 1012 may analyze and process the pixel values of the pixel points in the original video frame and output a plurality of target pixel point groups 1013.

[0284] The classification model may be obtained by training an initial machine learning model using third sample data. The initial machine learning model may include but is not limited to a fully convolutional neural network (FCN) and a full-resolution residual network (FRRN) . As shown in FIG. 10, the third sample data may include a plurality of sample original video frames, and each sample original video frame label is a sample pixel point group. The sample pixel point groups may correspond to a plurality of scenario categories in the sample original video, and a count of the plurality of scenario categories may be flexibly set as required.

[0285] In some embodiments, an output of the classification model 1012 may include the plurality of target pixel point groups of the original video frame and the grouping mark of each of the target pixel point groups. In this case, the label of the third sample data may also include a mark of the sample pixel point group. Different scenarios correspond to different grouping marks, i.e., different target pixel point groups correspond to different grouping marks. For example, a grouping mark corresponding to a sea region is 0, a grouping mark corresponding to a sky region is 1, a grouping mark corresponding to a ground region is 2, and a grouping mark corresponding to a plant region is 3.

[0286] In 1020, pixel point grouping information of the original video frame may be obtained based on the plurality of target pixel point groups.

[0287] Referring to the above, the encoding terminal may determine the one or more mask images of the original video frame based on the plurality of target pixel point groups. In some embodiments, the encoding terminal may count the count of the pixel points contained in each of the target pixel point groups, and determine the grouping mark of each of the target pixel point group based on the count. For example, the encoding terminal may sort the plurality of target pixel point groups according to the count of the pixel points contained, mark the sorted target pixel point groups in turn (e.g., mark as 0, 1, 2, 3, and 4) , and obtain the grouping mark of each of the target pixel point groups. Furthermore, the encoding terminal may obtain the one or more mask images of the original video frame based on the grouping mark of the plurality of target pixel point groups. For example, the encoding terminal may determine the pixel values (e.g., determine values of the grouping marks as the pixel values) of the pixel points in each of the target pixel point groups in the mask images based on the grouping mark of each of the target pixel point groups, thereby forming one or more mask images of the same size as the original video frame.

[0288] Furthermore, the encoding terminal may determine the one or more mask images as the pixel point grouping information of the original video frame; or determine the transformation information of the one or more mask images as the pixel point grouping information of the original video frame; or determine the quantization information of the one or more mask images as the pixel point grouping information of the original video frame; or determine the entropy encoding information of the one or more mask images as the pixel point grouping information of the original video frame. More descriptions may be found in the related descriptions of FIG. 7, which are not repeated here.

[0289] For example, taking the dimension of the original video frame being length*width*channel=768*512*3 as an example, after the original video frame is input into the classification model of FCN, the classification model classifies similar pixel points into one category based on the pixel value of the original video frame, and obtains three scenario categories of sky, airplane and lawn, thereby outputting three target pixel point groups, and a count of the pixel points contained in the three target pixel point groups is [176948, 117964, 98304] , respectively. After the encoding terminal sorts the three target pixel point groups in a descending order based on the count of pixel points, each of the target pixel point groups is assigned with a grouping label, i.e., [2, 1, 0] , respectively. Furthermore, the encoding terminal obtains the one or more mask images corresponding to the original video frame based on the grouping label, normalizes the one or more mask images, and inputs the one or more mask images into three 3*3 CNNs for transformation to obtain the transformation information, where convolution kernels of the three CNNs move with stride = 2 (i.e., downsampling by one time) . Furthermore, the encoding terminal performs inverse normalization on the transformation information, crop the transformation information to a range of [0, 3], and then quantize the transformation information to obtain the quantization information, and determine the quantization information as the pixel point grouping information.

[0290] It should be noted that the above description of the processes 800 and 1000 is only for illustration and description, and does not limit the scope of application of the present disclosure. For those skilled in the art, various modifications and changes can be made to the processes 800 and 1000 under the guidance of the present disclosure. However, these modifications and changes are still within the scope of the present disclosure.

[0291] FIG. 11 a flowchart illustrating an exemplary video decoding method according to some embodiments of the present disclosure. A process 1100 may be implemented by a decoding terminal. In some embodiments, the process 1100 may be performed by the processing device 110, the computing device 200 (e.g., the processor 210) , or the video encoding and decoding system 300 (e.g., the decoding module 320) . As shown in FIG. 11, in some embodiments, the process 1100 may include the following operations.

[0292] In 1110, pixel point grouping information of an original video frame corresponding to a decoded target video frame may be obtained from an encoding terminal.

[0293] It should be noted that in this operation, the decoded target video frame has not yet been obtained based on encoded data. Here, the original video frame corresponding to the decoded target video frame refers to an original video frame corresponding to a current target video frame.

[0294] Referring to the above, the pixel point grouping information may be obtained by the encoding terminal by performing pixel point grouping on the original video frame based on an image content of the original video frame. More descriptions regarding the pixel point grouping information and pixel point grouping may be found in the related descriptions of FIGs. 7-10, which are not repeated here.

[0295] As shown in FIG. 7, the pixel point information of the original video frame may be transmitted to the decoding terminal independently from the encoded data of the original video frame (i.e., the pixel point information of the original video frame and the encoded data of the original video frame are not fused) , or transmitted to the decoding terminal after the pixel point information of the original video frame and the encoded data of the original video frame are fused. In this embodiment, for ease of understanding, the following operations are described based on the pixel point information and the encoded data being transmitted independently as an example.

[0296] It can be understood that if the pixel point grouping information of the original video frame and the encoded data of the original video frame are fused and transmitted to the decoding terminal, the decoding terminal can obtain the pixel point grouping information of the original video frame by decoding received data.

[0297] In 1120, entropy decoding data may be obtained by performing entropy decoding on encoded data based on the pixel point grouping information.

[0298] Entropy decoding corresponds to entropy encoding, which is an inverse process of the entropy encoding. Similar to the entropy encoding, an entropy decoding mode may include but is not limited to variable length decoding and arithmetical decoding.

[0299] In some embodiments, the decoding terminal may obtain first context information by processing the pixel point grouping information through a first context model; generate a first prediction probability based on the first context information through a first probability model; and obtain the entropy decoding data by performing the arithmetical decoding on the encoded data based on the first prediction probability.

[0300] The first context model refers to a model configured to provide context information by the decoding terminal. For example, the first context model may be configured to provide the first probability model with the first context information related to the encoded data to guide the first probability model to perform probability prediction, such that the first probability model can more accurately predict the probability of each character.

[0301] The first context model may include a bidirectional Transformer network.

[0302] The first probability model may be configured to determine a prediction probability for each character in a character string to be decoded (e.g., the encoded data) for the arithmetical decoding. The first probability model may be a mapping table that records a correspondence between each character and the probability of the character. When the entropy decoding is performed, the encoding terminal may determine the prediction probability corresponding to each character by searching the mapping table to generate the first prediction probability. The first probability model may be an algorithm or a program code capable of calculating a probability value of each character in the character string to be encoded. In some embodiments, the first probability model may be a dynamic probability model.

[0303] In some embodiments, the first context model and the first probability model may be two modules of the same machine learning model. For example, the first generative machine learning model described in FIG. 5 may include an entropy decoding module, an inverse quantization module, and an inverse transformation module. The entropy decoding module may be configured to perform entropy decoding on the encoded data based on the pixel point grouping information to obtain the entropy decoding data. The inverse quantization module may be configured to perform inverse quantization on the entropy decoding data to obtain inverse quantization data. The inverse transformation module may be configured to perform inverse transformation on first fusion data to obtain the decoded target video frame corresponding to the original video frame. The entropy decoding module may include a network structure composed of the first context model and the first probability model, and an arithmetical decoding unit. The pixel point information of the original video frame may be input into the entropy decoding module. The first context model of the entropy decoding module may analyze and process the pixel point information and output the first context information, and input the first context information into the first probability model. The first probability model may determine the first prediction probability corresponding to the encoded data based on the first context information. The arithmetical decoding unit may further perform the arithmetical decoding on the encoded data based on the first prediction probability to obtain the entropy decoding data. The entropy decoding module, the inverse quantization module, and the inverse transformation module in the first generative machine learning model may be jointly trained. The joint training may be similar to the training process of the first generative machine learning model described above (e.g., FIG. 4 or FIG. 5) , which is not repeated here. In some embodiments, the entropy decoding module, the inverse quantization module, and the inverse transformation module in the first generative machine learning model may be a machine learning model for implementing the corresponding functions, or a partial structure of the machine learning model for implementing the corresponding functions.

[0304] In some embodiments, the first context model may include a logical determination module and a generation module. The logical determination module may be configured to sort the plurality of pixel point groups based on the pixel point grouping information, and determine the pixel value information corresponding to each of the pixel point groups in the encoded data. The generation module may be configured to process the pixel value information corresponding to the pixel points in each of the pixel point groups in the encoded data to obtain the first context information corresponding to each of the pixel point groups.

[0305] In some embodiments, the decoding terminal may sort (e.g., sort the pixel points in an ascending order based on the grouping marks) the plurality of pixel point groups (e.g., the plurality of target pixel point groups obtained from the encoding terminal) based on the pixel point grouping information; determine encoding information corresponding to each of the pixel point groups in the encoded data; and perform the arithmetical decoding on the encoding information based on an order of the plurality of pixel point groups.

[0306] Since there is a correspondence between the pixel points in the original video frame and the pixel points in the mask images, accordingly, the encoded data obtained based on the original video frame (or the second fusion data of the original video frame) also has a pixel point correspondence with the pixel point information of the original video frame.

[0307] Taking the sorting of the pixel points in an ascending order based on the grouping marks as an example, performing the arithmetic decoding based on the order of the plurality of pixel point groups may specifically include the following operations for an Nth pixel point group (N is a positive integer) among the plurality of pixel point groups.

[0308] In response to determining that N is 1, the decoding terminal may assign a fixed character probability to encoding information corresponding to a first pixel point group, and perform the arithmetic decoding on the encoding information corresponding to each of pixel points in the first pixel point group in the encoded data in turn based on the fixed character probability to obtain entropy encoding data corresponding to the first pixel point group. Alternatively, the decoding terminal may input the encoding information corresponding to the first pixel point group in the encoded data into the first probability model to obtain a first prediction probability corresponding to the first pixel point group, and perform the arithmetic decoding on the encoding information corresponding to each of the pixel points in the first pixel point group in the encoded data in turn based on the first prediction probability corresponding to the first pixel point group to obtain the entropy encoding data corresponding to the first pixel point group.

[0309] In response to determining that N is 2, the decoding terminal may obtain first context information corresponding to the first pixel point group by inputting, into a first context model, the entropy decoding information corresponding to the first pixel point group in the encoded data; generate a first prediction probability corresponding to a second pixel point group based on the first context information corresponding to the first pixel point group through the first probability model; and obtain the entropy encoding data corresponding to the second pixel point group by performing the arithmetical decoding on the encoding information corresponding to each of the pixel points in the second pixel point group in the encoded data based on the first prediction probability corresponding to the second pixel point group.

[0310] In response to determining that N is greater than 2, the decoding terminal may input entropy decoding data corresponding to an (N-1) th pixel point group into the first context model, the first context model determining first context information corresponding to the (N-1) th pixel point group based on first context information corresponding from the first pixel point group to an (N-2) th pixel point group and the entropy decoding data corresponding to the (N-1) th pixel point group; generate a first prediction probability corresponding to the Nth pixel point group based on the first context information corresponding to the (N-1) th pixel point group through the second probability model; and obtain the encoded data corresponding to the Nth pixel point group by performing the arithmetical decoding on the encoding information corresponding to each of the pixel points in the Nth pixel point group based on the first prediction probability corresponding to the Nth pixel point group.

[0311] In some embodiments, similar to the encoding terminal, in order to further improve the prediction accuracy of the first context information, first auxiliary inverse transformation information in an auxiliary bitstream may be introduced.

[0312] Referring to the above, the auxiliary bitstream may be sent to the decoding terminal together with the encoded data. In some embodiments, the decoding terminal may obtain first auxiliary entropy decoding information by performing first auxiliary entropy decoding on the auxiliary bitstream; obtain first auxiliary inverse quantization information by performing auxiliary inverse quantization on the first auxiliary entropy decoding information; and obtain the first auxiliary inverse transformation information by performing auxiliary inverse transformation on the first auxiliary inverse quantization information.

[0313] In some embodiments, the decoding terminal may sort (e.g., sort the pixel points in an ascending order based on the grouping marks) the plurality of pixel point groups (e.g., the plurality of target pixel point groups obtained from the encoding terminal) based on the pixel point grouping information; determine the encoding information corresponding to each of the pixel point groups in the encoded data; and determine the entropy decoding data based on the second auxiliary transformation information according to the order of the plurality of pixel point groups.

[0314] Taking the sorting of the pixel points in an ascending order based on the grouping marks as an example, determining the entropy decoding data based on the second auxiliary transformation information according to the order of the plurality of pixel point groups may specifically include the following operations for the Nth pixel point group among the plurality of pixel point groups.

[0315] In response to determining that N is 1, the decoding terminal may assign the fixed character probability to the encoding information corresponding to the first pixel point group, and perform the arithmetic decoding on the encoding information corresponding to each of the pixel points in the first pixel point group in the encoded data in turn based on the fixed character probability to obtain the entropy decoding data corresponding to the first pixel point group. Alternatively, the decoding terminal may input the encoding information corresponding to the first pixel point group in the encoded data into the first probability model to obtain the first prediction probability corresponding to the first pixel point group, and perform the arithmetic decoding on the encoding information corresponding to each of the pixel points in the first pixel point group in the encoded data in turn based on the first prediction probability corresponding to the first pixel point group to obtain the entropy decoding data corresponding to the first pixel point group. Alternatively, the decoding terminal may obtain initial first context information by inputting at least a portion of data of the first auxiliary inverse transformation information (e.g., the pixel value information corresponding to the first pixel point group in the first auxiliary inverse transformation information, or the first auxiliary inverse transformation information) into the first context model; generate the first prediction probability corresponding to the first pixel point group based on the initial first context information through the first probability model; and obtain the entropy decoding data corresponding to the first pixel point group by performing the arithmetic decoding on the encoding information corresponding to each of the pixel points in the first pixel point group in the encoded data in turn based on the first prediction probability corresponding to the first pixel point group.

[0316] In response to determining that N is 2, the decoding terminal may obtain the first context information corresponding to the first pixel point group by inputting, into the first context model, the entropy decoding data corresponding to the first pixel point group and at least the portion of the data of the first auxiliary inverse transformation information; generate a first prediction probability corresponding to the second pixel point group based on the first context information corresponding to the first pixel point group through the first probability model; and obtain entropy decoding data corresponding to the second pixel point group by performing the arithmetical decoding on encoding information corresponding to each of the pixel points in the second pixel point group in the encoded data based on the first prediction probability corresponding to the second pixel point group.

[0317] In response to determining that N is greater than 2, the decoding terminal may input entropy decoding data corresponding to the (N-1) th pixel point group and at least the portion of the data of the first auxiliary inverse transformation information into the first context model, the first context model determining first context information corresponding to the (N-1) th pixel point group based on first context information corresponding from the first pixel point group to the (N-2) th pixel point group, the entropy decoding data corresponding to the (N-1) th pixel point group, and at least the portion of the data of the first auxiliary inverse transformation information; generate the first prediction probability corresponding to the Nth pixel point group based on the first context information corresponding to the (N-1) th pixel point group through the first probability model; and obtain the entropy decoding data corresponding to the Nth pixel point group by performing the arithmetical decoding on the encoding information corresponding to the Nth pixel point group in the encoded data based on the first prediction probability corresponding to the Nth pixel point group.

[0318] At least the portion of the data of the second auxiliary inverse transformation information may include pixel value information corresponding to the pixel point groups in the first auxiliary inverse transformation information. For example, for the arithmetical decoding corresponding to the second pixel point group, the decoding terminal may obtain the first context information corresponding to the first pixel point group by inputting the entropy decoding data corresponding to the first pixel point group, and the pixel value information of the pixel point corresponding to each of the pixel points in the second pixel point group in the first auxiliary inverse transformation information into the first context model. As another example, for the arithmetical decoding corresponding to the (N-1) th pixel point group, the decoding terminal may input the entropy decoding data corresponding to the (N-2) th pixel point group, and the pixel value information corresponding to each of the pixel points the (N-1) th pixel point group in the first auxiliary inverse transformation information into the first context model. The first context model may determine the first context information corresponding to the (N-1) th pixel point group based on first context information corresponding from the first pixel point group to an (N-3) th pixel point group, the entropy decoding data corresponding to the (N-2) th pixel point group, and the pixel value information corresponding to each of the pixel points in the (N-1) th pixel point group in the first auxiliary inverse transformation information.

[0319] In some embodiments, in each decoding operation, all the first auxiliary inverse transformation information may be input into the first context model. In this case, each of the pixel point groups may share the same first auxiliary inverse transformation information. For example, for the arithmetical decoding corresponding to the (N-1) th pixel point group, the decoding terminal may input the entropy decoding data corresponding to the (N-1) th pixel point group and the first auxiliary inverse transformation information into the first context model.

[0320] In 1130, inverse quantization data may be obtained by performing inverse quantization on the entropy decoding data.

[0321] The inverse quantization is an inverse process of quantization. For example, a finite count of discrete values may be mapped to a continuous range of values by the inverse quantization.

[0322] In some embodiments, the decoding terminal may obtain a plurality of inverse quantization data corresponding to the plurality of pixel point groups by performing the inverse quantization on the entropy decoding data corresponding to each of the pixel point groups, and obtain the inverse quantization data corresponding to the encoded data by fusing the plurality of inverse quantization data.

[0323] In some embodiments, the decoding terminal may obtain the inverse quantization data by fusing the entropy decoding data corresponding to each of the pixel point groups and then performing the inverse quantization.

[0324] In 1140, first fusion data may be obtained by fusing the inverse quantization data with first reference information.

[0325] In some embodiments, a fusion mode may include but is not limited to splicing, addition, multiplication, weighted summation, etc. For example, the decoding module 320 may obtain the first fusion data by splicing the inverse quantization data and the first reference information. As another example, the decoding module 320 may obtain the first fusion data by performing the weighted summation on the inverse quantization data and the first reference information.

[0326] In 1150, the decoded target video frame may be obtained by performing an inverse transformation on the first fusion data.

[0327] The decoding terminal may obtain the decoded target video frame corresponding to the original video frame by performing the inverse transformation on the first fusion data. The inverse transformation is an inverse operation corresponding to the transformation of the encoding terminal, i.e., an inverse operation of the transformation. For example, the inverse transformation may be an inverse process of downsampling used by the encoding terminal, which is also referred to as upsampling.

[0328] Similar to the transformation, an upsampling mode of the inverse transformation may include but is not limited to bilinear upsampling, transformation upsampling based on a CNN, transformation upsampling based on an MLP neural network, etc.

[0329] It should be noted that the upsampling mode used by the decoding terminal corresponds to the downsampling mode used by the encoding terminal. For example, if the encoding terminal performs the transformation using the transformation downsampling based on the CNN, the decoding terminal performs the inverse transformation using the transformation upsampling based on the CNN.

[0330] It should be noted that the above description of the process 1100 is only for example and explanation, and does not limit the scope of application of the present disclosure. For those skilled in the art, various modifications and changes can be made to the process 1100 under the guidance of the present disclosure. However, these modifications and changes are still within the scope of the present disclosure. For example, the operation 1140 in the process 1100 may be omitted. In this case, in the operation 1150, the decoding terminal may obtain the decoded target video frame by performing the inverse transformation on the inverse quantization data. As another example, in the operations 1130 and 1150, the decoding terminal may obtain the decoded target video frame by performing the inverse quantization, the inverse transformation and other processing on the encoded data corresponding to each of the pixel point groups (e.g., the target pixel point groups) , and then fusing the processed data corresponding to each of the pixel point groups.

[0331] FIG. 12 a flowchart illustrating an exemplary video encoding method according to some embodiments of the present disclosure. A process 1200 may be implemented by an encoding terminal. In some embodiments, the process 1200 may be performed by the processing device 110, the computing device 200 (e.g., the processor 210) , or the video encoding and decoding system 300 (e.g., an encoding module 310) . As shown in FIG. 12, in some embodiments, the process 1200 may include the following operations.

[0332] In 1210, pixel point grouping information of an original video frame may be obtained by performing pixel point grouping on the original video frame based on an image content of the original video frame.

[0333] In some embodiments, the encoding terminal may perform pixel point grouping on the original video frame based on a correlation (e.g., a similarity between pixel values) of the pixel values of pixel points in the original video frame. Furthermore, the encoding terminal may mark the pixel points in the original video frame according to the grouping of the pixel points, thereby obtaining the pixel point grouping information of the original video frame.

[0334] In some embodiments, the encoding terminal may perform pixel point grouping on the original video frame based on the image content of the original video frame according to a preset continuous path. Specifically, the encoding module 310 may obtain a plurality of initial pixel point groups by traversing each of the pixel points in the original video frame in sequence according to the preset continuous path; obtain a plurality of target pixel point groups of the original video frame by adjusting pixel points in at least a portion of the plurality of initial pixel point groups. More descriptions regarding performing pixel point grouping according to the preset continuous path may be found in FIG. 8 and the related descriptions thereof.

[0335] In some embodiments, the encoding terminal may obtain the plurality of target pixel point groups of the original video frame by processing the pixel value of each of the pixel points in the original video frame through a classification model. More descriptions regarding performing pixel point grouping using the classification model may be found in FIG. 10 and the related descriptions thereof.

[0336] In some embodiments, the encoding terminal may assign corresponding grouping marks to the plurality of target pixel point groups, determine one or more mask images corresponding to the original video frame based on the grouping marks of the plurality of target pixel point groups, and obtain the pixel point grouping information of the original video frame based on the one or more mask images. More descriptions may be found in FIG. 7 and the related descriptions thereof (e.g., the operation 710) .

[0337] In 1220, second transformation data may be obtained by transforming the original video frame.

[0338] In some embodiments, the encoding terminal may perform nonlinear downsampling on the original video frame using the CNN so as to express a main feature of the original video frame using the more compact expression (e.g., a bitstream) , and obtain the second transformation data.

[0339] In 1230, second quantization data may be obtained by quantizing the second transformation data.

[0340] In some embodiments, the encoding terminal may obtain the second quantization data by mapping a continuous range of values corresponding to the second transformation data to a finite count of discrete values.

[0341] In 1240, encoded data may be obtained by performing entropy encoding on the second quantization data based on the pixel point grouping information.

[0342] In some embodiments, the encoding terminal may obtain second context information by processing the pixel point grouping information through a second context model; generate a second prediction probability based on the second context information through a second probability model; and obtain the encoded data by performing arithmetical encoding on the second quantization data based on the second prediction probability.

[0343] In some embodiments, the encoding terminal may sort (e.g., sort the pixel points in an ascending order based on the grouping marks) the plurality of pixel point groups (e.g., the plurality of target pixel point groups) based on the pixel point grouping information; determine pixel value information corresponding to each of the pixel point groups in the second quantization data; and determine the encoded data based on an order of the plurality of pixel point groups.

[0344] In some embodiments, the encoding terminal may obtain auxiliary transformation information by performing an auxiliary transformation on the second transformation data; obtain auxiliary quantization information by performing auxiliary quantization on the auxiliary transformation information; and obtain an auxiliary bitstream by performing auxiliary entropy encoding on the auxiliary quantization information. Furthermore, the encoding terminal may obtain second auxiliary entropy decoding information by performing second auxiliary entropy decoding on the auxiliary bitstream; obtain second auxiliary inverse quantization information by performing auxiliary inverse quantization on the second auxiliary entropy decoding information; and obtain second auxiliary inverse transformation information by performing an auxiliary inverse transformation on the second auxiliary inverse quantization information.

[0345] In some embodiments, the encoding terminal may sort (e.g., sort the pixel points in an ascending order based on the grouping marks) the plurality of pixel point groups (e.g., the plurality of target pixel point groups) based on the pixel point grouping information; determine the pixel value information corresponding to each of the pixel point groups in the second quantization data; and determine the encoded data based on the first auxiliary transformation information according to the order of the plurality of pixel point groups.

[0346] An entropy encoding mode of the second quantization data based on the pixel point grouping information may be similar to an entropy encoding mode of the first quantization data based on the pixel point grouping information. More descriptions may be found in FIG. 7 and the related descriptions thereof (e.g., the operation 750) .

[0347] In 1250, the encoded data may be sent to a decoding terminal.

[0348] The encoding terminal may transmit the encoded data obtained in the operation 1240 to the decoding terminal, such that the decoding terminal performs decoding and reconstruction based on the encoded data to obtain the decoded target video frame corresponding to the original video frame.

[0349] In some embodiments, a transformation module (configured to transform the original video frame to obtain the second transformation data) , a quantization module (configured to quantize the second transformation data to obtain the second quantization data) , a pixel point grouping information determination module (configured to perform pixel point grouping on the original video frame based on the image content of the original video frame to obtain the pixel point grouping information of the original video frame) , an entropy encoding module (configured to perform entropy encoding on the second quantization data based on the pixel point grouping information to obtain the encoded data, such as including an arithmetic encoding unit, the second context model, and the second probability model) , an auxiliary transformation module (configured for auxiliary transformation) , an auxiliary quantization module (configured for auxiliary quantization) , an auxiliary entropy encoding module (configured for auxiliary entropy encoding) , an auxiliary entropy decoding module (configured for auxiliary entropy decoding) , an auxiliary inverse quantization module (configured for auxiliary inverse quantization) , and an auxiliary inverse transform module (configured for auxiliary inverse transformation) of the encoding terminal may be trained separately or jointly. In some embodiments, each module of the encoding terminal may be a machine learning model for implementing a corresponding function, or a partial structure in the machine learning model for implementing a corresponding function. For example, each module of the encoding terminal may be a layer structure for implementing a corresponding function in the second generative machine learning model described above, and each module may correspond to one or more layers in the second generative machine learning model.

[0350] It should be noted that the above description of the process 1200 is only for example and explanation, and does not limit the scope of application of the present disclosure. For those skilled in the art, various modifications and changes can be made to the process 1200 under the guidance of the present disclosure. However, these modifications and changes are still within the scope of the present disclosure. For example, in the operation 1240, the encoding terminal may input the original video frame and the pixel point grouping information into a generative model composed of the second context model and the second probability model, and the generative model may determine the second context information based on the pixel point grouping information, obtain the second prediction probability based on the second context information, and guide the encoding process based on the second prediction probability, so as to obtain the encoded data.

[0351] FIG. 13 is a flowchart illustrating an exemplary video decoding method according to some embodiments of the present disclosure. A process 1300 may be implemented by a decoding terminal. In some embodiments, the process 1300 may be performed by the processing device 110, the computing device 200 (e.g., the processor 210) , or the video encoding and decoding system 300 (e.g., the decoding module 320) . As shown in FIG. 13, in some embodiments, the process 1300 may include the following operations.

[0352] In 1310, pixel point grouping information of an original video frame may be obtained from an encoding terminal.

[0353] Referring to the above, the pixel point grouping information may be obtained by performing pixel point grouping on the original video frame based on an image content of the original video frame. More descriptions regarding the pixel point grouping information and the pixel point grouping may be found in FIGs. 7-11 and the related descriptions thereof, which are not repeated here.

[0354] In 1320, entropy decoding data may be obtained by performing entropy decoding on encoded data based on the pixel point grouping information.

[0355] In some embodiments, the decoding terminal may obtain first context information by processing the pixel point grouping information through a first context model; generate a first prediction probability based on the first context information through a first probability model; and obtain the entropy decoding data by performing arithmetical decoding on the encoded data based on the first prediction probability.

[0356] In some embodiments, the decoding terminal may sort (e.g., sort the pixel points in an ascending order based on the grouping marks) a plurality of pixel point groups (e.g., the plurality of target pixel point groups) based on the pixel point grouping information; determine encoding information corresponding to each of the pixel point groups in the encoded data; and perform the arithmetical decoding on the encoding information based on an order of the plurality of pixel point groups.

[0357] Referring to the above, an auxiliary bitstream may be sent to the decoding terminal together with the encoded data. In some embodiments, the decoding terminal may obtain first auxiliary entropy decoding information by performing first auxiliary entropy decoding on the auxiliary bitstream; obtain first auxiliary inverse quantization information by performing auxiliary inverse quantization on the first auxiliary entropy decoding information; and obtain first auxiliary inverse transformation information by performing auxiliary inverse transformation on the first auxiliary inverse quantization information.

[0358] In some embodiments, the decoding terminal may sort (e.g., sort the pixel points in an ascending order based on the grouping marks) the plurality of pixel point groups (e.g., the plurality of target pixel point groups obtained from the encoding terminal) based on the pixel point grouping information; determine the encoding information corresponding to each of the pixel point groups in the encoded data; and determine the entropy decoding data based on the second auxiliary transformation information according to the order of the plurality of pixel point groups.

[0359] The operation 1320 may be similar to the operation 1120. More descriptions may be found in the related descriptions of the operation 1120, which are not repeated here.

[0360] In 1330, inverse quantization data may be obtained by performing inverse quantization on the entropy decoding data.

[0361] In some embodiments, the decoding terminal may obtain a plurality of inverse quantization data corresponding to the plurality of pixel point groups by performing the inverse quantization on the entropy decoding data corresponding to each of the pixel point groups, and obtain the inverse quantization data corresponding to the encoded data by fusing the plurality of inverse quantization data.

[0362] In some embodiments, the decoding terminal may obtain the inverse quantization data by fusing the entropy decoding data corresponding to each of the pixel point groups and then performing the inverse quantization.

[0363] In 1340, a decoded target video frame corresponding to the original video frame may be obtained by performing an inverse transformation on the inverse quantization data.

[0364] It should be noted that the upsampling mode used by the decoding terminal corresponds to the downsampling mode used by the encoding terminal. For example, if the encoding terminal performs the transformation using the transformation downsampling based on the CNN, the decoding terminal performs the inverse transformation using the transformation upsampling based on the CNN. More descriptions may be found in FIG. 11 and the related descriptions thereof, which are not repeated here.

[0365] In some embodiments, an inverse transformation module (configured to perform inverse transformation the inverse quantization data to obtain the decoded target video frame corresponding to the original video frame) , an inverse quantization module (configured to perform inversely quantization on the entropy decoding data to obtain the inverse quantization data) , an entropy decoding module (configured to perform entropy decoding on the encoded data based on the pixel point grouping information to obtain the entropy decoding data, such as including an arithmetic decoding unit, the first context model, and the first probability model) , an auxiliary entropy decoding module (configured for auxiliary entropy decoding) , the auxiliary inverse quantization module (configured for auxiliary inverse quantization) , and an auxiliary inverse transform module (configured for auxiliary inverse transformation) of the decoding terminal may be trained separately or jointly. In some embodiments, each module of the decoding terminal may be a machine learning model for implementing a corresponding function, or a partial structure in the machine learning model for implementing a corresponding function. For example, each module of the decoding terminal may be a layer structure in the first generative machine learning model for implementing a corresponding function, and each module may correspond to one or more layers in the first generative machine learning model.

[0366] In some embodiments, the modules (e.g., the pixel point grouping information determination module, the transformation module, the quantization module, the entropy encoding module, the auxiliary transformation module, the auxiliary quantization module, and the auxiliary entropy encoding module) of the encoding terminal and the modules (e.g., the inverse transformation module, the inverse quantization module, the entropy decoding module, the auxiliary entropy decoding module, the auxiliary inverse quantization module, and the auxiliary inverse transformation module) of the decoding terminal may be jointly trained. For example, training data containing a sample original video frame, or training data containing the sample original video frame and sample reference information may be input into the encoding terminal (e.g., input into the transformation module of the encoding terminal) , and the training data may be processed in sequence by each module of the encoding terminal to obtain sample encoding data (e.g., output by the entropy encoding module) and sample pixel point grouping information (e.g., output by the pixel point grouping information determination module) . Furthermore, the sample encoding data and the sample pixel point grouping information may be input into the decoding terminal (e.g., input into the entropy decoding module of the decoding terminal) , and the sample encoding data and the sample pixel point grouping information may be processed in sequence by each module of the decoding terminal to obtain a sample reconstructed video frame (e.g., output by the inverse transformation module) . During the training process, a first joint loss function D0 + λR0 may be constructed to perform joint parameter adjustment on each module of the encoding terminal and the decoding terminal, such that a value of the first joint loss function converges or is less than a first joint loss threshold, thereby completing the training of each module of the encoding terminal and the decoding terminal. In this case, the auxiliary entropy decoding module, the auxiliary inverse quantization module, and the auxiliary inverse transformation module obtained by the joint training may be provided at the encoding terminal and the decoding terminal simultaneously.

[0367] It should be noted that the above description of the process 1300 is only for example and explanation, and does not limit the scope of application of the present disclosure. For those skilled in the art, various modifications and changes can be made to the process 1300 under the guidance of the present disclosure. However, these modifications and changes are still within the scope of the present disclosure. For example, the decoding terminal may input the encoded data and the pixel point grouping information into a generative model composed of the first context model and the first probability model, and the generative model may determine the first context information based on the pixel point grouping information, determine the first prediction probability based on the first context information, and guide the decoding process of the encoded data based on the first prediction probability, so as to obtain the decoded target video frame.

[0368] FIG. 14 is a schematic diagram illustrating video encoding and decoding according to some embodiments of the present disclosure.

[0369] In some embodiments, as shown in FIG. 14, in response to determining that an original video frame is obtained, an encoding terminal may perform transformation, or transformation and quantization on the original video frame, and perform entropy encoding (e.g., arithmetical encoding) on processed data to obtain a main bitstream corresponding to the original video frame. The encoding terminal may perform pixel point grouping on the original video frame based on an image content of the original video frame to obtain a plurality of target pixel point groups of the original video frame, assign corresponding grouping marks to the plurality of target pixel point groups, and perform masking on the plurality of target pixel point groups in a mode matching an encoding order based on the grouping marks of the plurality of target pixel point groups to obtain a plurality of mask images corresponding to the original video frame. The encoding order may be correlated with a count of pixel points contained in the target pixel point groups. For example, the smaller the count, the higher the encoding order. It can be understood that one of the pixel points in the same target pixel point group is adjacent to at least one of the pixel points in another target pixel point group, thereby obtaining the plurality of target pixel point groups with continuous pixel arrangement, and achieving the integrity of the plurality of target pixel point groups.

[0370] Furthermore, the encoding terminal may transform the plurality of mask images, and perform entropy encoding on dimensionality reduction features to obtain a corresponding bitstream. The bitstream and the main bitstream of the original video frame may be sent to the decoding terminal together. Alternatively, the encoding terminal may transform and quantize the plurality of mask images, and input the plurality of mask images into the second context model provided at the encoding terminal, and the second context model may output the second context information by analyzing and processing the input data. The second probability model provided at the encoding terminal may determine the second prediction probability based on the second context information. The encoding terminal may adjust or re-encode the main bitstream corresponding to the original video frame based on the second prediction probability, and transmit an adjusted or re-encoded bitstream to the decoding terminal. The higher the second prediction probability, the fewer bits required for the corresponding character. Accordingly, the more accurate the probability prediction of the character by the second probability model, the higher the overall encoding performance of an encoder-decoder.

[0371] In some embodiments, the encoding terminal may obtain auxiliary transformation information by performing auxiliary transformation on the transformation data of the original video frame, and obtain the auxiliary bitstream by performing auxiliary entropy encoding (e.g., the arithmetical encoding) on the auxiliary transformation information. Optionally, the encoding may obtain auxiliary quantization information by performing auxiliary quantization on the auxiliary transformation information; and obtain the auxiliary bitstream by performing the auxiliary entropy encoding on the auxiliary quantization information. The auxiliary bitstream may be sent to the decoding terminal.

[0372] In some embodiments, as shown in FIG. 14, the encoding terminal may determine a fixed prediction probability using a fixed probability model, and obtain the auxiliary bitstream by performing an auxiliary entropy encoding operation based on the fixed prediction probability. Furthermore, the encoding terminal and / or the decoding terminal may obtain the auxiliary inverse transformation information (e.g., the second auxiliary inverse transformation information or the first auxiliary inverse transformation information) by performing auxiliary entropy decoding, auxiliary inverse quantization, and auxiliary inverse transformation on the auxiliary bitstream using the fixed probability model. The second auxiliary inverse transformation information may be input into the second context model, and the second context model may obtain the second context information using the second transformation data after the transformation + the second quantization data after the quantization of the original video frame, data after the transformation or data after the transformation + the quantization of the one or more mask images, and the second auxiliary inverse transformation information, thereby guiding the second probability model to perform probability prediction.

[0373] The fixed probability model is a simple entropy model, and parameters of the fixed probability model are obtained by offline learning. Obtaining the auxiliary bitstream or decoding the auxiliary bitstream through the fixed probability model can make at least one decoded video frame before the original video frame used as a reference video frame for generating the auxiliary bitstream, so as to assist the encoding and decoding prediction of the original video frame.

[0374] Corresponding to the encoding terminal, the decoding terminal may obtain the inverse quantization data by performing the entropy decoding and the inverse quantization on the bitstream corresponding to the one or more mask images and the main bitstream (e.g., the encoded data) corresponding to the original video frame, and input the inverse quantization data into the first context model provided at the decoding terminal. Optionally, the decoding terminal may obtain the first auxiliary inverse transformation information by performing the auxiliary entropy decoding, the auxiliary inverse quantization, and the auxiliary inverse transformation on the auxiliary bitstream using the fixed probability model, and input the first auxiliary inverse transformation information into the first context model. The first context model may obtain the first context information based on the inverse quantization data corresponding to the one or more mask images and the original video frame and the first auxiliary inverse transformation information, so as to guide the first probability model to perform probability prediction. The decoding terminal may perform decoding based on the first prediction probability output by the first probability model, and obtain the pixel points in all the reconstructed pixel point groups that match the encoding order in sequence to constitute the decoded target video frame corresponding to the original video frame.

[0375] The second context model and the first context model may be constructed based on a bidirectional Transformer network. Auxiliary features (e.g., the second auxiliary inverse transformation information or the first auxiliary inverse transformation information) obtained after the auxiliary inverse transformation may be directly used as the input data of the second context model and the first context model, such that the first / first context model can simultaneously refer to the decoded information and the enhanced information of the auxiliary bitstream to guide the first / first probability model to sequentially decode and reconstruct the plurality of target pixel points in groups with higher accuracy.

[0376] In some embodiments, for the complete encoding and decoding network shown in FIG. 14, an encoding and decoding branch of the auxiliary bitstream corresponding to the first and first context models may be frozen first, and other branches (e.g., an encoding and decoding branch of the bitstream corresponding to the original video frame, and an encoding and decoding branch of the bitstream corresponding to the mask images) of the first and first context models may be trained. After the first / first context model converge, the second auxiliary inverse transformation information / the first auxiliary inverse transformation information may be further used as the training input to enhance the richness of the training samples, and the encoding / decoding branches of the auxiliary bitstream of the first / first context model may be trained until convergence to obtain a complete first / first context model.

[0377] One or more embodiments of the present disclosure further provide an electronic device, including a memory (e.g., the storage device 120, and the memory 223) and a processor (e.g., the processing device 110, and the processor 210) coupled with each other. The memory may be configured to store program data (e.g., a computer program) . When the processor calls the program data, the video encoding and / or video decoding method in any of the above embodiments is implemented. More descriptions may be found in the detailed description of the method embodiments described above, which are not repeated here.

[0378] One or more embodiments of the present disclosure further provide a non-transitory computer-readable storage medium (e.g., the non-volatile storage medium 225) . The computer-readable storage medium may be configured to store program data (e.g., a computer program) . When the program data is executed by a processor (e.g., the processor 210) , the video encoding and / or video decoding method in any of the above embodiments is implemented. More descriptions may be found in the detailed description of the method embodiments described above, which are not repeated here.

[0379] Having thus described the basic concepts, it may be rather apparent to those skilled in the art after reading this detailed disclosure that the foregoing detailed disclosure is intended to be presented by way of example only and is not limiting. Various alterations, improvements, and modifications may occur and are intended to those skilled in the art, though not expressly stated herein. These alterations, improvements, and modifications are intended to be suggested by this disclosure and are within the spirit and scope of the exemplary embodiments of this disclosure.

[0380] Moreover, certain terminology has been used to describe embodiments of the present disclosure. For example, the terms “one embodiment, ” “an embodiment, ” and “some embodiments” mean that a particular feature, structure, or feature described in connection with the embodiment is included in at least one embodiment of the present disclosure. Therefore, it is emphasized and should be appreciated that two or more references to “an embodiment” or “one embodiment” or “an alternative embodiment” in various portions of this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or features may be combined as suitable in one or more embodiments of the present disclosure.

[0381] Furthermore, the recited order of processing elements or sequences, or the use of numbers, letters, or other designations therefore, is not intended to limit the claimed processes and methods to any order except as may be specified in the claims. Although the above disclosure discusses through various examples what is currently considered to be a variety of useful embodiments of the disclosure, it is to be understood that such detail is solely for that purpose and that the appended claims are not limited to the disclosed embodiments, but, on the contrary, are intended to cover modifications and equivalent arrangements that are within the spirit and scope of the disclosed embodiments. For example, although the implementation of various parts described above may be embodied in a hardware device, it may also be implemented as a software only solution, e.g., an installation on an existing server or mobile device.

[0382] Similarly, it should be appreciated that in the foregoing description of embodiments of the present disclosure, various features are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure aiding in the understanding of one or more of the various embodiments. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed subject matter requires more features than are expressly recited in each claim. Rather, claimed subject matter may lie in less than all features of a single foregoing disclosed embodiment.

[0383] In some embodiments, numbers describing the number of ingredients and attributes are used. It should be understood that such numbers used for the description of the embodiments use the modifier “about” , “approximately” , or “substantially” in some examples. Unless otherwise stated, “about” , “approximately” , or “substantially” indicates that the number is allowed to vary by ±20%. Correspondingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, and the approximate values may be changed according to the required features of individual embodiments. In some embodiments, the numerical parameters should consider the prescribed effective digits and adopt the method of general digit retention. Although the numerical ranges and parameters used to confirm the breadth of the range in some embodiments of the present disclosure are approximate values, in specific embodiments, settings of such numerical values are as accurate as possible within a feasible range.

[0384] For each patent, patent application, patent application publication, or other materials cited in the present disclosure, such as articles, books, specifications, publications, documents, or the like, the entire contents of which are hereby incorporated into the present disclosure as a reference. The application history documents that are inconsistent or conflict with the content of the present disclosure are excluded, and the documents that restrict the broadest scope of the claims of the present disclosure (currently or later attached to the present disclosure) are also excluded. It should be noted that if there is any inconsistency or conflict between the description, definition, and / or use of terms in the auxiliary materials of the present disclosure and the content of the present disclosure, the description, definition, and / or use of terms in the present disclosure is subject to the present disclosure.

[0385] Finally, it should be understood that the embodiments described in the present disclosure are only used to illustrate the principles of the embodiments of the present disclosure. Other variations may also fall within the scope of the present disclosure. Therefore, as an example and not a limitation, alternative configurations of the embodiments of the present disclosure may be regarded as consistent with the teaching of the present disclosure. Accordingly, the embodiments of the present disclosure are not limited to the embodiments introduced and described in the present disclosure explicitly.

Claims

1.A video decoding method, implemented by a decoding terminal, comprising:obtaining encoded data and first reference information from an encoding terminal, the first reference information including feature information of at least one decoded video frame before a target video frame; andobtaining a decoded target video frame based on the encoded data and the first reference information.2.The video decoding method of claim 1, wherein the obtaining the decoded target video frame based on the encoded data and the first reference information includes:obtaining the decoded target video frame by processing the encoded data and the first reference information through a first generative machine learning model.3.The video decoding method of claim 1, further comprising:obtaining prompt information from the encoding terminal, the prompt information reflecting at least one of following information related to the decoded target video frame: pixel value information, scaling information, timing information, segmented semantic information, and classified semantic information;the obtaining the decoded target video frame based on the encoded data and the first reference information includes: obtaining the decoded target video frame by processing the encoded data, the first reference information, and the prompt information through a first generative machine learning model.4.The video decoding method of claim 3, further comprising:determining first prompt data input by a user;the obtaining the decoded target video frame based on the encoded data and the first reference information includes: obtaining the decoded target video frame by processing the encoded data, the first reference information, the prompt information, and the first prompt data through the first generative machine learning model.5.The video decoding method of claim 4, wherein the first prompt data is configured to indicate an adjustment of an image content of the decoded target video frame during a process of obtaining the decoded target video frame.6.The video decoding method of any one of claims 1-5, further comprising:obtaining second prompt data from the encoding terminal;the obtaining the decoded target video frame based on the encoded data and the first reference information includes: obtaining the decoded target video frame by processing the encoded data, the first reference information, and the second prompt data, and / or the first prompt data through a first generative machine learning model.7.The video decoding method of claim 1, wherein the obtaining the decoded target video frame based on the encoded data and the first reference information includes:obtaining pixel point grouping information of an original video frame corresponding to the decoded target video frame from the encoding terminal, the pixel point grouping information being obtained by performing pixel point grouping on the original video frame based on an image content of the original video frame;obtaining entropy decoding data by performing entropy decoding on the encoded data based on the pixel point grouping information;obtaining inverse quantization data by performing inverse quantization on the entropy decoding data;obtaining first fusion data by fusing the inverse quantization data with the first reference information; andobtaining the decoded target video frame by performing an inverse transformation on the first fusion data.8.The video decoding method of claim 7, wherein the obtaining entropy decoding data by performing entropy decoding on the encoded data based on the pixel point grouping information includes:obtaining first context information by processing the pixel point grouping information through a first context model;generating a first prediction probability based on the first context information through a first probability model; andobtaining the entropy decoding data by performing arithmetical decoding on the encoded data based on the first prediction probability.9.The video decoding method of claim 7, wherein the obtaining entropy decoding data by performing entropy decoding on the encoded data based on the pixel point grouping information includes:sorting a plurality of pixel point groups based on the pixel point grouping information;determining encoding information corresponding to each of the plurality of pixel point groups in the encoded data; andperforming following operations on an Nth pixel point group of the plurality of pixel point groups based on an order of the plurality of pixel point groups, N being a positive integer;in response to determining that N is 2, obtaining first context information corresponding to a first pixel point group by inputting, into the first context model, entropy decoding data corresponding to the first pixel point group;generating a first prediction probability corresponding to a second pixel point group based on the first context information corresponding to the first pixel point group through a first probability model; andobtaining entropy decoding data corresponding to the second pixel point group by performing arithmetical decoding on encoding information corresponding to the second pixel point group based on the second prediction probability corresponding to the second pixel point group; andin response to determining that N is greater than 2, inputting entropy decoding data corresponding to an (N-1) th pixel point group into the first context model, the first context model determining first context information corresponding to the (N-1) th pixel point group based on first context information corresponding from the first pixel point group to an (N-2) th pixel point group, and the entropy decoding data corresponding to the (N-1) th pixel point group;generating a first prediction probability corresponding to the Nth pixel point group based on the first context information corresponding to the (N-1) th pixel point group through the first probability model; andobtaining entropy decoding data corresponding to the Nth pixel point group by performing arithmetical decoding on encoding information corresponding to the Nth pixel point group based on the second prediction probability corresponding to the Nth pixel point group.10.The video decoding method of claim 7, further comprising:obtaining an auxiliary bitstream from the encoding terminal;obtaining first auxiliary entropy decoding information by performing first auxiliary entropy decoding on the auxiliary bitstream;obtaining first auxiliary inverse quantization information by performing auxiliary inverse quantization on the first auxiliary entropy decoding information; andobtaining first auxiliary inverse transformation information by performing an auxiliary inverse transformation on the first auxiliary inverse quantization information.11.The video decoding method of claim 10, wherein the obtaining entropy decoding data by performing entropy decoding on the encoded data based on the pixel point grouping information includes:sorting a plurality of pixel point groups based on the pixel point grouping information;determining encoding information corresponding to each of the plurality of pixel point groups in the encoded data; andperforming following operations on an Nth pixel point group of the plurality of pixel point groups based on an order of the plurality of pixel point groups, N being a positive integer;in response to determining that N is 2, obtaining first context information corresponding to a first pixel point group by inputting, into a first context model, entropy decoding data corresponding to the first pixel point group and at least a portion data of the first auxiliary inverse transformation information;generating a first prediction probability corresponding to a second pixel point group based on the first context information corresponding to the first pixel point group through a first probability model; andobtaining entropy decoding data corresponding to the second pixel point group by performing arithmetical decoding on encoding information corresponding to the second pixel point group based on the first prediction probability corresponding to the second pixel point group; andin response to determining that N is greater than 2, inputting entropy decoding data corresponding to an (N-1) th pixel point group and at least the portion data of the first auxiliary inverse transformation information into the first context model, the first context model determining first context information corresponding to the (N-1) th pixel point group based on first context information corresponding from the first pixel point group to an (N-2) th pixel point group, entropy decoding data corresponding to the (N-1) th pixel point group, and at least the portion data of the first auxiliary inverse transformation information;generating a first prediction probability corresponding to the Nth pixel point group based on the first context information corresponding to the (N-1) th pixel point group through the first probability model; andobtaining entropy decoding data corresponding to the Nth pixel point group by performing arithmetical decoding on encoding information corresponding to the Nth pixel point group based on the first prediction probability corresponding to the Nth pixel point group.12.The video decoding method of any one of claims 8-11, wherein the first context model includes a bidirectional Transformer network.13.An image decoding method, implemented by a decoding terminal, comprising:obtaining pixel point grouping information of an original video frame from an encoding terminal, the pixel point grouping information being obtained by performing pixel point grouping on the original video frame based on an image content of the original video frame;obtaining entropy decoding data by performing entropy decoding on encoded data based on the pixel point grouping information;obtaining inverse quantization data by performing inverse quantization on the entropy decoding data; andobtaining a decoded target video frame corresponding to the original video frame by performing an inverse transformation on the inverse quantization data.14.A video encoding method, implemented by an encoding terminal, comprising:obtaining an original video frame and second reference information, the second reference information including feature information of at least one decoded video frame before the original video frame;obtaining encoded data based on the original video frame and the second reference information; andsending the encoded data to a decoding terminal.15.The video encoding method of claim 14, wherein the obtaining encoded data based on the original video frame and the second reference information includes:obtaining the encoded data by processing the original video frame and the second reference information through a second generative machine learning model.16.The video encoding method of claim 15, further comprising:generating, based on the original video frame and the second reference information, prompt information through the second generative machine learning model; orgenerating, based on the original video frame, the prompt information through the second generative machine learning model; the prompt information and the encoded data being sent to the decoding terminal together.17.The video encoding method of claim 16, wherein the prompt information reflects at least one of following information related to the original video frame: pixel value information, scaling information, timing information, segmented semantic information, and classified semantic information.18.The video encoding method of claim 15, further comprising:determining second prompt data input by a user, the second prompt data and the encoded data being sent to the decoding terminal together.19.The video encoding method of claim 18, wherein the second prompt data is configured to indicate an adjustment of an image content of the original video frame.20.The video encoding method of claim 14, wherein the obtaining, based on the original video frame and the second reference information, encoded data includes:obtaining pixel point grouping information of the original video frame by performing pixel point grouping on the original video frame based on an image content of the original video frame;obtaining second fusion data by fusing the original video frame and the second reference information;obtaining first transformation data by transforming the second fusion data;obtaining first quantization data by quantizing the first transformation data; andobtaining the encoded data by performing entropy encoding on the first quantization data based on the pixel point grouping information.21.The video encoding method of claim 20, wherein the pixel point grouping information and the encoded data are sent to the decoding terminal together.22.The video encoding method of claim 20, wherein the image content includes pixel values of pixel points in the original video frame, and the obtaining pixel point grouping information of the original video frame by performing pixel point grouping on the original video frame based on an image content of the original video frame includes:obtaining a plurality of initial pixel point groups by traversing each of the pixel points in the original video frame in turn based on a preset continuous path;obtaining a plurality of target pixel point groups of the original video frame by adjusting pixel points in at least a portion of the plurality of initial pixel point groups; andobtaining the pixel point grouping information of the original video frame based on the plurality of target pixel point groups; whereinthe adjusting pixel points in at least a portion of a plurality of initial pixel point groups includes:for two adjacent initial pixel point groups among the plurality of initial pixel point groups, migrating k pixel points satisfying a preset condition in a second initial pixel point group of the two adjacent initial pixel point groups to a first initial pixel point group of the two adjacent initial pixel point groups, the first initial pixel point group being located before the second initial pixel point group, and k being a positive integer.23.The video encoding method of claim 22, wherein the preset condition includes that in the first initial pixel point group and the second initial pixel point group of the two adjacent initial pixel point groups, a mean square difference between values of pixel points in the first initial pixel point group adjacent to the second initial pixel point group and a value of a kth pixel point in the second initial pixel point group is less than a mean square difference threshold.24.The video encoding method of claim 20, wherein the image content includes pixel values of pixel points in the original video frame, and the obtaining pixel point grouping information of the original video frame by performing pixel point grouping on the original video frame based on an image content of the original video frame includes:obtaining a plurality of target pixel point groups of the original video frame by processing the pixel values of the pixel points in the original video frame through a classification model; andobtaining the pixel point grouping information of the original video frame based on the plurality of target pixel point groups.25.The video encoding method of any one of claims 22-24, wherein the obtaining the pixel point grouping information of the original video frame based on the plurality of target pixel point groups includes:determining, based on the plurality of target pixel point groups, one or more mask images of the original video frame; andobtaining the pixel point grouping information of the original video frame by transforming and / or quantizing the one or more mask images.26.The video encoding method of claim 20, wherein the obtaining the encoded data by performing entropy encoding on the first quantization data based on the pixel point grouping information includes:obtaining second context information by processing the pixel point grouping information through a second context model;generating a second prediction probability based on the second context information through a second probability model; andobtaining the encoded data by performing arithmetical encoding on the first quantization data based on the second prediction probability.27.The video encoding method of claim 20, wherein the obtaining the encoded data by performing entropy encoding on the first quantization data based on the pixel point grouping information includes:sorting a plurality of pixel point groups based on the pixel point grouping information;determining pixel value information corresponding to each of the plurality of pixel point groups in the first quantization data; andperforming following operations on an Nth pixel point group of the plurality of pixel point groups based on an order of the plurality of pixel point groups, N being a positive integer;in response to determining that N is 2, obtaining second context information corresponding to a first pixel point group by inputting, into a second context model pixel value information corresponding to the first pixel point group in the first quantization data;generating a second prediction probability corresponding to a second pixel point group based on the second context information corresponding to the first pixel point group through a second probability model; andobtaining encoded data corresponding to the second pixel point group by performing arithmetical encoding on pixel value information corresponding to the second pixel point group in the first quantization data based on the second prediction probability corresponding to the second pixel point group; andin response to determining that N is greater than 2, inputting pixel value information corresponding to an (N-1) th pixel point group in the first quantization data into the second context model, the second context model determining second context information corresponding to the (N-1) th pixel point group based on second context information corresponding from the first pixel point group to an (N-2) th pixel point group and pixel value information corresponding to the (N-1) th pixel point group;generating a second prediction probability corresponding to the Nth pixel point group based on the second context information corresponding to the (N-1) th pixel point group through the second probability model; andobtaining encoded data corresponding to the Nth pixel point group by performing arithmetical encoding on pixel value information corresponding to the Nth pixel point group in the first quantization data based on the second prediction probability corresponding to the Nth pixel point group.28.The video encoding method of claim 20, further comprising:obtaining auxiliary transformation information by performing an auxiliary transformation on the first transformation data;obtaining auxiliary quantization information by performing auxiliary quantization on the auxiliary transformation information; andobtaining an auxiliary bitstream by performing auxiliary entropy encoding on the auxiliary quantization information, the auxiliary bitstream and the encoded data being sent to the decoding terminal together.29.The video encoding method of claim 28, further comprising:obtaining second auxiliary entropy decoding information by performing second auxiliary entropy decoding on the auxiliary bitstream;obtaining second auxiliary inverse quantization information by performing auxiliary inverse quantization on the second auxiliary entropy decoding information; andobtaining second auxiliary inverse transformation information by performing an auxiliary inverse transformation on the second auxiliary inverse quantization information.30.The video encoding method of claim 29, wherein the obtaining the encoded data by performing entropy encoding on the first quantization data based on the pixel point grouping information includes:sorting a plurality of pixel point groups based on the pixel point grouping information;determining pixel value information corresponding to each of the plurality of pixel point groups in the first quantization data; andperforming following operations on an Nth pixel point group of the plurality of pixel point groups based on an order of the plurality of pixel point groups, N being a positive integer;in response to determining that N is 2, obtaining second context information corresponding to a first pixel point group by inputting, into a second context model, pixel value information corresponding to the first pixel point group in the first quantization data and at least a portion data of the second auxiliary inverse transformation information;generating a second prediction probability corresponding to a second pixel point group based on the second context information corresponding to the first pixel point group through a second probability model; andobtaining encoded data corresponding to the second pixel point group by performing arithmetical encoding on pixel value information corresponding to the second pixel point group in the first quantization data based on the second prediction probability corresponding to the second pixel point group; andin response to determining that N is greater than 2, inputting pixel value information corresponding to an (N-1) th pixel point group in the first quantization data and at least the portion data of the second auxiliary inverse transformation information into the second context model, the second context model determining second context information corresponding to the (N-1) th pixel point group based on second context information corresponding from the first pixel point group to an (N-2) th pixel point group, pixel value information corresponding to the (N-1) th pixel point group, and at least the portion data of the second auxiliary inverse transformation information;generating a second prediction probability corresponding to the Nth pixel point group based on the second context information corresponding to the (N-1) th pixel point group through the second probability model; andobtaining encoded data corresponding to the Nth pixel point group by performing arithmetical encoding on pixel value information corresponding to the Nth pixel point group in the first quantization data based on the second prediction probability corresponding to the Nth pixel point group.31.The video encoding method of any one of claims 26-30, wherein the second context model includes a bidirectional Transformer network.32.An image encoding method, implemented by an encoding terminal, comprising:obtaining pixel point grouping information of an original video frame by performing pixel point grouping on the original video frame based on an image content of the original video frame;obtaining second transformation data by transforming the original video frame;obtaining second quantization data by quantizing the second transformation data;obtaining encoded data by performing entropy encoding on the second quantization data based on the pixel point grouping information; andsending the encoded data to a decoding terminal.33.A system, comprising:at least one storage device including a set of instructions; andat least one processor in communication with the at least one storage device, wherein when executing the set of instructions, the at least one processor is directed to perform the method of any one of claims 1-32.34.A non-transitory computer-readable storage medium, comprising at least one set of instructions, wherein when executed by one or more processors of a computing device, the at least one set of instructions causes the computing device to perform the method of any one of claims 1-32.

Citation Information

Patent Citations

  • Task-driven code stream structured image coding method

    CN110225341A

  • Improved entropy coding in image and video compression using machine learning

    CN113287306A

  • Semantic structured image coding and decoding method and system based on block mask

    CN115604490A

  • Image encoding and decoding method based on content generation, electronic equipment and storage medium

    CN119155452A

  • A method, an apparatus and a computer program product for video encoding and video decoding

    WO2023031503A1