Multi-path embedded video coding method and system based on generative artificial intelligence

Through the multi-path embedded video coding method of generative artificial intelligence, the problem of low storage and transmission efficiency of traditional video coding methods is solved, and personalized and customized transmission of high-definition video is realized, which is suitable for video encoding and decoding with different needs.

CN119583816BActive Publication Date: 2025-09-09INST OF ADVANCED TECH UNIV OF SCI & TECH OF CHINA +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411787829.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-09-09
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Traditional video encoding methods cannot effectively reduce storage occupancy and transmission information flow size, and cannot customize encoding according to the semantic information of the video and the needs of decoding users. In particular, it is difficult to efficiently transmit important video information in extremely low-bandwidth network communications.

Method used

A multi-path embedded video coding method based on generative artificial intelligence is adopted. Encoding and decoding are performed through prompt words and required parameters at the decoding and encoding ends. Generative AI is used to reconstruct the video, and feature extraction and semantic information coding are combined to achieve personalized and customized video transmission.

Benefits of technology

It achieves high-definition video transmission at extremely low bandwidth, reduces storage space and transmission information flow size, and can perform customized encoding and decoding according to the needs of the receiving end, supporting variable semantic expression for different groups of people and devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119583816B_ABST
    Figure CN119583816B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-path embedded video encoding method and system based on generative artificial intelligence. The method first encodes the input prompt word and requirement parameters and outputs them as a prompt word vector and a requirement parameter code stream; then decodes the prompt word vector and the requirement parameter code stream to output a feature vector and a feature map; extracts the feature vector and the feature map and the representation guidance vector to obtain multi-embedded feature information; reconstructs the obtained multi-embedded feature information through generative AI to obtain a video to be iterated; performs video quality evaluation on the video to be iterated; and generates a binary code stream for transmission by inputting the current prompt word vector and the requirement parameter code stream and the multi-embedded feature information. The present invention highly abstracts the important information features of transmission and reconstructs high-definition video at the receiving end, significantly reducing the storage space required for video encoding and the size of the transmission information stream, achieving higher-definition transmission under narrower bandwidth or making video encoding under extremely low bandwidth possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and video coding technology, and in particular to a technical solution for multi-path embedded video coding and video decoding based on generative artificial intelligence. Background Art

[0002] Reducing the amount of information flow during video encoding and storage is a key topic in video coding and decoding research. While many related technologies are widely used, further significantly reducing storage usage and transmission information flow remains a challenge. Currently, video coding and compression are becoming increasingly difficult, while improvements in compression efficiency are steadily decreasing. Traditional approaches to video coding and compression primarily rely on entropy coding, pixel-based correlation between previous and next frames, and intra-frame correlation, relying on pixel-level information and traditional statistical calculations to capture and characterize the motion characteristics of relevant pixels.

[0003] Therefore, traditional codec designs still require the storage and transmission of large amounts of motion information, relevant feature pixels, and key complete frame image data. Due to the limitations of information transmission theory, the performance of traditional codecs is also gradually reaching its limits. Furthermore, traditional video coding methods cannot customize encoding based on the semantic information and content of the video, nor can they customize decoding based on the needs of the decoding user.

[0004] Traditional video codecs struggle when applied to abstract or personalized expressions. Especially in extremely low-bandwidth networks, transmitting crucial and important video information from the sender to the receiver while minimizing resource usage is a challenge. Traditional methods are often time-consuming or even impossible to accomplish. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to transmit important and critical video information from a sending end to a receiving end while occupying as few resources as possible.

[0006] The present invention solves the above technical problems through the following technical means: a multi-path embedded video encoding method based on generative artificial intelligence, including encoding and decoding steps, wherein prompt words and required parameters from both the decoding end and the encoding end are present, and the encoding includes the following steps:

[0007] Step S101: Input the prompt word and requirement parameters of the decoding end and the prompt word and requirement parameters of the encoding end into the prompt word and requirement parameter encoding module (10), encode the input prompt word and requirement parameters and output them into a prompt word vector and requirement parameter code stream, and transmit them to the prompt word vector decoding module (20). At the same time, the output prompt word vector and requirement parameter code stream are also input into the semantic information encoding and code stream compression module (60), and are packaged as the input of this part into a binary code stream that can be transmitted;

[0008] Step S102, after obtaining the prompt word vector and the required parameter code stream, the prompt word vector decoding module (20) decodes them and outputs a feature vector and a feature map;

[0009] Step S103, after the feature vector and feature map are output, the feature vector, feature map and representation guidance vector are input into the video extraction module (30), and the multi-embedded feature information is obtained through the video extraction module (30);

[0010] Step S104, inputting the obtained multi-embedded feature information into the first generative AI decoding module (40), reconstructing the picture through generative AI and obtaining the video to be iterated;

[0011] Step S105 is used to receive the video to be iterated and evaluate its video quality. If the evaluation obtains corresponding indicator parameters and determines that the video does not need iteration, the iteration is completed. If the evaluation obtains corresponding indicator parameters and determines that the video still needs to be iterated;

[0012] Step S106, at this time, the semantic feature extraction and internal optimization iteration are completed, the current prompt word vector and the required parameter code stream and the multi-embedded feature information are input into the semantic information encoding and code stream compression module (60) to generate a binary code stream for transmission, and the encoder stops encoding.

[0013] As an optimized technical solution, step S101 specifically includes:

[0014] Inputting the decoding end prompt words and requirement parameters and the service (encoding end) prompt words and requirement parameters into the embedding layer of the prompt words and requirement parameter encoding module (10) for modal conversion;

[0015] Then input into the Encoder of the prompt word and requirement parameter encoding module (10), and output as prompt word vector and requirement parameter code stream, wherein the Encoder uses a basic learner based on the Transformer learner idea to perform feature extraction;

[0016] In the step S102, the prompt word vector decoding module (20) has a basic characterizer and can output a feature vector and a feature map.

[0017] As an optimized technical solution, the step S103 specifically includes:

[0018] The original video information is input into a video segmentation discriminator of a video extraction module (30), which reads the entire video or a very short segment of a real-time video and determines whether there is editing or drastic scene switching, thereby dividing the entire video into smaller video segments;

[0019] The smaller video segments after segmentation are input into the feature extraction module. In the feature extraction module, the expression and expression intensity of different feature extraction modules in the feature extraction module are partially selected according to the results of phase weight encoding. After feature extraction, each small video segment generates a corresponding multi-embedded feature information representation block. By stacking and splicing several multi-embedded feature information representation blocks, the multi-embedded feature information of the corresponding entire video is obtained.

[0020] As an optimized technical solution, in step S103, reading the entire video or a very short segment of the real-time video and determining whether there is editing or drastic scene switching, and dividing the entire video into smaller video segments includes:

[0021] By default, the video segmentation discriminator uses T seconds as a marking point, calculates the structural consistency parameters of the current frame and the previous frame, and judges that if the structural similarity between the two exceeds a certain fixed initial threshold, the similarity is considered acceptable. The current frame and the previous frame are continuous continuous pictures. If the current predetermined number of frames are all continuous pictures, they are all packaged into one group of video clips. If the current predetermined number of frames are not all continuous pictures, they are segmented and packaged into several groups of video clips based on the SSIM threshold, where the total number of each continuous video frame is less than or equal to the set slicing time x video frame rate. The entire video is divided into smaller video clips by using the segmentation discriminator, and the small video clips generated after segmentation are then input into the feature extraction module.

[0022] As an optimized technical solution, in step S103, in the feature extraction module, based on the result of phase weight encoding, partially selecting the expression and expression strength of different feature extraction modules in the feature extraction module specifically includes:

[0023] First, the video clips are input into the video content understanding module, which extracts features of semantic entities and background and foreground information within the video clips and generates representational prompt words for generative AI to complete video decoding and generate the video to be iterated;

[0024] At the same time, when the phase weight coding completes the weighted expression, feature extraction modules with different degrees of expression are selected according to the needs of the decoding end. If it is necessary to complete the pixel-level complete information expression, each feature extraction module will be fully called. If it is only necessary to express the most abstract semantic feature information video, only the representation prompt words generated by the video content understanding module will be used. For the remaining basic parameters, the default parameters are directly called in the first generative AI decoding module (40) without customization and pixel-level precise expression. If it is necessary to modify the video feature representation strength, video restoration degree and personalized expression degree according to the needs of the decoding end, the input prompt word information is expressed through phase weight coding and different feature extraction and feature descriptors are called. At the same time, different feature extraction and feature descriptors are shown with different expression strengths according to the different phase weight expression strengths.

[0025] As an optimized technical solution, step S104 specifically includes:

[0026] By decomposing the multi-embedded feature information into blocks, an embedded feature semantic information representation block is obtained. The representation prompt word is input into the representation prompt word embedding layer to complete the embedding operation of the prompt word to generate high-dimensional array information that can be read by the model. As long as the edge information feature map, motion information feature map, semantic feature map, semantic control feature map, semantic internal consistency feature map, semantic super-resolution feature map and personalized style feature vector exist, they are input into the auxiliary representation and guidance embedding layer accordingly. In the auxiliary representation and guidance embedding layer, the above-representable feature maps and feature vectors are uniformly encoded and embedded, and output as a high-dimensional array in a unified format to ensure that the input specifications and data structure are unified in the auxiliary representation hierarchical expression module. The basic control parameters are input into the parameter loading model to load customized parameter values. Among them, if the auxiliary representation and guidance embedding layer does not receive any information, it directly constructs information with the same data structure and fills it with the number 0. If some values ​​are missing in the parameter loading model, it is loaded with default parameters. If the representation prompt word embedding layer does not receive any information, the first generative AI decoding module (40) stops working.

[0027] After completing the data loading for the three components above, the data generated by the representation prompt word embedding layer is input into the representation prompt word core module of the generative AI video inference model. The data generated by the auxiliary representation and guidance embedding layers are then input into the auxiliary representation layered expression module of the generative AI video inference model to express their corresponding features based on different personalized features and representation strengths. Finally, the generative AI model generates the video to be iterated / reconstructed.

[0028] As an optimized technical solution, step S106 specifically includes:

[0029] The head and tail of the iteratively completed multi-embedded feature information are inserted into the head identification information and the tail identification information. The inserted identification information is mainly used to mark the head and tail of the information fragment. At the same time, the head contains basic description information of the semantic information of the multi-embedded semantic feature. After the information is recorded, the multi-embedded feature information representation block 1 to the semantic information representation block n are losslessly compressed. After the lossless compression is completed, the compressed multi-embedded feature information compressed representation block is input into the multi-embedded feature information compressed representation block similarity comparison mapping model. The model compares the similarity between the feature representation blocks and replaces the representation blocks with different representations and high similarity through transfer learning.

[0030] When the module is executed normally, several mapping blocks with smaller volumes are obtained. After the migration mapping is completed, the header identification information and the tail identification information are spliced ​​onto the output content and entropy coding is performed to encode a pure binary bit stream. When the corresponding binary code stream is obtained, the encoder part is executed.

[0031] As an optimized technical solution, the decoding step specifically includes:

[0032] Step S107, the decoder receives the transmitted binary code stream, inputs the binary code stream into the code stream decoding module (70), and decodes the multi-embedded feature information and the prompt word vector and the required parameter code stream, wherein the prompt word vector and the required parameter code stream can be defaulted;

[0033] In step S108, the obtained multi-embedded feature information, prompt word vector and required parameter code stream are input into the second generative AI decoding module (80), and the multi-embedded feature information is mainly used, and the prompt word vector and required parameter code stream are used as guidance and configuration parameters to generate videos with different degrees of personalized or specialized semantics required by the corresponding receiver. If the prompt word vector and required parameter code stream are missing, the original video is directly generated and expressed.

[0034] The present invention also provides a system corresponding to any of the above-mentioned multi-path embedded video encoding methods based on generative artificial intelligence, including an encoder and a decoder, wherein the encoder includes:

[0035] The requirement parameter encoding module (10) is used to receive the prompt word and requirement parameter from the decoding end and the prompt word and requirement parameter input from the encoding end, encode the input prompt word and requirement parameter and output them into a prompt word vector and a requirement parameter code stream, and transmit them to the prompt word vector decoding module (20). At the same time, the output prompt word vector and requirement parameter code stream are also input to the semantic information encoding and code stream compression module, and are packaged as the input of this part into a binary code stream that can be transmitted;

[0036] A prompt word vector decoding module (20) is used to decode the prompt word vector and the required parameter code stream after obtaining them, and output a feature vector and a feature map;

[0037] A video extraction module (30) is used to receive the feature vector and feature map output by the prompt word vector decoding module (20), and obtain multi-embedded feature information after processing;

[0038] A first generative AI decoding module (40) is configured to receive the multi-embedded feature information, reconstruct the image through generative AI, and obtain a video to be iterated;

[0039] A video quality comprehensive evaluation module (50) is used to receive a video to be iterated and perform video quality evaluation on it. If the evaluation obtains corresponding index parameters and determines that the video does not need to be iterated, the iteration is completed. If the evaluation obtains corresponding index parameters and determines that the video still needs to be iterated in the module,

[0040] A semantic information encoding and code stream compression module (60) is used to receive the current prompt word vector and the required parameter code stream and multiple embedded feature information after the semantic feature extraction and internal optimization iteration are completed, and generate a binary code stream for transmission;

[0041] The decoder comprises:

[0042] A code stream decoding module (70) is used to receive the binary code stream transmitted by the encoder and decode the multi-embedded feature information and the prompt word vector and the required parameter code stream, wherein the prompt word vector and the required parameter code stream can be defaulted;

[0043] The second generative AI decoding module (80) is used to generate videos with different degrees of personalized or specialized semantics required by the receiver based on the multi-embedded feature information, the prompt word vector and the required parameter code stream as guidance and configuration parameters. If the prompt word vector and the required parameter code stream are missing, the original video is directly generated and expressed.

[0044] The advantages of the present invention are: the present invention highly abstracts the important information features of transmission and reconstructs high-definition video at the receiving end, significantly reducing the storage space required for video encoding and the size of the transmission information stream, achieving higher-definition transmission under narrower bandwidth or making video encoding under extremely low bandwidth possible, and can perform customized encoding and decoding according to the needs of the decoding side. The different feature representation vectors used in this method can be used to perform variable semantic expression for different groups of people, different audiences, and even for machines or industrial equipment, effectively solving the problem that traditional encoding methods cannot customize feature expression according to needs, specifically including:

[0045] In terms of temporal information processing, traditional encoders use a keyframe and reference frame system to reference repeated information between frames within the temporal information to eliminate redundancy. The present invention uses an understanding-information extraction-reconstruction approach to read and understand key information such as background, foreground, and semantics within each continuous video segment, extract relevant features and edge information, and reconstruct and achieve pixel-level image restoration, select partial video feature expression, and personalized video reconstruction. At the same time, the mutual migration of feature blocks is used to further reduce information redundancy between video segments, further reducing storage and bandwidth usage.

[0046] In terms of personalized expression, traditional encoders perform pixel-level video restoration in a lossy or lossless manner, and are unable to change the video content, such as characters, semantics, style, duration, and rhythm, according to the needs of the sending and receiving ends. Using the present invention, personalized encoding of the video can be achieved by inputting prompt words at the sending and receiving ends. For example, if you need to delete the role of a certain character, you only need to input the corresponding prompt words at the sending end, and the encoding system will extract and discard the characteristics and semantic information of the character, and the decoded video will no longer contain the character. If you need to change the style of the video or adapt to younger users, you only need to input the corresponding prompt words, and the encoding system will adjust the video style characteristics and process inappropriate content (such as deletion, mosaic, blocking, etc.), thereby achieving personalized expression. Similarly, for personalized needs for video length and clarity, the above method can also be used to control the length and clarity of the video after encoding and decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 1 is a diagram showing the structure of an encoder of a multi-path embedded video coding system based on generative AI according to an embodiment of the present application;

[0048] Figure 2 4 is a decoder structure diagram of a multi-path embedded video coding system based on generative AI according to an embodiment of the present application;

[0049] Figure 3 This is a schematic diagram of the overall framework flow of the multi-path embedded video coding system based on generative AI in this application;

[0050] Figure 4 This is the internal structure diagram of the video extraction module of this application;

[0051] Figure 5 It is a diagram of the multi-embedded feature information structure of this application;

[0052] Figure 6 This is the internal structure diagram of the generative AI decoding module of this application;

[0053] Figure 7 This is the internal structure diagram of the video quality comprehensive evaluation module of this application;

[0054] Figure 8 This is the internal structure diagram of the video semantic optimization module of this application;

[0055] Figure 9 This is the internal structure diagram of the semantic information encoding and code stream compression module of this application;

[0056] Figure 10 This is the internal structure diagram of the multi-embedded feature optimization module of this application;

[0057] Figure 11 This is the internal structure diagram of the prompt word and requirement parameter encoding module of this application. DETAILED DESCRIPTION

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0059] like Figure 1 and Figure 2 As shown in FIG, a multi-path embedded video coding system based on generative AI according to an embodiment of the present application is divided into two parts: an encoder and a decoder.

[0060] The video encoder includes a prompt word and demand parameter encoding module 10, a prompt word vector decoding module 20, a video extraction module 30, a generative AI decoding module 40, a video quality comprehensive evaluation module 50, and a semantic information encoding and bitstream compression module 60. The output end of the prompt word and demand parameter encoding module 10 is simultaneously connected to the input ends of the prompt word vector decoding module 20 and the semantic information encoding and bitstream compression module 60. The output end of the prompt word vector decoding module 20 is connected to the generative AI decoding module 40 and the semantic information encoding and bitstream compression module 60 after passing through the video extraction module 30. The output end of the generative AI decoding module 40 is connected to the video quality comprehensive evaluation module 50. The input end of the prompt word and demand parameter encoding module 10 is used to receive the prompt word and demand parameter input by the user (decoding end) / service (encoding end), and the output end of the semantic information encoding and bitstream compression module 60 is used to output a binary bitstream to the video decoder.

[0061] The video decoder includes a code stream decoding module 70 and a generative AI decoding module 80 connected to each other. The code stream decoding module 70 is used to receive the binary code stream transmitted by the video encoder, and the generative AI decoding module 80 outputs a video reconstructed based on semantics.

[0062] See Figure 3 The method for encoding using the multi-path embedded video coding system based on generative AI of the present invention comprises the following steps:

[0063] 1. Coding:

[0064] The encoder's initial state and workflow are defined by the original semantic information video and the prompts and required parameters from the user (decoding end) / the service (encoding end). Four scenarios are defined: the original semantic information video exists, the prompts and required parameters from the user (decoding end) exist, and the prompts and required parameters from the service (encoding end) exist; the original semantic information video exists, the prompts and required parameters from the user (decoding end) exist, and the prompts and required parameters from the service (encoding end) do not exist; the original semantic information video exists, the prompts and required parameters from the user (decoding end) do not exist, and the prompts and required parameters from the service (encoding end) exist; the original semantic information video exists, the prompts and required parameters from the user (decoding end) do not exist, and the prompts and required parameters from the service (encoding end) exist. The original semantic information video exists, the prompts and required parameters from the user (decoding end) do not exist, and the prompts and required parameters from the service (encoding end) do not exist. The original semantic information video is required for this system; the prompts and required parameters from the user (decoding end) / the service (encoding end) are optional. The following description assumes that both the prompts and required parameters from the user (decoding end) / the service (encoding end) exist.

[0065] When both the prompt word and the required parameters from the user (decoding end) or the service (encoding end) are present, encoding includes the following steps:

[0066] Step S101: Input the user (decoding end) prompt words and required parameters and the service (sending encoding end) prompt words and required parameters. Figure 11 , specifically including:

[0067] The user (decoding end) prompt words and demand parameters and the service (encoding end) prompt words and demand parameters are input into the embedding layer of the prompt word and demand parameter encoding module 10 for modal conversion, that is, text information such as words and numbers are converted into feature maps through the embedding layer;

[0068] This is then input into the encoder of the prompt word and requirement parameter encoding module 10, which outputs the prompt word vector and the requirement parameter code stream. The encoder must perform feature extraction using a basic learner based on, but not limited to, a Transformer learner. The Transformer learner features positional encoding and a multi-head self-attention mechanism, enabling unified and efficient extraction of global and local feature information.

[0069] Step S102: After obtaining the prompt word vector and the required parameter code stream, they are input into the prompt word vector decoding module 20 for decoding. The module has a basic characterizer similar to the decoder architecture and can output feature vectors and feature maps.

[0070] Step S103: After the feature vector and feature map are output, they and the representation guidance vector are input to the video extraction module 30. Figure 4 The video extraction module 30 further includes a video segment discriminator, a phase weight encoding, and a feature extraction module. The feature vector, feature map, and representation guidance vector are input into the phase weight encoding part of the video extraction module 30, and the phase weight encoding guides the partial or complete expression of the feature extraction of the video clip.

[0071] In this step, the original video information is input into the video segmentation discriminator of the video extraction module 30. The video segmentation discriminator reads the entire video or a very short segment of the real-time video (sliced ​​according to real-time requirements) and determines whether there are any clips or drastic scene changes. This segmentation discriminator has adjustable parameters and default parameters. By default, the video segmentation discriminator uses T seconds as a marker (for example, 2 seconds. For videos with high real-time requirements, the marker can be set directly based on the number of frames, with a minimum of 1 frame). It calculates the structural consistency parameter between the current frame and the previous frame (the structural consistency parameter calculation method includes but is not limited to the structural similarity index SSIM (Structural Similarity Index Measurement, SSIM) and the multi-scale structural similarity index measure (MS-SSIM)). If the structural similarity between the two frames exceeds a certain fixed initial threshold (this threshold is not automatically updated, and here, the structural consistency calculation method SSIM is used as an example. The SSIM threshold includes but is not limited to values ​​that can be taken within the definition domain, such as 0.6 and 0.7. At the same time, "exceeding" here refers to the degree of convergence of the two images in structural consistency. If a parameter obtained by a certain calculation method indicates that the smaller the value, the higher the convergence, then "less than" here becomes "less than", and "not exceeding" in the following text becomes "greater than or equal to"), then the similarity is acceptable and the current frame and the previous frame are continuous. If the video is at 60 fps, 2 seconds corresponds to 120 frames. If all 120 frames are continuous, they are packaged into a single video segment. If the video is at 60 fps, 2 seconds corresponds to 120 frames. If not all 120 frames are continuous, they are segmented and packaged into several video segments based on the SSIM threshold. The total number of consecutive video frames in each segment must be less than or equal to the set slicing time multiplied by the video frame rate. A segmentation discriminator is used to segment the entire video into smaller segments, which are then fed into the feature extraction module.

[0072] In the feature extraction module, the expression and expression strength of different feature extraction modules in the feature extraction module will be partially selected based on the results of phase weight encoding. After the video clip is input into the feature extraction module, it will be input into the video content understanding module separately. The video content understanding module will complete feature extraction of semantic entities, background and foreground information in the video clip and generate representation prompt words for subsequent generative AI to complete video decoding and generate the video to be iterated. At the same time, under the weighted expression completed by phase weight encoding, feature extraction modules with different expression levels will be selected according to the needs of the user / decoding end. If pixel-level complete information expression is required, the edge information descriptor, motion information descriptor, semantic extraction module, semantic control feature extraction module, semantic internal consistency feature extraction module, semantic super-resolution feature extraction module, personalized style feature extraction module, video basic information reading and basic control parameter fine-tuning module will be fully called. If only the most abstract semantic feature information video needs to be expressed, there is no need to call the above modules. Only the video content understanding module needs to generate representation prompt words. For the remaining basic parameters, the default parameters are directly called in the generative AI decoding module 40 without customization and pixel-level precise expression. If it is necessary to modify the video feature representation strength, video restoration degree, and personalized expression degree according to the needs of the user / decoding end, the input prompt word information is expressed through phase weight encoding and different feature extraction and feature descriptors are called. At the same time, different feature extraction and feature descriptors are shown with different expression strengths according to the different phase weight expression strengths.

[0073] It should be noted that the optional expression modules and feature descriptors above represent different purposes and functions.

[0074] The edge information feature descriptor primarily extracts all edge contour information and features within the image. It then selects the strength of contour feature representation for different semantic entities based on the phase weight encoding results and generates an edge information feature map. The motion information descriptor consists of three subcomponents: a motion vector direction feature descriptor, an optical flow motion estimation descriptor, and a semantic trajectory and relative position descriptor. The motion vector direction feature descriptor primarily extracts motion vector and direction feature information for different semantic entities within a video clip and generates a motion vector direction feature map. With the support of phase weight encoding, it can also select the strength of motion vector direction features for different semantic entities based on the representation guidance vector, thereby purposefully adjusting or reducing the total amount of transmitted information. The optical flow motion estimation descriptor describes the optical flow motion information features of different foreground and background motion semantics within a video clip. This feature assists in expressing global motion information and controlling motion blur within the image. The semantic extraction module comprises foreground semantic entity extraction, background semantic entity extraction, and language and text extraction. It primarily extracts the semantics of foreground, background, and language entities within a video clip. The semantic entities of the foreground and background can be extracted and used for the subsequent video reconstruction and semantic-based super-resolution of the generative AI decoder. It can also be used for the compression of semantic redundancy before and after multiple video clips (mainly manifested in the compression of multiple embedded feature information representation blocks in step S130 to generate mapping blocks). The language and text entity semantic extraction module can extract the text content and its meaning feature information, which is used to restore the corresponding text with higher clarity and can be personalized and modified into other languages ​​and texts in different language environments according to the requirements of the representation guidance vector (such as Chinese to English, Japanese to Hebrew, etc.). The semantic control feature extraction module includes semantic trajectory and relative position descriptors, which are mainly used to accurately describe and guide the motion trajectory of local and specific semantics; through the expression of user / decoding end prompt words or representation guidance vectors under different intensities and semantic entities of phase weight encoding, the accurate motion trajectory description of the semantic entities and semantic entity sets of specific targets is achieved, and it is also used to ensure the consistency and personalized correction of the motion trajectory of semantic entities. The semantic internal consistency feature extraction module is mainly used to ensure the internal consistency of semantic entities between different frames generated by the generative AI decoder, and to maintain pixel-level consistency and structural consistency within the semantics. The semantic super-resolution feature extraction module is mainly used for image denoising and super-resolution based on semantic information. The module can simultaneously extract the pixel noise features and semantic entity noise features of the specified semantic entity area and generate corresponding feature maps. At the same time, it will also extract low-resolution feature maps for video key frames for subsequent super-resolution denoising and pixel-level video image restoration of the specified semantic entity in the generative AI decoder.The personalized style feature extraction module is mainly used to extract the style of the entire video clip and its style feature information. It can be used to accurately restore the style, contrast, dynamic range, highlight and dark details of the video; it can also be used to represent different style conversion features under the guidance of prompt word vectors for generating personalized style expression features and their corresponding picture details in the generative AI decoder. The video basic information reading and basic control parameter fine-tuning module is used to read basic information such as the number of video frames, frame rate, aspect ratio, color space, and control parameters for fine-tuning some semantic entities. The above control parameters and basic video information will be input into the generative AI decoder to control the consistency of the generated video basic parameter information.

[0075] Specifically, how to use the above-mentioned feature descriptors and feature maps to address inter-frame content redundancy requires the following: within a video clip, accurately restoring all semantic and motion information of the foreground and background in the original video requires invoking some or all of the following feature descriptors and modules: motion information descriptor, edge information feature descriptor, optical flow motion estimation descriptor, semantic control feature extraction module, semantic extraction module, language and text extraction module, and semantic internal consistency feature extraction module. By invoking these feature descriptors and modules, the video clip is abstracted and extracted into a record of the interaction of multiple foreground and background semantic entities. This is no longer concretely stored as multiple consecutive frames. By extracting and separating the entity semantics, it is no longer necessary to record, transmit, and restore them using pixels, but rather to directly restore the semantic entities and reconstruct the video clip. Therefore, the motion information of semantic entities recorded between frames is extracted and stored using the motion vector direction feature descriptor, optical flow motion estimation descriptor, semantic trajectory and relative position descriptor, foreground semantic entity extraction, and background semantic entity extraction. Furthermore, by utilizing the edge information feature descriptor, semantic understanding module, language and text extraction module, and semantic internal consistency feature extraction module to read and obtain the detailed features and necessary descriptive information of semantic entities, as well as their overlapping relationships, this effectively reduces inter-frame semantic entity information duplication. Thus, through bidirectional feature description, semantic entity understanding and extraction, and motion information extraction, redundant information between frames within a video clip is removed.

[0076] It's worth noting that the basic video information reading and basic control parameter fine-tuning module can also provide additional regulatory control. If the user / decoder needs to reduce a one-hour video to 30 minutes or expand it to two hours, this component will read the original video length and guide the expression of different feature extraction modules and descriptors based on feature importance levels. This will then constrain the generated video length in the subsequent video reconstruction module, gradually and iteratively achieving the goal of personalized duration control. If the user / decoder requires deleting a certain part, semantic meaning, segment, or character to reduce, maintain, or increase the total duration, this component will guide the expression of certain modules and control the overall duration of the reconstructed / iterated video. If the encoder requires real-time transmission, this component's control parameters can also control the specified number of encoded frames, with a minimum transmission size of one frame, to improve encoding efficiency and ensure encoding and transmission consistency.

[0077] After completing the video extraction of the video, multiple embedded feature information will be obtained. Among them, multiple embedding is reflected in the different representations and feature extraction modules that can be selected based on the needs of the user / decoding end. The multiple embedded feature information includes Figure 5 In the structure shown, after feature extraction, each short video segment will generate a corresponding multi-embedded feature information representation block. Stacking and splicing several multi-embedded feature information representation blocks will correspond to the multi-embedded feature information of the entire video.

[0078] Step S104: When multiple embedded feature information is obtained, part of the data is input into the generative AI decoding module 40 for generating a video reference to be iterated. Figure 6As shown. By decomposing the multiple embedded feature information into blocks, an embedded feature semantic information representation block is obtained. Specifically, in each embedded feature semantic information representation block, the representation prompt word is input into the representation prompt word embedding layer to complete the embedding operation of the prompt word to generate high-dimensional array information that can be read by the model. If the edge information feature map, motion information feature map, semantic feature map, semantic control feature map, semantic internal consistency feature map, semantic super-resolution feature map, and personalized style feature vector exist, they are input into the auxiliary representation and guidance embedding layer accordingly. In the auxiliary representation and guidance embedding layer, the above-representable feature maps and feature vectors are uniformly encoded and embedded, and output as a high-dimensional array in a unified format, ensuring unified input specifications and unified data structure in the auxiliary representation hierarchical expression module. Basic control parameters are input into the parameter loading model to load customized parameter values. If the auxiliary representation and guidance embedding layer does not receive any information, it directly constructs information with the same data structure and fills it with the number 0. If some values ​​are missing in the parameter loading model, the default parameters are loaded. If the representation prompt word embedding layer does not receive any information, the generative AI decoding module 40 stops working. After completing the data loading of the above three parts, the data generated by the representation prompt word embedding layer is input into the representation prompt word expression core module in the generative AI video reasoning model; the data generated by the auxiliary representation and guidance embedding layer is input into the auxiliary representation hierarchical expression module in the generative AI video reasoning model to express its corresponding features according to different personalized features and representation strengths. Finally, the generative AI model is used to generate the video to be iterated / reconstructed. Among them, the video to be iterated refers to the video used for iterative optimization by the encoder part, and the reconstructed video refers to the final video presented to the user by the decoder part. The generative AI video reasoning model is mainly responsible for generating key frames or video reasoning framework structures, including but not limited to generative base reasoning models based on Diffusion-like, VAE-like (VQ-VAE, AE), GAN-like, and multi-method fusion (Transformer-Diffusion, VAE-Diffusion, Transformer-GAN and GAN-Diffusion, etc.).

[0079] Step S105: After the AI ​​decoding module 40 completes the generation of the video to be iterated, it is input into the video quality comprehensive evaluation module 50. Figure 7 In the video quality comprehensive evaluation module 50, there are two required input items and two optional input items, namely: the original semantic information video and the video to be iterated; the user (receiving end) prompt words and required parameters and the service (encoding end) prompt words and required parameters.

[0080] The original semantic information video is input into the video image stacking embedding layer; this embedding layer stacks the frames and converts the stacked images into a high-dimensional array using an encoder-like base learner similar to the previous module description. The processing of the iterated video is similar to the processing of the original semantic information video: the video frames are stacked and input into the encoder-like base learner to convert the stacked images into a high-dimensional array.

[0081] When processing user (receiving end) prompt words and requirement parameters and service (encoding end) prompt words and requirement parameters, they are input into the corresponding embedding layer to convert them from text and numbers into high-dimensional vectors or high-dimensional arrays.

[0082] The core idea behind the above embedding layer processing is to use a unified type of base learner to convert modal information such as video, images, text, and numbers into unified modal data that can be processed, feature extracted, and analyzed. If user (receiving end) prompt words and requirement parameters and service (encoding end) prompt words and requirement parameters exist, the four high-dimensional arrays obtained from the embedding layer, including the user (receiving end) prompt words and requirement parameters, the service (encoding end) prompt words and requirement parameters, the original semantic information video, and the video to be iterated, are input into the reference quality assessment model. The reference quality assessment model includes pixel-level, feature-level, semantic-level, and segment-level gradient evaluation of the video image. Furthermore, a before-and-after comparison is performed based on the guidance expressed by the user / decoding end prompt words and requirement parameters after passing through the embedding layer, and a bidirectional quality assessment is performed on the original semantic information video and the video to be iterated. Evaluation criteria include, but are not limited to, brightness consistency, chromaticity consistency, color gamut consistency, contrast consistency, dynamic range consistency, pixel consistency, semantic consistency, style consistency, semantic recognizability consistency, emotional style consistency, language consistency, character movement consistency, semantic relative position relationship consistency, semantic overlap consistency, foreground consistency, background consistency, semantic target trajectory consistency, internal semantic consistency, internal semantic coherence consistency, inter-frame contrast consistency, inter-frame optical flow feature consistency, inter-frame brightness consistency, inter-frame smoothness consistency, and inter-frame semantic loss consistency. The evaluation results are abstracted and expressed as feature vectors and input into the comprehensive evaluation model. Similarly, a no-reference quality evaluation is performed separately for the video to be iterated. The high-dimensional array obtained by embedding the user (decoding end) / service (encoding end) prompt words and requirement parameters is input into the no-reference quality evaluation model. The evaluation metrics for the no-reference quality evaluation of the iterative video include, but are not limited to, brightness, chroma, color gamut, contrast, dynamic range, pixel resolution, image contrast, semantic resolution, style expression, semantic recognizability, emotional style expression, language clarity, character movement coherence, semantic relative position expression, semantic overlap, foreground clarity, background clarity, semantic target trajectory description, semantic internal coherence, inter-frame contrast, feature extraction difficulty, inter-frame optical flow feature clarity, inter-frame brightness and darkness variation, inter-frame smoothness, and inter-frame semantic loss. After completing the above no-reference evaluation, a feature vector of the no-reference quality evaluation result is obtained and input into a comprehensive evaluation large language model. This large model comprehensively analyzes the performance of the iterative video and the relevant feature representations of the original semantic information to produce a semantic guidance vector, a representation guidance vector, and a multi-embedded feature optimization guidance vector, totaling three vectors for the next round of video iteration.Among them, the semantic guidance vector mainly acts on the original semantic information video, and is used to optimize and improve the original semantic information, performing operations including but not limited to changing brightness, contrast, semantic clarity, etc. to improve the original semantic information video. The representation guidance vector mainly acts on the video understanding module, and is mainly used for tuning and adaptive strength expression of various feature extraction modules, feature descriptors, semantic understanding modules, etc. By optimizing feature extraction and semantic understanding, the expression effect and subjective and objective video quality of the video to be iterated are improved, and semantic-level and fragment-level optimization is completed for the iterated video. The multi-embedded feature optimization guidance vector mainly acts on the multi-embedded semantic features, and by fine-tuning the parameters, feature vectors, and feature maps of the multi-embedded semantics and its internal representation blocks, pixel-level optimization adjustment of the iterated video is achieved.

[0083] After the operation of the video quality comprehensive evaluation module 50 is completed, a judgment parameter will be generated, and it will be judged whether the iteration of the current video to be iterated is completed. If the module obtains the corresponding index parameter through evaluation and judges that the video still needs to be iterated, it will judge whether the iteration is completed or not. If the video to be iterated judges whether the iteration is completed or not, the relevant difference parameters are calculated and sent to the semantic optimization module 91, the representation guidance module 92, and the multi-embedded feature optimization module 93, and the iteration continues. When the semantic optimization module 91 receives the relevant parameters, it optimizes the original semantic information video. If the parameter is missing, no optimization is performed; when the representation guidance module 92 receives the relevant parameters, it optimizes the video extraction module 30. If the parameter is missing, no optimization is performed; when the multi-embedded feature optimization module 93 receives the relevant parameters, it optimizes the multi-embedded feature information content expressed. If the parameter is missing, no optimization is performed.

[0084] Specifically, the execution is as follows: semantic optimization module 91, characterization guidance vector and multi-embedded feature optimization module 93. The function of the characterization guidance vector has been described in S103 and will not be repeated here. Figure 8, input the original semantic information video into the video image stacking embedding layer, similar to step S105, stack the video images and perform the embedding operation to complete the modal conversion, and then input the Encoder after the embedding is completed. The Encoder here and the subsequent Decoder also include but are not limited to the Encoder-like structure and Decoder-like structure in the Transformer. By encoding the original semantic video frame by frame, a feature map set of all frames is obtained. The feature map set contains complete information for each video frame image. If the reverse operation is performed at this time, the original frame image will be directly generated. The feature map set and the semantic optimization guidance vector are input into the semantic information optimization model, and the model will complete the feature-level and semantic-level optimization of the original video based on the guidance vector and obtain the optimized feature map set. At this time, the optimized feature map set is input into the Decoder and the optimized original semantic information video is decoded. Specifically to the multi-embedded feature optimization module 93, as Figure 10 By inputting the multi-embedded feature optimization guidance vector into the multi-embedded feature optimization model, pixel-level and feature-level optimization is performed on the multi-embedded feature semantic representation blocks within the multi-embedded semantic features. The optimized multi-embedded feature information can be modified through internal feature maps and feature vectors to achieve better feature expression.

[0085] It should be noted that the Encoder-like learners and Decoder-like representations mentioned include but are not limited to the base learners split from the Encoder-Decoder architecture in the Transformer architecture, the Encoder / Decoder base learners in the Bidirectional Encoder Representations from Transformers (BERT) architecture based on the transformer architecture, and the Image Encoder / Decoder base learners in the Swin-Transformer architecture. As long as the Encoder / Decoder-like base learners and embedding layer network structures and Convolutional Neural Network (CNN)-like structures can uniformly convert different modal information into each other, they are all applicable to this patent proposal.

[0086] Step S106, when the video quality evaluation comprehensive module is completed and the parameters are judged to meet the conditions, the iteration is completed. At this time, the system enters the semantic information encoding and code stream compression module 60 to perform operations. When the system executes this step, the bidirectional iterative optimization of the original video and the video to be iterated is completed. Figure 9As shown. The head and tail of the iteratively completed multi-embedded feature information are inserted into the head identification information and the tail identification information. The inserted identification information is mainly used to mark the head and tail of the information fragment. At the same time, the header will also contain basic description information of the multi-embedded semantic feature semantic information, such as: the number of multi-embedded feature representation blocks, the total size of the multi-embedded feature representation block and the verification information. After the information is recorded, the multi-embedded feature information representation block 1 to the semantic information representation block n (n is the total number of representation blocks contained in a certain multi-embedded feature information) are losslessly compressed. Among them, the lossless compression method includes but is not limited to Huffman coding, sparse matrix compression using triple representation method and cross linked list method and other lossless compression methods. After completing the lossless compression, the compressed multi-embedded feature information compressed representation block 1 to multi-embedded feature information compressed representation block n are input into the multi-embedded feature information compressed representation block similarity comparison mapping model. The model can replace the representation blocks with different representations and high similarity by comparing the similarity between the feature representation blocks through transfer learning.

[0087] Specifically, this approach addresses the issue of redundant memory across multiple segments (resolving internal redundancy within multiple embedded feature information) using representation blocks and mapping blocks. If semantic entities and other information are highly similar across different segments, the representation block and mapping block are combined to replace the representation block with similar semantic entities and other information. This method uses the mapping block to perform operations such as addition, deletion, fine-tuning, rotation, inversion, and mirroring on the semantic entities and motion information within the representation block. The mapping and cross-reference scope covers all feature pointer blocks within a single transmission. If the entire video is processed in a single pass, full cross-reference is performed across all representation blocks within the entire video. If the video corresponding to a specific time segment is processed or transmitted in a single pass, cross-reference is performed within the representation blocks corresponding to the video segment within that time segment. In the extreme case, if only a single frame is transmitted at a time, the representation block and mapping block reference representation operations are not performed, and the image is directly entropy-coded into a binary bitstream. This effectively eliminates redundant information across multiple segments within the video, eliminating the need for repeated memory, storage, and transmission of redundant semantic entity and motion information.

[0088] When the module performs reverse input, if the header identification information is read to obtain a total of 2 feature representation blocks, but there are only multi-embedded feature information compression representation block 1 and mapping block 1, then the multi-embedded feature information compression representation block 1 and mapping block 1 are reversely input into the multi-embedded feature information compression representation block similarity comparison mapping model, and multi-embedded feature information compression representation block 1 and multi-embedded feature information compression representation block 2 will be obtained. When the module is executed normally, several mapping blocks with smaller volumes will be obtained. After the migration mapping is completed, the header identification information and the tail identification information are spliced ​​to the output content and entropy coding is performed to encode a pure binary bit stream (a digital stream consisting only of 0s and 1s). The binary bit stream obtained by entropy coding can be adapted to cross-platform applications of most communication platforms. It should be noted that step 106 can be executed in reverse. When step S106 is executed and the corresponding binary code stream is obtained, the encoder part is executed.

[0089] 2. Decoding, including:

[0090] Step S107: After the decoder has completely received all binary code streams of a video and verified them to be correct, it inputs the binary code stream into the code stream decoding module. Code stream decoding corresponds to Figure 9 The reverse operation. The code stream is first input into the entropy decoding module corresponding to the entropy coding and decoded to generate multi-embedded feature information with head and tail identifiers. At this time, some multi-embedded feature information compression representation blocks are missing in the multi-embedded feature information. In this case, the mapping block is needed to complete the mapping of the missing representation blocks. Specific examples are as follows:

[0091] Generally, when the multi-embedded feature information compression representation block n is missing, the multi-embedded feature information compression representation block n-1 and mapping block n-1 exist. In this case, the multi-embedded feature information compression representation block n is replaced by the multi-embedded feature information compression representation block n-1 and mapping block n-1. Generally, the multi-embedded feature information compression representation block n and mapping block n appear in pairs. There are also additional cases, such as the current multi-embedded feature information internally is: multi-embedded feature information compression representation block 1, mapping block 1, mapping block 1, mapping block 1, mapping block 6, multi-embedded feature information compression representation block 6, mapping block 6, multi-embedded feature information compression representation block 8. At this time, the actual restoration results are: multi-embedded feature information compression representation block 1, multi-embedded feature information compression representation block 2 (consisting of multi-embedded feature information compression representation block 1 and the first mapping block 1), multi-embedded feature information compression representation block 3 (consisting of multi-embedded feature information compression representation block 1 and the second mapping block 1), multi-embedded feature information compression representation block 4 (consisting of multi-embedded feature information compression representation block 1 and the third mapping block 1), multi-embedded feature information compression representation block 5 (consisting of multi-embedded feature information compression representation block 6 and the first mapping block 6), multi-embedded feature information compression representation block 6, multi-embedded feature information compression representation block 7 (consisting of multi-embedded feature information compression representation block 6 and the second mapping block 6), and multi-embedded feature information compression representation block 8.

[0092] Once the compressed representation block is fully mapped and restored, all internal representation blocks are decompressed, restoring the information within the representation block to its uncompressed state. Simultaneously, the multi-embedded feature information is verified using the basic descriptive information in the header identification information. Once the verification is correct, the header and tail identification information are discarded. Ultimately, this step yields complete multi-embedded feature information.

[0093] After the code stream decoding module completes its operation in step S108, it then inputs the complete multi-embedded feature information into the generative AI decoding module 80 for video reconstruction. This is similar to the operation in step S104 in the encoder and will not be repeated here. After this sub-step is completed, the semantically reconstructed video is restored. At this point, the decoder is complete.

[0094] At this point, the multi-path embedded video coding system based on generative AI has been fully implemented.

[0095] 1. This method uses a breakthrough video encoding and decoding system based on a generative AI model, which can significantly reduce the storage space and information flow required for storage and transmission.

[0096] 2. After adopting this method, different degrees of media source information expression can be controlled, including but not limited to lossy compression, variable semantic information, variable semantic target trajectory information, variable content body, variable target posture information, variable background information, semantic information encryption, high-density abstract feature expression, etc.

[0097] 3. This method enables customized media source encoding and decoding based on the individual needs of users (decoders) and services (encoders). When the receiving end does not require excessive redundant semantic and feature information, necessary semantic and feature information can be discarded and transmitted before transmission, further reducing the required information stream size.

[0098] 4. After using this method, unlike previous traditional video encryption methods, this method can achieve semantic-level feature selection encryption, and the encrypted information is more refined; this method helps to meet more detailed encryption requirements and provides an effective solution for media sources that need to be split and transmitted or low-bandwidth or ultra-low-bandwidth media encryption transmission.

[0099] 5. After using this method, the transmission integrity of the semantic feature expression information can be fed back and corrected according to the needs of the user (decoding end) / service (encoding end).

[0100] 6. After using this method, the information flow size of the transmitted content and the specifications and size of the media information generated after the information flow is decoded can be flexibly controlled according to the needs of the decoding end. Compared with the original media information, it can be decoded to produce media information with higher resolution, stronger readability, clearer semantics, and more satisfactory to the needs of the receiver.

[0101] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi-path embedded video coding method based on generative artificial intelligence, characterized by: The encoding process includes encoding and decoding steps, wherein the prompt words and required parameters from both the decoding end and the encoding end are present. The encoding process includes the following steps: Step S101: Input the prompt word and requirement parameters of the decoding end and the prompt word and requirement parameters of the encoding end into the prompt word and requirement parameter encoding module (10), encode the input prompt word and requirement parameters and output them into a prompt word vector and requirement parameter code stream, and transmit them to the prompt word vector decoding module (20). At the same time, the output prompt word vector and requirement parameter code stream are also input into the semantic information encoding and code stream compression module (60) as the input of this part and packaged into a binary code stream that can be transmitted; Step S102, after obtaining the prompt word vector and the required parameter code stream, the prompt word vector decoding module (20) decodes them and outputs a feature vector and a feature map; Step S103, after the feature vector and feature map are output, the feature vector, feature map and representation guidance vector are input into the video extraction module (30), and multiple embedded feature information is obtained through the video extraction module (30); Step S104, inputting the obtained multi-embedded feature information into the first generative AI decoding module (40), reconstructing the picture through generative AI and obtaining the video to be iterated; Step S105 is used to receive the video to be iterated and evaluate its video quality. If the evaluation obtains corresponding indicator parameters and determines that the video does not need iteration, the iteration is completed. If the evaluation obtains corresponding indicator parameters and determines that the video still needs to be iterated; Step S106, at this time, the semantic feature extraction and internal optimization iteration are completed, the current prompt word vector and the required parameter code stream and the multi-embedded feature information are input into the semantic information encoding and code stream compression module (60) to generate a binary code stream for transmission, and the encoder stops encoding.

2. The multi-path embedded video coding method based on generative artificial intelligence according to claim 1, wherein: The step S101 specifically includes: Inputting the prompt words and requirement parameters at the decoding end and the prompt words and requirement parameters at the encoding end into the embedding layer of the prompt words and requirement parameter encoding module (10) for modal conversion; Then input it into the Encoder of the prompt word and requirement parameter encoding module (10), and output it as the prompt word vector and the requirement parameter code stream, wherein the Encoder uses a basic learner based on the Transformer learner idea to perform feature extraction; In the step S102, the prompt word vector decoding module (20) has a basic characterizer and can output a feature vector and a feature map.

3. The multi-path embedded video coding method based on generative artificial intelligence according to claim 1, wherein: The step S103 specifically includes The original video information is input into the video segmentation discriminator of the video extraction module (30), which reads the entire video or a very short segment of the real-time video and determines whether there is editing or drastic scene switching, and divides the entire video into smaller video segments; The smaller video segments after segmentation are input into the feature extraction module. In the feature extraction module, the expression and expression intensity of different feature extraction modules in the feature extraction module are partially selected according to the results of phase weight encoding. After feature extraction, each small video segment generates a corresponding multi-embedded feature information representation block. By stacking and splicing several multi-embedded feature information representation blocks, the multi-embedded feature information of the corresponding entire video is obtained.

4. The multi-path embedded video coding method based on generative artificial intelligence according to claim 3, wherein: In step S103, the entire video or a very short segment of the real-time video is read and whether there is editing or drastic scene switching, and the entire video is divided into smaller video segments, including: By default, the video segmentation discriminator uses T seconds as a marking point, calculates the structural consistency parameters of the current frame and the previous frame, and judges that if the structural similarity between the two exceeds a certain fixed initial threshold, the similarity is considered acceptable. The current frame and the previous frame are continuous continuous pictures. If the current predetermined number of frames are all continuous pictures, they are all packaged into one group of video clips. If the current predetermined number of frames are not all continuous pictures, they are segmented and packaged into several groups of video clips based on the SSIM threshold, where the total number of each continuous video frame is less than or equal to the set slicing time x video frame rate. The entire video is divided into smaller video clips by using the segmentation discriminator, and the small video clips generated after segmentation are then input into the feature extraction module.

5. The multi-path embedded video coding method based on generative artificial intelligence according to claim 1, wherein: In step S103, in the feature extraction module, based on the result of phase weight encoding, partially selecting the expression and expression strength of different feature extraction modules in the feature extraction module specifically includes: First, the video clips are input into the video content understanding module, which extracts features of semantic entities and background and foreground information within the video clips and generates representational prompt words for generative AI to complete video decoding and generate the video to be iterated; At the same time, when the phase weight coding completes the weighted expression, feature extraction modules with different expression levels are selected according to the needs of the decoding end. If it is necessary to complete the pixel-level complete information expression, each feature extraction module will be fully called. If it is only necessary to express the most abstract semantic feature information video, only the representation prompt words generated by the video content understanding module will be used. For the remaining basic parameters, the default parameters are directly called in the first generative AI decoding module (40) without customization and pixel-level precise expression. If it is necessary to modify the video feature representation strength, video restoration degree and personalized expression degree according to the needs of the decoding end, the input prompt word information is expressed through phase weight coding and different feature extraction and feature descriptors are called. At the same time, different feature extraction and feature descriptors are shown with different expression strengths according to the different phase weight expression strengths.

6. The multi-path embedded video coding method based on generative artificial intelligence according to claim 5, wherein: In step S103, each feature extraction module and feature descriptor includes: Edge information descriptor, motion information descriptor, semantic extraction module, semantic control feature extraction module, semantic internal consistency feature extraction module, semantic super-resolution feature extraction module, personalized style feature extraction module, video basic information reading and basic control parameter fine-tuning module; The edge information feature descriptor is used to extract all edge contour information and features in the picture, and select the contour feature expression strength of different semantic entities according to the phase weight coding result and generate an edge information feature map; The motion information descriptor includes three sub-parts: a motion vector direction feature descriptor, an optical flow motion estimation descriptor, and a semantic trajectory and relative position descriptor. The motion vector direction feature descriptor is used to extract feature information of motion vectors and their directions of different semantic entities in a video clip and generate a motion vector direction feature map. With the support of phase weight coding, the expression strength of motion vector direction features of different semantic entities is selected according to the representation guidance vector, and the total amount of transmitted information is purposefully adjusted or reduced. The optical flow motion estimation descriptor is used to describe the optical flow motion information features of motion semantics in different foregrounds and backgrounds in a video clip. This feature is used to assist in expressing global motion information and motion blur control in the picture. The semantic extraction module includes foreground semantic entity extraction, background semantic entity extraction, and language and text extraction. It is used to extract the semantics of foreground semantic entities, background semantic entities, and language and text entities in video clips. The extracted foreground and background semantic entities are then used for video reconstruction and semantic-based super-resolution in the generative AI decoder. It is also used for compressing semantic redundancy in multiple video clips. The language entity semantic extraction module extracts the text content and its meaning feature information to restore the corresponding text with higher clarity and personalize it into other languages ​​in different language environments based on the requirements of the representation guidance vector; The semantic control feature extraction module includes semantic trajectory and relative position descriptors, which are used to accurately describe and guide the motion trajectory of local and specific semantics. By expressing the decoding end prompt words or representation guidance vectors under different intensities and semantic entities of phase weight encoding, it can accurately describe the motion trajectory of specific semantic entities and semantic entity sets. It is also used to ensure the consistency and personalized correction of the semantic entity motion trajectory. The semantic internal consistency feature extraction module is used to ensure the internal consistency of semantic entities between different frames generated by the generative AI decoder, maintaining pixel-level consistency and structural consistency within the semantics; The semantic super-resolution feature extraction module is used for image denoising and super-resolution based on semantic information. This module simultaneously extracts pixel noise features and semantic entity noise features of the specified semantic entity area and generates corresponding feature maps. It also extracts low-resolution feature maps from video keyframes for subsequent super-resolution denoising and pixel-level video image restoration of the specified semantic entity in the generative AI decoder. The personalized style feature extraction module is used to extract the style and style feature information of the entire video clip, accurately restore the style, contrast, dynamic range, highlight and shadow details of the video, and represent different style conversion features under the guidance of prompt word vectors, so as to generate personalized style expression features and corresponding picture details in the generative AI decoder; The video basic information reading and basic control parameter fine-tuning module is used to read the video frame number, frame rate, aspect ratio, color space basic information and some semantic entity fine-tuning control parameters. The control parameters and video basic information will be input into the generative AI decoder part to control the consistency of the generated video basic parameter information.

7. The multi-path embedded video coding method based on generative artificial intelligence according to claim 1, wherein: The step S104 specifically includes: By decomposing the multi-embedded feature information into blocks, an embedded feature semantic information representation block is obtained. The representation prompt word is input into the representation prompt word embedding layer to complete the embedding operation of the prompt word to generate high-dimensional array information that can be read by the model. As long as the edge information feature map, motion information feature map, semantic feature map, semantic control feature map, semantic internal consistency feature map, semantic super-resolution feature map and personalized style feature vector exist, they are input into the auxiliary representation and guidance embedding layer accordingly. In the auxiliary representation and guidance embedding layer, the above-representable feature maps and feature vectors are uniformly encoded and embedded, and output as a high-dimensional array in a unified format to ensure that the input specifications and data structure are unified in the auxiliary representation hierarchical expression module. The basic control parameters are input into the parameter loading model to load customized parameter values. Among them, if the auxiliary representation and guidance embedding layer does not receive any information, it directly constructs information with the same data structure and fills it with the number 0. If some values ​​are missing in the parameter loading model, it is loaded with the default parameters. If the representation prompt word embedding layer does not receive any information, the first generative AI decoding module (40) stops working. After completing the data loading of the above three parts, the data generated by the representation prompt word embedding layer is input into the representation prompt word expression core module in the generative AI video reasoning model; the data generated by the auxiliary representation and guidance embedding layer is input into the auxiliary representation hierarchical expression module in the generative AI video reasoning model to express its corresponding features according to different personalized features and representation strengths; finally, the generative AI model is used to generate the video to be iterated / reconstructed.

8. The multi-path embedded video coding method based on generative artificial intelligence according to claim 1, wherein: Step S106 specifically includes: The head and tail of the iteratively completed multi-embedded feature information are inserted into the head identification information and the tail identification information. The inserted identification information is used as the head and tail of the marking information segment. At the same time, the head contains basic description information of the semantic information of the multi-embedded semantic feature. After the information is recorded, the multi-embedded feature information representation block 1 to the semantic information representation block n are losslessly compressed. After the lossless compression is completed, the compressed multi-embedded feature information compressed representation block is input into the multi-embedded feature information compressed representation block similarity comparison mapping model. The model compares the similarity between the feature representation blocks and replaces the representation blocks with different representations and high similarity through transfer learning. When the module is executed normally, several mapping blocks with smaller volumes are obtained. After the migration mapping is completed, the header identification information and the tail identification information are spliced ​​onto the output content and entropy coding is performed to encode a pure binary bit stream. When the corresponding binary code stream is obtained, the encoder part is executed.

9. The multi-path embedded video coding method based on generative artificial intelligence according to claim 1, wherein: The decoding step specifically includes: Step S107, the decoder receives the transmitted binary code stream, inputs the binary code stream into the code stream decoding module (70), decodes the multi-embedded feature information and the prompt word vector and the required parameter code stream, wherein the prompt word vector and the required parameter code stream can be defaulted; Step S108, the obtained multi-embedded feature information and prompt word vector and required parameter code stream are input into the second generative AI decoding module (80), which mainly uses the multi-embedded feature information and the prompt word vector and required parameter code stream as guidance and configuration parameters to generate videos with different degrees of personalized or specialized semantics required by the corresponding receiver. If the prompt word vector and required parameter code stream are missing, the original video is directly generated and expressed.

10. A multi-path embedded video coding system based on generative artificial intelligence, characterized by: Includes an encoder and a decoder, where the encoder includes: The requirement parameter encoding module (10) is used to receive the prompt word and requirement parameter from the decoding end and the prompt word and requirement parameter input from the encoding end, encode the input prompt word and requirement parameter and output them into a prompt word vector and a requirement parameter code stream, and transmit them to the prompt word vector decoding module (20). At the same time, the output prompt word vector and requirement parameter code stream are also input into the semantic information encoding and code stream compression module, and are packaged as the input of this part into a binary code stream that can be transmitted; A prompt word vector decoding module (20) is used to decode the prompt word vector and the required parameter code stream after obtaining them, and output a feature vector and a feature map; The video extraction module (30) is used to receive the feature vector and feature map output by the prompt word vector decoding module (20), and obtain multi-embedded feature information after processing; A first generative AI decoding module (40) is configured to receive the multi-embedded feature information, reconstruct the image through generative AI, and obtain a video to be iterated; A video quality comprehensive evaluation module (50) is used to receive a video to be iterated and perform video quality evaluation on it. If the evaluation obtains corresponding indicator parameters and determines that the video does not need to be iterated, the iteration is completed. If the evaluation obtains corresponding indicator parameters and determines that the video still needs to be iterated in this module, The semantic information encoding and code stream compression module (60) is used to receive the current prompt word vector and the required parameter code stream and multiple embedded feature information after the semantic feature extraction and internal optimization iteration are completed, and generate a binary code stream for transmission; The decoder comprises: A code stream decoding module (70) is used to receive the binary code stream transmitted by the encoder and decode the multi-embedded feature information and the prompt word vector and the required parameter code stream, wherein the prompt word vector and the required parameter code stream can be defaulted; The second generative AI decoding module (80) is used to generate videos with different degrees of personalized or specialized semantics required by the corresponding receiver based on the multi-embedded feature information, the prompt word vector and the required parameter code stream as guidance and configuration parameters. If the prompt word vector and the required parameter code stream are missing, the original video is directly generated and expressed.

Citation Information

Patent Citations

  • Predictive video decoding and rendering based on artificial intelligence

    US20240244216A1

  • A task-oriented video semantic coding system

    WO2024114817A1