Video information processing method, multimedia information processing method and device

By analyzing the structural similarity and texture differences between the coded frame groups and the reference frame groups of video clips, and selecting appropriate encoding decisions, the problem of excessive time-consuming in the video encoding standard is solved, rapid encoding and reduction in calculations is achieved, and user experience is improved.

CN113709461BActive Publication Date: 2025-08-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110298455.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-19
Publication Date
2025-08-22
Estimated Expiration
2041-03-19

AI Technical Summary

Technical Problem

The existing video encoding standards take too long to choose the best partition mode, resulting in video compression processing extending user waiting time, large amount of calculations, and affecting user experience.

Method used

By obtaining the structural similarity difference between the coded frame group and the reference frame group of video clips, the coded frame group type is determined, and the encoding decision is selected based on the texture difference parameters, including static and dynamic early skip modes, reducing the encoding decision waiting time.

Benefits of technology

Quickly and accurately determine encoding decisions, improve video encoding speed, reduce calculation amount, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113709461B_ABST
    Figure CN113709461B_ABST
Patent Text Reader

Abstract

The present invention provides a video information processing method, a multimedia information processing method, an apparatus, an electronic device and a storage medium. The method comprises: determining the type of coding frame group in the video to be encoded according to the structural similarity difference between the current coding frame group and the reference coding frame group included in the video segment to be analyzed; determining the texture difference parameters of the current coding unit and the reference coding unit of the video segment to be analyzed; determining the coding decision that matches the video to be encoded based on the texture difference parameters of the current coding unit and the reference coding unit; processing the video to be encoded according to the determined coding decision to realize the encoding of the video to be encoded, thereby more quickly determining the encoding method for the video, reducing the waiting time for selecting the coding decision, and improving the speed of the video encoding process. At the same time, the computational amount of video information processing is saved, the computational amount of the device is reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video information processing technology, and in particular to a video information processing method, a multimedia information processing method, a device, an electronic device and a storage medium. Background Art

[0002] Video coding standards such as AVS3 (Audio Video Coding Standard), VVC (Versatile Video Coding), and future video coding standards have adopted more intra-frame and inter-frame coding tools, as well as loop filtering tools. When selecting the optimal division method for a coding unit, the encoder first determines the available division methods for that unit based on constraints. Selecting the optimal division method for a coding unit in a video requires iterating through each division method and making a rate-distortion optimization decision. This rate-distortion optimization decision-making process is time-consuming during the overall encoding process, hindering video compression and increasing user wait time. Summary of the Invention

[0003] In view of this, the embodiments of the present invention provide a video information processing method, a multimedia information processing method, an apparatus, an electronic device and a storage medium, which can quickly and accurately determine the coding decision that matches the video to be encoded through the type of coding frame group and the texture difference of the coding unit, and can also more quickly determine the encoding method for the video, reduce the waiting time for selecting the coding decision, and improve the speed of the video encoding process. At the same time, it saves the computational amount of video information processing, reduces the computational amount of the device, and improves the user experience.

[0004] The technical solution of the embodiment of the present invention is achieved as follows:

[0005] An embodiment of the present invention provides a video information processing method, including:

[0006] Acquire a video to be encoded in a video processing environment, and intercept a video segment to be analyzed in the video to be encoded;

[0007] determining the type of the coding frame group in the to-be-coded video according to a structural similarity difference between a current coding frame group and a reference coding frame group included in the to-be-coded video segment;

[0008] Determining a texture difference parameter between a current coding unit and a reference coding unit of the video segment to be analyzed;

[0009] When a texture difference parameter between the current coding unit and the reference coding unit is greater than or equal to a texture difference parameter threshold, determining that an encoding decision matching the video to be encoded is not to execute an early skip mode;

[0010] When a texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, and the coding frame group is of a first frame group type, determining that a coding decision matching the to-be-coded video is a static early skip mode;

[0011] When a texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, and the coding frame group is of a second frame group type, determining that a coding decision matching the to-be-coded video is a dynamic early skip mode;

[0012] The video to be encoded is processed according to the determined encoding decision to achieve encoding of the video to be encoded.

[0013] An embodiment of the present invention further provides a method for processing multimedia information, the method comprising:

[0014] Separating target audio and target video from multimedia information;

[0015] determining an encoding decision that matches the target video;

[0016] Determining a corresponding encoding method according to the encoding decision;

[0017] Processing the target video using the determined encoding method to achieve encoding of the target video;

[0018] The target video and the target audio that have been coded are encapsulated into new multimedia information to achieve compression of the multimedia information; wherein the coding decision is obtained as described above.

[0019] An embodiment of the present invention further provides a video information processing device, the device comprising:

[0020] An information transmission module is used to obtain a video to be encoded and intercept a video segment to be analyzed in the video to be encoded;

[0021] an information processing module, configured to determine the type of the coding frame group in the to-be-coded video according to a structural similarity difference between a current coding frame group and a reference coding frame group included in the to-be-coded video segment;

[0022] The information processing module is configured to determine a texture difference parameter between a current coding unit and a reference coding unit of the video segment to be analyzed;

[0023] The information processing module is configured to, when a texture difference parameter between the current coding unit and the reference coding unit is greater than or equal to a texture difference parameter threshold, determine that a coding decision matching the video to be coded is not to execute the early skip mode; when the texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, and the coding frame group is of a first frame group type, determine that a coding decision matching the video to be coded is a static early skip mode; and when the texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, and the coding frame group is of a second frame group type, determine that a coding decision matching the video to be coded is a dynamic early skip mode;

[0024] The information processing module is used to process the video to be encoded according to the determined encoding decision, so as to realize encoding of the video to be encoded.

[0025] In the above scheme,

[0026] The information processing module is configured to obtain the intra-frame prediction frame of the current coding frame group and the intra-frame prediction frame of the reference coding frame group;

[0027] The information processing module is configured to determine corresponding structural similarity differences based on the intra-frame prediction frame of the current coding frame group and the intra-frame prediction frame of the reference coding frame group;

[0028] The information processing module is used to determine a structural similarity difference threshold that matches the video processing environment;

[0029] The information processing module is configured to determine that the coding frame group in the to-be-coded video is a first type of coding frame group when the structural similarity difference is less than the structural similarity difference threshold;

[0030] The information processing module is configured to determine that the coding frame group in the to-be-coded video is a second type of coding frame group when the structural similarity difference is greater than or equal to the difference threshold.

[0031] In the above scheme,

[0032] The information processing module is used to obtain the brightness average value, brightness value variance, and brightness value covariance corresponding to the intra-frame prediction frame of the current coding frame group;

[0033] The information processing module is used to obtain the brightness average value, brightness value variance, and brightness value covariance corresponding to the intra-frame prediction frame of the reference coding frame group;

[0034] The information processing module is used to determine the structural similarity difference between the current coding frame group and the reference coding frame group based on the brightness average, brightness value variance, and brightness value covariance corresponding to the intra-frame prediction frame of the current coding frame group and the brightness average, brightness value variance, and brightness value covariance corresponding to the intra-frame prediction frame of the reference coding frame group.

[0035] In the above scheme,

[0036] The information processing module is used to determine a histogram difference parameter and a gradient difference parameter between the current coding unit and the reference coding unit;

[0037] The information processing module is configured to determine a texture difference parameter between the current coding unit and the reference coding unit of the video segment to be analyzed based on a histogram difference parameter and a gradient difference parameter between the current coding unit and the reference coding unit.

[0038] In the above scheme,

[0039] The information processing module is configured to determine the type of the corresponding coding frame group when the texture difference parameter between the current coding unit and the reference coding unit is greater than or equal to the texture difference parameter threshold;

[0040] The information processing module is configured to execute a static early skip mode when the coded frame group is of a first frame group type;

[0041] The information processing module is configured to execute a dynamic early skip mode when the coded frame group is of a second frame group type.

[0042] In the above scheme,

[0043] The information processing module is used to determine the number of extended quadtree structures and the number of binary tree structures in the current coding unit;

[0044] The information processing module is configured to mark the type of the current coding unit based on the number of extended quadtree structures and the number of binary tree structures in the current coding unit;

[0045] The information processing module is configured to adjust the texture difference parameter threshold according to the type of the current coding unit.

[0046] In the above scheme,

[0047] The information processing module is configured to determine the quadtree partitioning as a disabled partitioning mode when both the histogram difference and the structural difference of the quadtree partitioning that matches the video segment to be analyzed are smaller than the quadtree histogram difference threshold and the quadtree structural difference threshold;

[0048] The information processing module is configured to determine the binary tree partitioning as a disabled partitioning mode when both the histogram difference and the structural difference of the binary tree partitioning that matches the video segment to be analyzed are less than a preset binary tree histogram difference threshold and a binary tree structural difference threshold;

[0049] The information processing module is used to determine the extended quadtree partitioning as a disabled partitioning mode when the histogram difference and structural difference of the extended quadtree partitioning matching the video clip to be analyzed are both less than the preset extended quadtree histogram difference threshold and extended quadtree structural difference threshold.

[0050] In the above scheme,

[0051] The information processing module is used to obtain identification information of the video to be encoded, the video to be encoded corresponding to the encoding decision and the video information after encoding processing;

[0052] The information processing module is used to generate a target block based on the identification information of the video to be encoded, the encoding decision corresponding to the video to be encoded, and the video information after encoding processing, and add the target block to the blockchain network.

[0053] An embodiment of the present invention further provides a multimedia information processing device, comprising:

[0054] An information separation device, used for separating target audio and target video from multimedia information;

[0055] a video processing device for determining an encoding decision that matches the target video;

[0056] The video processing device determines a corresponding encoding method according to the encoding decision;

[0057] The video processing device is used to process the target video using the determined encoding method to achieve encoding of the target video;

[0058] The video processing device is used to encapsulate the target video and the target audio that have been coded into new multimedia information to achieve compression of the multimedia information.

[0059] An embodiment of the present invention further provides an electronic device, comprising:

[0060] a memory for storing executable instructions;

[0061] The processor is configured to implement any one of the video information processing methods described above, or the multimedia information processing method described above, when running the executable instructions stored in the memory.

[0062] An embodiment of the present invention also provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implements any of the video information processing methods described in the foregoing, or implements any of the point cloud data decoding methods described in the foregoing.

[0063] The embodiments of the present invention have the following beneficial effects:

[0064] The embodiment of the present invention obtains a video to be encoded in a video processing environment and intercepts a video segment to be analyzed in the video to be encoded; determines the type of the encoding frame group in the video to be encoded based on the structural similarity difference between the current encoding frame group and the reference encoding frame group included in the video segment to be analyzed; when it is determined that the encoding frame group in the video to be encoded is of the first frame group type, determines the texture difference parameters of the current encoding unit and the reference encoding unit of the video segment to be analyzed; determines the encoding decision that matches the video to be encoded based on the texture difference parameters of the current encoding unit and the reference encoding unit; processes the video to be encoded based on the determined encoding decision to achieve encoding of the video to be encoded, thereby enabling the encoding decision that matches the video to be encoded to be quickly and accurately determined based on the type of the encoding frame group and the texture difference of the encoding unit, and also more quickly determining the encoding method for the video, reducing the waiting time for selecting the encoding decision, and improving the speed of the video encoding process. At the same time, the computational workload of video information processing is saved, the computational workload of the device is reduced, and the user experience is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 A schematic diagram of a usage scenario of the multimedia information processing method provided by an embodiment of the present invention;

[0066] Figure 2 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention;

[0067] Figure 3 This is an optional flowchart of a video information processing process in an embodiment of the present invention;

[0068] Figure 4 An optional flowchart of a video information processing method provided by an embodiment of the present invention;

[0069] Figure 5 An optional flowchart of a video information processing method provided by an embodiment of the present invention;

[0070] Figure 6 Schematic diagram of SDIF changes between GOPs of different sequences in an embodiment of the present invention;

[0071] Figure 7An optional flowchart of a video information processing method provided by an embodiment of the present invention;

[0072] Figure 8 A schematic diagram of difference calculation in a video information processing method according to an embodiment of the present invention;

[0073] Figure 9 1 is a schematic diagram of the architecture of a video information processing device 100 provided by an embodiment of the present invention;

[0074] Figure 10 2 is a schematic diagram of the structure of a blockchain in a blockchain network 200 provided in an embodiment of the present invention;

[0075] Figure 11 2 is a functional architecture diagram of the blockchain network 200 provided in an embodiment of the present invention;

[0076] Figure 12 An optional flowchart of the multimedia information processing method provided by the embodiment of the present invention

[0077] Figure 13 Schematic diagram of data processing of a video processing method in an embodiment of the present invention. DETAILED DESCRIPTION

[0078] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0079] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0080] Before further explaining the embodiments of the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are subject to the following interpretations.

[0081] 1) Reference frame: In video encoding and decoding, a frame that serves as reference data for other frames and is used to reconstruct frames to obtain reference data between frames during the encoding / decoding process.

[0082] 2) SDK: The full name is Software Development Kit, which can be translated into software development kit. It is a collection of development tools for building application software for specific software packages, software frameworks, hardware platforms, operating systems, etc. In a broad sense, it includes a collection of related documents, examples and tools to assist in the development of a certain type of software.

[0083] 3) P frame: Inter-frame prediction frame, which can use intra-frame prediction and inter-frame prediction, and can make forward reference prediction coding decisions.

[0084] 4) B frame: Inter-frame prediction frame, which can use intra-frame prediction and inter-frame prediction, and can be forward, backward, and bidirectional reference prediction.

[0085] 5) I frame: Intra-frame prediction frame, which uses intra-frame information for prediction.

[0086] 6) Video codec standard: a certain agreed-upon rule for decoding video bitstreams.

[0087] 7) Video transcoding: This process converts a compressed video stream into another video stream to accommodate different network bandwidths, terminal processing capabilities, and user needs.

[0088] 8) Client: The carrier that implements specific functions in the terminal. For example, the mobile client (APP) is the carrier of specific functions in the mobile terminal, such as executing the function of online live broadcast (video streaming) or the function of playing online videos.

[0089] 9) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0090] 10) Mini Programs are programs developed using a front-end language (such as JavaScript) that implement services within a Hypertext Markup Language (HTML) page. These programs are downloaded by a client (such as a browser or any client with an embedded browser core) via a network (e.g., the internet) and interpreted and executed within the client's browser environment, eliminating the need for client-side installation. For example, by activating a mini program on a terminal through voice commands, a social networking client can download and run mini programs for various services, such as image editing and human eye image correction.

[0091] 11) Transaction: Equivalent to the computer term "transaction," a transaction includes operations that need to be submitted to a blockchain network for execution. It does not refer solely to transactions in a business context. Given the conventional use of the term "transaction" in blockchain technology, the embodiments of this invention follow this convention.

[0092] For example, the Deploy transaction is used to install a specified smart contract to a node in the blockchain network and prepare it for invocation; the Invoke transaction is used to append transaction records to the blockchain by invoking a smart contract and perform operations on the blockchain's state database, including update operations (including adding, deleting, and modifying key-value pairs in the state database) and query operations (i.e., querying key-value pairs in the state database).

[0093] 12) Blockchain: It is an encrypted, chained transaction storage structure formed by blocks.

[0094] For example, the header of each block can include the hash value of all transactions in the block, as well as the hash value of all transactions in the previous block, thereby achieving tamper-proof and anti-forgery of transactions in the block based on the hash value; newly generated transactions are filled into the block and, after consensus among nodes in the blockchain network, will be appended to the end of the blockchain to form a chain-like growth.

[0095] 13) Blockchain Network: A collection of nodes that incorporate new blocks into the blockchain through consensus.

[0096] 14) Ledger: A general term for the blockchain (also known as ledger data) and the state database synchronized with the blockchain.

[0097] Among them, the blockchain records transactions in the form of files in the file system; the state database records transactions in the blockchain in the form of different types of key (Key) value (Value) pairs to support fast queries on transactions in the blockchain.

[0098] 15) Smart Contracts: Also known as chain code or application code, these are programs deployed on nodes in a blockchain network. Nodes execute smart contracts invoked in received transactions to update or query key-value pairs in the ledger database.

[0099] 16) Consensus: This is a process within a blockchain network used to reach agreement on transactions within a block among multiple nodes involved. The agreed-upon block is appended to the end of the blockchain. Consensus mechanisms include Proof of Work (PoW), Proof of Stake (PoS), Delegated Proof-of-Stake (DPoS), and Proof of Elapsed Time (PoET).

[0100] Figure 1 Schematic diagram of the use scenario of the video information processing method provided by the embodiment of the present invention, see Figure 1 The terminals (including terminals 10-1 and 10-2) are provided with corresponding clients capable of performing different functions. The clients are terminals (including terminals 10-1 and 10-2) that use different service processes to obtain different video information from corresponding servers 200 via network 300 for browsing. The terminals are connected to server 200 via network 300. Network 300 can be a wide area network (WAN), a local area network (LAN), or a combination of the two, using a wireless link for data transmission. The types of videos obtained by the terminals (including terminals 10-1 and 10-2) from corresponding servers 200 via network 300 vary. For example, the terminals (including terminals 10-1 and 10-2) can obtain videos (i.e., videos containing video information or corresponding video links) from corresponding servers 200 via network 300, or they can obtain videos of different types (e.g., short videos or long videos) from corresponding servers 400 via network 300 for browsing. Servers 200 and 400 may store different types of videos. In some embodiments of the present invention, the processes of different types of videos stored in the server 200 may be written in software codes of different programming languages, and the code objects may be different types of code entities. For example, in software code written in C language, a code object may be a function. In software code written in JAVA language, a code object may be a class, and in iOS OC language, it may be a section of target code. In software code written in C++ language, a code object may be a class or a function. In this application, no distinction is made between the compilation environments of different types of videos.

[0101] Furthermore, when the server 200 sends or receives different types of videos to the terminals (terminal 10-1 and / or terminal 10-2) via the network 300, the video information needs to be compressed because the storage space occupied by the video information is large. As an example, the server 200 is configured to obtain a video to be encoded in a video processing environment and intercept a video segment to be analyzed from the video to be encoded; determine the type of the coding frame group in the video to be encoded based on the structural similarity difference between the current coding frame group and the reference coding frame group included in the video segment to be analyzed; when it is determined that the coding frame group in the video to be encoded is of the first frame group type, determine texture difference parameters between the current coding unit and the reference coding unit of the video segment to be analyzed; determine a coding decision that matches the video to be encoded based on the texture difference parameters between the current coding unit and the reference coding unit; and process the video to be encoded using the determined coding decision to achieve encoding of the video to be encoded.

[0102] The structure of the server of the embodiment of the present invention is described in detail below. The server can be implemented in various forms, such as a dedicated terminal with video information processing function, such as a gateway, or a server with video information processing function, such as the aforementioned Figure 1 Server 200 in. Figure 2 The schematic diagram of the structure of the electronic device provided in the embodiment of the present invention can be understood as follows: Figure 2 Only the exemplary structure of the server is shown, not all structures, and can be implemented as needed. Figure 2 Partial or complete structure shown.

[0103] The server provided in the embodiment of the present invention includes: at least one processor 201, a memory 202, a user interface 203 and at least one network interface 204. The various components in the electronic device 20 are coupled together via a bus system 205. It can be understood that the bus system 205 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 205 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 205 is not described in detail. Figure 2 Various buses are labeled as bus system 205.

[0104] The user interface 203 may include a display, a keyboard, a mouse, a trackball, a click wheel, keys, buttons, a touch pad or a touch screen.

[0105] It will be appreciated that the memory 202 may be volatile memory or non-volatile memory, or may include both. The memory 202 in this embodiment of the present invention can store data to support the operation of the terminal (e.g., 10-1). Examples of such data include any computer program used to operate on the terminal (e.g., 10-1), such as an operating system and application programs. The operating system includes various system programs, such as a framework layer, a core library layer, and a driver layer, which implement various basic services and handle hardware-based tasks. Application programs may include various application programs.

[0106] In some embodiments, the video information processing device provided by the embodiments of the present invention can be implemented using a combination of software and hardware. As an example, the video information processing device provided by the embodiments of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the video information processing method provided by the embodiments of the present invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0107] As an example of a video information processing device provided by an embodiment of the present invention being implemented by a combination of software and hardware, the video information processing device provided by an embodiment of the present invention can be directly embodied as a combination of software modules executed by the processor 201. The software module can be located in a storage medium, and the storage medium is located in the memory 202. The processor 201 reads the executable instructions included in the software module in the memory 202, and combines with the necessary hardware (for example, including the processor 201 and other components connected to the bus 205) to complete the video information processing method provided by the embodiment of the present invention.

[0108] As an example, the processor 201 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0109] As an example of a hardware implementation of the video information processing device provided in an embodiment of the present invention, the device provided in an embodiment of the present invention can be directly executed by a processor 201 in the form of a hardware decoding processor. For example, the video information processing method provided in an embodiment of the present invention can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0110] The memory 202 in the embodiment of the present invention is used to store various types of data to support the operation of the electronic device 20. Examples of such data include any executable instructions for operating on the electronic device 20, such as executable instructions. A program implementing the video information processing method according to the embodiment of the present invention may be included in the executable instructions.

[0111] In other embodiments, the video information processing device provided by the embodiment of the present invention can be implemented in software. Figure 2 The video information processing device 2020 stored in the memory 202 is shown. This device can be software in the form of a program or plug-in, and includes a series of modules. As an example of a program stored in the memory 202, the video information processing device 2020 can be included. The video information processing device 2020 includes the following software modules: an information transmission module 2081 and an information processing module 2082. When the software modules in the video information processing device 2020 are read into the RAM by the processor 201 and executed, the video information processing method provided by the embodiment of the present invention is implemented. The functions of each software module in the video information processing device 2020 are described below:

[0112] The information transmission module 2081 is used to obtain the video to be encoded and intercept the video segment to be analyzed in the video to be encoded.

[0113] The information processing module 2082 is configured to determine the type of the coding frame group in the to-be-coded video according to the structural similarity difference between the current coding frame group and the reference coding frame group included in the to-be-coded video segment.

[0114] The information processing module 2082 is configured to determine a texture difference parameter between a current coding unit and a reference coding unit of the video segment to be analyzed when it is determined that the coding frame group in the video to be encoded is of the first frame group type.

[0115] The information processing module 2082 is configured to determine a coding decision that matches the video to be encoded based on the texture difference parameter between the current coding unit and the reference coding unit.

[0116] The information processing module 2082 is configured to process the video to be encoded according to the determined encoding decision, so as to implement encoding of the video to be encoded.

[0117] according to Figure 2 In one aspect of the electronic device shown, the present application further provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of the computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the various embodiments and combinations of embodiments provided in the various optional implementations of the above-mentioned video information processing method.

[0118] Before introducing the video information processing method provided by this application, the video information processing process in the related art is first introduced, wherein: Figure 3 This is an optional flow chart of the video information processing process in an embodiment of the present invention. In the video compression process, such as VVC (Versatile Video Coding, universal video coding) and AVS3 (Audio Video Coding Standard 3, audio and video coding standard 3), HEVC (High Efficiency Video Coding, high performance video coding) encoding process, many mode decision processes are involved. Figure 3 Related video coding technologies, such as H.265 / HEVC (High Efficient Video Coding), H.266 / VVC (Versatile Video Coding), and AVS (Audio Video Coding Standard) (such as AVS3), use a hybrid coding framework to perform the following operations and processing on the input raw video signal:

[0119] 1. Block Partition Structure: The input image is divided into several non-overlapping processing units, each of which performs similar compression operations. This processing unit is called a CTU (Coding Tree Unit) or LCU (Large Coding Unit). Below the CTU, further subdivision can be performed to obtain one or more basic coding units, called CUs (Coding Units). Each CU is the most basic element in the encoding process. The following describes the various possible encoding methods for each CU.

[0120] 2. Predictive Coding: This includes methods such as intra-frame prediction and inter-frame prediction. The original video signal is predicted using a selected reconstructed video signal to generate a residual video signal. The encoder must select the most appropriate predictive coding mode for the current CU from among many possible modes and inform the decoder. Intra-frame prediction uses the predicted signal from a previously encoded and reconstructed region within the same image. Inter-frame prediction uses the predicted signal from a previously encoded image (called a reference image) that is different from the current image.

[0121] Taking the AVS3 coding standard as an example, a video frame is first divided into multiple Large Coding Units (LCUs). Each LCU is then recursively divided into smaller Coding Units (CUs). There are six possible CU splitting methods: non-split, vertical binary split, horizontal binary split, vertical extended binary split, horizontal extended quadtree split, and quadtree split. Selecting the optimal splitting method for a CU requires iterating through each splitting method and performing a rate-distortion optimization decision. This process is time-consuming during the encoding process, hindering video compression and increasing user latency.

[0122] Combine Figure 2 The electronic device 20 shown illustrates the video information processing method provided by the embodiment of the present invention, see Figure 4 , Figure 4 This is an optional flow chart of the video information processing method provided by the embodiment of the present invention. It can be understood that: Figure 4 The steps shown can be performed by various servers running video information processing devices, such as dedicated terminals, servers or server clusters with video information processing functions. Figure 4 The steps shown are explained.

[0123] Step 401: The video information processing apparatus obtains a video to be encoded in a video processing environment, and intercepts a video segment to be analyzed in the video to be encoded.

[0124] Among them, taking the videos uploaded by WeChat mini-program as an example, users can upload videos shot by the terminal or videos saved in the terminal through WeChat mini-program. The video information processing device encapsulated in the readable storage medium of the server can obtain these videos to be encoded through the communication link, and intercept videos with a fixed number of frames (for example, the first 10 frames of the video) or a fixed time (1 second of the first second of the video) as the video clip to be analyzed.

[0125] Step 402: The video information processing apparatus determines the type of the coding frame group in the to-be-coded video according to the structural similarity difference between the current coding frame group and the reference coding frame group included in the to-be-analyzed video segment.

[0126] In some embodiments of the present invention, determining the type of the coding frame group in the to-be-coded video based on the structural similarity difference between the current coding frame group and the reference coding frame group included in the to-be-coded video segment may be achieved by:

[0127] Obtain the intra-frame prediction frame of the current coding frame group and the intra-frame prediction frame of the reference coding frame group; determine the corresponding structural similarity difference based on the intra-frame prediction frame of the current coding frame group and the intra-frame prediction frame of the reference coding frame group; determine a structural similarity difference threshold that matches the video processing environment; when the structural similarity difference is less than the structural similarity difference threshold, determine that the coding frame group in the to-be-encoded video is a first type of coding frame group; when the structural similarity difference is greater than or equal to the difference threshold, determine that the coding frame group in the to-be-encoded video is a second type of coding frame group.

[0128] in, Figure 5An optional flow chart of a video information processing method provided for an embodiment of the present invention. Due to different types of videos to be encoded, they can be divided into two categories: dynamic scene videos and fixed scene videos. Dynamic scene videos can include film and television drama videos, mobile phone shooting videos and other types of videos. Fixed scene videos can include: network conference videos and other types of videos. When calculating the structural similarity difference, the brightness average, brightness value variance, brightness value covariance, and pixel dynamic range corresponding to the intra-frame prediction frames of the current coding frame group can be obtained; the brightness average, brightness value variance, and brightness value covariance corresponding to the intra-frame prediction frames of the reference coding frame group are obtained; based on the brightness average, brightness value variance, and brightness value covariance corresponding to the intra-frame prediction frames of the current coding frame group and the brightness average, brightness value variance, and brightness value covariance corresponding to the intra-frame prediction frames of the reference coding frame group, refer to Formula 1, Formula 2 and Formula 3 to determine the structural similarity difference between the current coding frame group and the reference coding frame group.

[0129] Formula 1

[0130] Formula 2

[0131] in, is the average brightness value in the Q image, is the average brightness value in the S image, is the variance of brightness in Q, is the variance of brightness in S, is the covariance of the bright spots in Q and S, , is a constant to maintain stability, L is the dynamic range of pixel values, k1 is 0.01, and k2 is 0.03.

[0132] Formula 3

[0133] In some embodiments of the present invention, when determining the coding frame group type of an image frame through the coding frame group, the SDIF between GOPs of the AVS3 standard reference sequence can be counted. In order to adapt to different video processing scenarios, the structural similarity difference threshold can also be adjusted based on different test sequences such as FourPeople, Johnny, KristenAndSara in the HEVC standard reference sequence. Figure 6 , Figure 6 FIG. 1 is a schematic diagram of SDIF changes between GOPs of different sequences in an embodiment of the present invention. Figure 6 As shown, by setting an adaptive threshold, dynamic scene videos and static scene videos in the video to be processed can be distinguished. The SDIF threshold can be set to 0.2 to match different video processing environments.

[0134] Step 403: Determine texture difference parameters between the current coding unit and the reference coding unit of the video segment to be analyzed.

[0135] In some embodiments of the present invention, determining the texture difference parameter between the current coding unit and the reference coding unit of the video segment to be analyzed may be achieved by:

[0136] Determine the corresponding current coding unit and reference coding unit in the video segment to be analyzed; determine the histogram difference parameter and gradient difference parameter of the current coding unit and the reference coding unit; and determine the texture difference parameter of the current coding unit and the reference coding unit of the video segment to be analyzed based on the histogram difference parameter and gradient difference parameter of the current coding unit and the reference coding unit. Figure 7 An optional flow chart of the video information processing method provided by an embodiment of the present invention, when the encoder encodes the current CU, it can search for the position of the CU in the reference frame, calculate the texture difference between the current CU and the reference CU, and then compare the degree of difference with the threshold to determine whether the EQT partitioning mode or the BT partitioning mode is prohibited. If the difference meets the difference threshold, it is determined whether the current GOP is a slowly changing GOP. If so, a static early skip mode is used to speed up the processing of the video image frame. Otherwise, a dynamic early skip mode is used to improve the processing of videos with large changes. If the difference does not meet the threshold, the normal encoding mode is used and the early skip mode is not executed. When the condition determines that the current CU disables EQT, the available options are QT, BT, and no partitioning; when the condition determines that the current CU disables BT, the available options are QT, EQT, and no partitioning.

[0137] Step 404: The video information processing apparatus determines a coding decision that matches the video to be coded based on the texture difference parameters between the current coding unit and the reference coding unit.

[0138] In some embodiments of the present invention, when the texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, determining the coding decision that matches the to-be-coded video based on the type of the coding frame group may be implemented in the following manner:

[0139] When the texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, the type of the corresponding coding frame group is determined; when the coding frame group is of the first frame group type, the static early skip mode is executed; when the coding frame group is of the second frame group type, the dynamic early skip mode is executed. Figure 8 , Figure 8The diagram of difference calculation of the video information processing method provided by the embodiment of the present invention is shown in FIG4 , wherein the calculation method of the histogram difference is shown in Formula 4 and Formula 5, and the calculation operator of the gradient difference is shown in FIG4 . Figure 8 shown.

[0140] Formula 4

[0141] Formula 5

[0142] in, They are the current CU and the reference CU, and is the number of each luma pixel in the two CUs.

[0143] Step 405: The video information processing apparatus processes the video to be encoded according to the determined encoding decision to implement encoding of the video to be encoded.

[0144] In some embodiments of the present invention, when adjusting the threshold, the number of extended quadtree structures and the number of binary tree structures in the current coding unit may also be determined; based on the number of extended quadtree structures and the number of binary tree structures in the current coding unit, the type of the current coding unit may be marked;

[0145] The texture difference parameter threshold is adjusted based on the type of the current coding unit. Data can be extracted from the reference sequences BQMall, BQTerrace, FourPeople, and Johnny used in the High Efficiency Video Coding (HEVC) standard. The extracted QPs are {27, 32, 38, 45}, and the Rancom-Access (RA) mode is used. A decision tree is used for training. In this solution, a size restriction is imposed, resulting in the following extracted features:

[0146]

[0147] in, and Indicates the number of EQT and BT currently used within the CU.

[0148] The labels for decision tree training are given by as well as Determine, referring to Formula 6 and Formula 7:

[0149] Formula 6

[0150] formula 7

[0151] if Less than the set , mark the CU as restricted partitioning, otherwise set it to unrestricted partitioning; if Less than the set , mark the CU as restricted partitioning, otherwise set it to unrestricted partitioning.

[0152] The embodiments of the present invention may be implemented in conjunction with cloud technology or blockchain network technology. Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network or local area network to achieve data computing, storage, processing, and sharing. It can also be understood as a general term for network technology, information technology, integration technology, management platform technology, and application technology based on cloud computing business models. The backend services of technical network systems require a large amount of computing and storage resources, such as video websites, image websites, and more portal websites. Therefore, cloud technology needs to be supported by cloud computing.

[0153] It should be noted that cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides these resources is called the "cloud." To users, the resources in the "cloud" appear to be infinitely scalable and can be accessed at any time, used on demand, and expanded at any time, with a pay-per-use fee. As a provider of cloud computing's basic capabilities, a cloud computing resource pool platform, often referred to as Infrastructure as a Service (IaaS), is established. Various types of virtual resources are deployed within the resource pool for external customers to choose from. The cloud computing resource pool primarily includes computing devices (which can be virtualized machines, including operating systems), storage devices, and network devices.

[0154] In some embodiments of the present invention, the identification information of the video to be encoded and the encoding decision corresponding to the video to be encoded can also be obtained; based on the identification information of the video to be encoded, the video to be encoded and the encoding decision corresponding to the video to be encoded, a target block is generated, and the target block is added to the blockchain network. Specifically, the identification information of the video to be encoded, the encoding decision corresponding to the video to be encoded and the encoded video information are sent to the blockchain network so that

[0155] The nodes of the blockchain network fill the identification information of the video to be encoded, the encoding decision corresponding to the video to be encoded, and the video information after encoding into a new block, and when a consensus is reached on the new block, the new block is appended to the end of the blockchain.

[0156] In the above solution, the method further includes:

[0157] Receive data synchronization requests from other nodes in the blockchain network; verify the permissions of the other nodes in response to the data synchronization requests; when the permissions of the other nodes are verified, control data synchronization between the current node and the other nodes to enable the other nodes to obtain the identification information of the video to be encoded, the encoding decision corresponding to the video to be encoded, and the video information after encoding.

[0158] In the above scheme, the method also includes: in response to a query request, parsing the query request to obtain a corresponding user identifier; obtaining permission information in a target block in the blockchain network based on the user identifier; verifying the matching of the permission information and the user identifier; when the permission information matches the user identifier, obtaining the corresponding identification information of the video to be encoded, the encoding decision corresponding to the video to be encoded, and the video information after encoding in the blockchain network; in response to the query request, pushing the obtained identification information of the video to be encoded, the encoding decision corresponding to the video to be encoded, and the video information after encoding to the corresponding client, so that the client obtains the corresponding identification information of the video to be encoded, the encoding decision corresponding to the video to be encoded, and the video information after encoding stored in the blockchain network.

[0159] Continue to see Figure 9 , Figure 9 It is a schematic diagram of the architecture of the video information processing device 100 provided by an embodiment of the present invention, including a blockchain network 200 (consensus nodes 210-1 to consensus nodes 210-3 are exemplarily shown), an authentication center 300, a business entity 400 and a business entity 500, which are described below respectively.

[0160] The blockchain network 200 is flexible and diverse, and can be, for example, a public blockchain, a private blockchain, or a consortium blockchain. For example, with a public blockchain, any business entity's electronic device, such as a user terminal or server, can access the blockchain network 200 without authorization. For example, with a consortium blockchain, a business entity can obtain authorization and its subordinate electronic devices (such as terminals / servers) can access the blockchain network 200, becoming client nodes in the blockchain network 200.

[0161] In some embodiments, client nodes may only function as observers of blockchain network 200, providing support for business entities initiating transactions (e.g., storing data on-chain or querying on-chain data). Client nodes may implement the functions of consensus node 210 of blockchain network 200, such as sorting, consensus services, and ledger functions, by default or selectively (e.g., depending on the specific business needs of the business entity). This allows the business entity's data and business processing logic to be migrated to blockchain network 200 to the greatest extent possible, ensuring the trustworthiness and traceability of data and business processing processes through blockchain network 200.

[0162] The consensus nodes in the blockchain network 200 receive transactions submitted by client nodes (for example, the client node 410 belonging to the business entity 400 shown in the previous embodiment, and the client node 510 belonging to the database operator system) from different business entities (for example, the business entity 400 and the business entity 500 shown in the previous embodiment), execute transactions to update the ledger or query the ledger, and various intermediate results or final results of the transaction execution can be returned to the client node of the business entity for display.

[0163] For example, the client node 410 / 510 can subscribe to events of interest in the blockchain network 200, such as transactions occurring in a specific organization / channel in the blockchain network 200, and the consensus node 210 pushes the corresponding transaction notification to the client node 410 / 510, thereby triggering the corresponding business logic in the client node 410 / 510.

[0164] The following uses an example of multiple business entities accessing a blockchain network to manage instruction information and business processes that match the instruction information to illustrate an exemplary application of the blockchain network.

[0165] See also Figure 9 Multiple business entities involved in the management process, such as business entity 400, which can be a video information processing device, and business entity 500, which can be a display system with video information processing capabilities, register with authentication center 300 to obtain their respective digital certificates. These digital certificates include the business entity's public key and the digital signature of authentication center 300 on the business entity's public key and identity information. These certificates are attached to the transaction along with the business entity's digital signature and are then sent to the blockchain network, where the blockchain network retrieves the digital certificate and signature from the transaction to verify the authenticity of the message (i.e., whether it has been tampered with) and the identity of the business entity sending the message. The blockchain network then performs identity verification, such as whether the entity has the authority to initiate transactions. Any client running on an electronic device (such as a terminal or server) under the control of a business entity can request access to blockchain network 200 and become a client node.

[0166] The client node 410 of the business entity 400 is used to obtain the identification information of the video to be encoded and the encoding decision corresponding to the video to be encoded; based on the identification information of the video to be encoded, the encoding decision corresponding to the video to be encoded and the video information after encoding processing, generate a target block, and add the target block to the blockchain network 200.

[0167] The corresponding instruction information and the coding decision that matches the instruction information are sent to the blockchain network 200. Business logic can be pre-configured in the client node 410. When a corresponding coding decision is formed, the client node 410 automatically sends the instruction information to be processed and the coding decision that matches the instruction information to the blockchain network 200. Alternatively, a business person of the business entity 400 can log in to the client node 410, manually package the instruction information, the coding decision that matches the instruction information, and the corresponding conversion process information, and send them to the blockchain network 200. During the transmission, the client node 410 generates a transaction corresponding to the update operation based on the instruction information, the coding decision that matches the instruction information, and the corresponding conversion coding decision. The transaction specifies the smart contract to be invoked to implement the update operation and the parameters to be passed to the smart contract. The transaction also carries the client node 410's digital certificate and a signed digital signature (for example, obtained by encrypting the transaction digest using the private key in the client node 410's digital certificate). The transaction is then broadcast to the consensus nodes 210 in the blockchain network 200.

[0168] When consensus node 210 in blockchain network 200 receives a transaction, it verifies the digital certificate and digital signature carried in the transaction. If verification is successful, it then determines whether business entity 400 has transaction authority based on the identity of the business entity 400 carried in the transaction. Either of these verifications will result in transaction failure. After successful verification, node 210 signs its own digital signature (for example, by encrypting the transaction digest using node 210-1's private key) and continues to broadcast the transaction within blockchain network 200.

[0169] After receiving a successfully verified transaction, consensus node 210 in blockchain network 200 inserts the transaction into a new block and broadcasts it. When consensus node 210 in blockchain network 200 broadcasts a new block, it performs a consensus process on the new block. If consensus is successful, the new block is appended to the end of its stored blockchain, and the state database is updated based on the transaction results. The transactions in the new block are executed: For transactions that submit and update pending instruction information, a coding decision matching the instruction information, and corresponding process trigger information, a key-value pair consisting of the instruction information, the coding decision matching the instruction information, and the corresponding process trigger information is added to the state database.

[0170] The business personnel of the business entity 500 logs in to the client node 510, enters the instruction information and the coding decision query request that matches the instruction information, and the client node 510 generates a transaction corresponding to the update operation / query operation based on the instruction information and the coding decision query request that matches the instruction information. The transaction specifies the smart contract that needs to be called to implement the update operation / query operation and the parameters passed to the smart contract. The transaction also carries the digital certificate of the client node 510 and the signed digital signature (for example, the transaction summary is encrypted using the private key in the digital certificate of the client node 510), and broadcasts the transaction to the consensus node 210 in the blockchain network 200.

[0171] The consensus node 210 in the blockchain network 200 receives the transaction, verifies the transaction, fills the block and reaches consensus, then appends the filled new block to the end of the blockchain stored in itself, updates the status database according to the transaction result, and executes the transaction in the new block: for the transaction submitted to update the manual recognition result corresponding to a certain coding decision data information, the key-value pair corresponding to the coding decision data information in the status database is updated according to the manual recognition result; for the transaction submitted to query a certain coding decision data information, the instruction information and the key-value pair corresponding to the coding decision matching the instruction information are queried from the status database, and the transaction result is returned.

[0172] It is worth noting that in Figure 9The process of directly uploading instruction information, the coding decision that matches the instruction information, and the corresponding process triggering information to the blockchain is exemplified in FIG. However, in other embodiments, when the instruction information and the coding decision that matches the instruction information are large in data volume, the client node 410 may upload the instruction information and the hash of the coding decision that matches the instruction information, and the corresponding instruction information and the hash of the coding decision that matches the instruction information to the blockchain in pairs, and store the instruction information, the coding decision that matches the instruction information, and the corresponding process triggering information in a distributed file system or database. After the client node 510 obtains the instruction information, the coding decision that matches the instruction information, and the corresponding process triggering information from the distributed file system or database, it may verify them in conjunction with the corresponding hash in the blockchain network 200, thereby reducing the workload of the uploading operation.

[0173] As an example of blockchain, see Figure 10 , Figure 10 This is a schematic diagram of the structure of the blockchain in the blockchain network 200 provided by an embodiment of the present invention. The header of each block can include the hash values ​​of all transactions in the block as well as the hash values ​​of all transactions in the previous block. The records of newly generated transactions are filled into the block and, after consensus among the nodes in the blockchain network, are appended to the end of the blockchain to form a chain growth. The chain structure between blocks based on hash values ​​ensures that transactions in the block are tamper-proof and anti-forgery.

[0174] The following describes an exemplary functional architecture of the blockchain network provided by an embodiment of the present invention. Figure 11 , Figure 11 2 is a functional architecture diagram of a blockchain network 200 provided in an embodiment of the present invention, including an application layer 201, a consensus layer 202, a network layer 203, a data layer 204, and a resource layer 205, which are described below respectively.

[0175] The resource layer 205 encapsulates the computing resources, storage resources, and communication resources of each node 210 in the blockchain network 200.

[0176] The data layer 204 encapsulates various data structures that implement the ledger, including the blockchain implemented as files in the file system, the key-value state database, and the existence proof (such as the hash tree of transactions in the block).

[0177] The network layer 203 encapsulates the functions of point-to-point (P2P) network protocol, data transmission mechanism and data verification mechanism, access authentication mechanism and business subject identity management.

[0178] Among them, the P2P network protocol realizes the communication between the nodes 210 in the blockchain network 200, the data propagation mechanism ensures the propagation of transactions in the blockchain network 200, and the data verification mechanism is used to achieve the reliability of data transmission between nodes 210 based on cryptographic methods (such as digital certificates, digital signatures, public / private key pairs); the access authentication mechanism is used to authenticate the identity of the business entity joining the blockchain network 200 according to the actual business scenario, and grant the business entity the right to access the blockchain network 200 when the authentication is passed; the business entity identity management is used to store the identity of the business entity allowed to access the blockchain network 200, as well as the permissions (such as the type of transactions that can be initiated).

[0179] The consensus layer 202 encapsulates the mechanism by which nodes 210 in the blockchain network 200 reach consensus on blocks (i.e., the consensus mechanism), as well as transaction management and ledger management functions. Consensus mechanisms include consensus algorithms such as POS, POW, and DPOS, and support pluggable consensus algorithms.

[0180] Transaction management is used to verify the digital signature carried in the transaction received by the verification node 210, verify the identity information of the business subject, and determine whether it has the authority to conduct the transaction based on the identity information (read relevant information from the business subject identity management); for business subjects that have obtained authorization to access the blockchain network 200, they all have digital certificates issued by the certification center. The business subject uses the private key in its own digital certificate to sign the submitted transaction, thereby declaring its legal identity.

[0181] Ledger management is used to maintain the blockchain and state database. Consensus-reached blocks are appended to the end of the blockchain. Transactions within the consensus-reached blocks are executed. If the transaction includes an update, the key-value pairs in the state database are updated. If the transaction includes a query, the key-value pairs in the state database are queried and the query results are returned to the client node of the business entity. Multiple query operations on the state database are supported, including: querying blocks by block vector number (e.g., transaction hash value); querying blocks by block hash value; querying blocks by transaction vector number; querying transactions by transaction vector number; querying business entity account data by account number (vector number); and querying the blockchain in a channel by channel name.

[0182] The application layer 201 encapsulates various services that can be implemented by the blockchain network, including transaction traceability, evidence storage, and verification.

[0183] The structure of the multimedia information processing device of the embodiment of the present invention is described in detail below. The multimedia information processing device can be implemented in various forms, such as a dedicated terminal with video information processing function, such as a gateway, or a multimedia information processing device with video information processing function, such as the aforementioned Figure 1Server 400 in.

[0184] Combine Figure 2 The electronic device shown illustrates the multimedia information processing method provided by the embodiment of the present invention, see Figure 12 , Figure 12 This is an optional flowchart of the multimedia information processing method provided by the embodiment of the present invention. It can be understood that: Figure 12 The steps shown can be performed by various servers running multimedia information processing devices, such as dedicated terminals with multimedia information processing functions, multimedia information processing devices, or multimedia information processing device clusters. Figure 12 The steps shown are explained.

[0185] Step 1201: The multimedia information processing device separates target audio and target video from the multimedia information.

[0186] in, Figure 13 This is a data processing diagram for the video processing method according to an embodiment of the present invention. The frame output buffer of each frame allocates additional space to store extra_info. This information includes, but is not limited to, the macroblock type, segmentation method, bits size, MV and reference frame, QP information, and frame type and size data for each macroblock. After the encoder receives a frame of frame_data, it also obtains this extra_info information, which can be used directly to guide the estimation of the complexity of the current frame. When calculating inter-frame complexity, refer to the following steps:

[0187] Step 1301: Calculate the structural similarity between the two preceding and following GOPs.

[0188] Step 1302: Compare with the structural similarity difference threshold that matches the video processing environment to determine whether it is less than the structural similarity difference threshold. If so, execute step 1304; otherwise, execute step 1303.

[0189] Step 1303: Confirm that the GOP changes dramatically.

[0190] Step 1304: Confirm that it is a slowly changing GOP.

[0191] Step 1305: Calculate the texture difference value between Ref.CU and Curr.CU.

[0192] Step 1306: Compare the texture difference value with the threshold to determine whether it is less than the texture difference threshold. If so, execute step 1308; otherwise, execute step 1307.

[0193] Step 1307: Do not perform early skip mode.

[0194] Step 1308: Determine whether it is a dynamic gop, if so, execute step 1309, otherwise execute step 1310.

[0195] Step 1309: Execute dynamic early skip mode.

[0196] Step 1310: Execute static early skip mode.

[0197] After the encoding decision is determined, execution continues at step 1202 .

[0198] Step 1202: The multimedia information processing apparatus determines an encoding decision that matches the target video.

[0199] Step 1203: The multimedia information processing apparatus processes the target video according to the determined encoding decision to implement encoding of the target video.

[0200] Step 1204: The multimedia information processing device transmits the encoded video.

[0201] In some embodiments of the present invention, tests were conducted using the low-delay (LD) encoding mode on the HPM-9.1 platform. The test server configuration can be 16.04.1-Ubuntu, 40 Intel(R), Xeon(R), Gold 6148 CPU @2.4GHz operating environment. The test data used in this embodiment is shown in Table 1. Due to the limited number of fixed-scene videos in the AVS3-specified test sequence, some fixed-scene test sequences from other standards were included in the test. The names of each sequence are shown in Table 1. In the above tests, the fast algorithm for limiting the number of equalization times (EQTs) prohibited the use of this method for LCUs with fewer than 5 EQTs. The encoding information and disparity data used to preset the disparity threshold primarily came from frames in the video sequences SlidShow, Johnny, KristenAndSara, and Vidyo4. The CU size restriction algorithm disabled the EQT partitioning method for CUs with a length-to-width product greater than 2048, and the encoding QPs used were 27, 32, 38, and 45, respectively. The HPM9.1-EQT encoding time in Table 1 is when all CUs are directly prohibited from using EQT partitioning, the HPM9.1-BT encoding time in Table 1 is when all CUs are directly prohibited from using BT partitioning, and the HPM9.1-EQT+HPM9.1-BT encoding time in Table 2 is when all EQT and BT partitioning are directly prohibited. Total time saved S T , EQT time saving S EQT , BT time saving S BT The calculation method of is shown in formula (8) and formula (9).

[0202]

[0203]

[0204]

[0205] Among them, Torg represents the average time and Tfast represents the shortest time

[0206] Table 1 Test performance of EQT and BT on HPM-9.1

[0207]

[0208] Table 2 Test performance of EQT and BT on HPM-9.1

[0209]

[0210] From the test data, it can be seen that by using the EQT and BT limiting algorithms, an overall time savings of 11% and 24% can be achieved respectively, and using the EQT+BT limiting algorithm can achieve an overall time savings of 31%. The BD-PSNR brought by the EQT, BT, and EQT+BT limiting algorithms are -0.011, -0.049, and -0.058 respectively, and the BD-BR brought are 0.362%, 1.677%, and 1.974% respectively. The above test data demonstrates the effectiveness of the video processing method provided by this application.

[0211] The embodiments of the present invention have the following beneficial effects:

[0212] The embodiment of the present invention obtains a video to be encoded in a video processing environment and intercepts a video segment to be analyzed in the video to be encoded; determines the type of the encoding frame group in the video to be encoded based on the structural similarity difference between the current encoding frame group and the reference encoding frame group included in the video segment to be analyzed; when it is determined that the encoding frame group in the video to be encoded is of the first frame group type, determines the texture difference parameters of the current encoding unit and the reference encoding unit of the video segment to be analyzed; determines the encoding decision that matches the video to be encoded based on the texture difference parameters of the current encoding unit and the reference encoding unit; processes the video to be encoded based on the determined encoding decision to achieve encoding of the video to be encoded, thereby enabling the encoding decision that matches the video to be encoded to be quickly and accurately determined based on the type of the encoding frame group and the texture difference of the encoding unit, and also more quickly determining the encoding method for the video, reducing the waiting time for selecting the encoding decision, and improving the speed of the video encoding process. At the same time, the computational workload of video information processing is saved, the computational workload of the device is reduced, and the user experience is improved.

[0213] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A video information processing method, characterized in that: The method comprises: Obtaining a video to be encoded, and intercepting a video segment to be analyzed from the video to be encoded; determining the type of the coding frame group in the to-be-coded video according to a structural similarity difference between a current coding frame group and a reference coding frame group included in the to-be-coded video segment; Determining a texture difference parameter between a current coding unit and a reference coding unit of the video segment to be analyzed; When a texture difference parameter between the current coding unit and the reference coding unit is greater than or equal to a texture difference parameter threshold, determining that an encoding decision matching the video to be encoded is not to execute an early skip mode; When a texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, and the coding frame group is of a first frame group type, determining that a coding decision matching the to-be-coded video is a static early skip mode; When a texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, and the coding frame group is of a second frame group type, determining that a coding decision matching the to-be-coded video is a dynamic early skip mode; The video to be encoded is processed according to the determined encoding decision to achieve encoding of the video to be encoded.

2. The method according to claim 1, characterized in that The determining the type of the coding frame group in the to-be-coded video according to the structural similarity difference between the current coding frame group and the reference coding frame group included in the to-be-coded video segment includes: Obtaining an intra-frame prediction frame of the current coding frame group and an intra-frame prediction frame of the reference coding frame group; determining corresponding structural similarity differences according to the intra-frame prediction frame of the current coding frame group and the intra-frame prediction frame of the reference coding frame group; Determining a structural similarity difference threshold that matches the video processing environment; When the structural similarity difference is less than the structural similarity difference threshold, determining that the encoding frame group in the to-be-encoded video is of the first frame group type; When the structural similarity difference is greater than or equal to the difference threshold, it is determined that the encoding frame group in the to-be-encoded video is of the second frame group type.

3. The method according to claim 2, characterized in that The determining, according to the intra-frame prediction frame of the current coding frame group and the intra-frame prediction frame of the reference coding frame group, corresponding structural similarity difference comprises: Obtaining a brightness average, brightness value variance, and brightness value covariance corresponding to the intra-frame prediction frame of the current coding frame group; Obtaining a brightness average, a brightness value variance, and a brightness value covariance corresponding to an intra-frame prediction frame of the reference coding frame group; Based on the brightness average, brightness variance, and brightness covariance corresponding to the intra-frame prediction frames of the current coding frame group and the brightness average, brightness variance, and brightness covariance corresponding to the intra-frame prediction frames of the reference coding frame group, the structural similarity difference between the current coding frame group and the reference coding frame group is determined.

4. The method according to claim 1, wherein The determining of the texture difference parameter between the current coding unit and the reference coding unit of the video segment to be analyzed includes: Determining a corresponding current coding unit and a reference coding unit in the video segment to be analyzed; Determining a histogram difference parameter and a gradient difference parameter between the current coding unit and a reference coding unit; Determine a texture difference parameter between the current coding unit and the reference coding unit of the video segment to be analyzed based on the histogram difference parameter and the gradient difference parameter between the current coding unit and the reference coding unit.

5. The method according to claim 1, wherein The method further comprises: Determining the number of extended quadtree structures and the number of binary tree structures in the current coding unit; Marking a type of the current coding unit based on the number of extended quadtree structures and the number of binary tree structures in the current coding unit; The texture difference parameter threshold is adjusted according to the type of the current coding unit.

6. The method according to claim 5, characterized in that The method further comprises: When the histogram difference and the structural difference of the quadtree partitioning that matches the video clip to be analyzed are both smaller than the quadtree histogram difference threshold and the quadtree structural difference threshold, the quadtree partitioning is determined as a disabled partitioning mode; When the histogram difference and the structural difference of the binary tree partition matching the video segment to be analyzed are both smaller than the preset binary tree histogram difference threshold and the binary tree structural difference threshold, the binary tree partition is determined as a disabled partitioning mode; When the histogram difference and structural difference of the extended quadtree partition matching the video clip to be analyzed are both smaller than the preset extended quadtree histogram difference threshold and extended quadtree structural difference threshold, the extended quadtree partition is determined as a disabled partitioning mode.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Obtaining identification information of the video to be encoded, encoding decision corresponding to the video to be encoded, and video information that has been encoded; Based on the identification information of the video to be encoded, the encoding decision corresponding to the video to be encoded, and the video information after encoding processing, a target block is generated, and the target block is added to the blockchain network.

8. A multimedia information processing method, characterized in that: The method comprises: Separating target audio and target video from multimedia information; Determining a coding decision that matches the target video, wherein the coding decision is obtained according to the method of any one of claims 1 to 7; Processing the target video according to the determined encoding decision to achieve encoding of the target video; The encoded target video and the target audio are encapsulated into new multimedia information.

9. A video information processing device, characterized in that: The device comprises: An information transmission module is used to obtain a video to be encoded and intercept a video segment to be analyzed in the video to be encoded; an information processing module, configured to determine the type of the coding frame group in the to-be-coded video according to a structural similarity difference between a current coding frame group and a reference coding frame group included in the to-be-coded video segment; The information processing module is configured to determine a texture difference parameter between a current coding unit and a reference coding unit of the video segment to be analyzed; The information processing module is configured to, when a texture difference parameter between the current coding unit and the reference coding unit is greater than or equal to a texture difference parameter threshold, determine that a coding decision matching the video to be coded is not to execute the early skip mode; when the texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, and the coding frame group is of a first frame group type, determine that a coding decision matching the video to be coded is a static early skip mode; and when the texture difference parameter between the current coding unit and the reference coding unit is less than the texture difference parameter threshold, and the coding frame group is of a second frame group type, determine that a coding decision matching the video to be coded is a dynamic early skip mode; The information processing module is used to process the video to be encoded according to the determined encoding decision, so as to realize encoding of the video to be encoded.

10. A multimedia information processing device, characterized in that: The device comprises: An information separation device, used for separating target audio and target video from multimedia information; A video processing device, configured to determine a coding decision matching the target video, wherein the coding decision is obtained according to the method of any one of claims 1 to 7; The video processing device is used to process the target video according to the determined encoding decision to achieve encoding of the target video; The video processing device is used to encapsulate the target video and the target audio that have been coded into new multimedia information.

11. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable instructions; The processor is configured to implement the video information processing method according to any one of claims 1 to 7, or the multimedia information processing method according to claim 8, when running the executable instructions stored in the memory.

12. A computer-readable storage medium storing executable instructions, characterized in that: When the executable instructions are executed by the processor, the video information processing method according to any one of claims 1 to 7 is implemented, or the multimedia information processing method according to claim 8 is implemented.

Citation Information

Patent Citations

  • Video information processing method and device and multimedia information processing method and device

    CN111294591A

  • Coding unit division method and device, encoder and storage medium

    CN111669602A

  • Video encoding method and device, equipment and storage medium

    CN112449182A