Video coding and decoding method and device for specifying hidden layer characteristic attribute and storage equipment
By specifying the syntax elements of hidden layer feature attributes in the code stream, the problem that the intelligent task network cannot use decoded features correctly is solved, and more accurate feature processing and task completion is achieved.
Patent Information
- Application Number
- CN202311530486.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-05-16
AI Technical Summary
In machine intelligence tasks, it is difficult for the existing technology to effectively specify and use the attributes of decoded features, resulting in the intelligent task network being unable to correctly use features to complete subsequent task inference.
The subsequent processing methods for decoding features are guided by adding or parsing syntax elements that specify hidden layer feature attributes in the code stream, including intelligent task network model, feature input location and feature extractor type.
It clearly specifies the ability and purpose of decoding features, provides accurate processing guidance for the receiver, improves the accuracy of specific tasks, and supports targeted task network re-optimization.
Smart Images

Figure CN120017851A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video encoding and decoding, and in particular to a video encoding and decoding method, apparatus and storage device for specifying hidden layer feature attributes. Background Art
[0002] In machine intelligence scenarios, some existing image and video encoding and decoding methods directly encode and decode images and videos, and then use the reconstructed images and videos to further complete machine intelligence tasks. However, there are also some methods that encode and decode feature data in the reasoning process of intelligent analysis tasks. In this framework, it is first necessary to use the feature extraction network of the task network to extract features from the image and video, and then use the image and video codec to encode and decode the features. Finally, the decoded features are used as the input of the machine intelligence task for reasoning to complete the machine intelligence task. Since the machine intelligence task networks are not the same, the size and content of the features to be encoded are also inconsistent. If the location where the feature data should be fed into the machine intelligence task network is not indicated in the bitstream, the machine intelligence task network may not be able to use the feature to complete subsequent task reasoning. Summary of the invention
[0003] The embodiment of the present invention provides a video encoding and decoding method and device for specifying hidden layer feature attributes, including the following aspects:
[0004] A first aspect of the present invention provides a video decoding method for specifying hidden layer feature attributes, comprising:
[0005] Parse the syntax element of the specified hidden layer feature attribute from the bitstream, where the hidden layer feature attribute includes any one or a combination of the following information:
[0006] A type or class of intelligent task network models corresponding to the decoding features;
[0007] The location of the decoded feature input into the intelligent task network model;
[0008] The type of feature extractor on the sender side.
[0009] Furthermore, the decoding method also includes:
[0010] The intelligent task network is called, and based on the hidden layer feature attributes, the hidden layer features are used to complete the reasoning or analysis of the intelligent task.
[0011] A second aspect of the present invention provides a video encoding method for specifying hidden layer feature attributes, comprising:
[0012] Add a syntax element specifying a hidden layer feature attribute in the bitstream, wherein the hidden layer feature attribute includes any one or a combination of the following information:
[0013] A type or class of intelligent task network models corresponding to the decoding features;
[0014] The location of the decoded feature input into the intelligent task network model;
[0015] The type of feature extractor on the sender side.
[0016] A third aspect of the present invention provides a video decoding device for specifying hidden layer feature attributes, comprising:
[0017] processor;
[0018] A memory for storing a code stream; and
[0019] One or more programs are used to implement the video decoding method with specified hidden layer feature attributes as described in the first aspect above.
[0020] A fourth aspect of the present invention provides a video encoding device for specifying hidden layer feature attributes, comprising:
[0021] processor;
[0022] A memory for storing a code stream; and
[0023] One or more programs are used to implement the video encoding method with specified hidden layer feature attributes as described in the second aspect above.
[0024] A fifth aspect of the present invention provides a storage device, which includes a code stream, which is decoded using the video decoding method for specifying hidden layer feature attributes as described in the first aspect and is used in an intelligent task network.
[0025] A sixth aspect of the present invention provides a storage device, which includes a code stream, which is encoded using the video encoding method for specifying hidden layer feature attributes as described in the second aspect above and is used in an intelligent task network.
[0026] The beneficial effects of this application are:
[0027] The syntax elements are used in the code stream to specify the attributes of the decoding features, which in turn guide the subsequent processing methods of the features. This attribute indicates the type of intelligent task that can be completed by the semantic information contained in the decoding features. It can also indicate that the decoding features can be directly used as input data of a certain layer in the middle of the intelligent task network and complete a specific task. It can also indicate that the decoding features can support a set of intelligent tasks and guide the receiving end to call the corresponding adapter to process the decoded features to use them to complete specific tasks. It can also indicate the type of feature extractor used by the sending end and guide the receiving end to use the feature extractor to redesign and train the backend network. The advantage of this method is that it can clearly define the capabilities and uses of the decoding features, providing accurate guidance for the receiving end to perform subsequent processing; at the same time, it can also provide constraints on the scope of semantic information for the receiving end to perform targeted task network re-optimization, which helps the receiving end use the decoding features to improve the accuracy of specific tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 Schematic diagram of the overall framework of the video decoding device for specifying hidden layer feature attributes of the present invention.
[0029] Figure 2 Schematic diagram of the network structure of the RCNN FPN intelligent task network used in the embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the specific technical scheme of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0032] An embodiment of the present invention shows a video decoding method for specifying hidden layer feature attributes, including any one or a combination of the following information:
[0033] A type or class of intelligent task network models corresponding to the decoding features;
[0034] The location of the decoded feature input into the intelligent task network model;
[0035] The type of feature extractor on the sender side.
[0036] The following will describe in detail the details of using latent features to complete reasoning or analysis of intelligent tasks based on the attributes of latent features in conjunction with specific examples.
[0037] In one embodiment, the receiving end parses out the syntax element specifying the attribute of the hidden layer feature, and the syntax element specifies a smart task network model corresponding to the decoding feature, and the decoding feature can be input into the smart task network model specified by the attribute of the hidden layer feature. For example, if there are multiple smart task network models, the syntax element specifying the hidden layer feature can be designed as follows:
[0038] Syntactic elements related to hidden feature attributes Descriptors task_module_idx u(n)
[0039] The grammatical element task_module_idx represents the serial number of the intelligent task network model that the decoding feature can input, and u(n) means that the grammatical element can be represented by an n-bit unsigned integer. For example, the meanings of different values of task_module_idx can be as follows: Figure 2 :
[0040]
[0041]
[0042] In another embodiment, the receiving end parses out the syntax element specifying the hidden layer feature attribute, and the syntax element specifies a class of intelligent task network models corresponding to the decoding feature. The intelligent class task network model may include several intelligent task network models of the same type of tasks but different network structures, and the decoding feature may be input into this class of different intelligent task network models. For example, if there are multiple types of tasks, the syntax element specifying the hidden layer feature may be designed as follows:
[0043] Syntactic elements related to hidden feature attributes Descriptors task_idx u(n)
[0044] The syntax element task_idx represents a type of task number that can be input for this feature, and u(n) means that this syntax element can be represented by an n-bit unsigned integer. For example, the meanings of different values of task_idx can be as follows:
[0045] task_idx Type of tasks represented 0 Object Detection 1 Instance Segmentation 2 Pose Estimation 3 Face Recognition … …
[0046] In another embodiment, the receiving end parses out a syntax element specifying a hidden layer feature attribute, and the syntax element specifies that the decoding feature is input to a certain intermediate layer in the intelligent task network model, and the decoding feature can be input to a certain layer in the intelligent task network indicated by the syntax element. For example, if the intelligent task network has multiple intermediate layers that can receive the decoding feature, the syntax element specifying the hidden layer feature can be designed as follows:
[0047] Syntactic elements related to hidden feature attributes Descriptors feature_input_position_idx u(n)
[0048] The syntax element feature_input_position_idx represents the sequence number of the decoded feature input to the different intermediate layers in the intelligent task network, and u(n) means that the syntax element can be represented by an n-bit unsigned integer. For example, the meaning of different values of feature_input_position_idx can be as follows, where each input intermediate layer refers to Figure 2 :
[0049] feature_input_position_id The intermediate layer position of the input task network represented by 0 Faster RCNN X101 FPN network res2 node 1 Faster RCNN X101 FPN network res3 node 2 Faster RCNN X101 FPN network res4 node 3 Faster RCNN X101 FPN network res5 node … …
[0050] In another embodiment, the receiving end parses out the syntax element specifying the hidden layer feature attribute, and the syntax element specifies that the decoding feature is simultaneously input into certain intermediate layers in the intelligent task network. The decoding feature can be input into certain layers in the intelligent task network indicated by the syntax element. Based on the previous input into a certain intermediate layer, if the decoding feature needs to be input into multiple intermediate layers, the sequence number indicating the decoding feature input into multiple intermediate layers can be added to the feature_input_position_idx syntax element, as shown below, where multiple intermediate layers are input. Figure 2 :
[0051]
[0052] In another embodiment, the receiving end parses out the syntax element specifying the attribute of the hidden layer feature, which specifies a set of intelligent tasks supported by the decoding feature, and instructs the receiving end to call the corresponding adapter to process the decoding feature to use it to complete the reasoning of the specific task. For example, if there are multiple adapters on the receiving end, the syntax element specifying the hidden layer feature can be designed as follows:
[0053] Syntactic elements related to hidden layer feature attributes Descriptors feature_adapter_idx u(n)
[0054] The syntax element feature_adapter_idx represents the adapter number that the receiving end should use, and u(n) indicates that the syntax element can be represented by an n-bit unsigned integer. For example, the meaning of different values of feature_adapter_id can be as follows:
[0055] feature_adapter_idx Adapter represented 0 Adapter 0 1 Adapter 1 2 Adapter 2 … …
[0056] In another embodiment, the receiving end parses out a syntax element that specifies the hidden layer feature attribute. The syntax element specifies the type of feature extractor used by the sending end. The receiving end can use the syntax element to redesign and retrain the backend network to improve task performance. For example, the syntax element can be designed as follows:
[0057] Syntactic elements related to hidden layer feature attributes Descriptors feature_extractor_idx u(n)
[0058] The syntax element feature_extractor_idx represents the feature extractor number corresponding to the feature to be encoded, and u(n) means that the syntax element can be represented by an n-bit unsigned integer. For example, the meanings of different values of feature_extractor_idx can be as follows, where each feature extractor refers to Figure 2 :
[0059]
[0060] The receiving end can obtain the corresponding feature extractor at the back end according to the syntax element of the specified hidden layer feature attribute, for example, the extraction network before the res2 node of the Faster RCNN X101 FPN network can be obtained, and the feature extractor can be used at the receiving end to redesign and train the back-end task network.
[0061] Another embodiment of the present invention shows a video encoding method for specifying hidden layer feature attributes, including:
[0062] A syntax element specifying hidden features is added to the bitstream, and the syntax element specifying hidden features specifies the attributes of the decoded hidden features; further, the attributes of the hidden features indicate the type of intelligent tasks that can be completed by the semantic information contained in the decoded features of the receiving end, and can also indicate that the decoded features of the receiving end can be directly used as input data of a certain layer in the middle of the intelligent task network and complete a specific task, and can also indicate that the decoded features of the receiving end can support a group of intelligent tasks, and guide the receiving end to call the corresponding adapter to process the decoded features to use them to complete specific tasks, and can also indicate the type of feature extractor used by the sending end, and guide the receiving end to use the feature extractor to redesign and train the back-end network. The advantage of this method is that the information that can clearly define the capabilities and uses of the decoding features is written into the bitstream at the sending end, providing accurate guidance for the receiving end to perform subsequent processing; at the same time, it can also provide constraints on the range of semantic information for the receiving end to perform targeted task network re-optimization, which helps the receiving end to use the decoding features to improve the accuracy of specific tasks.
[0063] In one embodiment, the transmitting end adds a syntax element specifying the attribute of the hidden layer feature and writes it into the bitstream. The syntax element specifies an intelligent task network model corresponding to the decoding feature, and the decoding feature can be input into the intelligent task network model specified by the attribute of the hidden layer feature. The specific syntax element design is similar to the corresponding decoding method embodiment.
[0064] In another embodiment, the transmitting end adds a syntax element specifying the hidden layer feature attribute and writes it into the bitstream. The syntax element specifies a class of intelligent task network models corresponding to the decoding feature. The intelligent class task network model may include several intelligent task network models with the same task but different network structures. The decoding feature may be input into this class of different intelligent task network models. The specific syntax element design is similar to the corresponding decoding method embodiment.
[0065] In another embodiment, the transmitting end adds a syntax element specifying the hidden layer feature attribute and writes it into the bitstream. The syntax element specifies that the decoding feature is input to a certain middle layer in the intelligent task network model. The decoding feature can be input to a certain layer in the intelligent task network indicated by the syntax element. The specific syntax element design is similar to the corresponding decoding method embodiment.
[0066] In another embodiment, the transmitting end adds a syntax element specifying the hidden layer feature attribute and writes it into the bitstream. The syntax element specifies that the decoding feature is simultaneously input into certain intermediate layers in the intelligent task network. The decoding feature can be input into certain layers in the intelligent task network indicated by the syntax element. The specific syntax element design is similar to the corresponding decoding method embodiment.
[0067] In another embodiment, the transmitting end adds a syntax element specifying the hidden layer feature attribute and writes it into the bitstream, the syntax element specifies a set of intelligent tasks supported by the decoding feature, and instructs the receiving end to call the corresponding adapter to process the decoding feature to use it to complete the reasoning of the specific task. The specific syntax element design is similar to the corresponding decoding method embodiment.
[0068] In another embodiment, the transmitting end adds a syntax element specifying the hidden layer feature attribute and writes it into the bitstream. The syntax element specifies the type of feature extractor used by the transmitting end. The receiving end can use the syntax element to redesign and retrain the backend network to improve the task performance. The specific syntax element design is similar to the corresponding decoding method embodiment.
[0069] Another embodiment of the present invention shows a video encoding device for specifying hidden layer feature attributes, such as Figure 1As shown, it includes a parser 101, a decoder 102 and a task network 103. The parser 101 parses the syntax elements of the specified hidden layer features from the code stream to be decoded to obtain the attributes of the hidden layer features, and other syntax elements are further input into the decoder 102 to realize the decoding of the hidden layer features to obtain the decoded features. The syntax elements that specify the hidden layer features specify the attributes of the decoded hidden layer features, which indicate the type of intelligent tasks that can be completed by the semantic information contained in the encoded features, and can also indicate that the encoded features can be directly used as input data of a certain layer in the middle of the intelligent task network and complete a specific task, and can also indicate that the encoded features can support a group of intelligent tasks, and guide the receiving end to call the corresponding adapter to process the decoded features to use them to complete specific tasks. According to the attribute information, the receiving end can call the relevant intelligent task network 103, and specify the input position of the decoded features in the task network 103 according to the input position in the syntax elements that specify the attributes of the hidden layer features.
[0070] Another embodiment of the present invention shows a video encoding device for specifying hidden layer feature attributes, including:
[0071] An encoder is used to add a syntax element specifying a hidden feature in a bitstream, wherein the syntax element specifying the hidden feature specifies the attribute of the decoded hidden feature, and the attribute of the hidden feature specifies the feature extraction network or the type of feature extraction network used to extract the hidden feature to be encoded.
[0072] In one embodiment, the attributes of the hidden layer features are also used by the receiving end to redesign and retrain the backend intelligent task network.
[0073] Another embodiment of the present invention further shows a decoding device, which includes a processor and a memory, and is used to execute the decoding method disclosed in the present invention.
[0074] Another type of embodiment of the present invention further shows a coding device, which includes a processor and a memory, and is used to execute the coding method disclosed in the present invention.
[0075] The above embodiments are only used to help understand the method and core idea of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A video decoding method for specifying hidden layer feature attributes, characterized in that: include: Parse the syntax element of the specified hidden layer feature attribute from the bitstream, where the hidden layer feature attribute includes any one or a combination of the following information: A type or class of intelligent task network models corresponding to the decoding features; The location of the decoded feature input into the intelligent task network model; The type of feature extractor on the sender side.
2. The video decoding method for specifying hidden layer feature attributes according to claim 1, characterized in that: Also includes: The intelligent task network is called, and based on the hidden layer feature attributes, the hidden layer features are used to complete the reasoning or analysis of the intelligent task.
3. A video encoding method for specifying hidden layer feature attributes, characterized in that: include: Add a syntax element specifying a hidden layer feature attribute in the bitstream, wherein the hidden layer feature attribute includes any one or a combination of the following information: A type or class of intelligent task network models corresponding to the decoding features; The location of the decoded feature input into the intelligent task network model; The type of feature extractor on the sender side.
4. A video decoding device for specifying hidden layer feature attributes, characterized in that: include: processor; A memory for storing a code stream; as well as One or more programs are used to implement the video decoding method of specifying hidden layer feature attributes as described in claim 1 or 2.
5. A video encoding device for specifying hidden layer feature attributes, characterized in that: include: processor; A memory for storing a code stream; as well as One or more programs are used to implement the video encoding method for specifying hidden layer feature attributes as described in claim 3.
6. A storage device, characterized in that: The storage device contains a code stream, which is decoded using the video decoding method for specifying hidden layer feature attributes as described in claim 1 or 2 and is used in an intelligent task network.
7. A storage device, characterized in that: The storage device contains a code stream, which is encoded using the video encoding method for specifying hidden layer feature attributes as claimed in claim 3 and is used in an intelligent task network.