Decoding device, encoding device, decoding method, and encoding method

By including modal information of image categories in the bitstream, the problem of insufficient information transmission between the encoding and decoding devices is solved, thereby achieving optimization of task processing and improvement of accuracy.

CN121909650APending Publication Date: 2026-04-21PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2024-09-19
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In the prior art, the encoding device fails to effectively convey the modal information of the image to the decoding device, which causes the decoding device to be unable to select the best task processing and reduces the execution accuracy of task processing.

Method used

The bitstream contains modal information representing the image category and is transmitted to the decoding device through the encoding device. The decoding device performs task processing based on this modal information.

Benefits of technology

It improves the execution accuracy of task processing performed by the decoding device, ensuring the optimization of task processing, including the accuracy of human vision and machine tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121909650A_ABST
    Figure CN121909650A_ABST
Patent Text Reader

Abstract

A decoding device is provided with: a circuit that acquires, from a bitstream, an image and a parameter associated with the image, the parameter including modal information indicating an image category of the image;
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a decoding device, an encoding device, a decoding method, and an encoding method. Background Technology

[0002] Patent Document 1 discloses an image processing system according to the background art. This image processing system includes an encoder (encoding device) and a decoder (decoding device). In the encoding device, various modalities of input images are input. The encoding device extracts features from the input images and inputs the extracted features to the decoding device. The decoding device performs an image parsing task based on the input features, thereby outputting a segmentation map.

[0003] However, in the background art, no research has been conducted on the transmission of modal information of an image from the encoding device to the decoding device.

[0004] Prior art literature

[0005] Patent documents

[0006] Patent Document 1: U.S. Patent Application Publication No. 2024 / 0046453 Summary of the Invention

[0007] The purpose of this disclosure is to provide a decoding device, encoding device, decoding method, and encoding method that can transmit modal information of an image from an encoding device to a decoding device, thereby improving the execution accuracy of task processing performed by the decoding device.

[0008] One aspect of the present disclosure relates to a decoding apparatus comprising circuitry and a memory connected to said circuitry, said circuitry acquiring an image from a bitstream and parameters associated with said image, said parameters including modal information representing the image category. Attached Figure Description

[0009] Figure 1 This is a simplified diagram illustrating the structure of an image processing system according to an embodiment of the present disclosure.

[0010] Figure 2 This diagram is a simplified representation of the circuitry of the encoding device.

[0011] Figure 3 This is a flowchart illustrating the processing performed by the circuitry of the encoding device.

[0012] Figure 4 It is a diagram that simplifies the structure of a bitstream.

[0013] Figure 5 This is a diagram illustrating the first example of the syntax related to setting modal information.

[0014] Figure 6 This is a diagram illustrating the first example of the syntax related to setting modal information.

[0015] Figure 7 This is the second example of a diagram illustrating the syntax related to setting modal information.

[0016] Figure 8 This is the second example of a diagram illustrating the syntax related to the setting of modal information.

[0017] Figure 9 This diagram is a simplified representation of the circuitry of a decoding device.

[0018] Figure 10 This is a flowchart illustrating the processing performed by the circuitry of the decoding device.

[0019] Figure 11 This diagram is a simplified representation of the circuitry of a decoding device.

[0020] Figure 12 This is a flowchart illustrating the processing performed by the circuitry of the decoding device.

[0021] Figure 13 This diagram is a simplified representation of the circuitry of the encoding device.

[0022] Figure 14 This is a simplified diagram illustrating the first example of the circuit structure of a decoding device.

[0023] Figure 15 This is a simplified diagram illustrating the second example of the circuit structure of a decoding device.

[0024] Figure 16 This is a diagram illustrating the first example of a multi-layered structure.

[0025] Figure 17 This is a diagram illustrating an example of the syntax related to setting modal information.

[0026] Figure 18 This is a diagram illustrating an example of the syntax related to setting modal information.

[0027] Figure 19 This is a diagram illustrating an example of the syntax related to setting modal information.

[0028] Figure 20 This is a diagram showing the second example of a multi-layered structure.

[0029] Figure 21 This diagram is a simplified representation of the circuitry of a decoding device.

[0030] Figure 22It is a diagram that simplifies the structure of a bitstream.

[0031] Figure 23 This is a diagram illustrating an example of the syntax related to setting modal information.

[0032] Figure 24 This is a diagram illustrating an example of the syntax related to setting modal information.

[0033] Figure 25 This is a diagram illustrating an example of image composition based on multiple components.

[0034] Figure 26 This is a diagram illustrating an example of the syntax related to setting modal information.

[0035] Figure 27 This is a diagram illustrating an example of the syntax related to setting modal information.

[0036] Figure 28 This diagram is a simplified representation of the circuitry of a decoding device.

[0037] Figure 29 This is a block diagram illustrating an example of the functional structure of the encoding section.

[0038] Figure 30 This is a diagram illustrating an example of the hierarchical structure of data in a stream.

[0039] Figure 31 This is a block diagram illustrating an example of the functional configuration of the decoding unit. Detailed Implementation

[0040] (The understanding that forms the basis of this disclosure)

[0041] The image processing system described in the background technology includes an encoding device and a decoding device. In the encoding device, input images of various modalities are taken as input. The encoding device extracts features from the input image and inputs the extracted features to the decoding device. The decoding device performs an image parsing task based on the input features, thereby outputting a segmentation map.

[0042] The tasks performed by the decoding device include both human vision and machine tasks. Human vision involves visual recognition or audiovisual processing based on dynamic images of operators or users. Machine tasks include a wide variety of tasks such as object detection, object tracking, object segmentation, action recognition, and pose estimation using AI models.

[0043] In the background art, modal information is not transmitted from the encoding device to the decoding device. Therefore, sometimes the optimal task processing or the optimal AI model corresponding to the image category of the input image is not selected, resulting in reduced execution accuracy of the task processing.

[0044] In order to solve this problem, the inventors came to the following understanding and thus conceived of this disclosure: by including modal information representing the image category in a bitstream and sending it from the encoding device to the decoding device, the decoding device performs task processing based on the modal information, thereby solving the above-mentioned problem.

[0045] The various methods disclosed herein will now be explained.

[0046] The decoding apparatus according to the first aspect of this disclosure includes a circuit and a memory connected to the circuit, the circuit acquiring an image from a bitstream and parameters associated with the image, the parameters including modal information representing the image category.

[0047] According to the first method, modal information representing the image category can be included in the bitstream and transmitted from the encoding device to the decoding device, thereby improving the execution accuracy of the task processing performed by the decoding device.

[0048] The decoding device involved in the second aspect of this disclosure, in the first aspect, may be at least one of the following image categories: visible light image, thermal image, infrared image, LiDAR image, RADAR image, or AI-generated image.

[0049] According to the second method, the decoding device can perform optimal task processing corresponding to visible light images, thermal images, infrared images, LiDAR images, RADAR images, or AI-generated images.

[0050] The decoding device involved in the third aspect of this disclosure, in the first or second aspect, may be such that the circuit performs task processing based on the image and switches the task processing based on the modal information.

[0051] According to the third method, the decoding device can perform optimal task processing corresponding to the image category, thereby improving the execution accuracy of the task processing performed by the decoding device.

[0052] The decoding device involved in the fourth aspect of this disclosure, in the third aspect, may be such that the task processing includes machine tasks.

[0053] According to the fourth method, the decoding device is able to perform the optimal machine task corresponding to the image category.

[0054] The decoding device involved in the fifth aspect of this disclosure, in the third or fourth aspect, may be such that the task processing includes human vision.

[0055] According to the fifth method, the decoding device is able to perform optimal human vision corresponding to the image category.

[0056] The decoding device involved in the sixth aspect of this disclosure, in any one of the first to fifth aspects, may be a circuit that performs task processing based on the image and switches the AI ​​model or image processing used in the task processing based on the modal information.

[0057] According to the sixth method, the decoding device can use the best AI model or the best image processing corresponding to the image category to perform task processing, which can improve the execution accuracy of the task processing performed by the decoding device.

[0058] In any of the first to sixth embodiments, the decoding device involved in the seventh aspect of this disclosure may include, in which the parameters further include transformation information for transforming the pixel values ​​of the image into sensor data values ​​or transforming the image into a sensed image.

[0059] According to the seventh method, the decoding device can transform the pixel values ​​of an image into sensor data values ​​or the image into a sensed image based on the transformation information.

[0060] In the seventh embodiment, the decoding apparatus involved in the eighth aspect of this disclosure may be a circuit that transforms the pixel values ​​of the image into sensor data values ​​or the image into the sensed image based on the transformation information, and performs task processing based on the sensor data values ​​or the sensed image.

[0061] According to the eighth method, the decoding device can perform task processing based on sensor data values ​​or sensed images by performing transformation processing before task processing.

[0062] The decoding apparatus involved in the ninth aspect of this disclosure, in the seventh aspect, may be a circuit that performs task processing based on the image and the transformation information.

[0063] According to the ninth method, the decoding device performs transformation processing during task processing, enabling it to perform task processing based on sensor data values ​​or sensed images.

[0064] In any of the 1st to 9th embodiments of this disclosure, the decoding apparatus may, during the acquisition of the parameters, involve the circuit acquiring the parameters from a given header region of the bitstream, wherein the given header region includes VUI or SEI.

[0065] According to method 10, the decoding device can easily obtain parameters from a given header region of the bitstream.

[0066] The decoding apparatus according to the 11th aspect of this disclosure, in any one of the 1st to 10th aspects, may be a multi-layer structure in which the bitstream has multiple image layers, and in the acquisition of the image, the circuit acquires multiple images of different image categories from the multiple image layers.

[0067] According to the 11th method, multiple images of different image categories can be sent from the encoding device to the decoding device, and the decoding device can appropriately obtain the multiple images from the bitstream.

[0068] In the 11th embodiment, the decoding apparatus according to the 12th embodiment of this disclosure may, during the acquisition of the parameters, involve the circuit acquiring multiple parameters associated with the multiple images from the header region of a specific image layer among the multiple image layers.

[0069] According to the 12th method, the decoding device is able to obtain, in summary, multiple parameters associated with multiple images of multiple image layers from the header region of a specific image layer.

[0070] In the 11th embodiment, the decoding apparatus according to the 13th embodiment of this disclosure may, during the acquisition of the parameters, involve the circuit acquiring multiple parameters associated with the multiple images from multiple header regions of the multiple image layers.

[0071] According to the 13th method, the decoding device is able to individually obtain each parameter associated with the image of each image layer from each header region of each image layer.

[0072] In any of the 1st to 10th embodiments, the decoding apparatus of the 14th embodiment of this disclosure may be such that the image is divided into multiple sub-images, and in the acquisition of the image, the circuit acquires the multiple sub-images of different image categories from the bitstream.

[0073] According to the 14th method, multiple sub-images of different image categories can be sent from the encoding device to the decoding device, and the decoding device can appropriately obtain the multiple sub-images from the bitstream.

[0074] In the 14th embodiment, the decoding apparatus involved in the 15th embodiment of this disclosure may, during the acquisition of the parameters, involve the circuit acquiring multiple parameters associated with the multiple sub-images from the header region of the image.

[0075] According to method 15, the decoding device is able to obtain, in summary, multiple parameters associated with multiple sub-images from the header region of the image.

[0076] In any of the 1st to 10th embodiments of this disclosure, the decoding apparatus may be such that the image is composed of multiple components, and in acquiring the image, the circuit acquires the multiple components of different image categories from the bitstream.

[0077] According to the 16th method, multiple components of different image categories can be sent from the encoding device to the decoding device, and the decoding device can appropriately obtain the multiple components from the bitstream.

[0078] In the 16th embodiment, the decoding apparatus according to the 17th embodiment of this disclosure may, during the acquisition of the parameters, involve the circuit acquiring multiple parameters associated with the multiple components from the header region of the image.

[0079] According to method 17, the decoding device is able to obtain multiple parameters associated with multiple components from the header region of the image.

[0080] In the 16th embodiment, the decoding apparatus according to the 18th embodiment of this disclosure may be configured to assign the modal information to one of the plurality of components, wherein, in the acquisition of the image, the circuit acquires the one component from the bitstream.

[0081] According to method 18, when only one component is the required image category, the decoding device can properly obtain that component from the bitstream by assigning modal information to that component.

[0082] In the 18th embodiment, the decoding apparatus according to the 19th embodiment of this disclosure may, in the acquisition of the image, acquire the other components as a dummy image when the modal information is not assigned to the other components that are different from the one component among the plurality of components.

[0083] According to method 19, the decoding device can reduce its processing load by acquiring other components of unassigned modal information as dummy images.

[0084] In any of the 1st to 19th embodiments, the decoding apparatus involved in the 20th aspect of this disclosure may be, in the acquisition of the image, wherein the circuit acquires a plurality of images of different image categories, and the circuit further performs at least one task processing based on the plurality of images.

[0085] According to the 20th method, it is possible to improve the execution accuracy of at least one task processing performed by the decoding device based on multiple images of different image categories.

[0086] In the 20th aspect, the decoding apparatus involved in the 21st aspect of this disclosure may be such that the plurality of images comprise visible light images and the at least one task processing comprises human vision.

[0087] According to the 21st method, by including visible light images in multiple images, human vision can be properly performed in the decoding device.

[0088] The encoding apparatus according to the 22nd aspect of this disclosure includes: a circuit; and a memory connected to the circuit, the circuit encoding an image and parameters associated with the image into a bitstream, the parameters including modal information representing the image category.

[0089] According to the 22nd method, the encoding device can include modal information representing the image category in the bitstream and transmit it to the decoding device. As a result, the execution accuracy of the task processing performed by the decoding device can be improved.

[0090] In the 22nd embodiment, the encoding device involved in the 23rd embodiment of this disclosure may be at least one of the following image categories: visible light image, thermal image, infrared image, LiDAR image, RADAR image, or AI-generated image.

[0091] According to the 23rd method, the decoding device can perform optimal task processing corresponding to visible light images, thermal images, infrared images, LiDAR images, RADAR images, or AI-generated images.

[0092] In the 22nd or 23rd embodiment, the encoding device involved in the 24th embodiment of this disclosure may be such that the circuit further generates the image based on sensor data values ​​or sensed images.

[0093] According to the 24th method, the encoding device generates an image based on sensor data values ​​or sensed images, and can appropriately send the image to the decoding device.

[0094] In the 24th embodiment, the encoding device involved in the 25th embodiment of this disclosure may include, in which the parameters further include transformation information for transforming the pixel values ​​of the image into the sensor data values, or transforming the image into the sensed image.

[0095] According to the 25th method, by including transformation information in the parameters, the pixel values ​​of an image can be transformed into sensor data values, or the image can be transformed into a sensed image by the decoding device.

[0096] In any of the 22nd to 25th embodiments, the encoding apparatus involved in the 26th aspect of this disclosure may be, in the encoding of the parameter, the circuit encodes the parameter into a given header region of the bit stream, the given header region including VUI or SEI.

[0097] According to method 26, parameters can be easily obtained from a given header region of the bitstream by a decoding device.

[0098] In any of the 22nd to 26th embodiments of the present disclosure, the bitstream has a multi-layer structure comprising multiple image layers, wherein in the encoding of the images, the circuit encodes multiple images of different image categories into the multiple image layers.

[0099] According to the 27th method, the encoding device can send multiple images of different image categories to the decoding device, and the decoding device can appropriately obtain the multiple images from the bitstream.

[0100] In the 27th embodiment, the encoding apparatus involved in the 28th embodiment of this disclosure may be, in the encoding of the parameters, the circuit encodes multiple parameters associated with the multiple images into a header region of a specific image layer among the multiple image layers.

[0101] According to method 28, the encoding device can encode multiple parameters associated with multiple images of multiple image layers in a summary manner into the header region of a specific image layer.

[0102] In the 27th embodiment, the encoding apparatus involved in the 29th embodiment of this disclosure may be, in the encoding of the parameters, the circuit encodes multiple parameters associated with the multiple images into multiple header regions of the multiple image layers.

[0103] According to method 29, the encoding device is able to encode each parameter associated with the image of each image layer individually into each header region of each image layer.

[0104] In any of the 22nd to 26th embodiments, the encoding apparatus involved in the 30th embodiment of this disclosure may be such that the image is divided into multiple sub-images, and in the encoding of the image, the circuit encodes the multiple sub-images of different image categories into the bitstream.

[0105] According to method 30, the encoding device can send multiple sub-images of different image categories to the decoding device, and the decoding device can appropriately obtain the multiple sub-images from the bitstream.

[0106] In the 30th embodiment, the encoding device involved in the 31st embodiment of this disclosure may be, in the encoding of the parameters, the circuit encodes multiple parameters associated with the multiple sub-images into the header region of the image.

[0107] According to method 31, the encoding device is able to encode multiple parameters associated with multiple sub-images in a summary manner into the header area of ​​the image.

[0108] The encoding apparatus involved in the 32nd aspect of this disclosure, in any of the 22nd to 26th aspects, may be such that the image is composed of multiple components, and in the encoding of the image, the circuit encodes the multiple components of different image categories into the bit stream.

[0109] According to the 32nd method, the encoding device can send multiple components of different image categories to the decoding device, and the decoding device can appropriately obtain the multiple components from the bitstream.

[0110] In the 32nd embodiment, the encoding apparatus involved in the 33rd embodiment of this disclosure may be, in the encoding of the parameters, the circuit encodes multiple parameters associated with the multiple components into the header region of the image.

[0111] According to the 33rd method, the encoding device is able to encode multiple parameters associated with multiple components into the header region of the image.

[0112] In the 32nd embodiment, the encoding apparatus involved in the 34th embodiment of this disclosure may be, in which modal information is assigned to one of the plurality of components, and in the encoding of the image, the circuit encodes the one component into the bit stream.

[0113] According to the 34th method, when only one component is the required image category, the encoding device can properly encode that component into the bit stream by assigning modal information to that component.

[0114] In the 34th embodiment, the encoding apparatus according to the 35th embodiment of this disclosure may, in the encoding of the image, encode the other components as dummy images when the modal information is not assigned to the other components that are different from the one component among the plurality of components.

[0115] According to method 35, encoding other components of unassigned modal information as dummy images by an encoding device can reduce the processing load of the encoding device.

[0116] In the decoding method according to the 36th aspect of this disclosure, the decoding device acquires an image from a bitstream and parameters associated with the image, the parameters including modal information representing the image category.

[0117] According to the 36th method, modal information representing the image category can be included in the bitstream and transmitted from the encoding device to the decoding device, thereby improving the execution accuracy of the task processing performed by the decoding device.

[0118] In the encoding method according to the 37th aspect of this disclosure, the encoding device encodes an image and parameters associated with the image into a bitstream, the parameters including modal information representing the image category.

[0119] According to method 37, the encoding device can include modal information representing the image category in the bitstream and transmit it to the decoding device. As a result, the execution accuracy of the task processing performed by the decoding device can be improved.

[0120] (Implementation of this disclosure)

[0121] Hereinafter, embodiments of the present disclosure will be described in detail using the accompanying drawings. Furthermore, elements labeled with the same reference numerals in different drawings represent the same or corresponding elements.

[0122] Furthermore, the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, constituent elements, steps, and order of steps shown in the following embodiments are examples and are not intended to limit this disclosure. In addition, any constituent elements in the following embodiments that are not described in the independent claims representing the highest-level concept are described as arbitrary constituent elements. Furthermore, in all embodiments, the contents can be substituted or combined. Moreover, these overall or specific methods can be implemented by a system, method, integrated circuit, computer program, or a recording medium such as a computer-readable CD-ROM, or by any combination of a system, method, integrated circuit, computer program, or recording medium.

[0123] Figure 1 This diagram illustrates a simplified structure of an image processing system according to an embodiment of the present disclosure. The image processing system includes an encoding device 1, a decoding device 2, and a transmission path NW.

[0124] In encoding device 1, image data D1 is input from an external device. The external device includes a camera or the like that captures moving images. The external device inputs the image data D1 of the captured moving images into encoding device 1.

[0125] Encoding device 1 generates a bitstream BS based on image data D1. Here, a bitstream refers to a sequence of digital data or a flow of digital data. A bitstream (or simply a stream) can be a single stream or composed of multiple streams at multiple levels. Furthermore, a bitstream can be transmitted via serial communication on a single transmission path or via packet communication on multiple transmission paths. Encoding device 1 sends the generated bitstream BS to decoding device 2 via transmission path NW. Decoding device 2 receives the bitstream BS.

[0126] Decoding device 2 decodes image data D1 based on bitstream BS, and performs task processing based on the decoded image data D1. Task processing includes human vision and machine tasks. Human vision involves visual recognition or audiovisual processing of dynamic images from operators or users. Machine tasks include various types of task processing such as object detection, object tracking, object segmentation, action recognition, or pose estimation using machine learning-based inference models, i.e., artificial intelligence (AI) models. The task processing unit performing human vision includes a display device such as a liquid crystal display (LCD) or an organic EL display. The task processing unit performing machine tasks includes an inferrer using AI.

[0127] The transmission path NW can be the Internet, a WAN (wide area network), a LAN (local area network), or any combination thereof. Ideally, the transmission path NW should be a dedicated network that ensures secure communication through access restrictions.

[0128] The encoding device 1 includes a circuit 11 and a memory 12 connected to the circuit 11. The circuit 11 is configured to include a processor such as a CPU. The memory 12 is configured to include any recording medium such as ROM, RAM, HDD, SSD, or semiconductor memory. The memory 12 stores the data processed by the circuit 11 or the data processed during the process.

[0129] The decoding device 2 includes a circuit 21 and a memory 22 connected to the circuit 21. The circuit 21 is configured to include a processor such as a CPU. The memory 22 is configured to include any recording medium such as ROM, RAM, HDD, SSD, or semiconductor memory. The memory 22 stores the data processed by the circuit 21 or the data processed during the process.

[0130] Figure 2 This diagram shows a simplified representation of the structure of the circuit 11 included in the encoding device 1. The circuit 11 includes an acquisition unit 31, a setting unit 32, an encoding unit 33, and a transmission unit 34.

[0131] Next, the coding unit 33 involved in this embodiment will be described. Figure 29 This is a block diagram illustrating an example of the functional configuration of the encoding unit 33 according to this embodiment. The encoding unit 33 encodes the image in blocks.

[0132] like Figure 29 As shown, the encoding unit 33 includes: a segmentation unit 102, a subtraction unit 104, a transform unit 106, a quantization unit 108, an entropy coding unit 110, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a block memory 118, a loop filter 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. Furthermore, the intra-frame prediction unit 124 and the inter-frame prediction unit 126 are configured as part of the prediction processing unit 125.

[0133] For example, Figure 29 The multiple components of the coding unit 33 shown are obtained through Figure 1 The circuit 11 and memory 12 shown are installed.

[0134] Circuit 11 is configured to include a processor such as a CPU. Circuit 11 can be a dedicated or general-purpose electronic circuit for encoding images, or it can be a collection of multiple electronic circuits. Alternatively, for example, circuit 11 can also implement... Figure 29 The coding unit 33 shown has multiple components, including the components other than those used for storing information.

[0135] The memory 12 can be a dedicated or general-purpose electronic circuit for storing information, or it can be a collection of multiple electronic circuits. The memory 12 can be externally connected to the circuit 11, or it can be built into the circuit 11. In addition, the memory 12 can be a disk or optical disk, or it can be a storage device or recording medium. Furthermore, the memory 12 can be non-volatile memory or volatile memory.

[0136] The memory 12 can store encoded images or the stream corresponding to the encoded images. Additionally, the memory 12 can also store a program for the processor to perform image encoding processing.

[0137] Alternatively, memory 12 can also be implemented. Figure 29 The encoding unit 33 shown here functions as one of the multiple components used for storing information. Specifically, the memory 12 can also be implemented... Figure 29 The block memory 118 and frame memory 122 shown below illustrate their functions. More specifically, the memory 12 can also store reconstructed images (specifically, reconstructed blocks or reconstructed images, etc.).

[0138] Furthermore, in the coding section 33, Figure 29 Some of the multiple components shown may be omitted from installation, and some of the multiple processes performed by these components may also be omitted from execution. Alternatively, Figure 29 A portion of the multiple components shown may also be mounted on other devices, and a portion of the multiple processes performed by the multiple components may also be performed by other devices.

[0139] Figure 3 This is a flowchart showing the process performed by the circuit 11 of the encoding device 1.

[0140] First, in step SP11, the acquisition unit 31 acquires image data D11, representing the processing object, i.e., image Q, input from an external device. Image data D11 is equivalent to... Figure 1 The image data D1 shown is shown.

[0141] Next, in step SP12, the setting unit 32 sets parameter P in association with image Q. Parameter P includes modal information D12 of image Q. Modal information D12 indicates the image category of image Q. Image category includes, for example, at least one of visible light image, thermal image, infrared image, LiDAR (light detection and ranging) image, RADAR (radio detection and ranging) image, or AI-generated image. The setting unit 32 can set parameter P by image analysis based on image data D11, or based on setting information input by the operator of encoding device 1.

[0142] Visible light images include natural images or RGB images, used to provide detailed color information for human vision or machine tasks. Thermal images include temperature distribution images, used for detecting organisms in dark environments. Infrared images include images captured using infrared cameras, used for imaging in low light. LiDAR images include distance images detected by laser-based range sensors, used to provide accurate distance information. RADAR images include distance images detected by radio-wave range sensors, used for ranging under arbitrary lighting conditions. AI-generated images include images generated by AI, or images obtained by appending AI-generated descriptive text to visible light images.

[0143] Next, in step SP13, the encoding unit 33 encodes the image Q represented by the image data D11 input from the acquisition unit 31 into the bit stream BS.

[0144] Next, in step SP14, the encoding unit 33 encodes the parameter P, which includes the modal information D12, input from the setting unit 32 into the bitstream BS. Here, encoding the parameter P into the bitstream BS can also be described as storing the parameter P in the bitstream BS, or simply storing the parameter P in the bitstream BS. Furthermore, the execution order of steps SP13 and SP14 can also be... Figure 3 Conversely, the examples can also be simultaneous.

[0145] Next, in step SP15, the transmitting unit 34 transmits the bit stream BS input from the encoding unit 33 to the decoding device 2 via the transmission path NW.

[0146] Figure 4 This diagram is a simplified representation of the structure of the bitstream BS. The bitstream BS has a header region 41 and a payload region 42. The encoding unit 33 stores the encoded data of the image Q in the payload region 42 and the encoded data of the parameter P associated with the image Q in the header region 41.

[0147] The encoding unit 33 can also encode the encoded data of parameter P into a given region 43 within the header region 41. The given region 43 can be VUI (video usability information) or SEI (Supplemental Enhancement Information). However, the given region 43 is not limited to VUI or SEI, and can also be VPS, SPS, PPS, PH, SH, APS, tile header, or system layer header, etc.

[0148] Figure 30 This is a diagram illustrating an example of the hierarchical structure of data in a stream. A stream, for example, contains video sequences. Figure 30 As shown in (A), a video sequence may include, for example, VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), SEI (Supplemental Enhancement Information), and multiple pictures.

[0149] VPS contains encoding parameters common to multiple layers in a dynamic image composed of multiple layers, and encoding parameters associated with the multiple layers or individual layers contained in the dynamic image.

[0150] The SPS contains parameters for the sequence, that is, the encoding parameters that the decoding device 2 refers to in order to decode the sequence. These encoding parameters can, for example, represent the width or height of the image. Furthermore, multiple SPSs can exist.

[0151] The PPS contains parameters for the images, that is, the encoding parameters that the decoding device 2 refers to in order to decode each image in the sequence. These encoding parameters may, for example, include a reference value for the quantization width used in the image decoding and a flag indicating the application of weighted prediction. Furthermore, multiple PPSs may exist. The SPS and PPS are sometimes simply referred to as a parameter set.

[0152] like Figure 30 As shown in (B), the image contains an image header and one or more slices. The image header contains encoding parameters referenced by the decoding device 2 for decoding the one or more slices.

[0153] like Figure 30 As shown in (C), the slice contains a slice header and one or more bricks. The slice header contains encoding parameters referenced by the decoding device 2 for decoding the one or more bricks.

[0154] like Figure 30 As shown in (D), a brick contains more than one CTU (Coding Tree Unit).

[0155] Alternatively, the image may not contain slices, but instead contain sets of tiles. In this case, the set of tiles contains more than one tile. Additionally, slices may also be included within the brick itself.

[0156] CTU is also known as a superblock or basic partitioning unit. For example... Figure 30 As shown in (E), the CTU includes a CTU header and one or more CUs (Coding Units). The CTU header contains encoding parameters referenced by the decoding device 2 for decoding the one or more CUs.

[0157] A CU can also be divided into multiple smaller CUs. Additionally, as... Figure 30As shown in (F), the CU contains a CU header, prediction information, and residual coefficient information. Prediction information is used to predict the CU. Residual coefficient information represents the prediction residuals. Furthermore, the CU is essentially the same as a PU (Prediction Unit) or TU (Transform Unit), but it can also contain multiple TUs smaller than the CU. Additionally, the CU can be processed at the level of each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. A VPDU is, for example, a fixed unit that can be processed in one stage during pipelined processing in hardware.

[0158] Furthermore, a flow may not have Figure 30 This is a subset of the multiple hierarchical levels shown. Furthermore, the order of these levels can be changed, and any level can be replaced with another level.

[0159] The image that is the object of the processing performed by the encoding device 1 or decoding device 2 at the current point in time is called the current image. If the processing is encoding, the current image is synonymous with the image to be encoded; if the processing is decoding, the current image is synonymous with the image to be decoded. Furthermore, the block (CU or a block of CU) that is the object of the processing performed by the encoding device 1 or decoding device 2 at the current point in time is called the current block. If the processing is encoding, the current block is synonymous with the block to be encoded; if the processing is decoding, the current block is synonymous with the block to be decoded.

[0160] Figure 5 , Figure 6 This is a diagram illustrating the first example of the syntax related to the setting of modal information D12. Figure 5 , 6 In the example shown, modal information D12 is represented by the value of the identifier of vui_modality_type contained in the VUI parameters.

[0161] like Figure 5 As shown, a value of 1 for the VUI parameter `vui_modality_info_present_flag` indicates the presence of `vui_modality_type`. A value of 0 for `vui_modality_info_present_flag` indicates the absence of `vui_modality_type`. Furthermore, a value of 0 for `vui_modality_info_present_flag` can also indicate that the value of `vui_modality_type` is 0.

[0162] like Figure 6As shown, a value of 0 for `vui_modality_type` indicates that image Q is a visible light image. A value of 1 indicates that image Q is a thermal image. A value of 2 indicates that image Q is an infrared image. A value of 3 indicates that image Q is a LiDAR image. A value of 4 indicates that image Q is a RADAR image. A value of 5 indicates that image Q is an AI-generated image. Other values ​​for `vui_modality_type` indicate a reserved box for future use. Furthermore, the value can be increased or decreased based on the image category of the object. Figure 6 The number of images defined in the document.

[0163] Figure 7 , Figure 8 This is the second example of a diagram illustrating the syntax related to the setting of modal information D12. Figure 7 , 8 In the example shown, modal information D12 is represented by the value of the identifier of mvi_modality_type contained in SEI messages such as machine_vision_indication.

[0164] like Figure 7 As shown, a value of 1 for `mvi_modality_info_present_flag` in the SEI message indicates the presence of `mvi_modality_type`. A value of 0 for `mvi_modality_info_present_flag` indicates the absence of `mvi_modality_type`. Furthermore, a value of 0 for `mvi_modality_info_present_flag` can also indicate that the value of `mvi_modality_type` is 0.

[0165] like Figure 8As shown, a value of 0 for `mvi_modality_type` indicates that image Q is a visible light image. A value of 1 indicates that image Q is a thermal image. A value of 2 indicates that image Q is an infrared image. A value of 3 indicates that image Q is a LiDAR image. A value of 4 indicates that image Q is a RADAR image. A value of 5 indicates that image Q is an AI-generated image. Other values ​​for `mvi_modality_type` indicate pre-defined boxes for future use. Furthermore, the value can be increased or decreased based on the image category of the object. Figure 8 The number of images defined in the document.

[0166] Generally, image data encoded into a bitstream (BS) consists of multiple components (image components). For example, a visible light image consists of multiple color components (R, G, B), or one luminance component (Y) and multiple chromatic aberration components (Cb, Cr). Depending on the image type, there are images composed of three image components, such as colored visible light images, as well as images composed of only one image component, such as monochrome visible light images (visible light images with only luminance components), thermal images, infrared images, LiDAR images, or RADAR images (hereinafter referred to as "single-component images").

[0167] As a first example, when the bitstream BS consists of three components, and the image Q is a single-component image, only the first component is encoded as a normal image into the bitstream BS, while the second and third components are encoded as dummy images. Encoding the dummy images is simpler than encoding the normal images. As a second example, when the bitstream BS consists of three components, and the image Q is a single-component image, the single-component image Q can be transformed into an image composed of three image components, and the first to third components are encoded as normal images into the bitstream BS. As a third example, when encoding the image Q is a single-component image, it can also be formatted as a bitstream BS consisting of only one component, and only the first component is encoded as a normal image into the bitstream BS.

[0168] Figure 9This diagram shows a simplified representation of the structure of the circuit 21 included in the decoding device 2. The circuit 21 includes: a receiving unit 51, a decoding unit 52, a switching unit 53A, and multiple n (n being a natural number of 2 or more) task processing units 541 to 542. n Task Processing Department 541-54 n The tasks performed include both human vision and machine tasks. Machine tasks include a wide variety of tasks that use AI models, such as object detection, object tracking, object segmentation, action recognition, or pose estimation.

[0169] Next, the decoding unit 52 according to this embodiment will be described. Figure 31 This is a block diagram illustrating an example of the functional configuration of the decoding unit 52 according to this embodiment. The decoding unit 52 decodes the encoded image, i.e., the stream, on a block-by-block basis.

[0170] like Figure 31 As shown, the decoding unit 52 includes: an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an adder unit 208, a block memory 210, a loop filter 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a segmentation determination unit 224. Furthermore, the intra-frame prediction unit 216 and the inter-frame prediction unit 218 are configured as part of the prediction processing unit 215.

[0171] For example, Figure 31 The decoding unit 52 shown has multiple components that are transmitted through Figure 1 The circuit 21 and memory 22 shown are installed.

[0172] Circuit 21 is configured to include a processor such as a CPU. Circuit 21 can be a dedicated or general-purpose electronic circuit for decoding streams, or it can be a collection of multiple electronic circuits. Alternatively, for example, circuit 21 can also implement... Figure 31 The decoding unit 52 shown has multiple components, including the components other than those used for storing information.

[0173] The memory 22 can be a dedicated or general-purpose electronic circuit for storing information, or it can be a collection of multiple electronic circuits. The memory 22 can be externally connected to the circuit 21, or it can be built into the circuit 21. Furthermore, the memory 22 can be a disk or optical disk, or it can be a storage device or recording medium. Additionally, the memory 22 can be either non-volatile or volatile.

[0174] The memory 22 can store the decoded stream or the decoded image. Additionally, the memory 22 can also store a program for the processor to perform the decoding process on the stream.

[0175] Alternatively, memory 22 can also be implemented Figure 31 The decoding unit 52 shown here has the function of storing information among its various constituent elements. Specifically, the memory 22 can also be implemented... Figure 31 The block memory 210 and frame memory 214 are shown to have specific functions. More specifically, the memory 22 can also store reconstructed images (specifically, reconstructed blocks or reconstructed images, etc.).

[0176] Furthermore, in the decoding unit 52, Figure 31 Some of the multiple components shown may be omitted from installation, and some of the multiple processes performed by these components may also be omitted from execution. Alternatively, Figure 31 A portion of the multiple components shown may also be mounted on other devices, and a portion of the multiple processes performed by the multiple components may also be performed by other devices.

[0177] Figure 31 The decoding unit 52 shown includes an inverse quantization unit 204, an inverse transform unit 206, an adder unit 208, a block memory 210, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, a prediction control unit 220, and a loop filter 212, which are respectively connected to... Figure 29 The encoding unit 33 shown includes the same inverse quantization unit 112, inverse transform unit 114, adder unit 116, block memory 118, frame memory 122, intra-frame prediction unit 124, inter-frame prediction unit 126, prediction control unit 128, and loop filter 120.

[0178] Figure 10 This is a flowchart showing the process performed by the circuit 21 of the decoding device 2.

[0179] First, in step SP21, the receiving unit 51 receives the bit stream BS sent by the encoding device 1 from the transmission path NW.

[0180] Next, in step SP22, the decoding unit 52 decodes the image Q based on the payload region 42 of the bitstream BS input from the receiving unit 51, thereby acquiring it. Furthermore, decoding may also include extraction. The decoding unit 52 outputs image data D21 of the image Q. Image data D21 is equivalent to... Figure 2 The image data shown is D11.

[0181] Next, in step SP23, the decoding unit 52 decodes the parameter P based on the header region 41 (or a given region 43 within the header region 41) of the bit stream BS input from the receiving unit 51, thereby obtaining the parameter P. The decoding unit 52 extracts the modal information D22 contained in the parameter P. The modal information D22 is equivalent to... Figure 2 The modal information D12 is shown. Furthermore, the execution order of steps SP22 and SP23 can also be... Figure 10 Conversely, the examples can also be simultaneous.

[0182] Next, in step SP24A, the switching unit 53A switches the task processing units 541 to 54 based on the modal information D22 input from the decoding unit 52. n The switching unit 53A maintains multiple image categories and multiple task processing units 541-54. n The correspondence is pre-set in a table of information (illustration omitted). Switching unit 53A refers to this table information to switch between multiple task processing units 541-542. n Among them, a task processing unit 541-54 is selected that corresponds to the image category represented by modal information D22. n .

[0183] Next, in step SP25, a task processing unit 541 to 542 selected by step SP24A... n Task processing is performed based on the image data D21 input from the decoding unit 52 via the switching unit 53A.

[0184] According to this embodiment, the encoding device 1 can include modal information D12 (D22) representing the image category of image Q in the bitstream BS and transmit it to the decoding device 2. As a result, the execution accuracy of the task processing performed by the decoding device 2 can be improved.

[0185] Furthermore, according to this embodiment, the decoding device 2 can perform optimal processing (task processing, etc.) corresponding to image categories such as visible light images, thermal images, infrared images, LiDAR images, RADAR images, or AI-generated images.

[0186] Furthermore, according to this embodiment, the decoding device 2 can perform optimal task processing corresponding to the image category, thereby improving the execution accuracy of the task processing performed by the decoding device 2.

[0187] Furthermore, according to this embodiment, the decoding device 2 is capable of performing the optimal machine task corresponding to the image category.

[0188] Furthermore, according to this embodiment, the decoding device 2 is capable of performing optimal human vision corresponding to the image category.

[0189] Furthermore, according to this embodiment, the decoding device 2 can easily decode the parameter P from a given region 43 of the bitstream BS.

[0190] Hereinafter, various modifications of the above-described embodiments will be described. These modifications can be arbitrarily combined and applied.

[0191] (Example 1)

[0192] Decoding device 2 can also switch between multiple AI models 551-55 used in task processing based on modal information D22. n .

[0193] Figure 11 This diagram shows a simplified representation of the circuit 21 included in the decoding device 2. The circuit 21 includes a receiving unit 51, a decoding unit 52, a switching unit 53B, and a task processing unit 54. The task processing unit 54 performs tasks including human vision or machine tasks. The task processing unit 54 has multiple n AI models 551 to 55. n AI Models 551-55 n This includes neural network models that have undergone machine learning. For example, if the task performed by the task processing unit 54 is object detection, AI models 551-55... n Convolutional neural network models, including ResNet or YOLO. AI Models 551-55 n Each of the setting parameters can be predefined in the decoding device 2, or can be transmitted from the encoding device 1 to the decoding device 2 through a different bit stream than the bit stream BS.

[0194] Figure 12 This is a flowchart showing the process performed by the circuit 21 of the decoding device 2.

[0195] Processing steps SP21 to SP23 and Figure 10 same.

[0196] Next, in step SP24B, the switching unit 53B switches AI models 551 to 55 based on the modal information D22 input from the decoding unit 52. n The switching unit 53B maintains multiple image categories and multiple AI models 551-55. n The correspondence is defined in a pre-set table (illustration omitted). The switching unit 53B refers to this table information to select from multiple AI models 551-552. n Choose an AI model that corresponds to the image category represented by modal information D22. (551-55) n .

[0197] Next, in step SP25, the task processing unit 54 uses an AI model 551-55 selected in step SP24B. n Task processing is performed based on image data D21 input from the decoding unit 52 via the switching unit 53B.

[0198] Furthermore, the switching unit 53B can also replace multiple AI models 551-55. n It can switch between multiple image processing methods. These multiple image processing methods include any image processing method such as filtering, segmentation, edge detection, or color analysis. Furthermore, the decoding device 2 can also switch between task processing units 541 to 542 based on modal information D22. n And AI models 551-55 n both sides.

[0199] According to this variation, the task processing unit 54 of the decoding device 2 can use the optimal AI model 551-55 corresponding to the image category of image Q. n Alternatively, the optimal image processing method can be used to perform the task processing. As a result, the execution accuracy of the task processing performed by the decoding device 2 can be improved.

[0200] (Second variation)

[0201] The encoding device 1 can also generate image Q based on sensor data values ​​or a sensed image. The parameter P can also include transformation information D13 (D23) for transforming the pixel values ​​of image Q into the aforementioned sensor data values ​​or transforming image Q into the aforementioned sensed image. Furthermore, the sensor data values ​​include the values ​​of data acquired from the sensor itself, or data columns obtained by arranging these values. The sensed image includes an image generated by arranging multiple sensor data values ​​into a matrix, or an image obtained by transforming that image. The sensed image can be a two-dimensional image or a three-dimensional image. Additionally, the pixel values ​​of the sensed image can be any number of bits, such as 8 bits, 10 bits, or 16 bits.

[0202] Figure 13 This diagram is a simplified representation of the structure of the circuit 11 included in the encoding device 1. The circuit 11 is relative to... Figure 2 The structure shown also has a transformation section 35.

[0203] The acquisition unit 31 acquires sensor data D10 input from an external device. The external device includes an image sensor, a temperature sensor, or a distance sensor, etc.

[0204] The transformation unit 35 generates image data D11 of image Q based on sensor data D10 by transforming the sensor data values ​​of sensor data D10 input from the acquisition unit 31 into pixel values ​​of image Q. Alternatively, instead of transforming the sensor data values ​​into pixel values, the transformation unit 35 may perform image transformation on the sensed image, thereby generating image data D11 of image Q based on the image data of the sensed image. Furthermore, image Q is an image to be encoded. An image to be encoded is an image obtained by transforming sensor data values ​​or a sensed image into an image that can be encoded. An image to be encoded is an image that can be recognized as an image by human vision. An image to be encoded is a two-dimensional image. The pixel values ​​of the image to be encoded are a number of bits conforming to standards such as 8 bits or 10 bits.

[0205] The transformations performed by the transformation unit 35 can also include transformations in units such as distance, temperature, or signal strength. Furthermore, the transformations performed by the transformation unit 35 can also include dimensional transformations, such as transformations from a three-dimensional image to a two-dimensional image. Additionally, the transformations performed by the transformation unit 35 can include coordinate system transformations, such as transformations from a three-dimensional polar coordinate system to a two-dimensional orthogonal coordinate system, or vice versa. Furthermore, the transformations performed by the transformation unit 35 can also include quantization transformations. Additionally, the transformations performed by the transformation unit 35 can include normalization-based data transformations. Normalization can include transformations using mathematical expressions or transformations using tabular information. Transformations using mathematical expressions can also include any method such as Min-Max Scaling, Z-score Normalization, or Log Transformation.

[0206] The transformation unit 35 inputs image data D11 to the encoding unit 33. Additionally, the transformation unit 35 inputs transformation information D13, used to inversely transform image data D11 into sensor data D10, to the setting unit 32. The transformation information D13 includes image sample information. This image sample information includes image sample interpretation, image sample representation type, angular resolution, or focal length, etc. The image sample interpretation defines the type of information used to represent image samples within an image. Image samples can represent signal strength such as distance, temperature, radio waves, radiation, or LiDAR pulses. The image sample representation type defines the method for transforming image sample values ​​into sensor data values. The angular resolution or focal length includes transformation parameters for coordinate system transformations, such as transformations from a two-dimensional orthogonal coordinate system to a three-dimensional polar coordinate system, or transformations from a two-dimensional orthogonal coordinate system to a three-dimensional orthogonal coordinate system.

[0207] The setting unit 32 sets parameters P in association with the image Q. Parameter P includes modal information D12 of the image Q. In addition, parameter P includes transformation information D13.

[0208] The encoding unit 33 encodes the image Q represented by the image data D11 input from the transformation unit 35 into the bit stream BS. Additionally, the encoding unit 33 encodes the parameter P, which includes modal information D12 and transformation information D13, input from the setting unit 32 into the bit stream BS.

[0209] The transmitting unit 34 transmits the bit stream BS input from the encoding unit 33 to the decoding device 2 via the transmission path NW.

[0210] Figure 14 This diagram is a simplified representation of the first example of the structure of the circuit 21 included in the decoding device 2. The circuit 21 is relative to... Figure 9 The structure shown also includes a transformation section 55.

[0211] The receiving unit 51 receives the bit stream BS sent by the encoding device 1 from the transmission path NW.

[0212] The decoding unit 52 decodes the image Q and the parameter P based on the bit stream BS input from the receiving unit 51. Furthermore, the decoding unit 52 extracts the modal information D22 and transform information D23 contained in the parameter P. The transform information D23 is equivalent to... Figure 13 The transformation information D13 is shown.

[0213] The transformation unit 55 transforms the pixel values ​​of the image Q represented by the image data D21 input from the decoding unit 52 into sensor data values ​​based on the transformation information D23, thereby generating sensor data D20 based on the image data D21. Sensor data D20 is equivalent to... Figure 13 The sensor data D10 is shown. Alternatively, the transformation unit 55 can transform the image Q into the aforementioned sensed image through image transformation based on the transformation information D23, instead of transforming pixel values ​​into sensor data values.

[0214] The transformations performed by the transformation unit 55 can also include transformations in units such as distance, temperature, or signal strength. Furthermore, the transformations performed by the transformation unit 55 can also include dimensional transformations such as transformations from a two-dimensional image to a three-dimensional image. Additionally, the transformations performed by the transformation unit 55 can include coordinate system transformations such as transformations from a two-dimensional orthogonal coordinate system to a three-dimensional polar coordinate system, or vice versa. Furthermore, the transformations performed by the transformation unit 55 can also include inverse quantization transformations. Additionally, the transformations performed by the transformation unit 55 can include data transformations based on inverse normalization. Inverse normalization can also include transformations using mathematical expressions or transformations using table information. Transformations using mathematical expressions can also include any method such as Min-Max Scaling, Z-score Normalization, or Log Transformation.

[0215] Based on the modal information D22 input from the decoding unit 52, the switching unit 53A switches the task processing units 541 to 54. n .

[0216] A task processing unit 541-54 selected by the switching unit 53A n Task processing is performed based on sensor data D20 input from the transformation unit 55 via the switching unit 53A or image data of the aforementioned sensed image.

[0217] Figure 15 This is a simplified diagram showing the second example of the structure of the circuit 21 included in the decoding device 2.

[0218] The receiving unit 51 receives the bit stream BS sent by the encoding device 1 from the transmission path NW.

[0219] The decoding unit 52 decodes the image Q and the parameter P based on the bit stream BS input from the receiving unit 51. Furthermore, the decoding unit 52 extracts the modal information D22 and transform information D23 contained in the parameter P.

[0220] Based on the modal information D22 input from the decoding unit 52, the switching unit 53A switches the task processing units 541 to 54. n .

[0221] Decoding unit 52 inputs image data D21 and transformation information D23 to a task processing unit 541-54 selected by switching unit 53A. n .

[0222] A task processing unit 541-54 selected by the switching unit 53A n Task processing is performed based on the input image data D21 and transformation information D23. This task processing unit 541-54... n Alternatively, sensor data D20 can be generated based on image data D21 by transforming the pixel values ​​of image Q represented by image data D21 into sensor data values ​​based on transformation information D23. This is a task processing unit 541-54. n Task processing can also be performed based on the generated sensor data D20. Alternatively, a task processing unit 541-54... n Image Q can also be transformed into the aforementioned sensed image through image transformation based on transformation information D23, and task processing can be performed based on the sensed image.

[0223] According to this variation, the encoding device 1 generates an image Q based on sensor data values ​​or sensed images, and can appropriately send the image Q to the decoding device 2.

[0224] Furthermore, according to this modified example, by including transformation information D13 (D23) in parameter P, the pixel values ​​of image Q can be transformed into sensor data values ​​or image Q can be transformed into a sensed image by the decoding device 2.

[0225] In addition, according to Figure 14 In the example configuration shown, the decoding device 2 can perform task processing based on sensor data values ​​or sensed images by performing transformation processing before task processing.

[0226] In addition, according to Figure 15 In the example configuration shown, the decoding device 2 can perform task processing based on sensor data values ​​or sensed images by performing transformation processing during task processing.

[0227] (3rd variation)

[0228] Bitstream BS can also have multiple m-layer image layers L1 to L1 (where m is a natural number greater than 2). m A multi-layered structure.

[0229] Figure 16 This is a diagram illustrating the first example of a multi-layered structure. In Figure 16 Only one access unit is shown. An access unit is the smallest unit of processing a temporal attribute, such as equivalent to one frame of a moving image. A bitstream (BS) is constructed by containing multiple temporally consecutive access units.

[0230] In the payload area 42 of the image layer L1, which is the lowest layer, the image Q, which serves as the main image, is stored. L1 The encoded data. In the payload area 42 of the second image layer L2, the image Q, serving as an auxiliary image, is stored. L2 The encoded data. Similarly, in the image layer L at the m-th layer. m The payload area 42 stores the image Q, which serves as an auxiliary image. Lm Encoded data.

[0231] Image Q L1 ~Q Lm They have mutually distinct image categories. For example, in a multi-layered structure with 3 layers and m=3, image Q... L1 For visible light images, image Q L2 This is an infrared image, image Q L3 This is a thermal image.

[0232] In image Q L1 ~Q Lm In cases where there is correlation (e.g., infrared images and thermal images), it is also possible to use image layers L1 to L2. m Interstitial reference image. On the other hand, in image Q L1 ~QLm In cases where there is no correlation between images (e.g., visible light images and distance images), images can also be processed in layers L1 to L2. m Do not refer to the image.

[0233] In the header region 41 of image layer L1, the data related to image Q is stored. L1 The associated parameter P was established. L1 The encoded data. Parameter P L1 Includes image Q L1 The modal information D12. Additionally, in the header region 41 of image layer L1, information related to image Q is also stored. L2 ~Q Lm The associated parameter P was established. L2 ~P Lm The encoded data. Parameter P L2 ~P Lm Includes image Q L2 ~Q Lm Modal information D12. To include all image layers L1 to L2... m Parameter P L1 ~P Lm The header region 41 stored in a specific image layer L1 can also use the SDI (scalability dimension information) _SEI message.

[0234] Figure 17 This is a diagram illustrating an example of the syntax related to the setting of modal information D12. Modal information D12 is represented by the value of the identifier sdi_aux_id[i] contained in the SDI_SEI message. When the value of sdi_aux_id[i] is 0, it indicates that in the i-th image layer L... i There is no image Q in the image that serves as an auxiliary image. Li When sdi_aux_id[i] has a value of 1 to 6, it indicates that in the i-th image layer L i exist Figure 17 The image category Q shown Li Furthermore, when sdi_aux_id[i] has a value of 7–127 or 160–255, it indicates a reserved box prepared for future use; when sdi_aux_id[i] has a value of 128–159, it indicates unspecified. Additionally, the number of boxes can be increased or decreased based on the image category of the object. Figure 17 The number of images defined in the document.

[0235] Figure 18 , Figure 19This is a diagram illustrating an example of the syntax related to setting modal information D12 without using the SDI_SEI message. Modal information D12 is represented as the value of the identifier mvi_modality_type[i] contained in SEI messages such as machine_vision_indication.

[0236] like Figure 18 , 19 As shown, the number of layers included in the multi-layer structure is specified by mvi_max_layers_minus1, and the modal information of the specified number of layers is summarized and described. Specifically, image layer L i The modal information is denoted as mvi_modality_type[i].

[0237] When mvi_modality_type[i] has a value of 0, it indicates that the image Q... Li The image category is visible light image. When mvi_modality_type[i] is 1, it indicates that the image Q... Li The image category is thermal image. When mvi_modality_type[i] has a value of 2, it indicates that the image Q... Li The image category is infrared image. When mvi_modality_type[i] has a value of 3, it indicates that the image Q... Li The image category is LiDAR image. When mvi_modality_type[i] has a value of 4, it indicates that the image is Q... Li The image category is RADAR image. When mvi_modality_type[i] has a value of 5, it indicates that the image is Q... Li The image category is AI-generated image. If the value of `mvi_modality_type[i]` is other than [prepared], it represents a pre-defined bounding box ensured for future use. Furthermore, the number of bounding boxes can be increased or decreased based on the image category of the object. Figure 19 The number of images defined in the document.

[0238] Figure 20 This is a diagram illustrating the second example of a multi-layered structure. Figure 20 Only one access unit is shown in the image.

[0239] In image layers L1 to L m The effective payload area is 42, which stores image Q. L1 ~Q Lm Encoded data.

[0240] In the header region 41 of image layer L1, the data related to image Q is stored. L1 The associated parameter P was established.L1 The encoded data. In the header region 41 of image layer L2, the data related to image Q is stored. L2 The associated parameter P was established. L2 The encoded data. Similarly, in image layer L m The header area 41 stores the image Q. Lm The associated parameter P was established. Lm The encoded data. Thus, each image layer L1~L m Parameter P L1 ~P Lm Stored in each image layer L1 to L2 m The header area is 41.

[0241] Each image layer L1~L m Parameter P L1 ~P Lm It can also be described in each image layer L1 to L2 m VUI or SEI. In Figure 5 , Figure 6 In the example shown, modal information D12 is represented by the value of the identifier of vui_modality_type contained in the VUI parameters. Figure 7 , Figure 8 In the example shown, modal information D12 is represented by the value of the identifier of mvi_modality_type contained in SEI messages such as machine_vision_indication. Furthermore, in order to include all image layers L1 to L2... m The SEI summary is stored in the header area 1, or you can use the SN (scalable nesting) _SEI message.

[0242] Figure 21 The diagram is a simplified representation of the structure of the circuit 21 included in the decoding device 2.

[0243] The decoding unit 52 decodes the image Q based on the payload region 42 of the multi-layered bitstream BS input from the receiving unit 51. L1 ~Q Lm Decoding is performed. The decoding unit 52 outputs image Q. L1 ~Q Lm Image data D211~D21 m .

[0244] Furthermore, the decoding unit 52, based on the header region 41 of the multi-layered bit stream BS input from the receiving unit 51, modifies the parameter P... L1 ~P Lm Decoding is performed. The decoding unit 52 extracts parameter P. L1 ~P LmThe modal information included is D221~D22 m .

[0245] The switching unit 53A is based on the modal information D221 to D22 input from the decoding unit 52. m For each image Q L1 ~Q Lm Switching Task Processing Units 541-54 n The switching unit 53A maintains multiple image categories and multiple task processing units 541-54. n The correspondence is defined in a pre-set table (illustration omitted). The switching unit 53A refers to this table information and switches between images Q for each image. L1 ~Q Lm From multiple task processing units 541 to 54 n The selection is based on modal information D221~D22. m The image category represents a task processing unit 541-54. n .

[0246] Task Processing Department 541-54 n Based on image data D211~D21 input from decoding unit 52 via switching unit 53A m To perform task processing.

[0247] In addition, Figure 21 The diagram shows a structure where one image data is input to one task processing unit 54, but it can also be configured to input multiple image data to one task processing unit 54. Specifically, the switching unit 53A can also maintain multiple image categories and multiple task processing units 541-54. n The correspondence is established in a pre-defined table, in which multiple images Q are associated with one task processing unit 54.

[0248] According to this variation, the encoding device 1 can encode multiple images Q of different image categories. L1 ~Q Lm The data is sent to decoding device 2, which can appropriately process the multiple images Q according to the bitstream BS. L1 ~Q Lm Decode it.

[0249] according to Figure 16 In the example configuration shown, the encoding device 1 is capable of encoding multiple image layers L1 to L2. m Multiple images Q L1 ~Q Lm Multiple parameters P were established for correlation. L1 ~P LmThe encoding is summarized and encoded into the header region 41 of a specific image layer L1. Furthermore, the decoding device 2 can summarize the encoding of multiple image layers L1 to L2 based on the header region 41 of the specific image layer L1. m Multiple images Q L1 ~Q Lm Multiple parameters P were established for correlation. L1 ~P Lm Decode it.

[0250] according to Figure 20 In the example configuration shown, the encoding device 1 is capable of encoding each image layer L1 to L2. m Image Q L1 ~Q Lm The parameters P that are associated have been established. L1 ~P Lm Encoded separately into each image layer L1 to L2 m Each header area 41. Additionally, the decoding device 2 can determine the image layers L1 to L2. m Each header region 41 is individually associated with each image layer L1 to L1. m Image Q L1 ~Q Lm The parameters P that are associated have been established. L1 ~P Lm Decode it.

[0251] (4th variation)

[0252] An image Q can also be segmented into multiple j (j is a natural number greater than 2) sub-images Q. S1 ~Q Sj .

[0253] Figure 22 This is a simplified diagram illustrating the structure of the bitstream BS. The encoding unit 33 encodes the sub-image Q... S1 ~Q Sj The encoded data is stored in payload area 42 and will be combined with sub-image Q. S1 ~Q Sj The associated parameter P was established. S1 ~P Sj The encoded data is stored in header area 41. Sub-image Q S1 ~Q Sj They have mutually distinct image categories. Parameter P S1 ~P Sj Contains sub-image Q S1 ~Q Sj The modal information D12. To include all sub-images Q... S1 ~Q Sj Parameter P S1 ~P SjThe summary is stored in header area 41, or the SN (scalable nesting) _SEI message can be used.

[0254] Figure 23 , Figure 24 This is a diagram illustrating an example of the syntax related to the setting of modal information D12. (And each sub-image Q) S1 ~Q Sj The associated modal information D12 is represented by the value of the mvi_modality_type identifier contained in the j-th machine_vision_indication in the SN_SEI message.

[0255] like Figure 24 As shown, when the value of mvi_modality_type is 0, it indicates that the sub-image Q... Sj The image category is visible light image. When the value of mvi_modality_type is 1, it indicates that the sub-image Q... Sj The image category is thermal image. When the value of mvi_modality_type is 2, it indicates that the sub-image Q... Sj The image category is infrared image. When the value of mvi_modality_type is 3, it indicates that the sub-image Q... Sj The image category is LiDAR image. With a value of 4 for mvi_modality_type, it indicates that the sub-image Q... Sj The image category is RADAR image. When the value of mvi_modality_type is 5, it indicates that the sub-image Q... Sj The image category is AI-generated image. If the value of `mvi_modality_type` is other than [prepared], it indicates a box prepared for future use. Furthermore, the number of boxes can be increased or decreased based on the image category of the object. Figure 24 The number of images defined in the document.

[0256] According to this variation, the encoding device 1 is able to encode multiple sub-images Q of different image categories. S1 ~Q Sj The data is sent to decoding device 2, which can appropriately process the multiple sub-images Q according to the bitstream BS. S1 ~Q Sj Decode it.

[0257] Furthermore, according to this modified example, the encoding device 1 is capable of encoding multiple sub-images Q. S1 ~Q Sj Multiple parameters P were established for correlation. S1 ~P SjThe encoding is then summarized and encoded into the header region 41 of image Q. Furthermore, according to this modified example, the decoding device 2 can summarize the encoding of multiple sub-images Q based on the header region 41 of image Q. S1 ~Q Sj Multiple parameters P were established for correlation. S1 ~P Sj Decode it.

[0258] (5th variation)

[0259] An image Q can also consist of multiple k components (k being a natural number greater than 2) Q. C1 ~Q Ck constitute.

[0260] Component Q C1 ~Q Ck These are the multiple components that make up an image Q. For example, a visible light image is composed of multiple color components (R, G, B), or a luminance component (Y) and multiple chromatic aberration components (Cb, Cr).

[0261] Figure 25 This shows that based on multiple components Q C1 ~Q Ck The diagram shows an example of the composition of image Q.

[0262] In the payload region 42 of the bitstream BS, the component Q corresponding to the first image component C1 is stored. C1 The encoded data, and the component Q corresponding to the second image component C2 C2 The encoded data, and similarly with the k-th image component C k The corresponding component Q Ck Encoded data.

[0263] Component Q C1 ~Q Ck They have mutually distinct image categories. For example, in a 3-component structure with k=3, component Q... C1 For visible light images, component Q C2 For infrared images, component Q C3 This is a thermal image.

[0264] In component Q C1 ~Q Ck In cases where there is correlation (e.g., infrared images and thermal images), it is also possible to compare image components C1 to C2. k Inter-reference image. On the other hand, in component Q C1 ~Q Ck In cases where there is no correlation (e.g., visible light images and distance images), it is also possible to use image components C1 to C2. k Do not refer to the image.

[0265] In the header area 41 of the bitstream BS, the component Q is stored. C1 ~Q Ck The associated parameter P was established. C1 ~P Ck The encoded data. Parameter P C1 ~P Ck Includes component Q C1 ~Q Ck Modal information D12.

[0266] Figure 26 , Figure 27 This is a diagram illustrating an example of the syntax related to the setting of modal information D12. Modal information D12 is represented as the value of the identifier mvi_modality_type[c] contained in SEI messages such as machine_vision_indication.

[0267] like Figure 27 As shown, when the value of mvi_modality_type[c] is 0, it indicates that the component Q... Ck The image category is visible light image. When mvi_modality_type[c] is 1, it indicates that the Q component... Ck The image category is thermal image. When the value of mvi_modality_type[c] is 2, it indicates that the Q component... Ck The image category is infrared image. When the value of mvi_modality_type[c] is 3, it indicates that the Q component... Ck The image category is LiDAR image. With mvi_modality_type[c] having a value of 4, it indicates that the Q component... Ck The image category is RADAR image. With mvi_modality_type[c] having a value of 5, it indicates that the Q component... Ck The image category is AI-generated image. If the value of `mvi_modality_type[c]` is other than [prepared], it indicates a pre-defined box intended for future use. Furthermore, the number of boxes can be increased or decreased based on the image category of the object. Figure 27 The number of images defined in the document.

[0268] Figure 28 The diagram is a simplified representation of the structure of the circuit 21 included in the decoding device 2.

[0269] The decoding unit 52 decodes component Q based on the payload region 42 of the multi-component structure bit stream BS input from the receiving unit 51. C1 ~Q Ck Decoding is performed. The decoding unit 52 outputs component Q.C1 ~Q Ck Image data D211~D21 k .

[0270] Furthermore, the decoding unit 52, based on the header region 41 of the multi-component structure bit stream BS input from the receiving unit 51, modifies the parameter P... C1 ~P Ck Decoding is performed. The decoding unit 52 extracts parameter P. C1 ~P Ck The modal information included is D221~D22 k .

[0271] The switching unit 53A is based on the modal information D221 to D22 input from the decoding unit 52. k For each component Q C1 ~Q Ck Switching Task Processing Units 541-54 n The switching unit 53A maintains multiple image categories and multiple task processing units 541-54. n The correspondence is defined in a pre-set table (illustration omitted). The switching unit 53A refers to this table information and selects each component Q... C1 ~Q Ck From multiple task processing units 541 to 54 n The selection is based on modal information D221~D22. k The image category represents a task processing unit 541-54. n .

[0272] Task Processing Department 541-54 n Based on image data D211~D21 input from decoding unit 52 via switching unit 53A k To perform task processing.

[0273] In addition, Figure 28 The diagram shows a structure where one image data is input to one task processing unit 54, but it can also be configured to input multiple image data to one task processing unit 54. Specifically, the switching unit 53A can also maintain multiple image categories and multiple task processing units 541-54. n The correspondence is established in a pre-defined table, in which multiple images Q are associated with one task processing unit 54.

[0274] In addition, depending on the image category, there are images composed of three image components, such as color visible light images, as well as images composed of only one image component, such as monochrome visible light images (visible light images with only brightness components), thermal images, infrared images, LiDAR images, or RADAR images (hereinafter referred to as "single-component images").

[0275] As an example, in the bitstream BS, there are 3 components Q. C1 ~Q C3 In the case of a color visible light image, the three components Q... C1 ~Q C3 All modal information D221 to D223 becomes a visible light image. On the other hand, in the case where the bitstream BS is composed of, for example, a monochromatic visible light image, an infrared image, and a thermal image, component Q... C1 The modal information D221 becomes a visible light image, and the component Q C2 The modal information D222 becomes an infrared image, and the component Q C3 The modal information D223 becomes a thermal image. In contrast, in the case where the bitstream BS consists, for example, only of infrared images, component Q... C1 The modal information D221 is converted into an infrared image, but the component Q is not processed. C2 Q C3 Assign modal information.

[0276] Encoding unit 33 will assign a component Q of modal information D221 C1 The image is encoded into a bitstream BS as usual. On the other hand, the encoding unit 33 encodes the unassigned modal information D222 to D22... k Other components Q C2 ~Q Ck The dummy image is encoded into the bitstream BS. The encoding process for dummy images can also be simpler than that for regular images.

[0277] In addition, the decoding unit 52 allocates a component Q of the modal information D221 to the bitstream BS. C1 Decoding is performed as a normal image. On the other hand, the decoding unit 52 decodes the unassigned mode information D222 to D22 according to the bitstream BS. k Other components Q C2 ~Q Ck Decoding is performed as a dummy image. The decoding process for a dummy image can be simpler than that for a regular image.

[0278] According to this variant, the encoding device 1 is able to encode multiple components Q of different image categories. C1 ~Q Ck The data is sent to decoding device 2, which can appropriately process the multiple components Q according to the bitstream BS. C1 ~Q Ck Decode it.

[0279] According to this modified example, the encoding device 1 is capable of encoding multiple components Q. C1 ~Q CkMultiple parameters P were established for correlation. C1 ~P Ck The header area 41 of the bitstream BS is encoded. Furthermore, the decoding device 2 can assign multiple components Q based on the header area 41 of the bitstream BS. C1 ~Q Ck Multiple parameters P were established for correlation. C1 ~P Ck Decode it.

[0280] According to this variation, in only one component Q C1 In the case of the required image category, for that one component Q C1 Assign modal information D221, thereby enabling the encoding device 1 to assign a component Q C1 After being properly encoded into the bitstream BS, the decoding device 2 can appropriately decode the component Q according to the bitstream BS. C1 Decode it.

[0281] According to this modified example, the encoding device 1 transmits the unassigned modal information D222~D22 k Other components Q C2 ~Q Ck Encoding the image as a dummy image reduces the processing load on encoding device 1. Furthermore, decoding device 2 uses the unassigned modal information D222~D22 k Other components Q C2 ~Q Ck Decoding the image as a dummy image can reduce the processing load on the decoding device 2.

[0282] Industrial availability

[0283] This disclosure is particularly useful for applications of image processing systems that include an encoding device for encoding an image into a bitstream and transmitting the bitstream, and a decoding device for decoding the image based on the received bitstream.

Claims

1. A decoding device, comprising: Circuit; and The memory is connected to the circuit. The circuit acquires the image from the bitstream and the parameters associated with the image. The parameters contain modal information representing the image category of the image.

2. The decoding device according to claim 1, wherein, The image categories include at least one of the following: visible light images, thermal images, infrared images, LiDAR images, RADAR images, or AI-generated images.

3. The decoding device according to claim 1, wherein, The circuit performs task processing based on the image. The circuit switches the task processing based on the modal information.

4. The decoding device according to claim 3, wherein, The task processing includes machine tasks.

5. The decoding device according to claim 3, wherein, The task processing includes human vision.

6. The decoding apparatus according to claim 1, wherein, The circuit performs task processing based on the image. The circuit switches between the AI ​​model or image processing used in the task processing based on the modal information.

7. The decoding device according to claim 1, wherein, The parameters also include transformation information for converting the pixel values ​​of the image into sensor data values, or converting the image into a sensed image.

8. The decoding apparatus according to claim 7, wherein, Based on the transformation information, the circuit transforms the pixel values ​​of the image into sensor data values, or transforms the image into the sensed image. The circuit performs task processing based on the sensor data values ​​or the sensed image.

9. The decoding apparatus according to claim 7, wherein, The circuit performs task processing based on the image and the transformation information.

10. The decoding apparatus according to claim 1, wherein, In acquiring the parameters, the circuit obtains the parameters from a given header region of the bitstream. The given header area contains either VUI or SEI.

11. The decoding apparatus according to claim 1, wherein, The bitstream has a multi-layer structure containing multiple image layers. In the acquisition of the image, the circuit acquires multiple images of different image categories from the multiple image layers.

12. The decoding apparatus according to claim 11, wherein, In acquiring the parameters, the circuit obtains multiple parameters associated with the multiple images from the header region of a specific image layer among the multiple image layers.

13. The decoding apparatus according to claim 11, wherein, In acquiring the parameters, the circuit obtains multiple parameters associated with the multiple images from multiple header regions of the multiple image layers.

14. The decoding apparatus according to claim 1, wherein, The image is segmented into multiple sub-images. In the acquisition of the image, the circuit acquires the plurality of sub-images of different image categories from the bitstream.

15. The decoding apparatus according to claim 14, wherein, In acquiring the parameters, the circuit obtains multiple parameters associated with the multiple sub-images from the header region of the image.

16. The decoding apparatus according to claim 1, wherein, The image is composed of multiple components. In the acquisition of the image, the circuit acquires the multiple components of different image categories from the bitstream.

17. The decoding apparatus according to claim 16, wherein, In acquiring the parameters, the circuit obtains multiple parameters associated with the multiple components from the header region of the image.

18. The decoding apparatus according to claim 16, wherein, The modal information is assigned to one of the plurality of components. In acquiring the image, the circuit acquires the component from the bitstream.

19. The decoding apparatus according to claim 18, wherein, In the absence of assigning modal information to other components that are different from the first component among the plurality of components, the circuit acquires the other components as a dummy image during the acquisition of the image.

20. The decoding apparatus according to claim 1, wherein, In the acquisition of the image, the circuit acquires multiple images of different image categories. The circuit also performs at least one task processing based on the plurality of images.

21. The decoding apparatus according to claim 20, wherein, The plurality of images includes visible light images. The at least one task processing involves human vision.

22. An encoding device comprising: Circuit; and The memory is connected to the circuit. The circuit encodes the image and the parameters associated with the image into a bitstream. The parameters contain modal information representing the image category of the image.

23. The encoding device according to claim 22, wherein, The image categories include at least one of the following: visible light images, thermal images, infrared images, LiDAR images, RADAR images, or AI-generated images.

24. The encoding device according to claim 22, wherein, The circuit also generates the image based on sensor data values ​​or sensed images.

25. The encoding device according to claim 24, wherein, The parameters also include transformation information for transforming the pixel values ​​of the image into the sensor data values, or transforming the image into the sensed image.

26. The encoding device according to claim 22, wherein, In the encoding of the parameters, the circuit encodes the parameters into a given header region of the bitstream. The given header area contains either VUI or SEI.

27. The encoding device according to claim 22, wherein, The bitstream has a multi-layer structure containing multiple image layers. In the encoding of the image, the circuit encodes multiple images of different image categories into the multiple image layers.

28. The encoding device according to claim 27, wherein, In the encoding of the parameters, the circuit encodes multiple parameters associated with the multiple images into the header region of a specific image layer among the multiple image layers.

29. The encoding device according to claim 27, wherein, In the encoding of the parameters, the circuit encodes multiple parameters associated with the multiple images into multiple header regions of the multiple image layers.

30. The encoding device according to claim 22, wherein, The image is segmented into multiple sub-images. In the encoding of the image, the circuit encodes the plurality of sub-images of different image categories into the bitstream.

31. The encoding device according to claim 30, wherein, In the encoding of the parameters, the circuit encodes multiple parameters associated with the multiple sub-images into the header region of the image.

32. The encoding device according to claim 22, wherein, The image is composed of multiple components. In the encoding of the image, the circuit encodes the multiple components of the image with different categories into the bit stream.

33. The encoding device according to claim 32, wherein, In the encoding of the parameters, the circuit encodes multiple parameters associated with the multiple components into the header region of the image.

34. The encoding device according to claim 32, wherein, The modal information is assigned to one of the plurality of components. In the encoding of the image, the circuit encodes the one component into the bit stream.

35. The encoding device according to claim 34, wherein, In the absence of assigning modal information to other components that are different from the first component among the plurality of components, the circuit encodes the other components as dummy images during the encoding of the image.

36. A decoding method, The decoding device acquires the image from the bitstream and the parameters associated with the image. The parameters contain modal information representing the image category of the image.

37. An encoding method, The encoding device encodes the image and the parameters associated with the image into a bitstream. The parameters contain modal information representing the image category of the image.

Citation Information

Patent Citations

  • Semi-supervised learning leveraging cross-domain data for medical imaging analysis

    US20240046453A1