Video compression system, video compression method, video decompression system, video decompression method, and program

The video compression and decompression systems control filters and quantization parameters block-by-block to address the challenge of optimizing image quality and bitrate in VLMs, achieving efficient compression and search performance.

WO2026116290A1PCT designated stage Publication Date: 2026-06-04NEC CORP

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NEC CORP
Filing Date
2025-11-25
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing video compression technologies struggle to finely control image quality and bitrate, particularly with Vision Language Models (VLMs), as they face a trade-off between resolution, data volume, and inference cost, making it difficult to optimize bitrate and quality, especially in low bitrate ranges.

Method used

A video compression system and method that employs a first filter group, an encoding unit, and a control network to control filters and quantization parameters block-by-block, and a video decompression system and method that controls second filters based on acquired information to enhance image quality and bitrate precision.

Benefits of technology

The system allows for precise control of image quality and bitrate, enabling high compression rates while maintaining recognition quality, especially for diverse and complex prompts, and reduces storage costs by optimizing video search systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025040915_04062026_PF_FP_ABST
    Figure JP2025040915_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention enables finer control over image quality and bit rate. This video compression system includes: a first filter group that includes one or more first filters to be applied to an input video; an encoding unit that encodes the video output from the first filter group; and a control network that, in accordance with the input video, outputs, to the first filter group, information for controlling the first filters to be applied to the video, for each of a plurality of blocks included in the video, and that, in accordance with the input video, outputs, to the encoding unit, information for controlling a quantization parameter in the encoding for each block.
Need to check novelty before this filing date? Find Prior Art

Description

Video Compression System, Video Compression Method, Video Expansion System, Video Expansion Method, and Program

[0001] The present disclosure relates to a video compression system, a video compression method, a video expansion system, a video expansion method, and a program.

[0002] As a related art, Patent Document 1 discloses a processing system. The processing system disclosed in Patent Document 1 preprocesses a part of a picture, that is, a block, using a filter before encoding a multimedia stream. In the processing system, a preprocessing module trains a model to predict the size and error of a block to be filtered by each filter based on the characteristics of the block. The preprocessing module uses this model to calculate the cost regarding the size and error when applying each filter to a predetermined block having specific characteristics. The preprocessing module applies the filter predicted to result in the best cost to the block. In Patent Document 1, the preprocessing module increases the compression rate of an image when the image is encoded using a predetermined quantization parameter with a predictable amount of quality loss.

[0003] Japanese Patent Translation of PCT International Publication No. 2022-534572

[0004] A Vision Language Model (VLM), which is a model for handling a composite of an image and text, is known. The VLM is used, for example, in video search. When a compressed video is used in video search or the like, more advanced compression control is considered desirable.

[0005] One object of the present disclosure is to provide a video compression system, a video compression method, a video expansion system, a video expansion method, and a program capable of more finely controlling image quality and bit rate.

[0006] A video compression system according to a first aspect of the present disclosure includes a first filter group including one or more first filters applied to an input video, an encoding unit that encodes the video output from the first filter group, and a control network that outputs information to the first filter group for controlling the first filters applied to each of a plurality of blocks contained in the video, according to the input video, and outputs information to the encoding unit for controlling the quantization parameters in the encoding for each of the blocks, according to the input video.

[0007] A video decompression system according to a second aspect of the present disclosure includes a decoding unit for decoding encoded video, a second filter group including one or more second filters applied to the decoded video, and a control unit that acquires information generated according to the encoded video for controlling a second filter applied to the decoded video for each of the multiple blocks included in the decoded video in the second filter group, and controls the second filter applied to the decoded video for each block based on the acquired information.

[0008] A third aspect of the present disclosure is a video compression method which includes a first filter group comprising one or more first filters, which applies the first filters to an input video, encodes the video output from the first filter group, controls the first filters applied to each of the multiple blocks contained in the video according to the input video, and controls the quantization parameters in the encoding for each block according to the input video.

[0009] A video decompression method according to a fourth aspect of the present disclosure involves decoding an encoded video, obtaining a second filter group including one or more second filters applied to the decoded video, and information for controlling the second filter group, which is generated according to the encoded video and includes one or more second filters applied to the decoded video, wherein the second filter group obtains information for controlling the second filter applied to the decoded video for each of a plurality of blocks included in the decoded video, and controls the second filter applied to the decoded video for each block based on the obtained information.

[0010] A program according to a fifth aspect of the present disclosure includes a first filter group comprising one or more first filters, which applies the first filters to an input video, encodes the video output from the first filter group, controls a control network, controls the first filters applied to each of the multiple blocks contained in the video according to the input video, and controls the quantization parameters in the encoding for each block according to the input video, causing the processor to perform these processes.

[0011] The video compression system, video compression method, video decompression system, video decompression method, and program related to this disclosure allow for more precise control of image quality and bitrate.

[0012] This is a block diagram showing a schematic configuration example of the video compression system related to this disclosure. This is a block diagram showing a schematic configuration example of the video decompression system related to this disclosure. This is a block diagram showing a configuration example of a video retrieval system including a compression device and a decompression device related to this disclosure. This is a block diagram showing a configuration example of a learning system used for learning a control network. This is a flowchart showing the operation procedure of the compression device. This is a flowchart showing the operation procedure of the decompression device. This is a block diagram showing a configuration example of a computer device.

[0013] Prior to describing the embodiments of this disclosure, the inventors will explain the matters they have considered. Currently, general VLMs are capable of recognition even at low bitrates because they perform inference at low resolution. In the future, it is conceivable that the resolution of the video handled by VLMs will increase, but there is a trade-off relationship between resolution or data volume and inference cost. On the other hand, in the low bitrate range, it is difficult to optimize bitrate and quality with gaze-area-based quantization parameter control alone. The inference accuracy of VLMs does not respond much to changes in quantization parameter values, but the bitrate responds sensitively to quantization parameter values. For this reason, it is difficult to finely control the quantization parameters, and it is difficult to optimize bitrate and quality. To lower the bitrate while maintaining the recognition quality of VLMs for diverse and complex prompts, advanced compression control is desired. Under these circumstances, the inventors have come up with the present invention.

[0014] The outline of this disclosure will now be explained. Figure 1 is a block diagram showing a schematic configuration example of a video compression system according to this disclosure. The video compression system 10 includes a first filter group 11, an encoding unit 12, and a control network 13.

[0015] The first filter group 11 includes one or more first filters applied to the input video. Here, "filter" can mean a functional unit that controls parameters or pixel values ​​for the video or frame images contained in the video. For example, the first filter may function as a functional unit that reduces the amount of information in the input video. The encoding unit 12 encodes the video output from the first filter group. The control network 13 outputs information to the first filter group 11 that controls the first filters applied to the video for each of the multiple blocks contained in the video, according to the input video. Here, "block" can mean each sub-region when the video or frame image is divided into multiple regions. The control network 13 also outputs information to the encoding unit 12 that controls the quantization parameters in encoding for each block, according to the input video.

[0016] In the video compression system 10 according to this disclosure, the first filter group 11 controls the first filter applied to the video block by block according to the input video, based on information output from the control network 13. The encoding unit 12 also controls the quantization parameters in encoding block by block according to the input video, based on information output from the control network 13. In this disclosure, since the first filter applied to the video can be controlled on a location-by-location basis, image quality and bitrate can be controlled more precisely.

[0017] Figure 2 is a block diagram illustrating a schematic configuration example of the video decompression system according to the present disclosure. The video decompression system 20 includes a decoding unit 21, a second filter group 22, and a control unit 23.

[0018] The decoding unit 21 decodes the encoded video. The second filter group 22 includes one or more second filters applied to the decoded video. The control unit 23 acquires information for controlling the second filter applied to the decoded video in the second filter group 22 for each of the multiple blocks contained in the decoded video that are generated according to the encoded video. Based on the acquired information, the control unit 23 controls the second filter applied to the decoded video for each block.

[0019] In the video decompression system 20 according to this disclosure, the second filter group 22 controls the second filter to be applied to the decoded video block by block, according to the video being decoded, based on the information acquired by the control unit 23. The encoded video has a first filter applied to it that is controlled on a location-by-location basis, for example, using the video compression system 10 described above. In this disclosure, since the second filter applied to the decoded video can be controlled on a location-by-location basis, the image quality and bitrate of the decoded video can be controlled more precisely.

[0020] The embodiments of this disclosure will be described in detail below with reference to the drawings. Note that the following description and drawings have been omitted and simplified as appropriate for clarity of explanation. Furthermore, in the following drawings, the same elements and similar elements are denoted by the same reference numerals, and redundant explanations have been omitted where necessary.

[0021] Figure 3 is a block diagram showing an example configuration of a video search system including a video compression system and a video decompression system according to the present disclosure. An embodiment will be described using Figure 3. The video search system 100 shown in Figure 3 has a compression device 110 and a video search device 130.

[0022] The video search system 100 can be used, for example, to search for desired video footage from video footage collected from multiple vehicles. For example, multiple vehicles each transmit driving video footage, which is video footage taken using cameras mounted on the vehicles, to a video collection device (not shown) via a wireless network. The video footage transmitted from each vehicle is divided into multiple video clips 200 according to predetermined criteria. The compression device 110 encodes the video clips 200 using a predetermined encoding scheme and stores the encoded video in the video storage unit 170. The video storage unit 170 uses high-speed storage such as a solid-state drive (SSD). The video search device 130 searches for desired video footage from the video footage stored in the video storage unit 170 in response to a prompt entered by the user. Here, the prompt can mean an instruction or search query entered by the user.

[0023] The compression device 110 includes a control network 111, a prefilter 112, and an encoder 113. Physically, the compression device 110 may be configured as a device having one or more memories and one or more processors. In the compression device 110, at least some of the functions of each part within the compression device 110 can be realized by one or more processors executing processing according to instructions read from one or more memories. The prefilter 112 and the encoder 113 may each be configured by dedicated hardware circuits. The compression device 110 corresponds to the video compression system 10 shown in Figure 1.

[0024] The prefilter 112 includes one or more filters applied to the video clip 200 encoded by the encoder 113. The frame image of the video clip 200 includes, for example, a plurality of blocks divided into a mesh-like structure with predetermined pixels. The prefilter 112 is configured to allow control over which filters are applied to each of the plurality of blocks included in the frame image of the video clip 200. The prefilter 112 includes, for example, a filter that controls the color of the frame image. The prefilter 112 is configured to allow control over the degree to which color is retained for each location in the frame image. More specifically, the prefilter 112 is configured to allow control over whether the values ​​of U and V are set to zero for each block when the pixel values ​​R'G'B' of each pixel in the frame image are converted to Y'UV. The prefilter 112 corresponds to the first filter group 11 shown in Figure 1.

[0025] The encoder 113 encodes the filtered video output from the pre-filter 112 while compressing it using a predetermined encoding method. The encoder 113 is configured so that the quantization parameters can be controlled for each block of the frame image. The encoder 113 stores the compressed and encoded video in the video storage unit 170. The encoder 113 may distribute the compressed and encoded video to the video retrieval device 130 via a network. The encoder 113 corresponds to the encoding unit 12 shown in Figure 1.

[0026] The control network 111 generates information to control the prefilter 112 and the encoder 113 according to the input video clip 200. In this embodiment, the control network 111 determines information to control the filter applied to the frame image in the prefilter 112 for each block. The control network 111 also determines information to control the quantization parameters in encoding for each block. The control network 111 outputs the information to control the prefilter 112 to the prefilter 112. The control network 111 also outputs information to control the encoder 113 to the encoder 113.

[0027] Furthermore, the control network 111 generates information to control the post-filter 152 of the decompression device 150 according to the input video clip 200. In this embodiment, the control network 111 determines information to control the filter applied to the frame image in the post-filter 152 for each block. The control network 111 stores the generated information to control the post-filter 152 in the video storage unit 170. The control network 111 may also distribute the information to control the post-filter 152 to the video retrieval device 130 via the network. The control network 111 includes a neural network such as a Transformer. The control network 111 is trained, for example, by reinforcement learning. The control network 111 corresponds to the control network 13 shown in Figure 1.

[0028] The video retrieval device 130 includes a recognition unit 132 and a decompression device 150. The decompression device 150 includes a decoder 151, a post-filter 152, and a control unit 153. Physically, the video retrieval device 130 and the decompression device 150 may each be configured as devices having one or more memories and one or more processors. In the video retrieval device 130 and the decompression device 150, at least some of the functions of each part in the video retrieval device 130 and at least some of the functions of each part in the decompression device 150 can be realized by one or more processors executing processing according to instructions read from one or more memories. In the decompression device 150, the decoder 151 and the post-filter 152 may each be configured by dedicated hardware circuits. The decompression device 150 may be included in the video retrieval device 130 or may be configured as a device independent of the video retrieval device 130. The decompression device 150 corresponds to the video decompression system 20 shown in Figure 2.

[0029] The decoder 151 decompresses or decodes the video compressed by the compressor 110. The post-filter 152 includes one or more filters applied to the decoded video, i.e., the video decompressed by the decoder 151. The post-filter 152 is configured so that the applied filter can be controlled for each of the multiple blocks contained in the frame image. The pre-filter 112 includes, for example, an image sharpening filter. The post-filter 152 is configured so that the degree of sharpening can be controlled for each location in the frame image. The decoder 151 corresponds to the decoding unit 21 shown in Figure 2. The post-filter 152 corresponds to the second filter group 22 shown in Figure 2.

[0030] The control unit 153 acquires control information for the post-filter 152 generated in the control network 111 from the video storage unit 170. According to the acquired control information, the control unit 153 controls the filters applied to the decoded video block by block in the post-filter. The control unit 153 corresponds to the control unit 23 shown in Figure 2.

[0031] The search prompt 131 is a prompt entered by the user during a video search. The search prompt 131 is written in natural language, for example. The search prompt 131 includes words that specify objects and situations to be recognized by the recognition unit 132. The recognition unit 132 acquires video stored in the video storage unit 170 via the expansion device 150 and generates recognition results for the acquired video. Based on the recognition results, the recognition unit 132 extracts video from the video storage unit 170 that corresponds to the objects and situations specified in the search prompt 131. The recognition unit 132 includes an artificial intelligence model, such as a VLM.

[0032] Figure 4 is a block diagram showing an example configuration of a learning system used for training the control network 111. The learning system 300 includes a pre-filter 301, an encoder 302, a decoder 303, a post-filter 304, a control network 305, a recognition unit 306, a recognition unit 307, an error acquisition unit 308, a data quantity acquisition unit 309, and a reward calculation unit 310. Physically, the learning system 300 can be configured as a device having one or more memories and one or more processors. In the learning system 300, at least some of the functions of each part within the learning system 300 can be realized by one or more processors executing processing according to instructions read from one or more memories.

[0033] The prefilter 301 includes one or more filters applied to the training video clip 330 encoded by the encoder 302. The prefilter 301 is configured so that the applied filter can be controlled for each of the multiple blocks contained in the frame image of the video clip. The encoder 302 encodes the filtered video output from the prefilter 301 while compressing it using a predetermined encoding scheme. The encoder 302 is configured so that the quantization parameters can be controlled for each block of the frame image.

[0034] The decoder 303 decompresses the video compressed by the encoder 302. The post-filter 304 includes one or more filters applied to the decoded video, i.e., the video decompressed by the decoder 303. The post-filter 304 is configured so that the applied filter can be controlled for each of the multiple blocks contained in the frame image. The pre-filter 301, encoder 302, decoder 303, and post-filter 304 correspond to the pre-filter 112, encoder 113, decoder 151, and post-filter 152 in the video retrieval system 100 shown in Figure 3, respectively. The pre-filter 301, encoder 302, decoder 303, and post-filter 304 may each be configured by dedicated hardware circuits.

[0035] The control network 305 generates information to control the pre-filter 301, encoder 302, and post-filter 304 according to the training video clip 330. The control network 305 includes, for example, a neural network such as a Transformer. The control network 305 determines, block by block, information to control the filter applied to the frame image in the pre-filter 301. The control network 305 also determines, block by block, information to control the quantization parameters in the encoding of the encoder 302. Furthermore, the control network 305 determines, block by block, information to control the filter applied to the frame image in the post-filter 304. The control network 305 outputs the information to control the pre-filter 301 to the pre-filter 301. The control network 111 also outputs information to control the encoder 302 to the encoder 302. The control network 111 also outputs information to control the post-filter 304 to the post-filter 304.

[0036] The learning prompt 350 includes a plurality of prompts used for training the control network 305. The learning prompt 350 includes a plurality of prompts that are assumed to be search prompts 131 input to the recognition unit 132 in the video retrieval device 130. The learning prompt 350 may be generated in response to the learning video clip 330, for example using generated artificial intelligence (AI). The learning prompt 350 may include a variety of prompts or complex prompts that may be possible for the learning video clip 330.

[0037] The recognition unit 306 generates recognition results for the learning video clip 330 based on the learning video clip 330 and the learning prompts 350. For example, the recognition unit 306 inputs the learning video clip 330 to the VLM and outputs the responses to multiple learning prompts 350 output from the VLM as recognition results.

[0038] The recognition unit 307 generates recognition results for the expanded video based on the expanded video output from the post-filter 304 and the learning prompts 350. For example, the recognition unit 307 inputs the expanded video to the VLM and outputs the responses to multiple learning prompts 350 output from the VLM as recognition results. Recognition units 306 and 307 correspond to the recognition unit 132 in the video search system 100 shown in Figure 3.

[0039] The error acquisition unit 308 acquires the error between the recognition result output from the recognition unit 306 and the recognition result output from the recognition unit 307. In other words, the error acquisition unit 308 uses the recognition result output from the recognition unit 306 as the correct answer data, and acquires the difference between the correct answer data and the recognition result output from the recognition unit 307 as the error. The error acquisition unit 308 uses an evaluation framework such as G-Eval to acquire the difference between the recognition result output from the recognition unit 306 and the recognition result output from the recognition unit 307. For example, the error acquisition unit 308 acquires the error of the response to the same prompt for each of the multiple learning prompts 350. The error acquisition unit 308 acquires the average of the acquired errors as the error for one learning video clip 330.

[0040] The error acquisition unit 308 may acquire video features of the VLM from the recognition unit 306 and the recognition unit 307, respectively, and acquire errors based on the acquired video features. For example, the error acquisition unit 308 may calculate the mean absolute error / (square root of mean squared error) of the video features for each patch, i.e., for each location, and acquire the calculated mean absolute error / (square root of mean squared error) as the error for each patch.

[0041] The data acquisition unit 309 acquires the data amount of the video compressed by the encoder 302. The reward calculation unit 310 calculates the reinforcement learning reward based on the error acquired by the error acquisition unit 308 and the data amount acquired by the data acquisition unit 309. For example, the reward calculation unit 310 calculates -(error + λ data amount) as the reward, where λ is any real number greater than 0. The control network 305 is learned by reinforcement learning by backpropagating the reward. The reward calculation unit 310 may calculate the reward for each patch based on the error of the VLM video features acquired for each patch, and backpropagate the calculated reward for each patch. The learned control network 305 is used as the control network 111 in the video retrieval system 100 shown in Figure 3.

[0042] Figure 5 is a flowchart showing the operation procedure of the compression device 110. The operation procedure of the compression device 110 corresponds to the video compression method. The compression device 110 acquires a video clip 200 (step A1). The control network 111 includes a neural network trained using, for example, the learning system 300 shown in Figure 4. The control network 111 determines the filter control in the pre-filter 112 and post-filter 152, and the quantization parameter values ​​in the encoder 113, block by block, according to the acquired video clip 200 (step A2).

[0043] The control network 111 outputs information to the prefilter 112 that controls the prefilter 112. The prefilter 112 applies a filter to the video clip 200 block by block based on the information transmitted from the control network 111 (step A3). For example, the control network 111 outputs information to the prefilter 112 that shows the distribution of the degree to which colors are retained for each location, i.e., block by block. In this case, the prefilter 112 controls the degree to which colors are retained for each block according to the information transmitted from the control network 111.

[0044] The control network 111 outputs information for controlling the encoder 113 to the encoder 113. Based on the information transmitted from the control network 111, the encoder 113 encodes, that is, compresses, the video clip to which the filter has been applied in the prefilter 112 (step A4). For example, the control network 111 outputs information indicating the distribution of quantization parameter values for each location, that is, for each block, to the encoder 113. In that case, the encoder 113 compresses each block of the video clip 200 with the quantization parameter value of each block according to the information transmitted from the control network 111.

[0045] The encoder 113 stores the encoded video clip 200, that is, the compressed video clip 200, in the video storage unit 170 (step A5). Also, the control network 111 stores filter control information, that is, information for controlling the postfilter 152 of the decompression device 150, in the video storage unit 170 (step A6). For example, the control network 111 stores information indicating the distribution of the degree of image sharpening for each location, that is, for each block, in the video storage unit 170.

[0046] FIG. 6 is a flowchart showing the operation procedure of the decompression device 150. The operation procedure of the decompression device 150 corresponds to the video decompression method. The decoder 151 acquires the compressed video clip from the video storage unit 170. The decoder 151 decodes, that is, decompresses, the compressed video clip (step B1).

[0047] The control unit 153 acquires the filter control information of the post-filter 152 from the video memory unit 170 (step B2). The control unit 153 outputs the filter control information to the post-filter 152. The post-filter 152 applies a filter to the extended video clip block by block according to the filter control information (step B3). For example, the video memory unit 170 stores information indicating the distribution of the degree of image sharpening for each block during video compression. The control unit 153 outputs the filter control information acquired from the video memory unit 170 to the post-filter 152. The post-filter 152 controls the degree of image sharpening for each block according to the filter control information. The video clip for which filter control is performed by the post-filter 152 is used for video search in the video search device 130.

[0048] In this embodiment, the control network 111 generates information for controlling the pre-filter 112, the encoder 113, and the post-filter 152 block by block. The control network 111 is trained in the learning system 300 so that the recognition error associated with video compression and the amount of data after compression are reduced. In this embodiment, the learned control network 111 is used to control the pre-filter 112 and the encoder 113 block by block in video compression. Also, the post-filter 152 applied after video decompression is controlled block by block using the information generated using the control network 111. In this embodiment, the degree of each of the pre-compression image processing and the post-decompression image processing can be controlled for each location, and finer control of image quality and bit rate can be achieved. In this embodiment, the video can be compressed at a high compression rate while suppressing the recognition error associated with compression. Therefore, while realizing good video search in the video search device 130, the storage cost of the video memory unit 170 can be suppressed.

[0049] In the above embodiment, an example in which the video compressed and stored in the video memory unit 170 is used for video search has been described. However, the present disclosure is not limited to this. The video compressed in the compression device 110 can also be used for other purposes.

[0050] Next, the physical configuration of the compression device 110, the decompression device 150, and the learning system 300 will be described. Figure 7 is a block diagram showing an example configuration of a computer device that can be used as the compression device 110, the decompression device 150, or the learning system 300. The computer device 500 has a processor 510 such as a Central Processing Unit (CPU), a storage unit 520, a Read Only Memory (ROM) 530, a Random Access Memory (RAM) 540, a communication interface (IF) 550, and a user interface 560.

[0051] The communication interface 550 is an interface for connecting the computer device 500 to a communication network via wired communication means or wireless communication means. The user interface 560 includes a display unit, such as a display. The user interface 560 also includes input units such as a keyboard, mouse, and touch panel.

[0052] The memory unit 520 is an auxiliary storage device capable of holding various types of data. The memory unit 520 does not necessarily have to be part of the computer device 500; it may be an external storage device or cloud storage connected to the computer device 500 via a network.

[0053] ROM 530 is a non-volatile memory device. For example, a semiconductor memory device such as a relatively small-capacity flash memory is used for ROM 530. The program executed by the processor 510 can be stored in the storage unit 520 or ROM 530. The storage unit 520 or ROM 530 stores various programs that realize the functions of the compression device 110, the decompression device 150, or the learning system 300.

[0054] The above program, when loaded into a computer, includes a set of instructions (or software code) for causing the computer to perform one or more of the functions described in the embodiments. The program may be stored in a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include RAM, ROM, flash memory, SSD or other memory technologies, Compact Disc (CD), digital versatile disc (DVD), Blu-ray® disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include, a temporary computer-readable medium or a communication medium that includes electrically, optically, acoustically, or otherwise propagating signals.

[0055] The RAM 540 is a volatile memory device. Various semiconductor memory devices such as Dynamic Random Access Memory (DRAM) or Static Random Access Memory (SRAM) can be used for the RAM 540. The RAM 540 may be used as an internal buffer for temporarily storing data, etc. The processor 510 loads a program stored in the memory unit 520 or ROM 530 into the RAM 540 and executes it. By executing the program, the functions of the compression device 110, the decompression device 150, or each part of the learning system 300 can be realized. The processor 510 may have an internal buffer that can temporarily store data, etc.

[0056] In the above embodiment, the compression device 110 and the decompression device 150 do not necessarily have to be configured as a single computer device. The compression device 110 and the decompression device 150 may each be configured using multiple physically separated devices. Similarly, the learning system 300 may also be configured using multiple physically separated devices.

[0057] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be made as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0058] Each drawing is merely illustrative to illustrate one or more embodiments. Each drawing may be associated with one or more other embodiments, rather than being associated with only one specific embodiment. As those skilled in the art will understand, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings, for example, to create embodiments not explicitly shown or described. Not all features or steps shown in any one drawing to illustrate an exemplary embodiment are necessarily required, and some features or steps may be omitted. The order of steps described in any of the drawings may be changed as appropriate.

[0059] Some or all of the above embodiments may also be described as follows, but are not limited to the following:

[0060] [Note 1] A video compression system comprising: a first filter group including one or more first filters applied to an input video; an encoding unit that encodes the video output from the first filter group; and a control network that outputs information to the first filter group for controlling the first filters applied to each of a plurality of blocks contained in the video, according to the input video, and outputs information to the encoding unit for controlling the quantization parameters in the encoding for each of the blocks, according to the input video.

[0061] [Note 2] The video compression system according to Note 1, wherein the control network is information for controlling a second filter group including one or more second filters applied to the decoded video obtained by decoding the encoded video, and further outputs information for controlling the second filters applied to the decoded video for each of the multiple blocks included in the decoded video.

[0062] [Note 3] The video compression system according to Note 2, wherein the encoded video and information for controlling the second group of filters are stored in a video storage unit.

[0063] [Note 4] The control network is generated by learning the recognition results for a plurality of prompts obtained by inputting decoded video obtained by decoding the encoded video into a recognition unit that performs recognition processing according to a prompt, an error index value indicating the error between the recognition results for the plurality of prompts and the correct recognition results for the plurality of prompts, and the amount of data of the encoded video, as described in any one of Notes 1 to 3.

[0064] [Note 5] The video compression system described in Note 4, wherein the correct recognition result is the recognition result for the plurality of prompts obtained by inputting the video before encoding into the recognition unit.

[0065] [Appendix 6] The video compression system according to Appendix 4 or 5, wherein the control network is learned based on a reward calculated according to the sum of the error index value and the amount of data.

[0066] [Note 7] A video decompression system comprising: a decoding unit for decoding encoded video; a second filter group including one or more second filters applied to the decoded video; and a control unit that acquires information generated according to the encoded video for controlling the second filter applied to the decoded video for each of the multiple blocks included in the decoded video in the second filter group, and controls the second filter applied to the decoded video for each block based on the acquired information.

[0067] [Note 8] The video decompression system according to Note 7, wherein the encoded video and the information for controlling the second filter group are stored in the video storage unit, and the control unit acquires the information for controlling the second filter group from the video storage unit.

[0068] [Note 9] The video decompression system according to Note 7 or 8, wherein the information controlling the second group of filters is generated using a control network that outputs to the first group of filters information to control the first filter applied to the video to be encoded for each of a plurality of blocks included in the video, according to the video to be encoded, and outputs to the encoding unit information to control the quantization parameters in the encoding for each of the blocks, according to the video to be encoded.

[0069] [Note 10] The control network is generated by learning the error index value indicating the error between the recognition results for a plurality of prompts and the correct recognition results for the plurality of prompts, and the amount of data of the encoded video, which are obtained by inputting the decoded video obtained by decoding the encoded video and applying the second filter to a recognition unit that performs recognition processing according to a prompt, and the recognition results for the correct answers for the plurality of prompts, as described in Note 9.

[0070] [Note 11] The video decompression system according to Note 10, wherein the correct recognition result is the recognition result for the plurality of prompts obtained by inputting the video before encoding into the recognition unit.

[0071] [Note 12] A video compression method comprising: a first filter group including one or more first filters, applying the first filters to an input video; encoding the video output from the first filter group; controlling a control network, controlling the first filters applied to each of the multiple blocks contained in the video according to the input video, and controlling the quantization parameters in the encoding for each block according to the input video.

[0072] [Note 13] A video decompression method comprising: decoding encoded video; a second filter group including one or more second filters applied to the decoded video; information for controlling the second filter group including one or more second filters applied to the decoded video, which is generated in accordance with the encoded video, wherein the second filter group controls the second filter applied to the decoded video for each of the plurality of blocks included in the decoded video; and controlling the second filter applied to the decoded video for each block based on the acquired information.

[0073] [Note 14] A program that causes a processor to perform the following processes: applying the first filters to an input video in a first filter group including one or more first filters; encoding the video output from the first filter group; controlling a control network; controlling the first filters applied to each of the multiple blocks contained in the video according to the input video; and controlling the quantization parameters in the encoding for each block according to the input video.

[0074] [Note 15] A program that decodes encoded video, includes a second filter group including one or more second filters applied to the decoded video, and controls the second filter group including one or more second filters applied to the decoded video, wherein the second filter group obtains information that controls the second filter applied to the decoded video for each of a plurality of blocks included in the decoded video, and causes the processor to execute a process to control the second filter applied to the decoded video for each block based on the obtained information.

[0075] Some or all of the elements (e.g., configuration and function) described in Appendices 2 through 6 that are dependent on Appendice 1 may also be dependent on Appendices 12 and 14 in the same way as those described in Appendices 2 through 6. Similarly, some or all of the elements (e.g., configuration and function) described in Appendices 8 through 11 that are dependent on Appendice 7 may also be dependent on Appendices 13 and 15 in the same way as those described in Appendices 8 through 11. Some or all of the elements described in any appendice may be applied to various hardware, software, recording means, systems, and methods for recording software.

[0076] This application claims priority based on Japanese Patent Application No. 2024-206330, filed on 27 November 2024, and incorporates all of its disclosures herein.

[0077] 10: Video compression system 11: First filter group 12: Encoding unit 13: Control network 20: Video decompression system 21: Decoding unit 22: Second filter group 23: Control unit 100: Video search system 110: Compressor 111: Control network 112: Pre-filter 113: Encoder 130: Video search device 131: Search prompt 132: Recognition unit 150: Decompressor 151: Decoder 152: Post-filter 153: Control unit 170: Video storage unit 200: Video clip 300: Learning system 301: Pre-filter 302: Encoder 303: Decoder 304: Post-filter 305: Control network 306, 307: Recognition unit 308: Error acquisition unit 309: Data volume acquisition unit 330: Learning video clip 350: Learning prompt 500: Computer device 510: Processor 520: Memory unit 530: ROM 540: RAM 550: Communication interface 560: User interface

Claims

1. A video compression system comprising: a first filter group including one or more first filters applied to an input video; an encoding unit that encodes the video output from the first filter group; and a control network that outputs information to the first filter group for controlling the first filters applied to each of a plurality of blocks contained in the video, according to the input video, and outputs information to the encoding unit for controlling the quantization parameters in the encoding for each of the blocks, according to the input video.

2. The video compression system according to claim 1, wherein the control network provides information for controlling a second filter group, which includes one or more second filters applied to a decoded video obtained by decoding the encoded video, and further outputs information for controlling a second filter applied to the decoded video for each of a plurality of blocks included in the decoded video.

3. The video compression system according to claim 2, wherein the encoded video and information for controlling the second group of filters are stored in a video storage unit.

4. The video compression system according to any one of claims 1 to 3, wherein the control network is generated by learning an error index value indicating the error between the recognition results for a plurality of prompts obtained by inputting a decoded video obtained by decoding the encoded video into a recognition unit that performs recognition processing according to a prompt, and the amount of data of the encoded video.

5. The video compression system according to claim 4, wherein the correct recognition result is the recognition result for the plurality of prompts obtained by inputting the video before encoding into the recognition unit.

6. The video compression system according to claim 4 or 5, wherein the control network is learned based on a reward calculated according to the sum of the error index value and the amount of data.

7. A video decompression system comprising: a decoding unit for decoding encoded video; a second filter group including one or more second filters applied to the decoded video; and a control unit that acquires information generated according to the encoded video for controlling the second filter applied to the decoded video for each of the multiple blocks included in the decoded video in the second filter group, and controls the second filter applied to the decoded video for each block based on the acquired information.

8. The video decompression system according to claim 7, wherein the encoded video and information for controlling the second group of filters are stored in a video storage unit, and the control unit obtains information for controlling the second group of filters from the video storage unit.

9. The video decompression system according to claim 7 or 8, wherein the information for controlling the second group of filters is generated using a control network that outputs to the first group of filters information to control the first filter applied to the video for each of a plurality of blocks contained in the video, according to the video being encoded, and outputs to the encoding unit information to control the quantization parameters in the encoding for each of the blocks, according to the video being encoded.

10. The video decompression system according to claim 9, wherein the control network is generated by learning an error index value indicating the error between the recognition results for a plurality of prompts obtained by inputting a decoded video obtained by decoding the encoded video and applying the second filter to a recognition unit that performs recognition processing according to a prompt, and the amount of data of the encoded video.

11. The video decompression system according to claim 10, wherein the correct recognition result is the recognition result for the plurality of prompts obtained by inputting the video before encoding into the recognition unit.

12. A video compression method comprising: a first filter group including one or more first filters, applying the first filters to an input video; encoding the video output from the first filter group; controlling a control network, controlling the first filters applied to each of the multiple blocks contained in the video according to the input video, and controlling the quantization parameters in the encoding for each block according to the input video.

13. A video decompression method comprising: decoding encoded video; a second filter group including one or more second filters applied to the decoded video; information for controlling the second filter group including one or more second filters applied to the decoded video, generated according to the encoded video, wherein the second filter group controls the second filter applied to the decoded video for each of the multiple blocks included in the decoded video; and controlling the second filter applied to the decoded video for each block based on the acquired information.

14. A program that causes a processor to perform the following processes: applying the first filters to an input video in a first filter group including one or more first filters; encoding the video output from the first filter group; controlling a control network; controlling the first filters applied to each of the multiple blocks contained in the video according to the input video; and controlling the quantization parameters in the encoding for each block according to the input video.

15. A program that decodes encoded video, includes a second filter group including one or more second filters applied to the decoded video, and controls the second filter group including one or more second filters applied to the decoded video, wherein the second filter group obtains information that controls the second filter applied to the decoded video for each of a plurality of blocks included in the decoded video, and causes a processor to execute a process to control the second filter applied to the decoded video for each block based on the obtained information.