Video processing system, video processing method, and program
The video processing system optimizes compression parameters using an estimation model to balance recognition accuracy and data volume, addressing the challenge of high compression degrading quality and accuracy in existing systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2025-11-18
- Publication Date
- 2026-06-04
AI Technical Summary
Existing video processing systems face challenges in achieving high recognition accuracy while reducing the data volume through video compression, as high compression ratios degrade video quality and recognition accuracy, and optimal parameter settings for balancing these factors are unclear.
A video processing system that includes a compression/decompression unit, recognition unit, error calculation unit, and model generation unit to estimate data volume and error index values, allowing for video compression that maintains recognition accuracy by determining optimal compression parameters using an estimation model.
The system effectively balances recognition accuracy and data volume reduction by optimizing compression parameters, reducing the amount of data while maintaining high recognition accuracy.
Smart Images

Figure JP2025040214_04062026_PF_FP_ABST
Abstract
Description
Image processing system, image processing method, and program
[0001] This disclosure relates to an image processing system, an image processing method, and a program.
[0002] As a related technology, Patent Document 1 discloses an image transmission system. In the image transmission system described in Patent Document 1, the imaging device captures images at a predetermined frame period and transmits the moving image data to the edge device. The edge device encodes the moving image data and transmits the encoded data to the server device.
[0003] The server device decodes the encoded data received from the edge device and generates a decoded image. The server device inputs the decoded image into a trained model and performs recognition processing on the decoded image. Based on the results of the recognition processing, the server device calculates the impact of each region of the decoded image on the recognition accuracy. The server device aggregates the impact of each region of the decoded image on the recognition accuracy for each encoded block. Based on the aggregated results, the server device generates a compression ratio map.
[0004] The server device transmits the generated compression ratio map to the edge device. The edge device uses the compression ratio map to encode each frame of the video data.
[0005] International Publication No. 2024 / 047734
[0006] Generally, increasing the compression ratio of video data reduces the amount of data, allowing more video data to be stored on storage. However, high compression ratios degrade the video quality of the decoded video, and the recognition accuracy of the recognition process performed on the decoded video deteriorates. If video is compressed with certain parameter settings, overall, either recognition accuracy is maintained but the data volume is large, or the data volume is reduced but recognition accuracy deteriorates. When performing recognition processing using decoded video, it is unclear how to set the video compression parameters for each video to reduce the amount of data while maintaining high recognition accuracy.
[0007] One of the purposes of this disclosure is to provide a video processing system, video processing method, and program that can achieve video compression that is compatible with maintaining a high level of recognition accuracy and reducing the amount of data.
[0008] A video processing system according to a first aspect of the present disclosure includes: a compression / decompression unit that compresses a video clip based on video compression parameters and decompresses the compressed video clip; a recognition unit that receives the decompressed video clip as input and outputs a recognition result of the input video clip; an error calculation unit that compares the recognition result of the decompressed video clip with a correct recognition result and calculates an error index value; and a model generation unit that generates a model for estimating the data amount of the compressed video clip and the error index value based on the video clip and the video compression parameters, wherein the model generation unit generates the model such that the difference between the data amount of the compressed video clip and the estimated data amount is small.
[0009] A video processing method according to a second aspect of the present disclosure compresses a video clip based on video compression parameters, decompresses the compressed video clip, generates a recognition result for the decompressed video clip, compares the generated recognition result for the decompressed video clip with a correct recognition result, calculates an error index value based on the result of the comparison, and generates a model such that the difference between the data amount of the compressed video clip estimated using a model that estimates the data amount of the compressed video clip and the error index value based on the video clip and the video compression parameters is reduced.
[0010] A program according to a third aspect of this disclosure compresses a video clip based on video compression parameters, decompresses the compressed video clip, generates a recognition result for the decompressed video clip, compares the generated recognition result for the decompressed video clip with a correct recognition result, calculates an error index value based on the result of the comparison, and causes a computer to perform the process of generating a model such that the difference between the data amount of the compressed video clip estimated using a model that estimates the data amount of the compressed video clip and the error index value based on the video clip and the video compression parameters is reduced.
[0011] The video processing system, video processing method, and program related to this disclosure are capable of achieving video compression that balances recognition accuracy and data volume reduction.
[0012] This is a block diagram showing a schematic configuration example of the video processing system related to this disclosure. This is a block diagram showing a configuration example of the video processing system related to this disclosure. This is a block diagram showing a configuration example of the learning system used to generate the estimation model. This is a schematic diagram illustrating the determination of video compression parameters. This is a flowchart showing the operation procedure of the learning system. This is a flowchart showing the operation procedure of the compression device. This is a block diagram showing a configuration example of the computer device.
[0013] Prior to describing embodiments of this disclosure, an overview of this disclosure will be given. Figure 1 is a block diagram showing an example of a schematic configuration of an image processing system according to this disclosure. The image processing system 10 includes a compression / decompression unit 11, a recognition unit 12, an error calculation unit 13, and a model generation unit 14.
[0014] The compression / decompression unit 11 compresses the video clip based on video compression parameters. Here, a video clip can mean, for example, the entirety of a set of video footage, or a portion of a set of video footage. The compression / decompression unit 11 also decompresses the compressed video clip. Decompressing a video clip is synonymous with decompressing a video clip. The recognition unit 12 receives the video clip decompressed by the compression / decompression unit 11 as input and outputs the recognition result of the input video clip. The error calculation unit 13 compares the recognition result of the decompressed video clip output from the recognition unit 12 with the correct recognition result. The error calculation unit 13 calculates an error index value based on the comparison result.
[0015] The model generation unit 14 generates a model that estimates the data volume and error index value of a compressed video clip based on the video clip and video compression parameters. The model generation unit 14 generates the model such that the difference between the data volume of the compressed video clip and the data volume estimated using the model is small. The model generation unit 14 also generates the model such that the difference between the error index value when recognition is performed using the compressed video clip and the error index value estimated using the model is small.
[0016] In this embodiment, the model generation unit 14 generates a model that estimates the data volume of the compressed video clip and the recognition error index value from the video clip to be compressed and the video compression parameters. By using the generated model, it is possible to estimate the data volume after compression and the recognition accuracy caused by the compression when compressing a video clip. By determining the video compression parameters using the recognition accuracy and data volume estimated using the model, it is possible to achieve video compression that reduces the amount of data while maintaining a recognition accuracy at the same level as the recognition accuracy of the video before compression.
[0017] The embodiments of this disclosure will be described in detail below with reference to the drawings. Note that the following description and drawings have been omitted and simplified as appropriate for clarity of explanation. Furthermore, in the following drawings, the same elements and similar elements are denoted by the same reference numerals, and redundant explanations have been omitted where necessary.
[0018] FIG. 2 is a block diagram showing a configuration example of a video processing system according to the present disclosure. In FIG. 2, the video processing system is configured as a video search system 100. An embodiment will be described using FIG. 2. The video search system 100 shown in FIG. 2 includes a compression device 110 and a video search device 130.
[0019] The video search system 100 can be used, for example, for a purpose of searching a desired driving video from driving videos collected from a plurality of vehicles. For example, each of the plurality of vehicles transmits, via a wireless network, a video captured using a camera mounted on the vehicle to a video collection device (not shown). The compression device 110 encodes the video in a predetermined encoding method and stores the encoded video in the video storage unit 150. In the present embodiment, it is assumed that the encoding of the video is synonymous with the compression of the video. For the video storage unit 150, a high-speed storage such as a solid-state drive (SSD) is used, for example. The video search device 130 searches for a desired video among the videos stored in the video storage unit 150 in response to a prompt input by the user. Here, the prompt can mean an instruction or a search query input by the user.
[0020] The video search device 130 includes a decoder 131 and a recognition unit 132. Physically, the video search device 130 can be configured as a device having one or more memories and one or more processors. In the video search device 130, at least a part of the functions of each unit in the video search device 130 can be realized by the one or more processors executing processing according to instructions read from the one or more memories. The decoder 131 may be configured by a dedicated hardware circuit.
[0021] The decoder 131 decompresses or decodes the video compressed by the compression device 110. The search prompt 133 is a prompt input by the user during video search. The search prompt 133 is described, for example, in natural language. The search prompt 133 includes words that specify objects and situations recognized by the recognition unit 132. The recognition unit 132 acquires the video stored in the video storage unit 150 via the decoder 131, and generates a recognition result for the search prompt 133 for the acquired video. Also, the recognition unit 132 may extract from the video storage unit 150 a video corresponding to the object and situation specified in the search prompt 133 based on the recognition result. The recognition unit 132 includes an artificial intelligence model such as, for example, a Vision Language Model (VLM).
[0022] The compression device 110 includes a video segmentation unit 111, an estimation model 112, a parameter determination unit 113, and an encoder 114. Physically, the compression device 110 may be configured as a device having one or more memories and one or more processors. In the compression device 110, at least part of the functions of each unit inside the compression device 110 can be realized by one or more processors executing processing according to instructions read from one or more memories. The encoder 114 may be constituted by a dedicated hardware circuit.
[0023] The video segmentation unit 111 divides the video to be compressed into a plurality of video clips. For example, the video segmentation unit 111 divides the video into N video clips where N is an integer of 2 or more. The video segmentation unit 111 divides, for example, the video transmitted from each vehicle into a plurality of video clips based on a predetermined criterion. The encoder 114 compresses and encodes each of the divided video clips. The encoder 114 is also called a video compression unit. The encoder 114 may set a fixation region in the frame image in the video clip, and compress the frame image so that the image quality of the fixation region is relatively higher than that of other regions. The encoder 114 stores the compressed video clip in the video storage unit 150. The encoder 114 may distribute the compressed and encoded video clip to the video search device 130 via a network.
[0024] Estimation model 112 is a model that estimates the data volume of a compressed video clip and the recognition error index value based on the video clip and video compression parameters. The error index value represents the recognition error in the recognition unit 132 of the video retrieval device 130. The error index value is also called the error score. Estimation model 112 estimates the data volume of the compressed video when the video clip is compressed in the encoder 114 based on certain parameters. Estimation model 112 also estimates the recognition error in the recognition unit 132 for a video clip obtained by decompressing a video clip that has been compressed in the encoder 114 based on certain parameters.
[0025] The parameter determination unit 113 determines the video compression parameters to be used to compress the video clips based on the divided video clips and the estimation model 112. For example, the parameter determination unit 113 estimates the data volume of the compressed video clip and the error index value for each candidate video compression parameter used to compress the video clip using the estimation model 112. Based on the estimated data volume of the video clip and the estimated error index value, the parameter determination unit 113 determines the video compression parameters to be used for encoding in the encoder 114 from the candidate video compression parameters. For example, the parameter determination unit 113 determines the video compression parameters to be used for encoding for each of the N video clips.
[0026] Figure 3 is a block diagram showing an example configuration of a learning system used to generate an estimation model 310. The learning system 300 includes a compression / decompression unit 301, a parameter setting unit 302, a recognition unit 303, a recognition unit 304, an error acquisition unit 305, and a model learning unit 306. Physically, the learning system 300 can be configured as a device having one or more memories and one or more processors. In the learning system 300, at least some of the functions of each part within the learning system 300 can be realized by one or more processors executing processing according to instructions read from one or more memories. The learning system 300 does not need to be configured as a separate system from the video retrieval system 100 shown in Figure 2, and some elements of the video retrieval system 100 may overlap with some elements of the learning system 300. The learning system 300 corresponds to the video processing system 10 shown in Figure 1.
[0027] The compression / decompression unit 301 compresses and encodes the learning video clip 330 using a predetermined encoding scheme. The compression / decompression unit 301 outputs the amount of data of the compressed learning video clip 330 as the measured data amount 360. The compression / decompression unit 301 decompresses the compressed learning video clip 330. The parameter setting unit 302 sets arbitrary video compression parameters as video compression parameters in the compression / decompression unit 301. The parameter setting unit 302 selects the video compression parameters to be set, for example, from a random distribution. The compression / decompression unit 301 may be configured by a dedicated hardware circuit. The compression / decompression unit 301 corresponds to the compression / decompression unit 11 shown in Figure 1. The compression / decompression unit 301 also corresponds to the encoder 114 and decoder 131 in the video retrieval system 100 shown in Figure 2. The compression / decompression unit 301 may be separated into a compression unit and a decompression unit. The compression portion of the compression / decompression unit 301 corresponds to the encoder 114, and the decompression portion of the compression / decompression unit 301 corresponds to the decoder 131.
[0028] The learning prompt 350 includes a plurality of prompts used to train the estimation model 310. The learning prompt 350 includes a plurality of prompts that are assumed to be search prompts 133 input to the recognition unit 132 in the video retrieval device 130. The learning prompt 350 may be generated in response to the training video clip 330, for example using generated artificial intelligence (AI). The learning prompt 350 may include a variety of prompts or complex prompts that may be possible for the training video clip 330.
[0029] The recognition unit 303 generates recognition results for the decompressed video clip based on the decompressed video clip output from the compression / decompression unit 301 and the learning prompts 350. For example, the recognition unit 303 inputs the decompressed video to the VLM and outputs the responses to multiple learning prompts 350 output from the VLM as recognition results.
[0030] The recognition unit 304 generates a recognition result of the training video clip 330 before compression, based on the training video clip 330 and the training prompt 350. For example, the recognition unit 304 inputs the training video clip 330 to the VLM and outputs the response to the training prompt 350 output from the VLM as the recognition result. Note that the learning system 300 does not necessarily need to have two recognition units 303 and 304; one recognition unit may be used as both the recognition unit 303 and 304. Recognition units 303 and 304 correspond to the recognition unit 12 shown in Figure 1. Also, recognition units 303 and 304 correspond to the recognition unit 132 in the video retrieval system 100 shown in Figure 2.
[0031] The error acquisition unit 305 acquires the error between the recognition result output from the recognition unit 303 and the recognition result output from the recognition unit 304. In other words, the error acquisition unit 305 uses the recognition result output from the recognition unit 304 as the correct answer data, and acquires the difference between the correct answer data and the recognition result output from the recognition unit 303 as the error. The error acquisition unit 305 acquires the error between the recognition result output from the recognition unit 303 and the recognition result output from the recognition unit 304 using, for example, an evaluation framework such as G-Eval. For example, the error acquisition unit 305 acquires the error of the response to the same prompt for each of the multiple learning prompts 350, and acquires the average of the acquired errors as the error for one learning video clip 330. The error acquisition unit 305 outputs the acquired error as the measured error score 370. The error acquisition unit 305 corresponds to the error calculation unit 13 shown in Figure 1.
[0032] The estimation model 310 is a model that estimates the data volume and error index value of a compressed video clip based on the video clip and video compression parameters. The estimation model 310 estimates the estimated data volume 380, which indicates the data volume of the compressed video clip, according to the training video clip 330 and the video compression parameters set by the parameter setting unit 302. The estimation model 310 also estimates the estimated error score 390 according to the training video clip 330 and the video compression parameters set by the parameter setting unit 302. The estimated error score 390 indicates the error in the result of the recognition process performed on the decoded video obtained by decoding the compressed video clip.
[0033] The estimation model 310 includes a relationship estimation network 311, a parameter extraction unit 312, and an estimation unit 313. The relationship estimation network 311 is a model that estimates the relationship between video compression parameters and the estimated data volume and estimated error score of the compressed video clip from the training video clip 330. For example, the relationship estimation network 311 estimates a function that estimates the estimated data volume from the independent variable, with the video compression parameters as the independent variable. The relationship estimation network 311 also estimates a function that estimates the estimated recognition error score from the independent variable, with the video compression parameters as the independent variable. The relationship estimation network 311 includes, for example, a combination of neural networks such as a Convolutional Neural Network (CNN), Transformer, and Multi-layer Perceptron (MLP), and a pre-trained Contrastive Language-Image Pre-Training (CLIP) Image Encoder.
[0034] The parameter extraction unit 312 extracts function parameter information for the function that estimates the estimated data amount from the relationship estimation network 311. The parameter extraction unit 312 also extracts function parameter information for the function that estimates the estimated error score from the relationship estimation network 311. The parameter extraction unit 312 extracts the coefficients of each term of a function expressed as a polynomial, for example, as function parameter information. The estimation unit 313 estimates the estimated data amount 380 and the estimated error score 390 based on the function parameter information extracted by the parameter extraction unit 312 and the video compression parameters set by the parameter setting unit 302. The estimation unit 313 estimates the estimated data amount 380 and the estimated error score 390 using, for example, MPL or a polynomial.
[0035] The model learning unit 306 generates an estimation model 310 such that the difference between the measured data amount 360 and the estimated data amount 380 becomes small and the difference between the measured error score 370 and the estimated error score becomes small. In the present embodiment, the model learning unit 306 learns the relational estimation network 311 of the estimation model 310 based on the measured data amount 360, the measured error score 370, the estimated data amount 380, and the estimated error score 390. The model learning unit 306 learns the relational estimation network 311 so that the difference between the measured data amount 360 and the estimated data amount 380 decreases. Further, the model learning unit 306 learns the relational estimation network 311 so that the difference between the measured error score 370 and the estimated error score 390 decreases. The model learning unit 306 is also called a model generation unit. The learned estimation model 310 is used as the estimation model 112 of the compression device 110 shown in FIG. 2. The model learning unit 306 corresponds to the model generation unit 14 shown in FIG. 1.
[0036] An explanation will be given of the determination of the video compression parameters of the video clip using the estimation model 112. FIG. 4 is a schematic diagram schematically showing the determination of the video compression parameters. The video segmentation unit 111 divides the video to be compressed into N video clips 200-1 to 200-N, for example, in a predetermined time unit. In the estimation model 112, with n being an integer from 1 to N, the video clip 200-n and the video compression parameters s n , r n , q n are input. Here, s n represents the resolution of the video clip, r n represents the frame rate of the video clip, and q n represents the bit rate of the video clip. The video compression parameters s n , r n , q n are selected, for example, from candidates for predetermined video compression parameters.
[0037] The estimation model 112 calculates, for the video clip 200-n, the estimated data amount R(s n , r n , q n ) and the estimated error score E(s n , rn ,q n The parameter determination unit 113 estimates the estimated data amount R(s) for each video clip. n ,r n ,q n ), and estimated error score E(s n ,r n ,q n The objective function is calculated from the following. For example, the parameter determination unit 113 calculates the objective function represented by the following formula, with λ being a real number greater than or equal to 0.
[0038] The parameter determination unit 113 determines the video compression parameters s that minimize the objective function. n ,r n ,q n This is determined as the video compression parameter for each video clip 200-n. The parameter determination unit 113 uses a predetermined algorithm, such as the gradient method, to determine the video compression parameter s that minimizes the objective function. n ,r n ,q n The parameter determination unit 113 searches for the video compression parameter s that minimizes the objective function. n ,r n ,q n Set the encoder 114 to the following parameters. The encoder 114 processes N video clips 200-1 to 200-N, each with a video compression parameter s. 1 ,r 1 ,q 1 ~s N ,r N ,q N Encode it using this method.
[0039] Next, the operation procedure will be explained. Figure 5 is a flowchart showing the operation procedure of the learning system 300. The operation procedure of the learning system 300 corresponds to the video processing method. The parameter setting unit 302 sets arbitrary video compression parameters to the compression / decompression unit 301 (step A1). The compression / decompression unit 301 compresses the learning video clip 330 based on the set video compression parameters (step A2). The compression / decompression unit 301 outputs the amount of data of the compressed learning video clip 330 as the measured data amount 360 (step A3).
[0040] The compression / decompression unit 301 decompresses the compressed learning video clip 330 (step A4). The recognition unit 303 performs recognition processing on the decompressed learning video clip (step A5). The recognition unit 304 also performs recognition processing on the learning video clip 330 before compression (step A6). The error acquisition unit 305 generates an actual error score 370 based on the recognition results of the recognition unit 303 and the recognition results of the recognition unit 304 (step A7).
[0041] The estimation model 310 estimates the estimated data volume 380 and the estimated error score 390 based on the training video clip 330 and the video compression parameters set in step A1 (step A8). In step A8, the relationship estimation network 311 of the estimation model 310 estimates the relationship between the video compression parameters and the estimated data volume and estimated error score of the compressed video clip from the training video clip 330. The parameter extraction unit 312 extracts function parameter information of the function that estimates the estimated data volume and estimated error score from the relationship estimation network 311. The estimation unit 313 estimates the estimated data volume 380 and the estimated error score 390 based on the function parameter information extracted by the parameter extraction unit 312 and the video compression parameters set in step A1.
[0042] The model learning unit 306 learns the estimation model 310 so that the difference between the measured data amount 360 and the estimated data amount 380 becomes smaller, and the difference between the measured error score 370 and the estimated error score becomes smaller (step A9). In step A9, the model learning unit 306 learns the relationship estimation network 311 so that the difference between the measured data amount 360 and the estimated data amount 380 decreases. The model learning unit 306 also learns the relationship estimation network 311 so that the difference between the measured error score 370 and the estimated error score 390 decreases. By repeatedly performing steps A1 to A9 for multiple training video clips 330, the estimation model 112 used in the compression device 110 is learned.
[0043] Figure 6 is a flowchart showing the operation procedure of the compression device 110. The video splitting unit 111 splits the video to be compressed into multiple video clips (step B1). The estimation model 112 estimates the amount of compressed data and the recognition error score for each video clip based on the split video clips and video compression parameters (step B2). The parameter determination unit 113 determines the video compression parameters to be used to compress each video clip based on the estimated amount of compressed data and the recognition error score (step B3). The encoder 114 compresses each video clip based on the determined video compression parameters (step B4).
[0044] In this embodiment, the estimation model 112 is generated in the learning system 300 such that the difference between the measured data amount 360 and the estimated data amount 380, and the difference between the measured error score 370 and the estimated error score 390 are reduced. In this embodiment, by using the estimation model 112, the data amount of the compressed video clip and the recognition error score can be estimated based on the video clip and video compression parameters. The parameter determination unit 113 determines the video compression parameters when compressing the video clip using the estimated data amount and error score. In this way, video compression that can maintain recognition accuracy while reducing the amount of data can be achieved.
[0045] If multiple video clips were compressed based on a single video compression parameter, the overall result would be either a large amount of data while maintaining recognition accuracy, or reduced data size but degraded recognition accuracy. In this embodiment, the parameter determination unit 113 calculates an objective function based on the estimated data size and estimated error score estimated for each video clip, and determines the video compression parameter that minimizes the objective function for each video clip. By doing so, compared to compressing each video clip based on a constant video compression parameter, it is possible to reduce the amount of data in the compressed video clips while reducing the recognition error.
[0046] In the above embodiment, R(s n ,r n ,q n) + λE(s n ,r n ,q n An example has been described in which the sum of ) is used as the objective function, but this disclosure is not limited to this. For example, the parameter determination unit 113 determines R(s) for each video clip. n ,r n ,q n ) + λE(s n ,r n ,q n The objective function is calculated as ) and the video compression parameter s that minimizes the objective function is calculated. n ,r n ,q n This can be determined for each video clip.
[0047] Furthermore, although the above embodiment describes an example in which compressed video stored in the video storage unit 150 is used for video retrieval, this disclosure is not limited thereto. Video compressed based on video compression parameters determined using the estimation model 112 can be used for other purposes as well. In that case, the learning system 300 uses the error in the recognition result for the application in which the video is used as the error score.
[0048] Next, the physical configuration of the compression device 110, the video retrieval device 130, and the learning system 300 will be described. Figure 7 is a block diagram showing an example configuration of a computer device that can be used as the compression device 110, the video retrieval device 130, or the learning system 300. The computer device 500 has a processor 510 such as a Central Processing Unit (CPU), a storage unit 520, a Read Only Memory (ROM) 530, a Random Access Memory (RAM) 540, a communication interface (IF) 550, and a user interface 560.
[0049] The communication interface 550 is an interface for connecting the computer device 500 to a communication network via wired communication means or wireless communication means. The user interface 560 includes a display unit, such as a display. The user interface 560 also includes input units such as a keyboard, mouse, and touch panel.
[0050] The memory unit 520 is an auxiliary storage device capable of holding various types of data. The memory unit 520 does not necessarily have to be part of the computer device 500; it may be an external storage device or cloud storage connected to the computer device 500 via a network.
[0051] ROM 530 is a non-volatile memory device. For example, a semiconductor memory device such as a relatively small-capacity flash memory is used for ROM 530. The program executed by the CPU 510 can be stored in the storage unit 520 or ROM 530. The storage unit 520 or ROM 530 stores various programs that realize the functions of the compression device 110, the video retrieval device 130, or the learning system 300.
[0052] The above program, when loaded into a computer, includes a set of instructions (or software code) for causing the computer to perform one or more of the functions described in the embodiments. The program may be stored in a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include RAM, ROM, flash memory, SSD or other memory technologies, Compact Disc (CD), digital versatile disc (DVD), Blu-ray® disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include, a temporary computer-readable medium or a communication medium that includes electrically, optically, acoustically, or otherwise propagating signals.
[0053] The RAM 540 is a volatile memory device. Various semiconductor memory devices such as Dynamic Random Access Memory (DRAM) or Static Random Access Memory (SRAM) can be used for the RAM 540. The RAM 540 may be used as an internal buffer for temporarily storing data. The CPU 510 loads a program stored in the memory unit 520 or ROM 530 into the RAM 540 and executes it. By executing the program, the CPU 510 can realize the functions of the compression device 110, the video retrieval device 130, or each part of the learning system 300. The CPU 510 may have an internal buffer that can temporarily store data.
[0054] In the above embodiment, the compression device 110 and the video retrieval device 130 do not necessarily have to be configured as a single computer device. The compression device 110 and the video retrieval device 130 may each be configured using multiple physically separated devices. Similarly, the learning system 300 may also be configured using multiple physically separated devices.
[0055] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be made as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0056] Each drawing is merely illustrative to illustrate one or more embodiments. Each drawing may be associated with one or more other embodiments, rather than being associated with only one specific embodiment. As those skilled in the art will understand, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings, for example, to create embodiments not explicitly shown or described. Not all features or steps shown in any one drawing to illustrate an exemplary embodiment are necessarily required, and some features or steps may be omitted. The order of steps described in any of the drawings may be changed as appropriate.
[0057] Some or all of the above embodiments may also be described as follows, but are not limited to the following:
[0058] [Note 1] A video processing system comprising: a compression / decompression unit that compresses a video clip based on video compression parameters and decompresses the compressed video clip; a recognition unit that receives the decompressed video clip as input and outputs a recognition result of the input video clip; an error calculation unit that compares the recognition result of the decompressed video clip with a correct recognition result and calculates an error index value; and a model generation unit that generates a model for estimating the data amount of the compressed video clip and the error index value based on the video clip and the video compression parameters, wherein the model generation unit generates the model such that the difference between the data amount of the compressed video clip and the estimated data amount is small.
[0059] [Note 2] The video processing system described in Note 1, wherein the recognition result of the decompressed video clip includes answers to a plurality of prompts obtained by inputting the decompressed video clip into a Vision Language Model (VLM) that outputs answers to prompts for an input video clip, and the correct recognition result includes answers to the plurality of prompts obtained by inputting the video clip before compression into the VLM.
[0060] [Note 3] The video processing system according to Note 1 or 2, wherein the model generation unit generates the model such that the difference between the amount of data of the compressed video clip and the estimated amount of data becomes small, and the difference between the calculated error index value and the estimated error index value becomes small.
[0061] [Appendix 4] The video processing system according to any one of Appendix 1 to 3, further comprising a parameter setting unit for setting video compression parameters used to compress the video clip in the compression / decompression unit.
[0062] [Note 5] The video processing system according to Note 4, wherein the parameter setting unit selects video compression parameters to be set in the compression / decompression unit from a randomly distributed set of video compression parameters.
[0063] [Appendix 6] The video processing system according to any one of Appendix 1 to 5, wherein the video compression parameters include at least one of the resolution, frame rate, and bitrate of the video clip.
[0064] [Note 7] The video processing system according to any one of Notes 1 to 6, wherein the model includes a relationship estimation network for estimating the relationship between the video compression parameters, the amount of data in the compressed video clip, and the error index value; a parameter extraction unit for extracting function parameter information for the estimated amount of data and function parameter information for the estimated error index value from the relationship estimation network; and an estimation unit for estimating the amount of data in the compressed video clip and the error index value based on the video clip, the video compression parameters used to compress the video clip in the compression / decompression unit, and the extracted function parameter information for the estimated amount of data and function parameter information for the estimated error index value.
[0065] [Note 8] The video processing system according to Note 7, wherein the model generation unit learns the relationship estimation network based on the amount of data of the compressed video clip, the estimated amount of data, the calculated error index value, and the estimated error index value.
[0066] [Note 9] The video processing system according to any one of Notes 1 to 8, further comprising: a video splitting unit that divides the video to be compressed into a plurality of video clips; a video compression unit that compresses each of the plurality of video clips; and a parameter determination unit that determines video compression parameters used for compressing the video clips based on the plurality of video clips and the model.
[0067] [Note 10] The video processing system according to Note 9, wherein the parameter determination unit estimates the amount of data of a compressed video clip and the error index value using the model for candidates of video compression parameters used to compress each video clip, and determines the video compression parameter to be used to compress the video clip from the candidates of video compression parameters based on the estimated amount of data of the video clip and the estimated error index value.
[0068] [Note 11] A video processing method that compresses a video clip based on video compression parameters, decompresses the compressed video clip, generates a recognition result for the decompressed video clip, compares the generated recognition result for the decompressed video clip with a correct recognition result, calculates an error index value based on the result of the comparison, and generates a model such that the difference between the data amount of the compressed video clip estimated using a model that estimates the data amount of the compressed video clip and the error index value based on the video clip and the video compression parameters is reduced.
[0069] [Note 12] A program that causes a computer to perform the following processes: compress a video clip based on video compression parameters; decompress the compressed video clip; generate a recognition result for the decompressed video clip; compare the generated recognition result for the decompressed video clip with a correct recognition result; calculate an error index value based on the result of the comparison; and generate a model that estimates the data amount of the compressed video clip and the error index value based on the video clip and the video compression parameters, such that the difference between the data amount of the compressed video clip estimated using the model and the data amount of the compressed video clip becomes small.
[0070] Some or all of the elements (e.g., configuration and function) described in Appendices 2 to 10 that are dependent on Appendice 1 may also be dependent on Appendices 11 and 12 in the same way as those described in Appendices 2 to 10. Some or all of the elements described in any appendice may be applicable to various hardware, software, recording means, systems, and methods for recording software.
[0071] This application claims priority based on Japanese Patent Application No. 2024-206329, filed on 27 November 2024, and incorporates all of its disclosures herein.
[0072] 10: Video processing system 11: Compression / decompression unit 12: Recognition unit 13: Error calculation unit 14: Model generation unit 100: Video search system 110: Compression device 111: Video splitting unit 112: Estimation model 113: Parameter determination unit 114: Encoder 130: Video search device 131: Decoder 132: Recognition unit 133: Search prompt 150: Video storage unit 200-1 to 200-N: Video clips 300: Learning system 301: Compression / decompression unit 302: Parameter setting unit 303, 304: Recognition unit 305: Error acquisition unit 306: Model learning unit 310: Estimation model 311: Relationship estimation network 312: Parameter extraction unit 313: Estimation unit 330: Learning video clips 350: Learning prompts 360: Actual data volume 370: Actual error score 380: Estimated data volume 390: Estimated error score 500: Computer device 510: Processor 520: Memory unit 530: ROM 540: RAM 550: Communication interface 560: User interface
Claims
1. A video processing system comprising: a compression / decompression unit that compresses a video clip based on video compression parameters and decompresses the compressed video clip; a recognition unit that receives the decompressed video clip as input and outputs a recognition result of the input video clip; an error calculation unit that compares the recognition result of the decompressed video clip with a correct recognition result and calculates an error index value; and a model generation unit that generates a model for estimating the data amount of the compressed video clip and the error index value based on the video clip and the video compression parameters, wherein the model generation unit generates the model such that the difference between the data amount of the compressed video clip and the estimated data amount is small.
2. The video processing system according to claim 1, wherein the recognition result of the decompressed video clip includes responses to a plurality of prompts obtained by inputting the decompressed video clip into a Vision Language Model (VLM) that outputs responses to prompts for an input video clip, and the correct recognition result includes responses to the plurality of prompts obtained by inputting the video clip before compression into the VLM.
3. The video processing system according to claim 1 or 2, wherein the model generation unit generates the model such that the difference between the amount of data of the compressed video clip and the estimated amount of data becomes small, and the difference between the calculated error index value and the estimated error index value becomes small.
4. The video processing system according to any one of claims 1 to 3, further comprising a parameter setting unit for setting video compression parameters used to compress the video clip in the compression / decompression unit.
5. The video processing system according to claim 4, wherein the parameter setting unit selects video compression parameters to be set in the compression / decompression unit from a randomly distributed set of video compression parameters.
6. The video processing system according to any one of claims 1 to 5, wherein the video compression parameters include at least one of the resolution, frame rate, and bitrate of the video clip.
7. The video processing system according to any one of claims 1 to 6, wherein the model includes a relationship estimation network for estimating the relationship between the video compression parameters, the amount of data in the compressed video clip, and the error index value; a parameter extraction unit for extracting function parameter information for the estimated amount of data and function parameter information for the estimated error index value from the relationship estimation network; and an estimation unit for estimating the amount of data in the compressed video clip and the error index value based on the video clip, the video compression parameters used to compress the video clip in the compression / decompression unit, and the extracted function parameter information for the estimated amount of data and function parameter information for the estimated error index value.
8. The video processing system according to claim 7, wherein the model generation unit learns the relationship estimation network based on the amount of data of the compressed video clip, the estimated amount of data, the calculated error index value, and the estimated error index value.
9. The video processing system according to any one of claims 1 to 8, further comprising: a video splitting unit for dividing video to be compressed into a plurality of video clips; a video compression unit for compressing each of the plurality of video clips; and a parameter determination unit for determining video compression parameters used for compressing the video clips based on the plurality of video clips and the model.
10. The video processing system according to claim 9, wherein the parameter determination unit estimates the amount of data of a compressed video clip and the error index value using the model for candidates of video compression parameters to be used to compress each video clip, and determines the video compression parameter to be used to compress the video clip from the candidates of video compression parameters based on the estimated amount of data of the video clip and the estimated error index value.
11. A video processing method comprising: compressing a video clip based on video compression parameters; decompressing the compressed video clip; generating a recognition result for the decompressed video clip; comparing the generated recognition result for the decompressed video clip with a correct recognition result; calculating an error index value based on the result of the comparison; and generating a model such that the difference between the data amount of the compressed video clip estimated using a model that estimates the data amount of the compressed video clip and the error index value based on the video clip and the video compression parameters is reduced.
12. A program that causes a computer to perform the following processes: compress a video clip based on video compression parameters; decompress the compressed video clip; generate a recognition result for the decompressed video clip; compare the generated recognition result for the decompressed video clip with a correct recognition result; calculate an error index value based on the result of the comparison; and generate a model that estimates the data amount of the compressed video clip and the error index value based on the video clip and the video compression parameters, such that the difference between the data amount of the compressed video clip estimated using the model and the data amount of the compressed video clip is small.