Information processing device, information processing method, and computer-readable non-transitory storage medium

The proposed solution addresses the performance degradation in AI video super-resolution systems by using a specialized coefficient database and a CNN that switches coefficients based on rendering condition changes, ensuring consistent and accurate inference across dynamic rendering conditions.

WO2025120882A1PCT designated stage expired Publication Date: 2025-06-12SONY GROUP CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/020808
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-08
Filing Date
2024-06-07
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing AI video super-resolution systems face performance degradation due to mismatched rendering resolutions between training and inference, particularly in dynamic environments like game engines, where rendering resolution changes frequently.

Method used

An information processing apparatus and method that utilizes a specialized coefficient database to store coefficients for various rendering conditions, with a CNN that switches these coefficients based on rendering condition changes and uses intermediate features as recurrent data for inference.

Benefits of technology

This configuration maintains high inference accuracy even when rendering conditions switch, by ensuring consistency in the format of recurrent data and using shared coefficient values for the final convolutional layer across different rendering conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024020808_12062025_PF_FP_ABST
    Figure JP2024020808_12062025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device includes a specialization coefficient database and a CNN. The specialization coefficient database stores a plurality of specialization coefficients in association with rendering conditions to be specialized. The plurality of specialization coefficients have a common coefficient value for a final convolution layer that is closest to an output layer. The CNN switches the specialization coefficients in accordance with the switching of the rendering conditions. The CNN performs inference relating to an input frame by using, as recurrent data, an intermediate feature amount that has been input to the final convolution layer.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and computer-readable non-transitory storage medium

[0001] The present invention relates to an information processing device, an information processing method, and a computer-readable non-transitory storage medium.

[0002] In general, the accuracy of AI inference increases as the characteristics of the input data used for inference and the input data used for learning become closer. If the rendering conditions of the input image used for inference (such as rendering resolution and the characteristics of the video scene to be specialized) are known, it is desirable to use coefficients (specialization coefficients) learned using data specialized for those rendering conditions. When the rendering conditions change between frames, it is desirable to also change the specialization coefficients in accordance with the change in rendering conditions.

[0003] For example, in real-time rendering of a game engine, the rendering resolution can change in various ways depending on the processing load. In AI video super-resolution that assumes only a specific rendering resolution, performance degradation occurs due to a mismatch between the rendering resolution during learning and the rendering resolution during inference. To solve this problem, a system configuration and learning method for AI video super-resolution that can accommodate various dynamic resolution changes (Dynamic Resolution: DR) are required.

[0004] Japanese Patent Application Laid-Open No. 2022-008037

[0005] There are two main architectural configurations for AI video super-resolution for real-time processing (online processing): the Image Recurrent model, which synthesizes the final output image into the next input, and the Feature Recurrent model, which synthesizes the intermediate features (intermediate features) that are the intermediate output of AI (especially a Convolutional Neural Network (CNN)) into the next input.

[0006] Typically, inference accuracy is higher when intermediate features with a large amount of information are used. However, while the Feature Recurrent model maintains high performance at a single rendering resolution, when models are trained independently for each rendering resolution, the features take on a format specific to each rendering resolution. As a result, when the rendering resolution changes dynamically, consistency cannot be maintained when switching models, and quality cannot be maintained when the rendering resolution is transitioned.

[0007] Therefore, the present disclosure proposes an information processing device, an information processing method, and a computer-readable non-transitory storage medium that are less likely to experience a decrease in inference accuracy when rendering conditions are switched.

[0008] According to the present disclosure, there is provided an information processing device including: a specialized coefficient database that stores multiple specialized coefficients, each of which has a common coefficient value in a final convolutional layer closest to an output layer, in association with a rendering condition that is the target of the specialized coefficients; and a CNN that switches the specialized coefficients in accordance with a switch in the rendering condition and performs inference on an input frame using intermediate features input to the final convolutional layer as recurrent data. The present disclosure also provides an information processing method in which information processing of the information processing device is executed by a computer, and a computer-readable non-transitory storage medium that stores a program that causes a computer to realize the information processing of the information processing device.

[0009] 1 is an explanatory diagram of high resolution AI in a game engine. FIG. 1 is an explanatory diagram of the architecture of high resolution AI for video. FIG. 1 is an explanatory diagram of the architecture of high resolution AI for video. FIG. 2 is a diagram explaining the characteristics of recurrent data according to training data. FIG. 3 is a diagram showing changes in inference accuracy when the format of recurrent data is not unified across two models. FIG. 4 is a diagram showing an example of the configuration of an information processing device of the present disclosure. FIG. 5 is an explanatory diagram of intermediate features. FIG. 6 is a diagram explaining control of the format of recurrent data by fixing the coefficients of the last block. FIG. 7 is a diagram showing the relationship between the filter of the final convolutional layer and feature frames. FIG. 8 is a diagram showing the flow of outputting "ch R". FIG. 9 is a diagram showing model switching according to a change in resolution. FIG. 10 is a diagram showing changes in inference accuracy when the format of recurrent data is unified across two models. FIG. 11 is a diagram showing an example of a method for learning specialization coefficients. FIG. 12 is a diagram showing an example of a method for learning specialization coefficients. FIG. 13 is a diagram showing an example of a processing flow during inference. FIG. 14 is a diagram showing an example of the hardware configuration of an information processing device.

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0011] The description will be given in the following order: [1. Problems of resolution changes in game engines] [2. Architecture of high-resolution AI for video] [3. Characteristics of recurrent data according to training data] [4. System configuration example of information processing device disclosed herein] [5. Control of recurrent data format by fixing coefficients of last block] [6. Improvement of consistency of recurrent data when switching models] [7. Method of learning specialized coefficients] [8. Processing flow during inference] [9. Example hardware configuration] [10. Effects]

[0012] In the present disclosure, a configuration is adopted in which the CNN specialization coefficients 41 (see FIG. 6 ) are switched in accordance with the switching of rendering conditions between frames. The specialization coefficients 41 are learned as coefficients specialized for specific rendering conditions. Examples of rendering conditions that can be specialized include rendering resolution and scene characteristics for which the rendered image is specialized (such as an explosion scene). In the following explanation, an example is described in which the renderer is used in real-time rendering of a game engine, and the specialization coefficients 41 are coefficients learned for each rendering resolution.

[0013] In the present disclosure, all specialized coefficients 41 to be switched share the coefficient values ​​of the final convolutional layer closest to the output layer (final convolutional layer 31L: see FIG. 6 ). This ensures the homogeneity of intermediate features acquired as recurrent data (data used for recursive input to the CNN) before and after switching of rendering conditions. Therefore, the inference accuracy of the CNN is less likely to decrease when switching rendering conditions. This will be explained in detail below.

[0014] [1. Problems with Resolution Changes in Game Engines] FIG. 1 is an explanatory diagram of high-resolution AI in a game engine.

[0015] In existing game engines, rendering resolution changes dynamically depending on factors such as the complexity of the scene. As a result, simply upconverting rendered images results in inconsistent quality. For still image support, a high-resolution AI (a model trained specifically for a specific rendering resolution) can be prepared for each rendering resolution, and the model can be switched for each rendering resolution. However, for video support, the next process is executed after receiving the results of the previous process. Therefore, when rendering resolution changes frequently, simply switching between high-resolution AI trained only for a specific rendering resolution results in a significant drop in quality.

[0016] 2. Architecture of High Resolution AI for Video FIGS. 2 and 3 are explanatory diagrams of the architecture of high resolution AI for video.

[0017] CNN processing for video generally uses a highly time-stable RNN (Recurrent Neural Network) structure. The RNN structure uses current information (input frame I C ) and past information, it is easy to maintain consistency in the output image. There are two ways to use past information. One is to use the output frame I, which is the inference image. O The other is the Image Recurrent model, which uses the feature output from the intermediate layer of the CNN (intermediate feature I H ) is synthesized into the next input and used as a Feature Recurrent model (Figure 3).

[0018] The Image Recurrent model is easy to learn robustly against changes in rendering resolution because the feedback information is an image. However, its performance is not high to begin with. The Feature Recurrent model is an image (output frame I O ) H Since it uses , the inference accuracy is higher than that of the Image Recurrent model. However, it is only after training that it becomes clear what features the CNN has extracted. Because there is randomness in the format in which the features are constructed, there is no guarantee that the features will be common between models trained for each rendering resolution. If an image with a rendering resolution different from that used during training is input, it is difficult to achieve high inference accuracy.

[0019] 3. Characteristics of recurrent data according to training data] Fig. 4 is a diagram illustrating characteristics of recurrent data according to training data. The left side of Fig. 4 is a diagram illustrating a training process using training data with resolution A. The right side of Fig. 4 is a diagram illustrating a training process using training data with resolution B.

[0020] CNN uses the most recent intermediate feature I H is used as recurrent data to infer the next frame. H contains multiple feature frames FE with different properties. During training, the output frame I OThe coefficients of the filter 32 (see FIG. 9) in the convolution layer 31 are learned so that the difference between the target image and the target image becomes small.

[0021] The formats of the feature frame FE and the filter 32 are automatically acquired through learning. The format refers to the type of feature and the order of the convolution operation. If no constraints are placed on learning, the formats may differ between models. Therefore, if two models are simply joined together to accommodate changes in rendering resolution, the format of the filter 32 will not match the format of the input feature frame FE, and inference performance will not be guaranteed.

[0022] 5 is a diagram showing changes in inference accuracy when the recurrent data format is not unified between the two models. When switching from a model with resolution A to a model with resolution B, the format of the feature frame FE output from the model before the switch (resolution A) does not match the format of the filter 32 of the model after the switch (resolution B), and accurate inference results cannot be obtained. As a result, color shifts and the like occur when switching models, and the PSNR (Peak Signal-to-Noise Ratio) drops significantly.

[0023] In view of the above circumstances, the present disclosure proposes a method for standardizing the format of recurrent data. In the present disclosure, the coefficient values ​​of the final convolutional layer 31L of each model are standardized. This standardizes the format of the filter of the final convolutional layer 31L for each model. By standardizing the filter format, the format of the recurrent data input to the final convolutional layer 31L is also standardized. The above information processing is realized, for example, by an information processing device 1 shown in FIG. 6 .

[0024] 6 is a diagram showing an example of the configuration of an information processing device 1 according to the present disclosure. The information processing device 1 includes a renderer 10, a scaler 20, a CNN 30, and a specialized coefficient database 40.

[0025] The renderer 10 performs rendering based on the rendering settings. The rendering settings include conditions for specializing the rendering as rendering conditions. For example, if the rendering is specialized for a specific rendering resolution, the value of the rendering resolution is acquired as the rendering condition. The renderer 10 outputs an image of one frame of a video as a rendered image. The block diagram in FIG. 6 explains the operation for one frame of a video.

[0026] The scaler 20 converts the rendered image or the image scaled to fit the architecture of the CNN 30 into an input frame I. C Scaling refers to the process of expanding the number of pixels in a rendered image. If the architecture of the CNN 30 accepts a fixed number of pixels, while the number of pixels in a rendered image changes depending on how it is used, the scaler 20 expands the number of pixels at a desired magnification so as to obtain a number of pixels that matches the specifications of the CNN 30. The method of expanding the number of pixels may be selected arbitrarily, such as bilinear or bicubic.

[0027] As described above, the renderer 10 may change the rendering resolution depending on the processing load. In this case, the scaler 20 scales the rendered image to fit the architecture of the CNN 30.

[0028] The specialized coefficient database 40 holds specialized coefficients 41 of the pre-trained CNN 30. The specialized coefficient database 40 stores the multiple specialized coefficients 41 in association with the rendering conditions that are the targets of the specialization. The multiple specialized coefficients 41 share the coefficient values ​​of the final convolutional layer (final convolutional layer 31L) that is closest to the output layer.

[0029] When the renderer 10 detects a change in the rendering conditions based on the rendering settings, it outputs a switching instruction CM for the specialized coefficients 41 to the specialized coefficient database 40. Based on the switching instruction CM, the specialized coefficient database 40 selects the specialized coefficients 41 corresponding to the changed rendering conditions and sends them to the CNN 30.

[0030] The CNN 30 switches the specialization coefficient 41 in accordance with the switching of the rendering conditions. The CNN 30 converts the image of the current frame acquired from the scaler 20 into the input frame I C The CNN 30 acquires the input frame I C and the intermediate feature I obtained in the inference of a past frame (for example, the frame immediately before the current frame). H and based on the input frame I C Make inferences regarding.

[0031] The CNN 30 includes multiple convolutional layers 31 (“Conv 1 "~"Conv last CNN 30 includes a final convolutional layer 31L ("Conv last The intermediate feature I H as recurrent data and input frame I C The CNN 30 outputs the image obtained by the inference (inference image) as an output frame I O The CNN 30 outputs the intermediate feature I input to the final convolutional layer 31L in the inference of the current frame. H to the next input frame I C This is obtained as recurrent data for inference.

[0032] FIG. 7 shows the intermediate feature I H FIG.

[0033] Intermediate feature I H indicates the internal feature of the CNN 30. In FIG. 7, C is input to the CNN 30, and the feature z n,l Here, "n" represents the number of dimensions, and "l" represents the layer number. The CNN 30 is expressed with four layers. The final output of the CNN 30 is z 1,3 ~z 64,3 7, the number of dimensions is 64, but the number of dimensions is determined by the configuration of the CNN 30 and does not necessarily have to be 64.

[0034] In the RNN structure, the inference accuracy is improved by using such multidimensional features rather than recursively using RGB images as input. In the RNN structure of the present disclosure, the intermediate feature I H is the multidimensional feature (z 1,l ~z n,l In practice, it is desirable to use the feature immediately before the final convolutional layer 31L (the output of the third layer in the example of FIG. 7 ), which is closer to the output, i.e., which is a feature that better represents the training data.

[0035] 6 , the CNN 30 switches the specialization coefficients 41 when the rendering resolution is switched, as a switching of the rendering conditions. When the rendering conditions are switched, the CNN 30 switches the intermediate feature I H The CNN 30 acquires the recurrent data and the input frame I generated based on the rendering conditions after the rendering conditions are switched. C and based on the input frame I C Make inferences regarding.

[0036] Intermediate feature I H The motion compensation used for the motion vector can be applied to the motion vector. The motion vector has a vector related to the spatial movement of the object between frames. The motion compensation is performed as a process of predicting the post-motion data from the pre-motion data. The post-motion compensation intermediate feature I H By inputting this to the CNN 30, it is possible to ensure consistency between frames.

[0037] 5. Control of the Format of Recurrent Data by Fixing the Coefficients of the Last Block FIG. 8 is a diagram for explaining control of the format of recurrent data by fixing the coefficients of the last block LB.

[0038] The last block LB refers to a group of layers of a neural network from the final convolutional layer 31L onward. In the present disclosure, learning is performed for each rendering resolution with the coefficients of the last block LB fixed as pre-trained coefficients. The coefficients of the model for each rendering resolution obtained by learning are acquired as specialized coefficients 41.

[0039] For example, each specialized coefficient 41 uses a group of images rendered under the corresponding rendering conditions as training data. The group of neural network layers from the final convolutional layer 31L onward is referred to as the last block LB. Each specialized coefficient 41 is obtained by learning the coefficients of layers other than the last block LB while fixing the coefficient value of the last block LB. The commonality of the last block LB is achieved by reusing the coefficient value of the last block LB obtained in learning at a certain rendering resolution when learning at another rendering resolution.

[0040] 9 is a diagram showing the relationship between the filters of the final convolutional layer 31L and the feature frame FE. The final convolutional layer 31L includes a filter for "ch R" ("filter A"), a filter for "ch G" ("filter B"), and a filter for "ch B" ("filter C"). "ch R", "ch G", and "ch B" refer to the channels for red (R), green (G), and blue (B), respectively. The output frame I O includes "ch R", "ch G" and "ch B".

[0041] FIG. 10 is a diagram showing the flow of outputting "ch R." FIG. 10 shows an example in which eight channels FE-1 to FE-8 are input as a feature frame FE. Eight filters 32-1 to 32-8 constituting "filter A" are applied to each channel. The convolution results for each channel are finally added together and output as "ch R." If the types and order of filters 32-1 to 32-8 are fixed, the types and order of channels FE-1 to FE-8 are also uniquely constrained. As a result, the format of the feature frame FE is unified.

[0042] 6. Improving Consistency of Recurrent Data When Switching Models FIGS. 11 and 12 are diagrams showing switching of models in response to a change in resolution.

[0043] 11 and 12, the rendering resolution is switched between frames 3 and 5. In frame 3, the model with resolution A ("Model A") switches to a model with resolution B ("Model B"), and in frame 5, the model with resolution B ("Model B") switches to a model with resolution C ("Model C").

[0044] When the rendering resolution is switched, the CNN 30 switches the coefficients to specialized coefficients 41 specialized for the rendering resolution after the switch. H is acquired as recurrent data, and inference is performed using the specialization coefficient 41 after switching.

[0045] For example, in FIG. P,Q " is the P-th input frame I generated at resolution Q. C and "Out P " is the Pth frame output frame I O The CNN 30 receives the second input frame I 2,A The CNN 30 performs inference on the image obtained by inference using a specialized coefficient 41 specialized for resolution A. 2 Output as

[0046] In the third frame, the CNN 30 switches the coefficients to specialized coefficients 41 specialized for resolution B. The CNN 30 uses the intermediate feature I obtained in the inference of the second frame. H is acquired as recurrent data, and the third input frame I 3,B Inference is made regarding intermediate feature I His acquired using the specialization coefficient 41 specialized for resolution A. Therefore, normally, a format mismatch occurs between the CNN 30 to which the specialization coefficient 41 specialized for resolution B is applied. However, in the present disclosure, the intermediate feature I H Since the format is unified, there is no format mismatch when switching rendering resolution.

[0047] 13 is a diagram showing the change in inference accuracy when the recurrent data format is unified between two models. As described above, when switching from a model with resolution A to a model with resolution B, no format mismatch occurs. Therefore, a correct inference result is obtained. As a result, there is no decrease in PSNR when switching between models.

[0048] 14 to 16 are diagrams showing an example of a method for learning a specialization coefficient. The learning process of this example is applied to VSR (Video Super Resolution). t}, {GT t}} n " indicates the training dataset for each rendering resolution. "n" is a label indicating the rendering resolution. "t" is a label indicating the frame number. "I" indicates the input to the model. "GT" indicates the ground truth data. "Out" indicates the output of the model (inferred image).

[0049] The learning process includes three steps ("STEP 0", "STEP 1", and "STEP 2"). STEP 0 is the preparation of learning data and architecture (model). In STEP 0, a learning dataset and Feature Recurrent Video AI architectures (number of rendering resolutions + 1) are prepared.

[0050] STEP 1 is the learning of the specialization coefficient 41 for a specific rendering resolution. In STEP 1, a rendering resolution (e.g., m) to be learned is arbitrarily selected. Then, the data set {{I t}, {GT t}} m Using the model VSR m The model VSR after learning is m The coefficients of the last block LB are transferred as fixed coefficients to the model of another rendering resolution.

[0051] STEP 2 is the learning of the specialization coefficient 41 for another rendering resolution (for example, n). In STEP 2, the learned model VSR m The coefficients of the last block LB of the model VSR of the other rendering resolution to be learned are extracted as fixed coefficients. n As a result, the last block LB is fixed to the fixed coefficient and the unlearned model VSR' is obtained. n The model VSR' is obtained. n For the dataset {{I t}, {GT t}} n The trained model VSR'' n is obtained.

[0052] 8. Processing Flow During Inference FIG. 17 is a diagram showing an example of a processing flow during inference.

[0053] The renderer 10 performs rendering based on the rendering settings. The scaler 20 scales the rendered image as needed. The CNN 30 uses an input I whose rendering resolution changes depending on the rendering settings. t is received (step S11).

[0054] At the start of inference (first frame), the intermediate feature I H Therefore, the intermediate feature I H For example, the input I 1 If the rendering resolution of label i, then model VSR'' as CNN30. i Model VSR'' is used. i The model VSR'' acquires dummy data in which all parameters of each feature frame FE are set to blank (data indicating 0). iis the input I 1 Inference is performed based on the dummy data (step S12). In Fig. 17, "O" indicates the output of the model (inferred image). "RFM" indicates the intermediate feature I obtained from the intermediate layer of the model. H Shows.

[0055] Assume that the rendering resolution is switched at the (t+1)th frame. The input I t+1 When the rendering resolution of label j, we use CNN30 as the model VSR'' j Model VSR'' is used. j is the intermediate feature RFM, which is the output of the previous process. t receives the current frame input I t+1 and intermediate feature RFM t Inference is made based on this (step S13).

[0056] Intermediate feature RFM t is the model before the switch VSR'' i Since this is the inference result, the model after switching should be VSR'' j However, in the method of the present disclosure, this format mismatch is resolved, so that problems such as color shift do not occur in the inference results. j outputs highly accurate inference results. t} can be output (step S14).

[0057] 9. Example of Hardware Configuration FIG. 18 is a diagram illustrating an example of the hardware configuration of the information processing device 1. As shown in FIG.

[0058] The information processing of the information processing device 1 is realized by, for example, a computer 1000. The computer 1000 has a CPU (Central Processing Unit) 1100, a RAM (Random Access Memory) 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0059] The CPU 1100 operates and controls each component based on a program (program data 1450) stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the program stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0060] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 starts up, as well as programs that depend on the hardware of the computer 1000 .

[0061] The HDD 1400 is a non-transitory computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the HDD 1400 is a recording medium that records an information processing program according to an embodiment as an example of program data 1450.

[0062] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0063] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display device, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.

[0064] For example, when the computer 1000 functions as the information processing device 1 according to the embodiment, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200 to realize the functions of the aforementioned components. The information processing program, various models, and various data according to the present disclosure are stored in the HDD 1400. The CPU 1100 reads and executes program data 1450 from the HDD 1400. Alternatively, the CPU 1100 may acquire these programs from another device via an external network 1550.

[0065] [10. Effects] The information processing device 1 has a specialized coefficient database 40 and a CNN 30. The specialized coefficient database 40 stores a plurality of specialized coefficients 41 in association with the rendering conditions that are the subject of the specialized processing. The specialized coefficients 41 share the coefficient value of the final convolutional layer 31 that is closest to the output layer. The CNN 30 switches the specialized coefficients 41 in accordance with the switching of the rendering conditions. The CNN 30 calculates the intermediate feature I input to the final convolutional layer 31. H as recurrent data and input frame I CIn the information processing method of the present disclosure, the processing of the information processing device 1 is executed by a computer 1000. The computer-readable non-transitory storage medium of the present disclosure stores a program that causes the computer 1000 to realize the processing of the information processing device 1.

[0066] According to this configuration, regardless of which specialization coefficient 41 is used, the intermediate feature quantity I H The intermediate feature I, which is the recurrent data, is obtained. H Since the homogeneity of the images is maintained, the accuracy of inference is less likely to decrease when the rendering conditions are changed.

[0067] In the inference process using the specialization coefficients 41 before the rendering conditions are switched, the CNN 30 calculates the intermediate feature I H The CNN 30 acquires the recurrent data and the input frame I generated based on the rendering conditions after the rendering conditions are switched. C and based on the input frame I C Make inferences regarding.

[0068] According to this configuration, the intermediate feature I output before the specialization coefficient 41 is switched H is used for the inference of the CNN 30 after switching the specialization coefficient 41. Normally, the input intermediate feature I H However, in the configuration of the present disclosure, the intermediate feature I H Therefore, the accuracy of inference is less likely to decrease.

[0069] A group of images rendered under the corresponding rendering conditions is used as learning data for learning each specialized coefficient 41. If the group of layers of the neural network from the final convolutional layer 31 onwards is referred to as a last block LB, each specialized coefficient 41 is obtained by learning the coefficients of layers other than the last block LB while fixing the coefficient values ​​of the last block LB.

[0070] According to this configuration, the final convolution layer 31 can reliably acquire multiple common specialized coefficients 41.

[0071] The CNN 30 switches the specialization coefficient 41 when the rendering resolution is switched as a switching of the rendering condition.

[0072] According to this configuration, the specialization coefficient 41 can be switched in response to switching of the rendering resolution.

[0073] The information processing device 1 includes a scaler 20. The scaler 20 scales a rendering image in accordance with the architecture of the CNN 30 to generate an input frame I C Output as

[0074] This configuration allows for highly accurate inference results to be obtained.

[0075] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0076] [Additional Notes] The present technology may also be configured as follows: (1) An information processing device comprising: a specialized coefficient database that stores a plurality of specialized coefficients, in which coefficient values ​​of a final convolutional layer closest to an output layer are commonized, in association with the rendering condition that became the target of the specialization; and a CNN that switches the specialized coefficients in accordance with a switch in the rendering condition, and performs inference on an input frame by using intermediate features input to the final convolutional layer as recurrent data. (2) The information processing device described in (1), wherein the CNN acquires, as the recurrent data, the intermediate features to be input to the final convolutional layer in an inference process that uses the specialized coefficients before the rendering condition is switched, and performs inference on the input frame based on the recurrent data and the input frame generated based on the rendering condition after the rendering condition is switched. (3) The information processing device according to (1) or (2), wherein each specialized coefficient is obtained by using a group of images rendered under the corresponding rendering condition as learning data, and learning coefficients of layers other than the last block, which is a group of layers of the neural network from the final convolutional layer onwards, while fixing coefficient values ​​of the last block. (4) The information processing device according to any one of (1) to (3), wherein the CNN switches the specialized coefficients by using a switch in rendering resolution as a switch in the rendering condition. (5) The information processing device according to (4), further comprising a scaler that outputs an image obtained by scaling a rendered image in accordance with an architecture of the CNN as the input frame. (6) An information processing method executed by a computer, comprising: storing a plurality of specialized coefficients, in which coefficient values ​​of a final convolutional layer closest to an output layer are common, in association with the rendering conditions that are the subject of the specialization; switching the specialized coefficients in accordance with switching of the rendering conditions; and performing inference regarding an input frame by using intermediate features input to the final convolutional layer as recurrent data.(7) A computer-readable non-transitory storage medium storing a program that causes a computer to perform the following: storing multiple specialized coefficients, each having a common coefficient value of the final convolutional layer closest to the output layer, in association with the rendering conditions that are the subject of the specialization; switching the specialized coefficients in accordance with a switch in the rendering conditions; and performing inference regarding an input frame by using intermediate features input to the final convolutional layer as recurrent data.

[0077] 1 Information processing device 20 Scaler 30 CNN 31 Convolution layer 40 Specialized coefficient database 41 Specialized coefficient I C Input Frame I H Intermediate feature LB Last block

Claims

1. An information processing device having: a specialized coefficient database that stores a plurality of specialized coefficients, in which the coefficient values ​​of a final convolutional layer closest to an output layer are commonized, in association with the rendering conditions that were the subject of the specialization; and a CNN that switches the specialized coefficients in accordance with a switch in the rendering conditions, and performs inference regarding an input frame by using intermediate features input to the final convolutional layer as recurrent data.

2. The information processing device of claim 1, wherein the CNN acquires the intermediate features to be input to the final convolutional layer as the recurrent data in an inference process using the specialization coefficients before the rendering conditions are switched, and performs inference regarding the input frame based on the recurrent data and the input frame generated based on the rendering conditions after the rendering conditions are switched.

3. The information processing device according to claim 1, wherein each specialized coefficient is obtained by using a group of images rendered under the corresponding rendering conditions as learning data, and learning the coefficients of layers other than the last block, which is a group of layers of the neural network after the final convolutional layer, while fixing the coefficient values ​​of the last block.

4. The information processing device according to claim 1, wherein the CNN switches the specialization coefficient when a rendering resolution is changed as a switching of the rendering condition.

5. The information processing device according to claim 4, further comprising a scaler that outputs an image obtained by scaling a rendering image in accordance with an architecture of the CNN as the input frame.

6. An information processing method executed by a computer, comprising: storing a plurality of specialization coefficients, the coefficient values ​​of which of a final convolutional layer closest to an output layer are common, in association with the rendering conditions that are the subject of the specialization; switching the specialization coefficients in accordance with switching of the rendering conditions; and performing inference regarding an input frame by using intermediate features input to the final convolutional layer as recurrent data.

7. A computer-readable non-transitory storage medium storing a program that causes a computer to perform the following operations: storing a plurality of specialization coefficients, in which the coefficient values ​​of the final convolutional layer closest to the output layer are common, in association with the rendering conditions that are the subject of the specialization; switching the specialization coefficients in accordance with switching of the rendering conditions; and performing inference regarding the input frame by using the intermediate features input to the final convolutional layer as recurrent data.

Citation Information

Patent Citations

  • Data processing apparatus, data processing method and data processing program

    JP2020119312A