Information processing system, information processing method, and computer-readable non-transitory storage medium

WO2026204723A1PCT designated stage Publication Date: 2026-10-01SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/010917
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-19
Publication Date
2026-10-01

Smart Images

  • Figure JP2026010917_01102026_PF_FP_ABST
    Figure JP2026010917_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing system comprises a style feature extraction unit and a style conversion network. The style feature extraction unit extracts a style feature quantity from style information, which represents the appearance of an image. The style conversion network acquires input data in which different viewpoint images are assigned to respective channels. The style conversion network converts, on the basis of the result of attention processing in the channel direction, which represents the direction in which the channels are merged, the style of each of the viewpoint images into the style indicated by the style feature quantity.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method, and computer-readable non-temporary storage medium

[0001] The present invention relates to an information processing system, an information processing method, and a computer-readable non-temporary storage medium.

[0002] 3D assets are created by reconstructing images taken from various viewpoints (viewpoint images). Recapturing a desired scene is time-consuming, so it is necessary to edit existing 3D assets to obtain 3D assets with different styles (sky color, brightness, weather, etc.) (style-transformed assets).

[0003] Leon A. Gatys et al (2015), “A Neural Algorithm of Artistic Style”

[0004] To obtain a style-transformed asset, it is necessary to render an existing 3D asset from various viewpoints, style-transform each of the numerous viewpoint images generated, and then reconstruct the 3D model. However, simply style-transforming each viewpoint image results in a set of style-transformed images with low correlation between viewpoints. Therefore, performing 3D reconstruction using such a set of style-transformed images will not yield a high-quality style-transformed asset.

[0005] Therefore, this disclosure proposes an information processing system, an information processing method, and a computer-readable non-temporary storage medium that are capable of acquiring high-quality style conversion assets.

[0006] According to this disclosure, an information processing system is provided, comprising: a style feature extraction unit that extracts style features from style information representing the appearance of an image; and a style conversion network that acquires input data in which different viewpoint images are assigned to each channel, and converts the style of each viewpoint image to the style indicated by the style features based on the result of attention processing in the channel direction representing the direction in which the channels are joined. Furthermore, according to this disclosure, an information processing method is provided in which the information processing of the information processing system is performed by a computer, and a computer-readable non-temporary storage medium is provided that stores a program that enables the computer to implement the information processing of the information processing system.

[0007] This figure shows an example of the configuration of the information processing system disclosed herein. This figure shows an example of the configuration of the style conversion unit. This figure shows an example of a conventional style conversion network. This figure shows an example of a style conversion image generated by a conventional style conversion network. This figure shows the behavior of Spatial Attention. This figure shows the behavior of Spatial Attention. This figure shows the behavior of Spatial-view attention. This figure shows the behavior of Spatial-view attention. This figure shows a modified example of the style conversion network. This figure shows another example of the configuration of the information processing system. This figure shows an example of a processing flow including viewpoint generation processing. This figure shows an example of the hardware configuration of the information processing system.

[0008] Embodiments of the present disclosure will be described in detail below with reference to the drawings. In each of the following embodiments, the same parts will be denoted by the same reference numerals, and redundant descriptions will be omitted.

[0009] The explanation will proceed in the following order: [1. Information Processing System Configuration] [2. Style Conversion Processing] [2-1. Spatial-View Attention] [2-2. Comparison with Conventional Style Conversion Networks] [3. Modified Style Conversion Networks] [4. Viewpoint Generation Processing] [5. Hardware Configuration Examples]

[0010] [1. Configuration of the Information Processing System] Figure 1 is a diagram showing an example of the configuration of the information processing system 1 of the present disclosure. The information processing system 1 includes a rendering unit 10, a style conversion unit 20, a 3D reconstruction unit 30, and an asset database 40.

[0011] The rendering unit 10 renders 3D assets AS from various viewpoints based on viewpoint information VW. The asset database 40 stores the 3D assets AS. 3D assets AS refer to material data used when creating 3D images. 3D assets AS include data of 3D models used in computer graphics, and 3D data scanned from real space, such as 3D Gaussian splatting.

[0012] 3D Gaussian splatting is a technique that reconstructs a 3D scene based on 2D images acquired from multiple viewpoints, generating an image from a new viewpoint. For information on 3D data generation methods using 3D Gaussian splatting, please refer to Non-Patent Document A below.

[0013] [Non-patent Document A] Thomas Muller et al (2023), “3D Gaussian Splatting for Real-Time Radiance Field Rendering.”

[0014] Specifically, point cloud data of objects in the space is acquired using 2D images of the space to be scanned, taken from various viewpoints. The point cloud data is calculated using Structure from Motion (SfM) technology. Using this point cloud data as initial values, Gaussian parameters, represented by the number and density of Gaussian distributions, are optimized according to the complexity of the scene. The final Gaussian parameters are stored as the information below.

[0015] • Mean value (μ): Indicates the center position of the Gaussian. • Covariance matrix: Determines the shape and directionality of the Gaussian and reflects the structure within the scene. • Opacity: Indicates the transparency of each Gaussian and represents the depth and overlap of the scene. • Color (RGB value): Holds the color information of each Gaussian and reproduces the colors of the scene.

[0016] The viewpoint information VW defines information about the position and orientation of each viewpoint to be rendered. For example, the rendering unit 10 projects the Gaussian parameters held by the 3D asset AS onto a 2D plane based on the viewpoint's position and orientation information. This enables the rendering process of the 3D asset AS to be executed.

[0017] The rendering unit 10 obtains viewpoint image VIs corresponding to each viewpoint by rendering the 3D asset AS. The rendering unit 10 generates multi-channel image data by assigning a different viewpoint image VI to each channel. The rendering unit 10 outputs the generated multi-channel image data as data for style conversion processing.

[0018] The style conversion unit 20 acquires the multi-channel image data generated by the rendering unit 10 as input data IN. The style conversion unit 20 performs style conversion processing on the input data IN based on the style information ST. Style information ST refers to information that represents the appearance of the image as desired or specified by the user. For example, style information ST may be presented as a text prompt, a reference image, a visual context indicating the direction of conversion, etc.

[0019] The style conversion unit 20 converts the style of each channel's viewpoint image VI to the style indicated by the style information ST through style conversion processing. The style conversion unit 20 generates multi-channel image data by assigning the multiple viewpoint images VI after style conversion to different channels, and outputs this as output data OUT for 3D reconstruction processing.

[0020] The 3D reconstruction unit 30 acquires the style-transformed viewpoint image VI included in the output data OUT as a style-transformed image CI. The 3D reconstruction unit 30 reconstructs the 3D scene using multiple style-transformed image CIs corresponding to each channel and generates a style-transformed asset CA. A style-transformed asset CA refers to a new 3D asset obtained by style-transforming the original 3D asset AS based on style information ST. As a reconstruction method, methods such as 3D Gaussian splatting can be used.

[0021] [2. Style Conversion Process] Figure 2 shows an example of the configuration of the style conversion unit 20.

[0022] The style conversion unit 20 includes a style feature extraction unit 21 and a style conversion network 22. The style feature extraction unit 21 extracts style feature quantities SF from style information ST. The style feature extraction unit 21 outputs the extracted style feature quantities SF to the intermediate layer of the style conversion network 22.

[0023] For example, style information ST is presented as text prompts such as "clear skies" or "cloudy." As a method for extracting style features SF from text prompts, known language-image feature transformation methods used in GPT (Generative Pretrained Transformer) (registered trademark), such as CLIP (Contrastive Language-Image Pre-training), can be employed.

[0024] The style conversion network 22 acquires a set of multi-channel viewpoint images from the rendering unit 10 as input data IN. The input data IN is composed of multi-channel image data, with a different viewpoint image VI assigned to each channel. Based on the results of attention processing in the channel direction, the style conversion network 22 converts the style of each viewpoint image VI into the style indicated by the style feature SF.

[0025] In this configuration, multiple viewpoint images VI are multi-channelized and input to the style conversion network 22. Since style conversion is performed while simultaneously referencing information from multiple viewpoints, a set of style-converted images is obtained in which the correlation between viewpoints is maintained. As a result, the accuracy of 3D reconstruction is improved, and a high-quality style-converted asset CA is obtained.

[0026] For example, the style transformation network 22 has a multi-layer structure including a Convolution Layer 23 and a Spatial-View Attention Layer 24. The viewpoint images rendered from multiple viewpoints are all simultaneously input to the Convolution Layer 23 as input data IN.

[0027] Convolution Layer 23 extracts intermediate features from the input data IN by performing a convolution operation on the input data IN using a CNN (Convolution Neural Network). Style transfer network 22 replaces some of the extracted intermediate features with style features SF in order to perform style transfer.

[0028] Spatial-View Attention Layer 24 performs integrated attention processing in both the channel and spatial directions on intermediate features, where some features have been replaced by style features (SF). "Channel direction" refers to the direction in which multiple channels corresponding to multiple viewpoint images (VI) are combined. "Performing integrated attention processing in both the channel and spatial directions" means performing Channel Attention and Spatial Attention in parallel.

[0029] When the style feature vector SF is input to the style transfer network 22, attention processing is repeatedly performed multiple times by the alternately stacked Convolution Layer 23 and Spatial-View Attention Layer 24. The style transfer network 22 simultaneously obtains a group of style-transformed images as a result of performing style transfer on each input viewpoint image VI.

[0030] [2-1. Spatial-View Attention] An attention mechanism refers to a mechanism that dynamically specifies where to focus attention in input data. Information about correlation between data can be obtained by performing attention processing.

[0031] As attention processing for images, Channel Attention and Spatial Attention are conceivable. Channel Attention is attention processing that detects correlation in the channel direction. Spatial Attention is attention processing that detects correlation for each location within an image.

[0032] In the present disclosure, each channel corresponds to a rendering viewpoint. In the example of FIG. 2, viewpoint numbers (1 to N) are added after the codes of viewpoint images VI and style converted images CI. Correlation between channels is grasped as correlation between viewpoints. Therefore, Channel Attention can also be described as View Attention.

[0033] The Spatial-View Attention Layer 24 performs calculation while mutually referencing data of each multi-channeled viewpoint image VI. Therefore, the Spatial-View Attention Layer 24 can implement Spatial-View Attention processing which combines Channel Attention (View Attention) and Spatial Attention. According to this configuration, a style converted image group in which both the correlation between viewpoints and the correlation between objects within an image are favorably retained can be obtained.

[0034] Hereinafter, the operation of the Spatial-View Attention Layer 24 will be specifically described. It is assumed herein that the intermediate feature amount X output by the Convolution Layer 23 is four-dimensional tensor data of N×H×W×C. At this time, "N" indicates the number of image data in the viewpoint direction. "H" indicates the height of the feature amount. "W" indicates the width of the feature amount. "C" indicates the number of channels of the feature map.

[0035] <1: Patch Division and Patch Embedding> The input intermediate feature X can be expressed as X∈R^(N×H×W×C). From this, each feature of N pieces of image data in the viewpoint direction is divided into patches of fixed size. Patch Embedding is performed with the linearly transformed features of these patches as one-dimensional vectors, thereby generating a token sequence Z_pt.

[0036] Z_pt = PatchEmbed(X) Z_pt∈R^(M×D) M = N×P P: the number of patches per viewpoint (for example, for 16×16 patches, P=(H×W) / 16^2. D: the number of dimensions of patch embedding.

[0037] <2: Position Embedding> In order to incorporate spatial and viewpoint information, a process of adding position embedding to the patch embedding token sequence is performed. Z = Z_pt + E_pos E_pos: learnable position embedding matrix.

[0038] <3: Generation of Query, Key and Value> For each token sequence Z, a process of applying a learnable linear transformation to generate query (Q), key (K) and value (V) is performed. Q = Z*W_Q K = Z*W_K V = Z*W_V Z: input token sequence (dimension N×D). W_Q, W_K, W_V: learnable weight matrices (dimension: D×d_k). Q, K, V: query, key and value matrices (dimension: M×d_k).

[0039] <4: Calculation of Spatio-Temporal Attention> A process is performed in which the inner product of query Q and key K is calculated, scaling and softmax are applied to obtain attention weights, and then the weights are applied to value V.

[0040] Attention(Q, K, V) = softmax((Q*K^T) / sqrt(d_k))*V Q*K^T: scaled inner product (measures the similarity between each token). sqrt(d_k): scaling factor. softmax: a function that converts attention scores into a probability distribution. Attention(Q, K, V): attention output (dimension: M×d_k).

[0041] <5: Multi-head attention> This process calculates h independent attention heads, combines them, and applies a linear transformation.

[0042] MultiHead(Q,K,V) = Concat(head_1,...,head_h) * W_O head_i = Attention(Q_i,K_i,V_i) head_i: Attention output calculated by the i-th head. Concat: Concatenates the outputs of all heads in the dimensional direction (dimension: N × (h * d_k)). W_O: Linear transformation matrix to return the combined output to the original dimension (dimension: (h * d_k) × D). MultiHead(Q,K,V): Final output of multi-head attention (dimension: N × D).

[0043] <6: Normalization and Residual Connection> The output of the multi-head attention is added to the original input token sequence (residual connection), and normalization is applied.

[0044] Z' = LayerNormal(Z + MultiHead(Q, K, V)) Z: Input token sequence (same as <2> above). MultiHead(Q, K, V): Output of multi-head attention calculated in <5> above. Z + MultiHead(Q, K, V): Residual connection. LayerNormal: Function to normalize the distribution of features. Z': Normalized token sequence (input to the next layer).

[0045] <7: Conversion to Spatial Structure> In order to convert to the format expected by Convolution Layer 23 (a 4-dimensional tensor), the token sequence is rearranged into its original spatial structure.

[0046] Z′conv=Reshape(Z′, shape=(N,H,W,D))∈R^(N×H×W×D)

[0047] [2-2. Comparison with Conventional Style Conversion Networks] Figure 3 shows an example of a conventional style conversion network 22C. Figure 3 shows the configuration proposed in Non-Patent Document B below.

[0048] [Non-Patent Literature B] Alexey Dosovitzky et al (2020), “An Image is Word 16x16 Words: Transformers for Image Recognition at Scale.”

[0049] The conventional style transfer network 22C performs style transfer processing for each viewpoint image VI. The Convolution Layer 28 extracts intermediate features from a single viewpoint image VI. The Spatial Attention Layer 29 replaces some of the extracted intermediate features with style features SF and performs style transfer processing based on the results of Spatial Attention.

[0050] The style transfer network 22C repeats the above process for the number of input viewpoint images VI. In the example in Figure 3, the number of viewpoint images VI is N. The style transfer network 22C repeats the extraction of intermediate features, replacement of intermediate features, and style transfer process N times to generate N style-transformed images CI.

[0051] In conventional style transfer networks 22C, style transfer processing is performed for each viewpoint image VI. However, unlike the method disclosed herein, it does not perform style transfer processing simultaneously on multiple multi-channel viewpoint images VI. As a result, a group of style-transformed images with low correlation between viewpoints is output. Even if 3D reconstruction is performed using such a group of style-transformed images, a high-quality style-transformed asset CA cannot be obtained.

[0052] For example, the Spatial Attention Layer 29 in Figure 3 can learn spatial dependencies within an image, similar to the Spatial-View Attention Layer 24 of this disclosure shown in Figure 2, but it cannot learn dependencies in the viewpoint direction. This is because there is no connection between the Spatial and View layers. Specifically, in the patch splitting process described in <1> above, the intermediate features are set to Xs ∈ R^(H × W × C), so Self Attention does not work for viewpoint direction information. As a result, there is a problem in that a set of style-transformed images with high correlation between viewpoints cannot be output.

[0053] In contrast, as already explained, according to the operation of the Spatial-View Attention Layer 24 in this disclosure, Self Attention is applied to all tokens that span the viewpoint direction and the spatial axis. Therefore, it becomes possible to simultaneously model the dependencies between space and viewpoint direction, and to output a set of style-transformed images with high correlation between viewpoints.

[0054] Figure 4 shows an example of a style-transformed image CI generated by a conventional style-transformed network 22C. Figure 4 shows a viewpoint image VI viewed from "Viewpoint 1". 1 And viewpoint image VI as seen from "Viewpoint 2" 2 This is shown. Viewpoint image VI 1 and viewpoint image VI 2 These are each converted to a style that indicates "cloudy" weather. As a result, the image of the entire scene looks cloudy, and the style-converted image CI 1 CI 2 It will be acquired as follows.

[0055] In conventional style transfer processing, feature extraction is performed from a single viewpoint image (VI), and image features such as weather and time of day are controlled within a non-linear feature space. Therefore, when the viewpoint changes, the intermediate features input to Spatial Attention Layer 29 also change, potentially resulting in different shapes and colors even for the same subject. This is what is meant by "outputting a set of style-transformed images with low correlation between viewpoints."

[0056] On the other hand, the 3D reconstruction in the 3D reconstruction unit 30 is based on the premise that images of the "same subject" taken from various viewpoints are input, and it reconstructs the 3D representation by detecting parts that have the same shape and color among the viewpoint images VI. Therefore, if a group of style-transformed images with low correlation between viewpoints, as in the conventional method, is input, the procedure for detecting parts that have the same shape and color among the viewpoint images VI may fail, potentially reducing the accuracy of the 3D reconstruction.

[0057] In contrast, the style conversion process of this disclosure takes a group of images rendered from multiple viewpoints as input in a multi-channel format, performs style conversion on each individual rendered image, and outputs the result in the same multi-channel format. This structure allows for simultaneous reference of information from multiple viewpoints, and when combined with self-attention processing, it becomes possible to output a group of style-converted images with high correlation between viewpoints. Furthermore, by using a group of style-converted images with high correlation between viewpoints, the accuracy of reconstruction in the subsequent 3D reconstruction unit 30 is improved.

[0058] This point will be explained in more detail. Figures 5 and 6 illustrate the behavior of Spatial Attention. Here, a single viewpoint image VI and its intermediate features are divided into multiple fixed-size patch PTs (Image Patches). In Spatial Attention, the single viewpoint image VI is divided spatially, and the individual patch PTs obtained by the division are linearly transformed. The features obtained by linearly transforming the patch PTs are patch-embedded as one-dimensional vectors, and a token sequence is generated (see Figure 5).

[0059] By performing the processes described in <2> to <7> above on the generated token sequence, the feature quantities after Spatial Attention are obtained (see Figure 6). This is the result of calculating the degree of relevance between tokens in the spatial direction using Self Attention, and similar regions are represented as grouped feature quantities. These feature quantities are then output from the Spatial-Attention Layer 29, after which the final style-transformed image CI is obtained.

[0060] At this time, for example in a sky region, due to the effect of grouping similar regions, the sky region is converted with similar colors that are "spatially consistent".

[0061] However, although such Spatial Attention can ensure spatial consistency during style conversion within a single image, it cannot refer to the relevance for images in the view direction. Therefore, there arises a problem that consistency in the view direction cannot be ensured in style conversion.

[0062] The Spatial-View Attention Layer 24 according to the present disclosure is designed in view of the above problems. FIGS. 7 and 8 are diagrams showing the behavior of spatial-view attention.

[0063] In Spatial-View Attention, each viewpoint image VI (viewpoint image VI 1 , viewpoint image VI 2 ) is divided in the spatial direction, and each viewpoint patch PT obtained by the division 1 , PT 2 is linearly transformed. The feature quantity obtained by linearly transforming each viewpoint patch PT 1 , PT 2 is subjected to Patch Embedding as a one-dimensional vector, and a token sequence is generated (see FIG. 7).

[0064] By performing the processes according to the above <2> to <7> on the generated token sequence, a feature quantity after Spatial-View Attention is obtained (see FIG. 8). This is a result of tokens mutually calculating relevance through Self-Attention in both the spatial direction and the view direction, and is expressed as a feature quantity in which similar regions across both the spatial direction and the view direction are grouped together. Then, the feature quantity is output from the Spatial-View Attention Layer 24, after which a final group of style-converted images is obtained.

[0065] In this case, for example, in the sky area, similar regions are grouped together in both the spatial and viewpoint directions, resulting in the sky area being transformed into similar colors that are "spatially consistent and consistent across viewpoints."

[0066] [3. Modified Style Transfer Network] Figure 9 shows a modified style transfer network 22.

[0067] The style transfer network 50 in this example includes an Encoder 51, a Bottleneck U-net 52, and a Decoder 53. The Encoder 51 encodes the input data IN into the feature space. The Bottleneck U-net 52 performs style transfer on the encoded input data IN features.

[0068] Bottleneck U-net 52 is a neural network with a Bottleneck structure. Bottleneck U-net 52 has a multilayer structure including a Convolution Layer 23 and a Spatial-View Attention Layer 24 as shown in Figure 2. For example, Bottleneck U-net 52 is configured as a neural network that constitutes the intermediate part of a latent diffusion model.

[0069] Decoder 53 decodes the results of the style transfer processing performed in the feature space into the image space. As a result of the decoding, a multi-channel style transfer image CI (style transfer image group) is generated as output data OUT for 3D reconstruction.

[0070] In this example, the input data IN is encoded and output to the subsequent processing stage. With this configuration, style transformation (Convolution, Spatial-View Attention) is performed in the feature space. Therefore, it is possible to achieve higher accuracy in style transformation than when style transformation is performed in the image space.

[0071] [4. Viewpoint Generation Process] Figure 10 shows another example of the configuration of the information processing system. The difference in this example from the configuration in Figure 1 is that the information processing system 2 has a viewpoint generation unit 60. The following explanation will focus on the differences from the configuration in Figure 1.

[0072] The viewpoint generation unit 60 acquires multiple style-converted images CI from the style conversion unit 20. The viewpoint generation unit 60 applies interpolation processing to the acquired multiple style-converted images CI to generate at least one interpolated image that can be seen from an intermediate position between viewpoints (intermediate viewpoint) (viewpoint generation processing). The 3D reconstruction unit 30 uses the multiple style-converted images CI acquired from the style conversion unit 20 and the at least one interpolated image generated by the viewpoint generation unit 60 to generate a 3D asset (style-converted asset CA) with the converted style.

[0073] Any interpolation process can be used, as long as it can create an interpolated image of an intermediate viewpoint that is not present in the input, while maintaining the three-dimensional features obtained from each style-transformed image CI. For example, one method is to perform correspondence point detection using two style-transformed image CIs for two specific viewpoints, perform motion estimation, and output an interpolated image. Alternatively, as shown in Non-Patent Document C below, a video generation model that generates videos from pre-trained images can be used to generate a video of camera work connecting two viewpoints, and this video can be output as an interpolated image.

[0074] [Non-patent Document C] Dejia Xu et al (2024), “CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation.”

[0075] The following describes an example of an information processing method that includes viewpoint generation processing. Figure 11 is a diagram showing an example of a processing flow that includes viewpoint generation processing.

[0076] The rendering unit 10 acquires viewpoint information VW and 3D asset AS (step S11). In the example in Figure 10, information on the position and orientation of K viewpoints is acquired as viewpoint information VW. The rendering unit 10 generates K viewpoint images VI based on the information of the specified K viewpoints.1 ~VI K This is generated (step S12).

[0077] The style conversion unit 20 converts K viewpoint images VI 1 ~VI K Perform style conversion processing on K style converted images CI 1 ~CI K The viewpoint generation unit 60 generates (step S13). The viewpoint generation unit 60 sets (N-K) intermediate viewpoints and generates interpolated images corresponding to each intermediate viewpoint. The viewpoint generation unit 60 treats the interpolated images as new style conversion images CI as seen from the intermediate viewpoints.

[0078] The viewpoint generation unit 60 uses the newly generated (N-K) style conversion image CI (interpolated image) to process the K style conversion image CI obtained from the style conversion unit 20. 1 ~CI K By adding this, a total of N style conversion image CI 1 ~CI N The 3D reconstruction unit 30 obtains the N style conversion images CI obtained from the viewpoint generation unit 60 (step S14). 1 ~CI N A style conversion asset CA is generated based on this (step S15).

[0079] This configuration generates high-quality style-transformed assets (CAs) using numerous style-transformed image CIs (Composition Indicators) that include interpolated images. The addition of interpolated images reduces the number of style-transformed image CIs required to generate the CA. Therefore, the computational load of the style-transformed process can be reduced.

[0080] The advantages of adjusting the number of style-converted images (CIs) through viewpoint generation processing are as follows: The style conversion unit 20 can convert the style of each viewpoint image (VI) while maintaining the correlation of viewpoint directions through the operation of the style conversion network 22. However, if the number of input and output images becomes very large, the processing cost increases, and it becomes difficult to maintain the correlation of all viewpoint directions. Therefore, it is desirable to minimize the number of input and output images as much as possible.

[0081] On the other hand, the 3D reconstruction unit 30 requires images rendered from as many viewpoints as possible in order to perform 3D reconstruction with higher accuracy. Therefore, it is desirable to reduce the number of input and output images, that is, to increase the number of style conversion images CI as much as possible.

[0082] The viewpoint generation unit 60 is applied in consideration of the above-mentioned problems. First, the style conversion unit 20 generates K style-converted image CIs with high-precision correlation between viewpoints at low cost. Next, the viewpoint generation unit 60 generates an interpolated image from the K style-converted image CIs while maintaining three-dimensional features, increasing the number of style-converted image CIs from K to N. Then, high-precision 3D reconstruction is performed using the increased number of style-converted image CIs from K to N, generating a style-converted 3D asset. This simultaneously solves the problems of the style conversion unit 20 and the 3D reconstruction unit 30.

[0083] [5. Hardware Configuration Example] Figure 12 shows an example of the hardware configuration of information processing systems 1 and 2.

[0084] Information processing systems 1 and 2 can be implemented by a computer 1000 as shown in Figure 12. The computer 1000 includes a processing circuit 1100, RAM 1200, ROM 1300, secondary storage device 1400, communication interface 1500, input / output interface 1600, display unit 1700, camera unit 1800, microphone 1900, and speaker 2000. The various parts of the computer 1000 are connected by a bus 1050.

[0085] The processing circuit 1100 operates based on a program stored in the ROM 1300 or secondary storage device 1400, and controls each part. For example, the processing circuit 1100 loads the program stored in the ROM 1300 or secondary storage device 1400 into the RAM 1200 and executes processing corresponding to various programs.

[0086] ROM 1300 stores boot programs such as the BIOS (Basic Input Output System) executed by the processing circuit 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.

[0087] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily records programs executed by the processing circuit 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that records programs for each process of the information processing systems 1 and 2 according to the embodiments of this disclosure, which are examples of program data 1450.

[0088] The communication interface 1500 is an interface for the computer 1000 to connect to the external network 1550. For example, the processing circuit 1100 can receive data from other devices or transmit data it has generated to other devices via the communication interface 1500.

[0089] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the processing circuit 1100 receives data from input devices such as a microphone 1900 or a touch panel via the input / output interface 1600. The processing circuit 1100 also transmits data to output devices such as a display unit 1700 or a speaker 2000 via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Discs), magneto-optical recording media such as MOs (Magneto-Optical Discs), tape media, magnetic recording media, or semiconductor memory.

[0090] The display unit 1700 is an interface for displaying information processed by the computer 1000. The display unit 1700 is, for example, a liquid crystal display or an organic electroluminescent display (OLED display). Alternatively, the display unit 1700 may be a touch panel display device or an image projection device.

[0091] The camera unit 1800 is an interface for the computer 1000 to capture images. The microphone 1900 is an interface for the computer 1000 to capture sound. The speaker 2000 is an interface for the computer 1000 to output processed sound. The various parts of the computer 1000 are connected by the bus 1050. Each interface does not necessarily have to be located inside the computer 1000, but may be located outside the computer 1000 via a network or the like. Furthermore, each part of the computer 1000 may be controlled by a circuit different from the processing circuit 1100. For example, the display unit 1700 may be controlled not by the processing circuit 1100, but by a circuit dedicated to display processing provided within the display unit 1700.

[0092] For example, when computer 1000 functions as an information processing system 1 or 2 according to the embodiments of this disclosure, the processing circuit 1100 of computer 1000 functions as various detection and control units (e.g., rendering unit 10, style conversion unit 20, 3D reconstruction unit 30, viewpoint generation unit 60) included in the information processing system 1 or 2 by executing a program loaded onto RAM 1200. The secondary storage device 1400 stores the information processing program and various data according to this disclosure. The processing circuit 1100 reads and executes program data 1450 from the secondary storage device 1400, but as another example, these programs may be obtained from other devices via an external network 1550. That is, the secondary storage device 1400 is not limited to being inside computer 1000, but may be located outside computer 1000. The processing circuit 1100 is an example of an integrated circuit, and CPU, MPU, GPU, APU, ASIC, and FPGA can all be considered integrated circuits.

[0093] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.

[0094] [Note] This technology can also be configured as follows: (1) An information processing system comprising: a style feature extraction unit that extracts style features from style information representing the appearance of an image; and a style transformation network that acquires input data in which different viewpoint images are assigned to each channel, and transforms the style of each viewpoint image into the style indicated by the style features based on the result of attention processing in the channel direction representing the direction in which the channels are joined. (2) The information processing system according to (1) above, wherein the style transformation network comprises: a Convolution Layer that extracts intermediate features from the input data; and a Spatial-View Attention Layer that performs integrated attention processing in the channel direction and spatial direction on the intermediate features in which some features have been replaced by the style features. (3) The information processing system according to (1) or (2) above, wherein the style conversion network comprises an Encoder for encoding the input data into a feature space, and a Decoder for decoding the result of the style conversion processing performed in the feature space into an image space. (4) The information processing system according to any one of (1) to (3) above, comprising a 3D reconstruction unit that acquires the viewpoint image after style conversion as a style conversion image and generates a 3D asset using a plurality of style conversion images corresponding to each channel. (5) The information processing system according to (4) above, comprising a viewpoint generation unit that generates at least one interpolated image visible from an intermediate position between viewpoints by applying interpolation processing to the plurality of style conversion images, wherein the 3D reconstruction unit generates the 3D asset having the converted style using the plurality of style conversion images and the at least one interpolated image. (6) A computer-based information processing method comprising: extracting style features from style information representing the appearance of an image; obtaining input data to which different viewpoint images are assigned for each channel; and converting the style of each viewpoint image to the style indicated by the style features based on the result of attention processing of the channel direction representing the direction in which the channels are joined.(7) The information processing method according to (6) above, comprising: extracting intermediate features from the input data; and performing integrated attention processing in the channel direction and spatial direction on the intermediate features from which some features have been replaced by the style features. (8) The information processing method according to (6) or (7) above, comprising: encoding the input data into a feature space; and decoding the result of the style transformation processing performed in the feature space into an image space. (9) The information processing method according to any one of (6) to (8) above, comprising: acquiring the viewpoint image after style transformation as a style transformation image; and generating a 3D asset using a plurality of style transformation images corresponding to each channel. (10) The information processing method according to (9) above, comprising: applying interpolation processing to the plurality of style transformation images to generate at least one interpolated image visible from an intermediate position between viewpoints; and generating the 3D asset having the transformed style using the plurality of style transformation images and the at least one interpolated image. (11) A computer-readable non-temporary storage medium that stores a program that causes a computer to perform the following: extract style features from style information representing the appearance of an image, acquire input data to which different viewpoint images are assigned for each channel, and convert the style of each viewpoint image to the style indicated by the style features based on the result of attention processing in the channel direction representing the direction in which the channels are joined.

[0095] 1,2 Information Processing System 21 Style Feature Extraction Unit 22,50 Style Conversion Network 23 Conversion Layer 24 Spatial-View Attention Layer 30 3D Reconstruction Unit 51 Encoder 53 Decoder CA Style Conversion Asset (3D Asset) CI Style Conversion Image IN Input Data SF Style Feature Quantity ST Style Information VI Viewpoint Image

Claims

A style feature extraction unit that extracts style features from style information that describes the appearance of an image, A style transformation network that acquires input data in which different viewpoint images are assigned to each channel, and transforms the style of each viewpoint image into the style indicated by the style feature based on the result of channel direction attention processing which represents the direction in which the channels are joined, An information processing system having   The aforementioned style conversion network is A Convolution Layer extracts intermediate features from the aforementioned input data, A Spatial-View Attention Layer performs integrated attention processing in the channel direction and spatial direction on the intermediate features, which have been partially replaced by the aforementioned style features. The information processing system according to claim 1, having the following features.   The aforementioned style conversion network is An Encoder that encodes the aforementioned input data into a feature space, A Decoder that decodes the result of the style transformation process performed in the feature space into the image space, The information processing system according to claim 1, having the following features.   The system includes a 3D reconstruction unit that acquires the viewpoint image after style conversion as a style conversion image and generates a 3D asset using multiple style conversion images corresponding to each channel. The information processing system according to claim 1.   The viewpoint generation unit generates at least one interpolated image visible from an intermediate position between viewpoints by applying interpolation processing to the aforementioned multiple style-transformed images. The 3D reconstruction unit generates the 3D asset having the converted style using the plurality of style conversion images and the at least one interpolation image. The information processing system according to claim 4.   Style features are extracted from style information that describes the appearance of an image. We acquire input data to which different viewpoint images are assigned to each channel. Based on the result of channel direction attention processing representing the direction in which the channels are joined, the style of each viewpoint image is converted to the style indicated by the style feature. An information processing method performed by a computer, which includes the ability to perform the following actions.   Intermediate features of the aforementioned input data are extracted, For the intermediate features, some of which have been replaced by the aforementioned style features, an integrated attention process is performed in the channel direction and spatial direction. The information processing method according to claim 6, wherein the method is as follows:   The input data is encoded into the feature space, The result of the style transformation process performed in the feature space is decoded into the image space. The information processing method according to claim 6, wherein the method is as follows:   The aforementioned viewpoint image after style conversion is obtained as the style-converted image. It generates 3D assets using multiple style-transformed images corresponding to each channel. The information processing method according to claim 6, wherein the method is as follows:   By applying interpolation processing to the aforementioned multiple style-transformed images, at least one interpolated image visible from an intermediate position between viewpoints is generated. Using the plurality of style conversion images and the at least one interpolation image, the 3D asset having the converted style is generated. The information processing method according to claim 9, wherein the method is as follows:   Style features are extracted from style information that describes the appearance of an image. We acquire input data to which different viewpoint images are assigned to each channel. Based on the result of channel direction attention processing representing the direction in which the channels are joined, the style of each viewpoint image is converted to the style indicated by the style feature. A computer-readable, non-temporary storage medium that stores a program that enables a computer to perform a certain action.