Automatically generated shader masks and parameters

The system addresses memory and processing challenges in computer simulations by combining grayscale images with colors using gradient descent to generate shader masks and parameters, enhancing graphics rendering efficiency.

JP2025533838AActive Publication Date: 2025-10-09SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025519646
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-05
Filing Date
2023-10-03
Publication Date
2025-10-09
Estimated Expiration
2043-10-03

AI Technical Summary

Technical Problem

Computer simulations, such as computer games, face challenges with increasing memory space and processing time requirements for shading operations as graphics become more sophisticated.

Method used

A system that combines grayscale images with colors using gradient descent to render a final color image, reducing memory usage and processing time by employing machine learning models and shaders to generate shader masks and parameters.

Benefits of technology

Reduces memory requirements and processing time for shading operations by generating shader masks and parameters efficiently, enabling improved graphics rendering in computer simulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533838000001
    Figure 2025533838000001
  • Figure 2025533838000002
    Figure 2025533838000002
  • Figure 2025533838000003
    Figure 2025533838000003
Patent Text Reader

Abstract

The graphics shader (200) takes two grayscale images (called "masks") (300, 302) and four colors to generate a full-color image. The logic behind the shader is two-fold: first, separating the colors from the image (408) allows for a wide variety of images (e.g., changing one color to get a brick wall of a different color); and second, the two grayscale masks take up less memory space than a full-color image. A script using differentiable programming and gradient descent "finds" the masks and colors for the target image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates generally to automatically generated shader masks and parameters. [Background technology]

[0002] As understood herein, computer simulations, such as computer games, use shaders, which are software programs, to paint game objects with colors and textures, and as understood herein, as game graphics become more sophisticated, memory space and processing time for shading operations become increasingly important. Summary of the Invention

[0003] Thus, the device includes at least one computer storage that is not a transitory signal, the at least one computer storage including instructions executable by the at least one processor to receive the first grayscale image and the second grayscale image, combine the grayscale image with a plurality of colors to render a test image, modify the test image using gradient descent, and output a final color image based at least in part on a loss indication associated with the gradient descent.

[0004] The first grayscale image and the second grayscale image may be based on a common image, i.e., may be two different grayscale versions of the same image, or may not be based on a common image.

[0005] In some examples, the instructions may be executable to combine the grayscale image with four colors to render a test image. The instructions may be embodied in a machine learning (ML) model and / or a shader. In a non-limiting embodiment, the instructions may be executable to output a final color image to a computer simulation for displaying the final color image during playback of the computer simulation.

[0006] In another aspect, an apparatus includes at least one processor programmed with instructions to identify at least a first grayscale image and a second grayscale image, combine the grayscale image and at least one color to render a test image, modify the test image using a gradient descent method, and output a final color image based at least in part on a loss indication associated with the gradient descent method.

[0007] In another aspect, a method includes receiving a first grayscale image, receiving a second grayscale image, and receiving at least one color. The method includes outputting a test image based at least in part on the grayscale image and the color. The method further includes applying a gradient descent method to modify the test image by minimizing a loss function until a final image is generated, and outputting the final image to a computer simulation.

[0008] The details of the present application, both as to its structure and operation, can best be understood in reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram of an exemplary system in accordance with the principles of the present invention; [Figure 2] 1 illustrates an exemplary hardware architecture. [Figure 3] 1 illustrates an exemplary software architecture. [Figure 4] 1 illustrates exemplary logic consistent with the principles of the present invention. [Figure 4A] Offer alternative representations. [Figure 5] Screenshots are presented that graphically illustrate the principles of the present invention. [Figure 6] Screenshots are presented that graphically illustrate the principles of the present invention. [Figure 7] Screenshots are presented that graphically illustrate the principles of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0010] The present disclosure generally relates to computer ecosystems, including aspects of consumer electronics (CE) device networks, such as, but not limited to, computer gaming networks. The systems herein may include server and client components, which may be connected via a network to enable data exchange between the client and server components. The client components may include one or more computing devices, including game consoles such as the Sony PlayStation®, or game consoles manufactured by Microsoft, Nintendo, or other manufacturers, extended reality (XR) headsets such as virtual reality (VR) headsets, augmented reality (AR) headsets, portable televisions (e.g., smart TVs, Internet-enabled TVs), portable computers such as laptops and tablet computers, and smartphones and other mobile devices, including additional examples described below. These client devices may operate in a variety of operating environments. For example, some of the client computers may employ the Linux® operating system, a Microsoft operating system, or a Unix® operating system, or an operating system manufactured by Apple or Google, or a Berkeley Software Distribution (BSD) OS, including derivatives of BSD, to name a few. These operating environments may be used to run one or more browsing programs, such as a Microsoft, Google, or Mozilla browser, or other browser program capable of accessing websites hosted by Internet servers as described below. Also, an operating environment according to the principles of the present invention may be used to run one or more computer game programs.

[0011] Servers and / or gateways may be used that may include one or more processors that execute instructions that configure the server to send and receive data over a network such as the Internet. Alternatively, clients and servers may be connected via a local intranet or virtual private network. The server or controller may be instantiated by a game console such as a Sony PlayStation®, a personal computer, or the like.

[0012] Information may be exchanged between the client and the server over a network. For this purpose and for security, the server and / or client may include firewalls, load balancers, temporary storage, and proxies, as well as other network infrastructure for reliability and security. One or more servers may form an apparatus that implements a method for providing a secure community, such as an online social website or gamer network, to network members.

[0013] The processor may be a single-chip processor or a multi-chip processor capable of executing logic through various lines such as address lines, data lines, and control lines, as well as registers and shift registers. Processors including digital signal processors (DSPs) may be circuit embodiments.

[0014] Components included in one embodiment may be used in other embodiments in any suitable combination, for example, any of the various components described herein and / or illustrated in the figures may be combined, interchanged, or excluded from other embodiments.

[0015] A "system having at least one of A, B, and C" (and similarly "a system having at least one of A, B, or C" and "a system having at least one of A, B, and C") includes a system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together.

[0016] 1 , an exemplary system 10 is shown, which may include one or more of the exemplary devices according to the principles of the present invention, as referenced above and further described below. A first of the exemplary devices included in system 10 is a consumer electronics (CE) device, such as an audio-video device (AVD) 12, including, but not limited to, a theater display system that may be projector-based, or an Internet-enabled TV with a TV tuner (equivalently, a set-top box that controls the TV). Alternatively, AVD 12 may also be a computer-controlled Internet-enabled (“smart”) phone, a tablet computer, a notebook computer, a head-mounted device (HMD) and / or a headset, such as smart glasses or a VR headset, other computer-controlled wearable devices, a computer-controlled Internet-enabled music player, a computer-controlled Internet-enabled headphones, a computer-controlled Internet-enabled implantable device, such as an implantable skin device, or the like. In any event, it should be understood that AVD12 is configured to implement the principles of the present invention (e.g., to communicate with other CE devices to implement the principles of the present invention, to execute the logic described herein, and to perform any other functions and / or operations described herein).

[0017] Thus, AVD 12 may be established with some or all of the components shown to implement such inventive principles. For example, AVD 12 may include one or more touch-enabled displays 14, which may be implemented with high-resolution or ultra-high-resolution flat screens of "4K" or higher. Touch-enabled display(s) 14 may include, for example, a capacitive or resistive touch-sensing layer with a touch-sensing electrode grid consistent with the inventive principles.

[0018] The AVD 12 may also include one or more speakers 16 for outputting audio in accordance with the principles of the present invention, and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to and controlling the AVD 12. The exemplary AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, a WAN, or a LAN, under the control of one or more processors 24. Thus, the interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as a mesh network transceiver. It should be understood that the processor 24 controls the AVD 12 to implement the principles of the present invention, including other elements of the AVD 12 described herein, such as controlling the display 14 to present images and receiving input from the display 14. Furthermore, it should be noted that the network interface 20 may be a wired or wireless modem or router, or other suitable interface, such as a wireless telephone transceiver or the Wi-Fi transceiver described above.

[0019] In addition to the above, AVD 12 may also include one or more input and / or output ports 26, such as a High-Definition Multimedia Interface (HDMI®) port or a Universal Serial Bus (USB) port for physically connecting to another CE device, and / or a headphone port for connecting headphones to AVD 12 to provide a user with audio from AVD 12 via headphones. For example, input port 26 may be connected, wired or wirelessly, to a cable or satellite source 26 a of audio-video content. Thus, source 26 a may be a separate or integrated set-top box or satellite receiver. Alternatively, source 26 a may be a game console or disc player containing content. If implemented as a game console, source 26 a may include some or all of the components described below in connection with CE device 48.

[0020] AVD 12 may further include one or more computer memory / computer-readable storage media 28, such as non-transitory disk-based or solid-state storage, which in some cases may be embodied as a standalone device within the AVD's chassis, or as a personal video recording device (PVR) or video disc player internal or external to the AVD's chassis for playing AV programs, or as a removable storage medium or server as described below. In some embodiments, AVD 12 may also include a position or location receiver, such as, but not limited to, a cellular telephone receiver 30, a GPS receiver 30, and / or an altimeter 30, configured to receive geographic location information from satellites or cellular towers and provide the information to processor 24 and / or in conjunction with processor 24 to determine the altitude at which AVD 12 is located.

[0021] Continuing with the description of the AVD 12, in some embodiments, the AVD 12 may include one or more cameras 32, which may be a thermal imaging camera, a digital camera such as a webcam, an IR sensor, an event-based sensor, and / or a camera integrated into the AVD 12 and may be controllable by the processor 24 to collect photographs / images and / or videos in accordance with the principles of the present invention. The AVD 12 may also include a Bluetooth transceiver 34 and other NFC elements 36 for communicating with other devices using Bluetooth and / or near field communication (NFC) technology, respectively. An exemplary NFC element may be a radio frequency identification (RFID) element.

[0022] Furthermore, AVD 12 may include one or more auxiliary sensors 38 that provide input to processor 24. For example, one or more of auxiliary sensors 38 may include one or more pressure sensors that form a layer of touch-enabled display 14 itself, and may be, but are not limited to, piezoelectric pressure sensors, capacitive pressure sensors, piezo-resistive wire strain gauges, optical pressure sensors, electromagnetic pressure sensors, etc. Examples of other sensors include pressure sensors, motion sensors such as accelerometers, gyroscopes, cyclometers, or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or cadence sensors, event-based sensors, gesture sensors (e.g., for sensing gesture commands). Thus, sensors 38 may be implemented by one or more motion sensors such as individual accelerometers, gyroscopes, and magnetometers, and / or an inertial measurement unit (IMU), which typically includes a combination of accelerometers, gyroscopes, and magnetometers to determine the position and orientation of AVD 12 in three dimensions, or by an event-based sensor such as an event detection sensor (EDS). An EDS consistent with the present disclosure provides an output indicative of a change in light intensity sensed by at least one pixel of the light-sensing array. For example, if the light sensed by the pixel is decreasing, the output of the EDS may be −1, and if it is increasing, the output of the EDS may be +1. An output binary signal of 0 may indicate no change in light intensity below a certain threshold.

[0023] The AVD 12 may also include an over-the-air (OTA) TV broadcast port 40 for receiving over-the-air (OTA) TV broadcasts, which provides input to the processor 24. In addition to the above, it should be noted that the AVD 12 may also include an infrared (IR) transmitter 42 and / or an IR receiver 42 and / or an IR transceiver 42, such as an Infrared Data Association (IRDA) device. A battery (not shown) may be provided for powering the AVD 12, which may be a kinetic energy harvester that converts kinetic energy into electrical power to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field-programmable gate array 46 may also be included. One or more haptic / vibration generators 47 may be provided for generating haptic signals that can be sensed by a person holding or interacting with the device. Thus, the haptic generator 47 may use an electric motor connected to an eccentric and / or unbalanced weight to vibrate all or part of the AVD 12 via the motor's rotatable shaft, which may rotate under the control of the motor (which may be controlled by a processor such as processor 24), thereby producing vibrations of various frequencies and / or amplitudes, as well as simulations of forces in various directions.

[0024] A light source such as a projector, such as an infrared (IR) projector, may also be included.

[0025] In addition to AVD 12, system 10 may include one or more other CE device types. In one embodiment, first CE device 48 may be a computer game console that can be used to transmit computer game audio and video to AVD 12 via commands sent directly to AVD 12 and / or through a server, as described below, while second CE device 50 may include similar components to first CE device 48. In the illustrated embodiment, second CE device 50 may be configured as a computer game controller operated by a player or as a head-mounted display (HMD) worn by a player. The HMD may include a head-up transparent or opaque display to present AR / MR content or VR content (more generally, extended reality (XR) content), respectively. The HMD may be configured as a glasses-type display or as a large VR-type display sold by a computer game console manufacturer.

[0026] In the example shown, only two CE devices are shown, but it will be understood that fewer or more devices may be used. The devices herein may implement some or all of the components shown with respect to AVD 12. Any of the components shown in the figures below may incorporate some or all of the components shown in the AVD 12 example.

[0027] Referring now to the aforementioned at least one server 52, the at least one server 52 includes at least one server processor 54, at least one tangible computer-readable storage medium 56, such as disk-based or solid-state storage, and at least one network interface 58, which, under the control of the server processor 54, enables communication with other exemplary devices over the network 22 and may, indeed, facilitate communication between the server and client devices in accordance with the principles of the present invention. It should be noted that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface, such as, for example, a wireless telephone transceiver.

[0028] Thus, in some embodiments, server 52 may be an entire internet server or server "farm" and may include and execute "cloud" functionality such that devices of system 10 may access the "cloud" environment through server 52, in an exemplary embodiment such as a network gaming application. Alternatively, server 52 may be implemented by one or more game consoles or other computers in the same room or nearby as the other devices shown.

[0029] The components shown in the figures below may include some or all of the components shown herein. Any user interfaces (UIs) described herein may be integrated and / or extended, and UI elements may be mixed and matched between UIs.

[0030] The principles of the present invention may employ a variety of machine learning models, including deep learning models. Machine learning models consistent with the principles of the present invention may employ a variety of algorithms trained using methods including supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature learning, self-learning, and other forms of learning. Examples of such algorithms that may be implemented by computer circuitry include one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and a type of RNN known as a long-short-term memory (LSTM) network. Support vector machines (SVMs) and Bayesian networks may also be considered examples of machine learning models. In addition to the types of networks described above, the models herein may be implemented by classifiers.

[0031] Thus, as understood herein, performing machine learning may include accessing a model, training the model with training data, and enabling the model to process additional data and make inferences. Thus, an artificial neural network / artificial intelligence model trained through machine learning may include an input layer, an output layer, and multiple hidden layers therebetween, which are configured and weighted to make inferences regarding a preferred output.

[0032] FIG. 2 illustrates a graphics shader 200 that may be executed by a GPU 202 to shade graphics from a source 204 of computer game graphics, such as a computer game console or server, for presentation on a display 206 .

[0033] 3, shader 200 takes in two grayscale images (called "masks") 300, 302 and four colors 304, 306, 308, 310 and generates a full-color image 312 from the input. The four colors 304, 306, 308, 310 may be, for example, red, green, blue, and yellow. The logic behind shader 200 is two-fold. First, separating the colors from the image allows for a wide variety of images (e.g., changing one color to get a brick wall of a different color), and second, the two grayscale masks take up less memory space than a full-color image.

[0034] A script using differentiable programming and gradient descent "finds" the mask and color of the target image. Figure 4 shows exemplary logic such a script may implement.

[0035] Beginning at block 400, two grayscale masks 300, 302 shown in FIG. 3 are generated along with four colors 304, 306, 308, 310 shown in FIG.

[0036] Proceeding to block 402, the logic generates a color image based on the grayscale masks 300, 302 and the four colors 304, 306, 308, 310. Proceeding next to block 404, the loss between the current image and a target image, which may be defined by the user, is determined. Based on the determined loss, if the determined loss is acceptably small or if a predetermined number of iterations has been reached, as indicated at decision diamond 406, the final grayscale mass and color are output at block 410. However, if the determined loss is not acceptably small or if the predetermined number of iterations has not been reached, the logic proceeds from decision diamond 406 to block 408.

[0037] At block 408, the grayscale mask and color are modified by applying gradient descent, and the logic loops back to block 402 to generate an updated image.

[0038] Gradient descent uses calculus to take a loss as input and identify ways to modify an image to lower the loss. This technique can be used in stochastic gradient descent, which can be used as an extension of the backpropagation algorithm used in ML model training, such as the ML models described herein. Stochastic gradient descent adds a probabilistic characteristic to the direction of the update. Weights may be used to calculate derivatives.

[0039] 4A shows that image parameters are generated in block 420. The parameters can be two grayscale masks and four color pixel values, represented by three values ​​for each RGB image. So, if the grayscale mask is 2000 x 20000 pixels, the total number of parameters is 2000 x 2000 x 2 + 4 x 3 parameters. Generating an image from parameters means performing the same calculations that a shader does. In pseudocode: a=lerp(color_0, color_1, mask_0) b=lerp(color_2, color_3, mask_0) generated_image = lerp(a, b, mask_1)

[0040] Proceeding to block 422, a loss is obtained. The loss is the difference between the generated image and the target image (the difference can be either absolute or mean squared error). The loss is then backpropagated in block 424 using gradient descent to modify the parameters of block 420. The user can determine the number of times the cycle is repeated, for example, 1000 to 6000 cycles.

[0041] This is further illustrated by Figures 5 and 6. In Figure 5, two grayscale masks 500, 502 are applied to the previous result and another mask to render a final color image 504, shown enlarged at 506. In Figure 6, two colors 600, 602 are combined with the final color 504 of Figure 5 to render a new color image 604.

[0042] FIG. 7 shows two grayscale masks 700, 702, which are essentially two different grayscale versions of the same image, that are combined with four colors 704 (red, black, yellow, and blue in the example shown) to render a final color image 712.

[0043] The above tools and techniques may be provided to end-user game computing devices, such as computer game consoles, so that end-user game players can use the tools described herein in-game (i.e., as part of playing a computer game) to create and / or modify game objects for each other.

[0044] Although particular embodiments are shown and described in detail herein, it should be understood that the subject matter encompassed by this invention is limited only by the scope of the claims.

Claims

1. 1. A device comprising at least one computer storage device that is not a transitory signal, The at least one computer storage includes instructions, the instructions being executed by the at least one processor to: receiving a first grayscale image and a second grayscale image; combining the grayscale image with a plurality of colors to render a test image; modifying the test image using gradient descent; outputting a final color image based at least in part on a loss indication associated with the gradient descent method; and A device capable of running

2. The device of claim 1 , wherein the first grayscale image and the second grayscale image are based on a common image.

3. The device of claim 1 , wherein the first grayscale image and the second grayscale image are not based on a common image.

4. The instruction: combining the grayscale image with four colors to render the test image; The device of claim 1 , capable of executing:

5. The device of claim 1 , wherein the instructions are embodied in a machine learning (ML) model.

6. The device of claim 1 , wherein the instructions are embodied in a shader.

7. The device of claim 1 , wherein the instructions are executable to output the final color image to a computer simulation for displaying the final color image during playback of the computer simulation.

8. The device of claim 1 comprising the at least one processor.

9. 1. An apparatus comprising at least one processor, The at least one processor is programmed with instructions, the instructions comprising: identifying at least a first grayscale image and a second grayscale image; combining the grayscale image with at least one color to render a test image; modifying the test image using gradient descent; outputting a final color image based at least in part on a loss indication associated with the gradient descent method; and A device that performs the above.

10. The apparatus of claim 9 , wherein the first grayscale image and the second grayscale image are based on a common image.

11. The apparatus of claim 9 , wherein the first grayscale image and the second grayscale image are not based on a common image.

12. The instruction: combining the grayscale image with four colors to render the test image; The apparatus of claim 9 , capable of performing the following:

13. The apparatus of claim 9 , wherein the instructions are embodied in a machine learning (ML) model.

14. The apparatus of claim 9 , wherein the instructions are embodied in a shader.

15. 10. The apparatus of claim 9, wherein the instructions are executable to output the final color image to a computer simulation for displaying the final color image during playback of the computer simulation.

16. receiving a first grayscale image; receiving a second grayscale image; receiving at least one color; outputting a test image based at least in part on the grayscale image and the color; applying gradient descent to modify the test images by minimizing a loss function until a final image is generated; outputting the final image to a computer simulation; A method comprising:

17. The method of claim 16 , wherein the first grayscale image and the second grayscale image are based on a common image.

18. The method of claim 16 , wherein the first grayscale image and the second grayscale image are not based on a common image.

19. The method of claim 16 , wherein the method is embodied in a machine learning (ML) model.

20. The method of claim 16 , wherein the method is performed by a shader.

Citation Information

Patent Citations

  • Apparatus and method for color image fusion

    US20020015536A1

  • Method and apparatus for automatically collecting terrain source data for display during flight simulation

    US20030190589A1

  • Edge-Aware Bilateral Image Processing

    US20170132769A1

  • Method and system for image enhancement

    US20180089812A1

  • Colorization of vector images

    US20190355154A1