Soft shadow image rendering method and apparatus, and device, storage medium and program product

By combining the generation of shadow masks and point cloud data, the problems of low efficiency and poor effect in soft shadow image rendering are solved, and a high-efficiency and realistic rendering effect is achieved.

WO2025241701A1PCT designated stage Publication Date: 2025-11-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2025/085629
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-20
Filing Date
2025-03-28
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing technologies consume high computational resources and are inefficient in soft shadow image rendering, resulting in harsh shadow effects and poor rendering quality.

Method used

By generating shadow masks and point cloud data, and combining the shadow masks to generate soft shadow masks, image rendering is performed, improving rendering efficiency and enhancing shadow effects.

Benefits of technology

It reduces performance overhead, improves rendering efficiency, and enhances the realism and smoothness of shadow effects by taking into account the three-dimensionality of objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025085629_27112025_PF_FP_ABST
    Figure CN2025085629_27112025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a soft shadow image rendering method and apparatus, and an electronic device, a computer-readable storage medium and a computer program product. The method comprises: acquiring an image to be rendered that comprises a shadow area; on the basis of the shadow area in said image, generating a shadow mask for said image; acquiring point cloud data of said image, and on the basis of the point cloud data and the shadow mask, generating a soft shadow mask for said image; and on the basis of the soft shadow mask, performing image rendering on said image to obtain a soft shadow image of said image.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, and apparatus for rendering soft shadow image, storage medium, and program product

[0001] Cross-reference to Related Applications

[0002] Embodiments of the present application are based on and claim priority from Chinese Patent Application No. 202410627683.7 filed on May 20, 2024, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0003] The present application relates to the technical field of computers, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for rendering a soft shadow image. BACKGROUND

[0004] In related technologies, when a soft shadow image is rendered, a corresponding shadow mask is generated, and then a shadow image is reconstructed using a graphics geometry sample and the shadow mask. However, such a method requires a large amount of computing resources, such as using sufficient samples to design three dimensions and colors, which not only has a high performance overhead, but also has a low rendering efficiency when the soft shadow image is rendered. In addition, the method of reconstructing a shadow image using only a graphics geometry sample and a shadow mask can result in a harsh shadow effect of the generated soft shadow image, which reduces the rendering effect when the soft shadow image is rendered. Therefore, related technologies have low rendering efficiency and poor rendering effect when a soft shadow image is rendered. SUMMARY

[0005] Embodiments of the present application provide a method, device, electronic device, computer-readable storage medium, and computer program product for rendering a soft shadow image, which can improve the efficiency and rendering effect when a soft shadow image is rendered.

[0006] The technical solution of the embodiments of the present application is as follows:

[0007] The embodiments of the present application provide a method for rendering a soft shadow image. The method is executed by an electronic device, and includes:

[0008] Obtaining a to-be-rendered image including a shadow region;

[0009] Generating a shadow mask of the to-be-rendered image based on the shadow region in the to-be-rendered image;

[0010] Obtaining point cloud data of the to-be-rendered image, and generating a soft shadow mask of the to-be-rendered image based on the point cloud data and the shadow mask;

[0011] perform image rendering on the to-be-rendered image based on the soft shadow mask to obtain a soft shadow image of the to-be-rendered image.

[0012] The embodiment of the present application provides a rendering device of a soft shadow image, and the device comprises:

[0013] The first obtaining module is configured to obtain a to-be-rendered image comprising a shadow area.

[0014] The generating module is configured to generate a shadow mask of the to-be-rendered image based on the shadow area in the to-be-rendered image.

[0015] The second obtaining module is configured to obtain point cloud data of the to-be-rendered image, and generate a soft shadow mask of the to-be-rendered image based on the point cloud data and the shadow mask.

[0016] The rendering module is configured to perform image rendering on the to-be-rendered image based on the soft shadow mask to obtain a soft shadow image of the to-be-rendered image.

[0017] The embodiment of the present application provides an electronic device, which comprises:

[0018] The memory is configured to store computer executable instructions or computer programs.

[0019] The processor is configured to execute the computer executable instructions or computer programs stored in the memory, so as to realize the rendering method of the soft shadow image provided in the embodiment of the present application.

[0020] The embodiment of the present application provides a computer readable storage medium, which comprises computer executable instructions or computer programs, and the processor executes the computer executable instructions or computer programs, so as to realize the rendering method of the soft shadow image provided in the embodiment of the present application.

[0021] The embodiment of the present application provides a computer program product, which comprises computer executable instructions or computer programs, and the processor executes the computer executable instructions or computer programs, so that the electronic device executes the rendering method of the soft shadow image provided in the embodiment of the present application.

[0022] The embodiment of the present application has the following beneficial effects:

[0023] First, a shadow mask of the to-be-rendered image including a shadow area is generated, then a soft shadow mask of the to-be-rendered image is generated based on point cloud data of the to-be-rendered image and the shadow mask, and finally, the to-be-rendered image is rendered based on the soft shadow mask to obtain a soft shadow image of the to-be-rendered image. In this way, the soft shadow image is obtained by rendering based on the generated soft shadow mask of the to-be-rendered image, which not only reduces the performance overhead, but also improves the rendering efficiency when the soft shadow image is rendered. At the same time, since the generation of the shadow is mostly related to the stereoscopic nature of the object, the point cloud data of the to-be-rendered image is also used in the soft shadow image rendering process, that is, the stereoscopic nature of the object in the to-be-rendered image in the three-dimensional environment is considered. Compared with the related art scheme of reconstructing the shadow image only by using the graphic geometry sample and the shadow mask, the problem of the generated soft shadow image being relatively harsh is solved, and the rendering effect when the soft shadow image is rendered is improved. BRIEF DESCRIPTION OF DRAWINGS

[0024] FIG. 1 is an architecture schematic diagram of a soft shadow image rendering system 100 provided by an embodiment of the present application;

[0025] FIG. 2 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;

[0026] FIG. 3 is a flow schematic diagram of a soft shadow image rendering method provided by an embodiment of the present application;

[0027] FIG. 4 is a flow schematic diagram of generating a shadow mask of a to-be-rendered image provided by an embodiment of the present application;

[0028] FIG. 5 is a schematic diagram of a target neighborhood range of a pixel point provided by an embodiment of the present application;

[0029] FIG. 6 is a schematic diagram of a to-be-rendered image and a depth image corresponding to the to-be-rendered image provided by an embodiment of the present application;

[0030] FIG. 7 is a model structure schematic diagram of a depth image generation model provided by an embodiment of the present application;

[0031] FIG. 8 is a flow schematic diagram of a soft shadow image rendering method provided by an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical scheme and advantages of the present application clearer, the embodiments of the present application will be described in detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by a person of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.

[0033] In the following description, "some embodiments" are described with reference to a subset of possible embodiments, but it is understood that "some embodiments" can be a same subset or different subset of all possible embodiments and can be combined with each other as long as there is no conflict.

[0034] In the following description, the terms "first\second\third" are merely used to distinguish similar objects, and do not represent a specific order of the objects. It is understood that the "first\second\third" can be interchanged in a specific order or sequence as long as it is allowed, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application is for the purpose of describing embodiments of this application only and is not intended to be limiting of this application.

[0036] Before further describing the embodiments of the present application, the terms and phrases involved in the embodiments of the present application are explained, and the terms and phrases involved in the embodiments of the present application are applicable to the following explanations.

[0037] 1) Client, also known as user end, refers to a program corresponding to a server to provide local services for users. Except for some application programs that can only run locally, it is generally installed on a common client and needs to be run in cooperation with a server, that is, there needs to be a corresponding server and service program in the network to provide corresponding services, so that a specific communication connection needs to be established between the client and the server to ensure the normal operation of the application program.

[0038] 2) Render Target (RT), a video memory buffer for rendering pixels, that is, a memory area. Such a memory area can exist simultaneously, that is, multiple render targets can exist simultaneously.

[0039] 3) Depth image, depth refers to the Z coordinate of a pixel in the view camera coordinate system, which is a unit of space. Depth maps can be obtained for any image. A depth map is a two-dimensional image that stores the depth values of all pixels in a single view. The depth map is a unit of space, such as millimeters. The significance of the depth map is to express the three-dimensional results of image matching in a more ordered and less storage space.

[0040] 4) Virtual light source, a graphical element that provides lighting effects in a virtual scene, obtained by simulating real light sources in a virtual scene.

[0041] 5) Virtual scene, a virtual scene displayed (or provided) by an application when running on a terminal. The virtual scene can be a simulated environment of the real world, a semi-simulated and semi-fictional virtual environment, or a purely fictional virtual environment. The virtual scene can be any of a two-dimensional virtual scene, a 2.5-dimensional virtual scene, or a three-dimensional virtual scene, and the embodiments of the present application do not limit the dimension of the virtual scene. For example, the virtual scene can include a sky, a land, an ocean, etc., the land can include desert, city, and other environmental elements, and a user can control a virtual object to move in the virtual scene.

[0042] 6) Virtual light, a light in a virtual scene emitted by a virtual light source used to illuminate the virtual scene. The virtual light includes direct light and indirect light. The direct light is emitted by the virtual light source, reflected by a virtual light point, and then reflected to a virtual camera. The indirect light is emitted by the virtual light source, reflected at least once, and then reflected to the virtual camera through the virtual light point.

[0043] 7) Virtual camera, a component simulating the function of a real camera. The virtual camera determines the viewing angle of the virtual scene by simulating the shooting function of the real camera, defines the position and direction of the virtual object controlled by the user in the virtual environment, and determines the scene content of the virtual scene observed by the virtual object controlled by the user.

[0044] 8) Soft shadows, soft and gradual shadow effects produced at the edges of objects due to factors such as the size of the light source, the scattering of light, and the blocking of objects. Unlike hard shadows, soft shadows have softer and smoother boundaries, and are closer to the shadow effects under real lighting conditions.

[0045] 9) Hard shadows, clear and well-defined shadow effects produced when there is no obstruction between the object and the light source under light illumination. The characteristic of hard shadows is that the edges of the shadow area are clear and sharp, without gradual or blurred effects, and look like the edges of an object completely blocking the light source.

[0046] 10) Soft shadow mask, usually a grayscale image or a transparency image, used to describe the shadow effect produced when light is incident on the surface of an object. The soft shadow mask can be used to enhance the realism and fidelity of the image, making the lighting effect look more natural and soft. The soft shadow mask usually contains transparency or grayscale information of the shadow on the surface of the object, reflecting the position, intensity of the light source, and the relative position relationship between the objects. By applying the soft shadow mask, the lighting effect can produce a gradually lightening or fading effect on the surface of the object, rather than a simple black and white shadow.

[0047] 11) Potential Shadow Area, refers to an area that may be covered by a shadow under the illumination of a light ray, that is, an area that may appear a shadow effect relative to other objects or surfaces under the position of a light source and the direction of light propagation. When light propagates from a light source to an object surface, if the light is blocked by other objects so that the light cannot directly reach a certain area, then this area has the potential to appear a shadow (Potential Shadow). The consideration of the Potential Shadow Area can help simulate the lighting effect in the real world, making the rendering of objects more realistic and lifelike. By analyzing and processing the Potential Shadow Area, the propagation path of the light can be accurately simulated in the rendering, and the appropriate shadow effect can be determined according to the geometry of the object and the position of the light source, which helps to enhance the realism and stereoscopic effect of the rendered image, making the lighting effect more natural and realistic.

[0048] The application scenarios of the rendering method of the soft shadow image provided in the embodiments of the present application are described below. For any scheme that needs to render a soft shadow image and display it, the rendering method of the soft shadow image provided in the embodiments of the present application can be used to render a soft shadow image.

[0049] In some embodiments, the soft shadow image can be a game image, that is, the rendering method of the soft shadow image provided in the embodiments of the present application can be applied to a game scenario, and the soft shadow image can be a game image including virtual objects, virtual props, virtual scenes in games, etc.

[0050] In some embodiments, the soft shadow image can be a film and television production image, that is, the rendering method of the soft shadow image provided in the embodiments of the present application can be applied to a film and television production scenario, and the soft shadow image can be an image in a science fiction or fantasy film and television work, or an image in a two-dimensional animation or three-dimensional animation, etc.

[0051] In some embodiments, the soft shadow image can be a design image, that is, the rendering method of the soft shadow image provided in the embodiments of the present application can be applied to an architectural and interior design scenario, and the soft shadow image can be an architectural design scheme display image or an interior design display image, etc.

[0052] In some embodiments, the soft shadow image can be a promotion image, that is, the rendering method of the soft shadow image provided in the embodiments of the present application can be applied to a promotion scenario, and the soft shadow image can be a product promotion image, a poster design image, etc. Specific settings can be made according to actual use requirements, which are not described here.

[0053] Referring to FIG. 1, which is an architecture diagram of a soft shadow image rendering system 100 provided by an embodiment of the present application, to implement a soft shadow image rendering scenario (for example, the embodiment of the present application can be applied in a game scenario, that is, the soft shadow image rendering scenario can be that, in the process of a user playing a game, a server obtains a game scenario image including a shadow region in the game; based on the shadow region in the game scenario image, a shadow mask of the game scenario image is generated; point cloud data of the game scenario image is obtained, and based on the point cloud data and the shadow mask, a soft shadow mask of the game scenario image is generated; based on the soft shadow mask, image rendering is performed on the game scenario image to obtain a soft shadow image of the game scenario image), a terminal (exemplarily shown as a terminal 400) is connected to a server 200 through a network 300, which can be a wide area network or a local area network, or a combination of the two, the terminal 400 is configured to be used by a user to use a client 401 to display on a display interface (exemplarily shown as a display interface 401-1), and the terminal 400 and the server 200 are connected to each other through a wired or wireless network.

[0054] In the embodiment, the server 200 is configured to obtain a to-be-rendered image including a shadow region; based on the shadow region in the to-be-rendered image, generate a shadow mask of the to-be-rendered image; obtain point cloud data of the to-be-rendered image, and based on the point cloud data and the shadow mask, generate a soft shadow mask of the to-be-rendered image; based on the soft shadow mask, perform image rendering on the to-be-rendered image to obtain a soft shadow image of the to-be-rendered image; and send the soft shadow image of the to-be-rendered image to the terminal 400.

[0055] The terminal 400 is further configured to receive the soft shadow image of the to-be-rendered image, and based on the display interface, display the soft shadow image.

[0056] In some embodiments, the server 200 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a set-top box, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, and a mobile device (for example, a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable game device, a smart speaker, and a smart watch), but is not limited thereto. The terminal device and the server can be connected directly or indirectly through wired or wireless communication, which is not limited in the embodiment of the present application.

[0057] Referring to FIG. 2, which is a structural schematic diagram of an electronic device provided by the embodiments of the present application, in actual application, the electronic device can be the server 200 or the terminal 400 shown in FIG. 1. Referring to FIG. 2, the electronic device shown in FIG. 2 includes at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together by a bus system 440. It can be understood that the bus system 440 is configured to realize the connection and communication between the components. The bus system 440 includes not only a data bus, but also a power bus, a control bus and a status signal bus. However, for the purpose of clear illustration, all the buses are marked as the bus system 440 in FIG. 2.

[0058] The processor 410 can be an integrated circuit chip with signal processing capability, such as a general purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc., wherein the general purpose processor can be a microprocessor or any conventional processor.

[0059] The user interface 430 includes one or more output devices 431 enabling presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons and controls.

[0060] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, and the like. The memory 450 optionally includes one or more storage devices remotely located from the processor 410 in a physical location.

[0061] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0062] In some embodiments, the memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, which are exemplarily illustrated below.

[0063] The operating system 451 includes system programs configured to handle various basic system services and perform hardware-related tasks, such as a framework layer, a core library layer, a driver layer, and the like, configured to implement various basic services and handle hardware-based tasks;

[0064] The network communication module 452 is configured to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), and the like;

[0065] The presentation module 453 is configured to enable presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, and the like) associated with the user interface 430 (e.g., a user interface configured to operate peripheral devices and display content and information);

[0066] The input processing module 454 is configured to detect and interpret one or more user inputs or interactions from the input devices 432.

[0067] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software, and FIG. 2 shows a soft shadow image rendering apparatus 455 stored in the memory 450, which can be in the form of programs and plug-ins and the like, including the following software modules: a first acquisition module 4551, a generation module 4552, a second acquisition module 4553, and a rendering module 4554, which are logical, and thus can be combined or further split according to the functions implemented. The functions of the various modules will be described below.

[0068] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software, and FIG. 2 shows a soft shadow image rendering apparatus 455 stored in the memory 450, which can be in the form of programs and plug-ins and the like, including the following software modules: a first acquisition module 4551, a generation module 4552, a second acquisition module 4553, and a rendering module 4554, which are logical, and thus can be combined or further split according to the functions implemented. The functions of the various modules will be described below.

[0069] In some embodiments, the terminal or the server can implement the rendering method of the soft shadow image provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or a software module in an operating system; can be a native application program (APP), i.e., a program that needs to be installed in an operating system to run, such as an instant messaging APP, a web browser APP; can also be a mini-program, i.e., a program that only needs to be downloaded into a browser environment to run; and can also be a mini-program that can be embedded into any APP. In summary, the above computer program can be any form of application program, module or plug-in.

[0070] Based on the above description of the rendering system and the electronic device of the soft shadow image provided in the embodiments of the present application, the rendering method of the soft shadow image provided in the embodiments of the present application is described below. In actual implementation, the rendering method of the soft shadow image provided in the embodiments of the present application can be implemented by a terminal or a server alone, or by a terminal and a server cooperatively. For example, the rendering method of the soft shadow image provided in the embodiments of the present application is executed by the server 200 in FIG. 1 alone. Referring to FIG. 3, FIG. 3 is a flowchart of the rendering method of the soft shadow image provided in the embodiments of the present application. The steps shown in FIG. 3 will be described below in combination with FIG. 3.

[0071] In step 101, the server acquires a to-be-rendered image including a shadow region.

[0072] In actual implementation, the to-be-rendered image can be pre-stored locally, can be acquired from the outside world (such as the Internet), or can be acquired in real time, for example, acquired in real time by a shooting device. In this regard, the embodiments of the present application do not make any limitation.

[0073] It should be noted that the to-be-rendered image can be an image corresponding to a virtual scene in a game, can be an image corresponding to an animation scene or an animation scene in video playing, can be a movie image in a science fiction or fantasy theme movie in a movie production scene, or can be an image in a two-dimensional animation or a three-dimensional animation; or can be a building design scheme display image in a building and interior design scene, an interior design display image, or a product promotion image, a poster design image in a promotion scene, and the like. In this regard, the embodiments of the present application do not make any limitation.

[0074] In step 102, a shadow mask of the to-be-rendered image is generated based on the shadow region in the to-be-rendered image.

[0075] It should be noted that after the to-be-rendered image is acquired, the shadow area in the to-be-rendered image is detected, and a shadow mask of the to-be-rendered image is generated based on the detected shadow area in the to-be-rendered image. The shadow mask is a binary image, which is used to distinguish the shadow area and the non-shadow area in the to-be-rendered image. For example, the pixel value of the shadow area in the to-be-rendered image is marked as 1, and the pixel value of the non-shadow area is marked as 0, or the pixel value of the shadow area is marked as 1, and the pixel value of the non-shadow area is marked as 0. The embodiments of the present application do not limit this.

[0076] In actual implementation, the process of detecting the shadow area in the to-be-rendered image to obtain the shadow area in the to-be-rendered image can be that, first, the to-be-rendered image is subjected to grayscale processing to obtain a grayscale image of the to-be-rendered image, and then the grayscale image is subjected to denoising processing to obtain a target grayscale image; then, a preset grayscale threshold is acquired, and for each pixel point in the target grayscale image, the pixel point is compared with the grayscale threshold, and based on the comparison result, a pixel point less than the grayscale threshold is selected from the plurality of pixel points as a target pixel point, and then a region formed by the target pixel point is taken as the shadow area; or, the shadow area in the to-be-rendered image is detected by color space conversion to obtain the shadow area in the to-be-rendered image, which can be that the to-be-rendered image is converted from a current color space (such as an RGB color space) to a target color space (such as an HSV, Lab, or other color space), and then the to-be-rendered image in the target color space is analyzed to obtain color information of the to-be-rendered image, so as to detect the shadow area by using the color information. For example, when detecting the shadow area of a natural scene, the image is converted from the RGB color space to the HSV color space, and the shadow area is detected by using the color information; or, the shadow area can also be recognized by analyzing the texture features of the to-be-rendered image. For example, when detecting the shadow area of an indoor scene, the shadow area is recognized by analyzing the texture features of the image. The way of detecting the shadow area in the to-be-rendered image to obtain the shadow area in the to-be-rendered image is various, and the embodiments of the present application do not limit this.

[0077] In actual implementation, referring to FIG. 4, FIG. 4 is a flow diagram of generating a shadow mask of a to-be-rendered image according to an embodiment of the present application. Based on FIG. 4, the process of generating the shadow mask of the to-be-rendered image based on the shadow area in the to-be-rendered image, that is, step 102, can be implemented by the following steps.

[0078] In step 1021, the brightness value of each pixel point in the to-be-rendered image is acquired.

[0079] In actual implementation, the image to be rendered is converted into image data in RGB format, color values of each pixel in the image to be rendered on R, G and B color channels, i.e., red, green and blue color channels, are obtained, and then based on the image data in RGB format, a brightness value of each pixel in the image to be rendered is determined, i.e., for each pixel, the following processing is performed: color values of the pixel on the three color channels are summed to obtain a sum result, a ratio of the sum result to a third constant is obtained, and the ratio is taken as the brightness value of the pixel, i.e.:

[0080] V = (R + G + B) / 3 …… Equation (1);

[0081] wherein V indicates the brightness value of the pixel, R indicates the color value of the pixel on the red color channel, G indicates the color value of the pixel on the green color channel, B indicates the color value of the pixel on the blue color channel, and 3 is a third constant preset.

[0082] In step 1022, for each pixel in the image to be rendered, the following processing is performed to obtain a mask value of the pixel: the pixel is dilated to obtain a local maximum brightness value of the pixel, and the pixel is eroded to obtain a local minimum brightness value of the pixel; and based on the local maximum brightness value and the local minimum brightness value, the mask value of the pixel is determined.

[0083] In actual implementation, before the pixel is dilated to obtain the local maximum brightness value of the pixel, a target neighborhood range of the pixel is obtained, and the target neighborhood range includes at least one adjacent pixel, where the adjacent pixel is used to indicate a pixel adjacent to the pixel.

[0084] It should be noted that the target neighborhood range herein is preset, for example, it can be a region formed by 8 adjacent pixels around the pixel, or a region formed by 15 adjacent pixels around the pixel, and the present embodiment does not limit this.

[0085] For example, referring to FIG. 5, which is a schematic diagram of a target neighborhood range of a pixel provided by the present embodiment, based on FIG. 5, for 9 pixels indicated by the dashed box 501, when the target neighborhood range indicates a region formed by 8 adjacent pixels around the pixel, 8 white points around the black point indicated by 5011 are adjacent pixels of the black point, i.e., the region formed by the white points is the target neighborhood range of the black point; and for 16 pixels indicated by the dashed box 502, when the target neighborhood range indicates a region formed by 15 adjacent pixels around the pixel, 15 white points around the black point indicated by 5021 are adjacent pixels of the black point, i.e., the region formed by the white points is the target neighborhood range of the black point.

[0086] Then, after obtaining the target neighborhood range of the pixel point, the process of performing the dilation processing on the pixel point to obtain the local maximum brightness value of the pixel point can be: obtaining the brightness values of each adjacent pixel point; selecting a maximum brightness value from the at least one brightness value, and taking the maximum brightness value as the local maximum brightness value of the pixel point, that is:

[0087] M(i, j) = max V(i+m, j+n) …… Equation (2);

[0088] wherein V is the brightness value of the pixel point, i and j are the row coordinate and column coordinate of the pixel point, K is the size of the kernel, that is, the number of adjacent pixel points, and m and n are the offset of the kernel, that is, the sliding range of the kernel on the image.

[0089] It should be noted that the dilation processing is a mathematical morphological dilation transformation. Through the dilation processing, the brightness value of each pixel will be replaced by the maximum brightness value in its surrounding neighborhood, so that the brightness values in the image become brighter or more prominent. Specifically, the dilation processing usually involves a structure element (also known as a kernel or a template), which defines the shape and size of the dilation operation. The structure element slides on the image (usually from left to right and from top to bottom), and at each position, the neighborhood range included by the structure element is determined. Then, the brightness value of each pixel point in the neighborhood range is obtained, and the maximum brightness value is selected from the brightness values of multiple pixel points as the brightness value of the neighborhood range. Then, the pixel at the center position in the neighborhood range is determined, and the selected brightness value is assigned to the pixel at the center position of the structure element, that is, the pixel at the center position in the neighborhood range. In this way, after the dilation processing, the value of each pixel on the brightness channel of the original image is the maximum brightness value in its neighborhood. In this way, the brightness changes in the image can be emphasized or the details of some local regions can be highlighted.

[0090] The process of performing the erosion processing on the pixel point to obtain the local minimum brightness value of the pixel point can be: obtaining the brightness values of each adjacent pixel point; selecting a minimum brightness value from the brightness values of each adjacent pixel point, and taking the minimum brightness value as the local minimum brightness value of the pixel point, that is:

[0091] m(i, j) = min V(i+m, j+n) …… Equation (5);

[0092] wherein V is the brightness value of the pixel point, i and j are the row coordinate and column coordinate of the pixel point, K is the size of the kernel, that is, the number of adjacent pixel points, and m and n are the offset of the kernel, that is, the sliding range of the kernel on the image.

[0093] It should be noted that the erosion processing is also a mathematical morphological erosion transform, which is used to extract the local minimum brightness value region in the image. Through the erosion processing, the brightness value of each pixel will be replaced by the minimum brightness value in its surrounding neighborhood. Specifically, the erosion processing also generally involves a structure element as described above, which defines the shape and size of the erosion operation. The structure element slides on the image (usually from left to right and from top to bottom), and at each position, the neighborhood range included by the structure element is determined. Then, the brightness value of each pixel in the neighborhood range is obtained, and the minimum brightness value is selected from the brightness values of multiple pixels as the brightness value of the neighborhood range. Then, the pixel at the center position in the neighborhood range is determined, and the selected brightness value is assigned to the pixel at the center position of the structure element, i.e., the pixel at the center position in the neighborhood range. In this way, after the erosion processing, the value of each pixel in the brightness channel of the original image is the minimum brightness value in its neighborhood. In this way, it can also be used to emphasize the brightness change in the image or highlight some local area details.

[0094] By applying the above embodiment, for each pixel point, the maximum brightness value of the pixel points in the target neighborhood range of the pixel point is taken as the local maximum brightness value of the pixel point after the dilation processing, and the minimum brightness value of the pixel points in the target neighborhood range of the pixel point is taken as the local minimum brightness value of the pixel point after the erosion processing. In this way, the accuracy of the selected local maximum brightness value and the local minimum brightness value is ensured, so that the rendering effect of the soft shadow image can be improved in the subsequent process.

[0095] In actual implementation, after the local maximum brightness value and the local minimum brightness value are determined, the process of determining the mask value of the pixel point based on the local maximum brightness value and the local minimum brightness value can be that the ratio of the local minimum brightness value and the local maximum brightness value of the pixel point is obtained; when the ratio is less than a target ratio, a first constant is determined as the mask value of the pixel point; when the ratio is greater than or equal to the target ratio, a second constant is determined as the mask value of the pixel point; wherein the first constant and the second constant are different.

[0096] It should be noted that the target ratio can be a pre-set value such as 0.5, and the first constant and the second constant can also be pre-set, for example, the first constant can be 0 and the second constant can be 1. The embodiments of the present application do not limit this.

[0097] Exemplarily, for the process of determining the mask value of the pixel point based on the local maximum brightness value and the local minimum brightness value, that is:

[0098] Wherein, m(i, j) is the local minimum brightness value of the pixel point, M(i, j) is the local maximum brightness value of the pixel point, T is a target ratio value set in advance, i and j are the row coordinate and column coordinate of the pixel point.

[0099] In this way, the mask value of the pixel point is determined based on the ratio of the local minimum brightness value and the local maximum brightness value of the pixel point; in this way, the mask value of the pixel point is accurately determined, thereby significantly improving the analysis accuracy of the image, and the shadow area in the image can be effectively detected in the subsequent process.

[0100] In step 1023, the shadow mask of the image to be rendered is obtained based on the mask value of each pixel point.

[0101] In actual implementation, as described above, the shadow mask is a binary image, and therefore, after the mask value of each pixel point, i.e. 0 or 1, is determined, the binary image corresponding to the image to be rendered, i.e. the shadow mask, can be obtained.

[0102] By applying the above embodiment, the shadow mask is determined based on the local maximum brightness value of the pixel point obtained by the dilation processing and the local minimum brightness value of the pixel point obtained by the erosion processing, which is helpful to distinguish the brightness change area in the image to be rendered, thereby facilitating the distinction between the shadow area and the non-shadow area in the image to be rendered, and further improving the rendering effect of the soft shadow image in the subsequent process.

[0103] In step 103, the point cloud data of the image to be rendered is obtained, and a soft shadow mask of the image to be rendered is generated based on the point cloud data and the shadow mask.

[0104] In actual implementation, the process of obtaining the point cloud data of the image to be rendered can be as follows: obtaining a depth image of the image to be rendered and the coordinates of each pixel point in the image to be rendered; converting the coordinates of each pixel point in the image to be rendered based on the depth image to obtain the three-dimensional coordinates of each pixel point; and obtaining the point cloud data of the image to be rendered based on the three-dimensional coordinates of each pixel point in the image to be rendered.

[0105] It should be noted that the depth image (Depth Image), also known as depth map (Depth Map), is a special image, and each pixel value of the image represents the distance or depth information of the corresponding point in the scene to the virtual camera, wherein the size of the pixel value represents the distance of the corresponding point to the virtual camera; and the point cloud data is used to indicate the surface shape and structure of the three-dimensional object in the three-dimensional scene corresponding to the image to be rendered, and is a three-dimensional data set composed of a large number of discrete points, each point usually carries its coordinates (X, Y, Z) in three-dimensional space and other attribute information such as color (R, G, B), normal vector, intensity, etc.

[0106] By using the above embodiment, the depth image of the to-be-rendered image is used to perform coordinate conversion on the to-be-rendered image to obtain the three-dimensional coordinates of each pixel point, so that the point cloud data of the to-be-rendered image is obtained based on the three-dimensional coordinates of each pixel point. In this way, the geometric structure of the three-dimensional scene, that is, the three-dimensional model, can be reconstructed through the depth image and the coordinate conversion, so as to facilitate obtaining the three-dimensional coordinates of each pixel point, and then the point cloud data of the to-be-rendered image can be accurately obtained.

[0107] In some embodiments, for the process of obtaining the depth image corresponding to the to-be-rendered image, a spatial coordinate system corresponding to the to-be-rendered image is constructed, and in the spatial coordinate system, the height value, that is, the depth value, of each pixel point in the to-be-rendered image is determined; based on the height value, the depth image of the to-be-rendered image is generated. For example, referring to FIG. 6, which is a schematic diagram of a to-be-rendered image and a depth image corresponding to the to-be-rendered image according to an embodiment of the present application, based on FIG. 6, 601 shows a to-be-rendered image, and 602 shows a depth image corresponding to the to-be-rendered image. First, the to-be-rendered image as shown in 601 is obtained, and a spatial coordinate system corresponding to the to-be-rendered image is constructed, and based on the to-be-rendered image, the height value of each pixel point in the to-be-rendered image in the spatial coordinate system is determined. The height values of the pixel points in the gray part of the image content indicated by the dashed box 6021 are the same, and different from the height values of the pixel points in the black part indicated by 602, so that based on the height values of the pixel points, the depth image corresponding to the to-be-rendered image as shown in 602 is generated.

[0108] In other embodiments, the process of obtaining the depth image corresponding to the to-be-rendered image can also be implemented by a model. A pre-trained depth image generation model is obtained, the to-be-rendered image is input into the depth image generation model, the depth value of each pixel point in the to-be-rendered image is obtained, and then based on the depth value of each pixel point, the depth image of the to-be-rendered image is obtained.

[0109] It should be noted that, referring to FIG. 7, which is a model structure diagram of a depth image generation model according to an embodiment of the present application, based on FIG. 7, the depth image generation model includes a feature extraction layer and a depth value prediction layer. After the to-be-rendered image is input into the depth image generation model, the image features of the to-be-rendered image are extracted through the feature extraction layer, then the depth value prediction is performed based on the image features of the to-be-rendered image through the depth value prediction layer, the depth value of each pixel point in the to-be-rendered image is obtained, and then based on the depth value of each pixel point, the depth image of the to-be-rendered image is obtained.

[0110] In actual implementation, the depth image generation model can use different architectures and loss functions, the size of the depth image is the same as that of the image to be rendered, i.e., HxW, where H is the height of the image and W is the width of the image; each pixel value D(i,j) in the depth image represents the distance from the i-th row and j-th column pixel in the image to the virtual camera corresponding to the image to be rendered. Wherein, the formula representation of the depth image generation model can be:

[0111] D = f(I, θ) …… Equation (9);

[0112] Wherein, I is the input image to be rendered, D is the output depth image, f is the function corresponding to the depth image generation model, and θ is the model parameter of the depth image generation model.

[0113] In actual implementation, for the process of converting the coordinates of each pixel point in the image to be rendered based on the depth image to obtain the three-dimensional coordinates of each pixel point, obtaining the intrinsic matrix of the virtual camera corresponding to the image to be rendered, wherein the intrinsic matrix is generally a 3x3 matrix, which is used to indicate the focal length, principal point and distortion parameters of the virtual camera, and can generally be obtained by the calibration process of the virtual camera; obtaining the depth value of each pixel point in the depth image; combining the intrinsic matrix and the depth value of each pixel point, the coordinates of each pixel point in the image to be rendered are converted to obtain the three-dimensional coordinates of each pixel point.

[0114] It should be noted that the calibration process of the virtual camera refers to the process of determining the internal parameters (such as focal length, principal point, distortion coefficient, etc.) and external parameters (such as position and direction) of the virtual camera to ensure that the virtual camera can accurately simulate the imaging characteristics of the real camera, that is, to ensure that the image generated by the virtual camera is consistent with the image captured by the real camera in terms of geometry and optical characteristics.

[0115] In this way, the three-dimensional coordinates of the pixel points are determined in combination with the intrinsic matrix of the virtual camera and the depth values of the pixel points; in this way, the accuracy of the determined three-dimensional coordinates of the pixel points is improved, thereby helping to improve the rendering effect of the soft shadow image in the subsequent process.

[0116] In actual implementation, in combination with the intrinsic matrix and the depth value of each pixel point in the depth image, the process of converting the coordinate of each pixel point in the image to be rendered to obtain the three-dimensional coordinate of each pixel point can be that, the coordinate of each pixel point in the image to be rendered is multiplied by the corresponding depth value to obtain the three-dimensional coordinate of each pixel point in the coordinate system corresponding to the virtual camera, then the three-dimensional coordinate of each pixel point in the coordinate system corresponding to the virtual camera is multiplied by the inverse matrix of the intrinsic matrix of the virtual camera to obtain the three-dimensional coordinate of each pixel point in the world coordinate system; the three-dimensional coordinate of each pixel point in the world coordinate system is taken as the converted three-dimensional coordinate, that is, the three-dimensional coordinate of each pixel point, and the specific formula is as follows:

[0117] wherein, K -1 is the inverse matrix of the camera intrinsic matrix, D(i, j) is the depth value of the pixel point, i and j are the row coordinate and the column coordinate of the pixel point, is the coordinate of each pixel point in the image to be rendered, that is, the two-dimensional coordinate.

[0118] In some embodiments, after obtaining the image to be rendered including the shadow area, the coordinate of the pixel point in the image to be rendered can also be normalized, and here, the coordinate of each pixel point is normalized to obtain the normalized coordinate of each pixel point in the image to be rendered; based on the depth image and the normalized coordinate, the coordinate of each pixel point in the image to be rendered is converted to obtain the three-dimensional coordinate of each pixel point.

[0119] It should be noted that the process of normalizing the coordinate of each pixel point to obtain the normalized coordinate of each pixel point in the image to be rendered is also:

[0120] wherein, i is the horizontal coordinate of the pixel point, j is the vertical coordinate of the pixel point, H is the height of the image, and W is the width of the image. In this way, the pixel coordinate (i, j) in the image to be rendered is converted into the normalized homogeneous coordinate (x, y, 1), wherein the range of x and y is [-1, 1], so that the influence of the size and resolution of the image on the coordinate transformation can be eliminated, and the coordinate transformation only depends on the intrinsic matrix of the camera.

[0121] In actual implementation, for each pixel in the to-be-rendered image based on the three-dimensional coordinates, the process of obtaining the point cloud data corresponding to the to-be-rendered image, based on the three-dimensional coordinates, obtains a target matrix, and takes the target matrix as the point cloud data corresponding to the to-be-rendered image, for example, stores the three-dimensional coordinates in an N*M matrix, where N=H*W is the total number of pixels in the to-be-rendered image, and M is the dimension number of the three-dimensional coordinates. In this way, the point cloud data of the three-dimensional scene corresponding to the to-be-rendered image is obtained, which can be used for subsequent shadow generation and rendering operations and the like.

[0122] In actual implementation, after obtaining the point cloud data and the shadow mask, a soft shadow mask can be generated based on the point cloud data and the shadow mask. The process of generating the soft shadow mask of the to-be-rendered image based on the point cloud data and the shadow mask can be that the shadow intensity of each pixel in the to-be-rendered image is predicted based on the point cloud data and the shadow mask to obtain the shadow intensity value of each pixel; and the soft shadow mask of the to-be-rendered image is obtained based on the shadow intensity value of each pixel.

[0123] It should be noted that the shadow intensity of each pixel is the transparency of the pixel. As described above, the soft shadow mask is usually a gray image or a transparency image. Therefore, after determining the shadow intensity value of each pixel, that is, the transparency, the transparency image corresponding to the to-be-rendered image, that is, the soft shadow mask, can be obtained.

[0124] It should be noted that the process of predicting the shadow intensity of each pixel in the to-be-rendered image based on the point cloud data and the shadow mask to obtain the shadow intensity value of each pixel can be implemented by a shadow intensity prediction model. The shadow intensity prediction model includes a point cloud encoder and a shadow predictor. First, a point cloud encoder is used to extract features from the point cloud data to obtain point cloud features, that is:

[0125] F p = p(P, θ p ) …… formula (13);

[0126] Where P is the input point cloud data, p is the point cloud encoder, and θ p is the parameter of the point cloud encoder.

[0127] Then, a shadow predictor is used to predict the shadow intensity value of each pixel based on the point cloud features and the shadow mask, that is:

[0128] Where q is the shadow predictor, θ q is the parameter of the shadow predictor, S is the shadow mask, and F p is the point cloud features extracted from the point cloud data.

[0129] By applying the above embodiments, the shadow intensity value of each pixel point is predicted based on the point cloud data and the shadow mask, so as to obtain the soft shadow mask of the to-be-rendered image based on the shadow intensity value of each pixel point; in this way, the shadow intensity value of each pixel point can be predicted with high precision by using the point cloud data and the shadow mask, and a high-quality soft shadow mask is generated, so that a more realistic soft shadow effect can be generated in the subsequent process.

[0130] In actual implementation, in addition to generating the soft shadow mask directly based on the point cloud data and the shadow mask, the soft shadow mask can also be generated based on other parameters. Next, taking specific different other parameters as examples, the process of generating the soft shadow mask is described.

[0131] In some embodiments, the soft shadow mask can also be generated in combination with the irradiation direction of the virtual light emitted by the virtual light source. After obtaining the to-be-rendered image including the shadow area, the to-be-rendered image can also be parsed to obtain the irradiation direction of the virtual light source corresponding to the to-be-rendered image; thus, the process of generating the soft shadow mask of the to-be-rendered image based on the point cloud data and the shadow mask can be generating the soft shadow mask of the to-be-rendered image based on the point cloud data, the shadow mask, and the irradiation direction of the virtual light source.

[0132] It should be noted that the formation of the shadow is related to the light, and based on this, the to-be-rendered image is parsed to obtain the virtual light source corresponding to the to-be-rendered image, wherein the virtual light source is a graphical element for providing lighting effects in a virtual scene and is obtained by simulating a real light source in a virtual scene; then the to-be-rendered image is analyzed based on the virtual light source to obtain the irradiation direction of the virtual light source corresponding to the to-be-rendered image, so that the soft shadow mask is generated in combination with the irradiation direction of the virtual light emitted by the virtual light source corresponding to the to-be-rendered image, which enriches the data used when generating the soft shadow mask, can not only improve the accuracy of the generated soft shadow mask, but also improve the rendering effect of the subsequent soft shadow image, that is, a realistic soft shadow effect can be generated, thereby enhancing the light and shadow effect and atmosphere of the soft shadow image.

[0133] It should be noted that the process of generating the soft shadow mask of the to-be-rendered image based on the point cloud data, the shadow mask, and the irradiation direction of the virtual light source can also be implemented by using a shadow intensity prediction model. As shown in formula (13), first, a point cloud encoder is used to extract features of the point cloud data to obtain point cloud features, and then a shadow predictor is used to predict the shadow degree value of each pixel point based on the point cloud features, the shadow mask, and the irradiation direction of the virtual light source, that is:

[0134] wherein q is the shadow predictor, and θ qis a parameter of the shadow predictor, S is a shadow mask, F p is a point cloud feature obtained by feature extraction on the point cloud data, and L is an illumination direction of the virtual light source.

[0135] It should be noted that after obtaining the shadow intensity prediction value of each pixel point, i.e., the shadow intensity value, the soft shadow mask of the to-be-rendered image can be obtained based on the shadow intensity value of each pixel point. Meanwhile, the illumination direction of the virtual light source in the above process is obtained analytically, and in addition to this, the illumination direction of the virtual light source can also be preset, for which the embodiments of the present application do not make any limitation.

[0136] In some embodiments, the soft shadow mask can also be generated in combination with the direct light component and the scattered light component corresponding to the to-be-rendered image. After obtaining the to-be-rendered image including the shadow area, the to-be-rendered image can also be analyzed to obtain the direct light component and the scattered light component corresponding to the to-be-rendered image. Thus, the process of generating the soft shadow mask of the to-be-rendered image based on the point cloud data and the shadow mask can be generating the soft shadow mask of the to-be-rendered image based on the point cloud data, the shadow mask, the direct light component, and the scattered light component.

[0137] It should be noted that the to-be-rendered image is first analyzed to obtain the direct light component and the scattered light component corresponding to the to-be-rendered image, i.e., the direct light component and the scattered light component are separated from the to-be-rendered image. Here, different image processing techniques are used to achieve this, for example, using the ratio between direct light and scattered light or predicting direct light and scattered light by training a deep learning model, for which the embodiments of the present application do not make any limitation.

[0138] The direct light component refers to light rays directly emitted from the virtual light source to the surface of an object and reflected or transmitted, i.e., light rays that directly reach the surface of an object without being blocked or scattered by other objects, usually including light rays from the main light source (such as the virtual light source) and responsible for producing bright spots, highlights, and other strong reflection effects on the surface of an object. The scattered light component refers to the phenomenon of random reflection of light rays on the surface of an object due to the roughness or material properties of the object. The propagation direction of the scattered light is uniform, and there is no specific light source directly contributing. The scattered light mainly causes a dull and soft lighting effect on the surface of an object, increasing the overall brightness of the object.

[0139] By applying the above embodiments, the soft shadow mask of the to-be-rendered image is generated in combination with the direct light component and the scattered light component. In this way, a high-quality soft shadow mask can be generated, which helps to improve the rendering effect of the soft shadow image in the subsequent process.

[0140] In actual implementation, after obtaining the direct light component and the scattered light component, based on the point cloud data, the shadow mask, the direct light component and the scattered light component, a process of generating a soft shadow mask of the image to be rendered can be that, based on the direct light component and the scattered light component, a potential shadow area in the image to be rendered is determined; based on the point cloud data, the shadow mask and the potential shadow area in the image to be rendered, a shadow intensity of each pixel point in the image to be rendered is predicted to obtain a shadow intensity value of each pixel point; and based on the shadow intensity value of each pixel point, the soft shadow mask of the image to be rendered is obtained.

[0141] It should be noted that the potential shadow area refers to an area that can be covered by a shadow under light irradiation, that is, under the position of the light source and the propagation direction of the light, the area can be blocked by other objects or surfaces to appear a shadow effect. When the light propagates from the light source to the surface of the object, if the light is blocked by other objects, the light cannot directly reach a certain area, and then the area has a potential shadow possibility.

[0142] Therefore, in the rendering process of the soft shadow image, considering the potential shadow area, not only can the light effect in the real world be simulated to make the rendering of the object more realistic and lifelike, but also the propagation path of the light can be accurately simulated in the rendering, and the appropriate shadow effect can be determined according to the geometric shape of the object and the position of the light source, which helps to enhance the realism and stereoscopic effect of the soft shadow rendering image, and makes the light effect more natural and lifelike.

[0143] In actual implementation, based on the direct light component and the scattered light component, a process of determining the potential shadow area in the image to be rendered can be that, for each pixel point in the image to be rendered, the following processing is performed: based on the direct light component and the scattered light component, a first brightness value of the pixel point in a direct light dimension corresponding to the direct light component and a second brightness value of the pixel point in a scattered light dimension corresponding to the scattered light component are obtained; based on the first brightness value and the second brightness value, a target pixel point is selected from a plurality of pixel points included in the image to be rendered; wherein the first brightness value of the target pixel point is less than a first brightness value threshold, and the second brightness value of the target pixel point is greater than a second brightness value threshold, and the second brightness value threshold is greater than the first brightness value threshold; and a region where the target pixel point is located in the image to be rendered is determined as the potential shadow area in the image to be rendered.

[0144] For example, if a pixel point is dark in the direct light component, that is, the pixel point is dark when irradiated by the direct light, for example, the brightness value of the pixel point when irradiated by the direct light is less than or equal to a brightness threshold, and the pixel point is bright in the scattered light component, that is, the pixel point is bright when irradiated by the scattered light, for example, the brightness value of the pixel point when irradiated by the scattered light is greater than the brightness threshold, then the pixel point is a potential shadow area, that is, the pixel point is a target pixel point.

[0145] It should be noted that the first brightness value of the pixel point in the direct light dimension corresponding to the direct light component refers to the brightness value of the pixel point when it is irradiated by direct light, and the second brightness value of the pixel point in the scattered light dimension corresponding to the scattered light component refers to the brightness value of the pixel point when it is irradiated by scattered light; at the same time, the first brightness value threshold and the second brightness value threshold can be pre-set.

[0146] In actual application, the first brightness value of the pixel point under the direct light component is determined, and the second brightness value of the pixel point under the scattered light component is determined, so as to determine the potential shadow area in the image based on the first brightness value and the second brightness value; in this way, the potential shadow area in the image to be rendered can be accurately determined, so that the subsequent rendering effect can be further improved when the image is rendered based on the potential shadow area in the subsequent process.

[0147] In actual implementation, the process of generating the soft shadow mask of the image to be rendered based on the point cloud data, the shadow mask, the direct light component and the scattered light component is similar to the process of generating the soft shadow mask of the image to be rendered based on the point cloud data and the shadow mask, and the process of generating the soft shadow mask of the image to be rendered based on the point cloud data, the shadow mask and the irradiation direction of the virtual light source, that is, the shadow degree of each pixel point in the image to be rendered is predicted based on the point cloud data, the shadow mask, the direct light component and the scattered light component, to obtain the shadow degree value of each pixel point; the soft shadow mask corresponding to the image to be rendered is obtained based on the shadow degree value of each pixel point. For this, the embodiments of the present application do not make redundant description.

[0148] In step 104, the image to be rendered is rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered.

[0149] In actual implementation, after the soft shadow mask is obtained, the image to be rendered can be directly rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered, or the image to be rendered can be restored to obtain a shadow-free image, so that the image to be rendered is rendered based on the soft shadow mask and the shadow-free image to obtain a soft shadow image of the image to be rendered.

[0150] In actual implementation, after the image to be rendered is rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered, the shadow area in the image to be rendered can also be shadow-eliminated based on the shadow mask to obtain a shadow-free image; thus, the process of rendering the image to be rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered can be to fuse the soft shadow mask with the shadow-free image to obtain a fused image, that is: In actual implementation, the process of generating the soft shadow mask of the image to be rendered based on the point cloud data, the shadow mask, the direct light component and the scattered light component is similar to the process of generating the soft shadow mask of the image to be rendered based on the point cloud data and the shadow mask, and the process of generating the soft shadow mask of the image to be rendered based on the point cloud data, the shadow mask and the irradiation direction of the virtual light source, that is, the shadow degree of each pixel point in the image to be rendered is predicted based on the point cloud data, the shadow mask, the direct light component and the scattered light component, to obtain the shadow degree value of each pixel point; the soft shadow mask corresponding to the image to be rendered is obtained based on the shadow degree value of each pixel point. For this, the embodiments of the present application do not make redundant description.

[0148] In step 104, the image to be rendered is rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered.

[0149] In actual implementation, after the soft shadow mask is obtained, the image to be rendered can be directly rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered, or the image to be rendered can be restored to obtain a shadow-free image, so that the image to be rendered is rendered based on the soft shadow mask and the shadow-free image to obtain a soft shadow image of the image to be rendered.

[0150] In actual implementation, after the image to be rendered is rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered, the shadow area in the image to be rendered can also be shadow-eliminated based on the shadow mask to obtain a shadow-free image; thus, the process of rendering the image to be rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered can be to fuse the soft shadow mask with the shadow-free image to obtain a fused image, that is:

[0151] wherein, is an input soft shadow mask, is an input shadow-free image, F f is a fused feature representation, f is a fusion network, θ f is a parameter of the fusion network.

[0152] Then, image rendering is performed on the fused image to obtain a soft shadow image corresponding to the to-be-rendered image, that is,

[0153] wherein, is an output soft shadow image, F f is a fused feature representation, r is an image generation network, θ r is a parameter of the image generation network.

[0154] It should be noted that the process of removing shadows from the shadow region in the to-be-rendered image based on the shadow mask to obtain a shadow-free image can be implemented based on a model, such as a diffusion model, or can be implemented by other means, and the present application does not limit this.

[0155] In some embodiments, when the process of removing shadows from the shadow region in the to-be-rendered image based on the shadow mask to obtain a shadow-free image is implemented based on a model, the process of restoring the shadow region in the to-be-rendered image based on the shadow mask to obtain a shadow-free image can be, based on the shadow mask, determining the shadow region in the to-be-rendered image; based on the shadow region, performing at least one denoising processing on the to-be-rendered image to obtain a shadow-free image.

[0156] Among them, for the process of performing at least one denoising processing on the to-be-rendered image to obtain a shadow-free image, it can be implemented based on a diffusion model, wherein a pre-trained diffusion model is obtained, and then based on the diffusion model, at least one denoising processing is performed on the to-be-rendered image to obtain a shadow-free image.

[0157] By applying the above embodiments, at least one denoising processing is performed on the to-be-rendered image based on the shadow region to obtain a shadow-free image; in this way, through denoising processing, the shadow region in the image can be removed with high precision, which is convenient for subsequent image processing and analysis.

[0158] It should be noted that the number of denoising processes is pre-set and associated with the diffusion model; the diffusion model can convert any image restoration task into a mapping process from a Gaussian noise image to a target image, the core idea of the model is to gradually diffuse the latent representation of the image to Gaussian noise, then use a reverse diffusion network to gradually recover the latent representation of the image from Gaussian noise, and finally through a decoder, the latent representation is converted into a target image.

[0159] In this way, by recovering the shadow-free image from the shadow image, and then applying the shadow-free image in the subsequent soft shadow image generation process, compared with directly generating a soft shadow image based on the shadow image, not only the clarity and realism of the soft shadow image are improved, but also the rendering effect of the soft shadow image is improved.

[0160] In some embodiments, for the process of obtaining a shadow-free image by at least once denoising the image to be rendered, when performing image feature processing, an anchor stripe self-attention mechanism or a window self-attention mechanism can be used, so that not only the long-distance dependency relationship and local detail information of the image can be captured, while maintaining the efficiency of calculation and memory, or the local context information and global perception ability of the image can be captured, while reducing the overhead of calculation and memory; moreover, efficient and explicit image hierarchical structure modeling is achieved, and the accuracy in the image feature processing process is improved.

[0161] It should be noted that for the use process of the anchor stripe self-attention mechanism, the image to be rendered can be first divided into multiple stripes, and then self-attention operation is performed on each stripe while using anchor points to constrain the information flow between the stripes, wherein the stripe represents the distribution of the anchor points on the input sequence, i.e. the interval and position between the anchor points; and the anchor points are pre-selected to guide the self-attention model to focus on a specific position. In this way, the model can more effectively capture the long-distance dependency relationship in the input sequence, and can avoid the problem of excessive consumption of computing resources caused by too long sequence.

[0162] And for the use process of the window self-attention mechanism, when performing image feature processing, the image can also be divided into multiple windows, and the self-attention mechanism can be applied to each window. In this way, the model can more focused on the key areas in the input data, thereby improving the quality and accuracy of the generated image.

[0163] By using the above-mentioned embodiments of the present application, first, a shadow mask of a to-be-rendered image including a shadow area is generated, then a soft shadow mask of the to-be-rendered image is generated based on point cloud data of the to-be-rendered image and the shadow mask, and finally, image rendering is performed on the to-be-rendered image based on the soft shadow mask to obtain a soft shadow image of the to-be-rendered image. In this way, the soft shadow image is obtained by performing image rendering based on the generated soft shadow mask of the to-be-rendered image, which not only reduces the performance overhead, but also improves the rendering efficiency when rendering the soft shadow image. At the same time, since the generation of shadows is mostly related to the stereoscopic nature of objects, the point cloud data of the to-be-rendered image is also used in the soft shadow image rendering process in the present application, that is, the stereoscopic nature of objects in the to-be-rendered image in a three-dimensional environment is considered. Compared with the related art scheme of reconstructing a shadow image only by using a graphics geometry sample and a shadow mask, the problem of the generated soft shadow image being relatively harsh is solved, and the rendering effect when rendering the soft shadow image is improved.

[0164] In the following, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0165] In the related art, when performing soft shadow image rendering, a corresponding shadow mask is mostly generated first, and then a shadow image is reconstructed by using a graphics geometry sample and the shadow mask. However, such a method needs a large amount of calculation, such as using sufficient samples to design three dimensions and colors, and it is difficult to realize real-time shadow reconstruction. Not only is the performance overhead extremely high, but also the efficiency when rendering the soft shadow image is relatively low. At the same time, the way of reconstructing a shadow image only by using a graphics geometry sample and a shadow mask will also cause the generated soft shadow image to be relatively harsh, which reduces the rendering effect when rendering the soft shadow image.

[0166] Based on this, the embodiments of the present application provide a rendering method of a soft shadow image, which can recover a shadow-free image from a shadow image, and then generate a realistic soft shadow effect according to a three-dimensional point cloud and a light source direction. The main innovations of the method are as follows:

[0167] First, the method uses a diffusion process-based generative model to convert an arbitrary image restoration task into a mapping problem from a Gaussian noise image to a target image, thereby realizing effective extraction and reconstruction of the latent representation of the image.

[0168] Second, the method utilizes the cross-scale similarity and anisotropic features of the image, and realizes efficient and explicit modeling of the image hierarchy through an anchor stripe self-attention mechanism and a window self-attention mechanism, thereby improving the quality and details of the image.

[0169] Thirdly, the method uses a shadow image decomposition model to decompose the shadow image into direct light component and scattered light component, so as to realize the detection and segmentation of the shadow area, and provide the basis for the subsequent soft shadow generation.

[0170] Fourthly, the method uses a shadow generation module to generate a soft shadow mask according to the three-dimensional point cloud and the light source direction, so as to realize the simulation and rendering of the soft shadow effect.

[0171] Referring to FIG. 8, FIG. 8 is a flowchart of the soft shadow image rendering method provided by the embodiment of the present application. Based on FIG. 8, the soft shadow image rendering method provided by the embodiment of the present application is realized through steps 801 to 806. The soft shadow image rendering method provided by the embodiment of the present application is divided into four steps, i.e. shadow detection, three-dimensional projection, diffusion generation and soft shadow generation.

[0172] The shadow detection process is to use a simple shadow detection algorithm based on color information to convert the image to HSV color space, so as to obtain the brightness value of each pixel point in the image, and then use a maximum minimum value filter to identify the shadow area and the non-shadow area, and set the pixel value of the shadow area to 1 and the pixel value of the non-shadow area to 0, so as to obtain the shadow mask.

[0173] In actual implementation, the purpose of shadow detection is to identify the shadow area in the image according to the color information of the image and generate a shadow mask. Here, a simple shadow detection algorithm based on a maximum minimum value filter is used. This algorithm has the following advantages: first, it does not need to know the position and direction of the light source in advance, nor does it need to segment or classify the image; second, it can handle different lighting conditions and shadow types, including hard shadow, soft shadow and self-shadow; third, the calculation amount is low and suitable for real-time application. The basic idea of this algorithm is that the brightness value of the shadow area is usually lower than that of the surrounding non-shadow area, so the shadow area can be judged by comparing the local maximum value and the minimum value of the image. The algorithm includes the following steps: first, convert the image to HSV color space to obtain the brightness value of each pixel point in the image; then define the kernel size k of the maximum minimum value filter, generally take an odd number such as k = 15; then use the maximum value filter to perform dilation operation on the brightness channel to obtain the local maximum value matrix M as shown in formulas (2), (3) and (4); then use the minimum value filter to perform erosion operation on the brightness channel to obtain the local minimum value matrix m as shown in formulas (5), (6) and (7); finally, calculate the shadow mask as shown in formula (8).

[0174] And the three-dimensional projection process, that is, using a deep learning-based method to estimate the depth map of the scene from the two-dimensional image; then, according to the depth map and the camera parameters, each pixel in the two-dimensional image is projected to the corresponding pixel point in the three-dimensional space. In this way, the point cloud representation of the three-dimensional scene can be obtained.

[0175] In actual implementation, the purpose of three-dimensional projection is to correspond each pixel in the image (to be rendered image) to a point in the three-dimensional space according to the two-dimensional image and the camera parameters, so as to obtain the point cloud representation (point cloud data) of the three-dimensional scene. This is an important step because it can convert the information of the three-dimensional image into the information of the three-dimensional space, providing a basis for subsequent shadow generation and rendering operations. Here, a deep learning-based method is used to estimate the depth map of the scene from the two-dimensional image, and then a projection transformation is used to convert the two-dimensional coordinates into three-dimensional coordinates. The process includes the following steps: first, depth map estimation, using a depth estimation network, inputting a two-dimensional image, and outputting the depth value of each pixel. Among them, the depth estimation network is a deep convolutional neural network that can learn the depth information of the scene from the two-dimensional image without any other supervision signal or prior knowledge. The depth estimation network can use different architectures and loss functions, such as U-Net, ResNet, BerHu, etc. The size of the depth map is the same as that of the two-dimensional image, that is, HxW, where H is the height of the image and W is the width of the image. Each pixel value D(i,j) in the depth map represents the distance from the i-th row and j-th column pixel in the image to the camera, in meters, as shown in equation (9).

[0176] Second, two-dimensional coordinate normalization. In order to facilitate subsequent projection transformation, the pixel coordinates (i,j) in the two-dimensional image need to be converted to normalized homogeneous coordinates (x,y,1) as shown in equations (11) and (12); where the range of x and y is [-1,1]. In this way, the size and resolution of the image can be eliminated, and the projection transformation only depends on the intrinsic matrix of the camera.

[0177] Third, three-dimensional coordinate calculation. As shown in equation (10), according to the intrinsic matrix K of the camera, the two-dimensional coordinates (x,y,1) and the depth value D(i,j) can be converted to three-dimensional coordinates (X,Y,Z,1), where X, Y and Z represent the coordinates of the horizontal, vertical and depth directions in the three-dimensional space, respectively, in meters. Wherein, the intrinsic matrix K of the camera is a 3x3 matrix, which contains the focal length, principal point and distortion parameters of the camera, which can generally be obtained by the camera calibration process.

[0178] Fourth, point cloud generation. Store the three-dimensional coordinates (X, Y, Z, 1) as an N x 4 matrix, where N = H x W is the total number of pixels in the image. In this way, the point cloud representation of the three-dimensional scene is obtained, which can be used for subsequent shadow generation and rendering operations. Among them, the point cloud is a commonly used three-dimensional data structure, which can represent any shape and object in three-dimensional space, and can also be processed and analyzed, such as plane fitting, surface reconstruction, feature extraction, etc.

[0179] And the diffusion generation process is to recover the shadow-free image from the shadow image using the latent space diffusion model-based method. This method utilizes the cross-scale similarity and anisotropy characteristics of the image, and realizes efficient and explicit image hierarchical structure modeling through the anchor stripe self-attention mechanism and window self-attention mechanism. This method can handle large-size real images and has achieved excellent results on various image restoration tasks.

[0180] In actual implementation, the purpose of diffusion generation is to recover the shadow-free image from the shadow image, that is, to remove the shadow effect in the image, making the image clearer and more natural. Here, a latent space diffusion model-based method is used, which utilizes the cross-scale similarity and anisotropy characteristics of the image, and realizes efficient and explicit image hierarchical structure modeling through the anchor stripe self-attention mechanism and window self-attention mechanism. This method can handle large-size real images and has achieved excellent results on various image restoration tasks. The latent space diffusion model can convert any image restoration task into a mapping problem from a Gaussian noise image to a target image. The core idea of this model is to gradually diffuse the latent representation of the image to Gaussian noise, and then use a reverse diffusion network to gradually recover the latent representation of the image from Gaussian noise, that is:

[0181] Where z0 is the latent representation of the image, zt is the diffusion state at step t, βt is the diffusion coefficient at step t, and ∈t is the noise term at step t.

[0182] Finally, through a decoder, the latent representation is converted into a target image, that is:

[0183] Where, is the output shadow-free image, g is a decoder network, φ is the parameter of the decoder network, and z0 is the latent representation of the image.

[0184] It should be noted that the reverse diffusion network can be a deep convolutional neural network, the core idea of which is to use a residual connection U-Net structure to extract and fuse features for each diffusion state, and then use a prediction module to predict the residual of the next diffusion state according to the current diffusion state and the noise term, so as to realize the process of reverse diffusion.

[0185] In actual implementation, for the anchor stripe self-attention mechanism. This mechanism is a method of image feature extraction and fusion based on self-attention, which can capture long-distance dependencies and local details of the image while maintaining high efficiency of computation and memory. The core idea of this mechanism is to divide the image into multiple stripes, and then perform self-attention operation on each stripe while using anchor points to constrain the information flow between stripes, thereby realizing cross-stripe feature fusion.

[0186] In actual implementation, for the window self-attention mechanism. This mechanism is a method of image feature extraction and fusion based on self-attention, which can capture local context information and global perception ability of the image while reducing the overhead of computation and memory. The core idea of this mechanism is to divide the image into multiple windows, and then perform self-attention operation on each window while using a learnable position encoding to enhance the feature representation within the window, thereby realizing feature fusion within the window.

[0187] The process of soft shadow generation is to use a neural network-based method to generate realistic soft shadow effects based on shadow-free images, shadow masks, and three-dimensional point clouds. This method uses a shadow image decomposition model to decompose the shadow image into direct light components and scattered light components, then uses a shadow generation module to generate a soft shadow mask based on three-dimensional point clouds (i.e., point cloud representation) and light source direction (i.e., direct light component and scattered light component), and finally uses a shadow fusion module to fuse the soft shadow mask and the shadow-free image to obtain the final soft shadow image.

[0188] In actual implementation, the purpose of soft shadow generation is to generate realistic soft shadow effects based on shadow-free images, shadow masks, and three-dimensional point clouds. Here, a neural network-based method is used, which uses a shadow image decomposition model to decompose the shadow image into direct light components and scattered light components, then uses a shadow generation module to generate a soft shadow mask based on three-dimensional point clouds and light source direction, and finally uses a shadow fusion module to fuse the soft shadow mask and the shadow-free image to obtain the final soft shadow image.

[0189] For the shadow image decomposition model. This model is a deep convolutional neural network that decomposes a shadow image into direct light and diffuse light components. The core idea of this model is to use an encoder-decoder structure to extract features from the shadow image and reconstruct it, while detecting and segmenting the shadow region based on the shadow mask, thereby achieving the decomposition of the shadow image.

[0190] For the shadow generation module. This module is a deep convolutional neural network that generates a soft shadow mask based on a three-dimensional point cloud and light source direction. The core idea of this module is to use a point cloud encoder to extract features from the three-dimensional point cloud, and then use a shadow predictor to predict the shadow degree of each point based on the point cloud features and light source direction, thereby achieving the generation of a soft shadow mask.

[0191] For the shadow fusion module. This module is a deep convolutional neural network that fuses the soft shadow mask and the shadow-free image to obtain the final soft shadow image. The core idea of this module is to use a fusion network to extract features from the soft shadow mask and the shadow-free image and fuse them, and then use a reconstruction network (image generation network) to reconstruct and refine the fused features, thereby achieving the generation of a soft shadow image, as shown in equations (16) and (17).

[0192] By applying the above embodiments of the present application, a shadow mask of a to-be-rendered image including a shadow region is first generated, then a soft shadow mask of the to-be-rendered image is generated based on point cloud data of the to-be-rendered image and the shadow mask, and finally the to-be-rendered image is rendered based on the soft shadow mask to obtain a soft shadow image of the to-be-rendered image. In this way, the soft shadow image is obtained by rendering the to-be-rendered image using the generated soft shadow mask, which not only reduces the performance overhead, but also improves the rendering efficiency when rendering the soft shadow image. At the same time, since the generation of shadows is mostly related to the stereoscopic nature of objects, the point cloud data of the to-be-rendered image is also utilized in the soft shadow image rendering process, i.e., the stereoscopic nature of objects in the to-be-rendered image in a three-dimensional environment is considered. Compared with the related art scheme that only uses geometric samples and a shadow mask to reconstruct a shadow image, the problem of harsh shadow effects of the generated soft shadow image is solved, and the rendering effect when rendering the soft shadow image is improved.

[0193] The following continues to illustrate an exemplary structure of the soft shadow image rendering device 455 provided by the embodiments of the present application as a software module. In some embodiments, as shown in FIG. 2, the software module stored in the soft shadow image rendering device 455 of the memory 450 can include:

[0194] The first acquisition module 4551 is configured to acquire a to-be-rendered image including a shadow region.

[0195] The generating module 4552 is configured to generate a shadow mask of the image to be rendered based on the shadow region in the image to be rendered.

[0196] The second obtaining module 4553 is configured to obtain point cloud data of the image to be rendered, and generate a soft shadow mask of the image to be rendered based on the point cloud data and the shadow mask.

[0197] The rendering module 4554 is configured to perform image rendering on the image to be rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered.

[0198] In some embodiments, the generating module 4552 is further configured to obtain a brightness value of each pixel point in the image to be rendered; and perform the following processing for each pixel point in the image to be rendered to obtain a mask value of the pixel point: perform dilation processing on the pixel point to obtain a local maximum brightness value of the pixel point, and perform erosion processing on the pixel point to obtain a local minimum brightness value of the pixel point; determine the mask value of the pixel point based on the local maximum brightness value and the local minimum brightness value; and obtain the shadow mask corresponding to the image to be rendered based on the mask values of the pixel points.

[0199] In some embodiments, the apparatus further includes a third obtaining module configured to obtain a target neighborhood range of the pixel point, the target neighborhood range including at least one adjacent pixel point, the adjacent pixel point being used to indicate a pixel point adjacent to the pixel point; and the generating module 4552 is further configured to obtain a brightness value of each adjacent pixel point; select a maximum brightness value from at least one of the brightness values, and take the maximum brightness value as a local maximum brightness value of the pixel point; and select a minimum brightness value from the brightness values of the adjacent pixel points, and take the minimum brightness value as a local minimum brightness value of the pixel point.

[0200] In some embodiments, the generating module 4552 is further configured to obtain a ratio of the local minimum brightness value and the local maximum brightness value of the pixel point; determine a first constant as the mask value of the pixel point when the ratio is less than a target ratio; and determine a second constant as the mask value of the pixel point when the ratio is greater than or equal to the target ratio; wherein the first constant and the second constant are different.

[0201] In some embodiments, the second obtaining module 4553 is further configured to obtain a depth image of the image to be rendered, and coordinates of each pixel point in the image to be rendered; convert the coordinates of each pixel point in the image to be rendered based on the depth image to obtain three-dimensional coordinates of each pixel point; and obtain point cloud data of the image to be rendered based on the three-dimensional coordinates of each pixel point in the image to be rendered.

[0202] In some embodiments, the second obtaining module 4553 is further configured to construct a spatial coordinate system corresponding to the image to be rendered, and determine height values of each pixel point in the image to be rendered in the spatial coordinate system; and generate a depth image of the image to be rendered based on the height values.

[0203] In some embodiments, the second obtaining module 4553 is further configured to obtain an intrinsic matrix of a virtual camera corresponding to the image to be rendered, the intrinsic matrix being used to indicate a focal length, a principal point and distortion parameters of the virtual camera; obtain a depth value of each pixel point in the depth image; and convert the coordinates of each pixel point in the image to be rendered in combination with the intrinsic matrix and the depth value of each pixel point to obtain three-dimensional coordinates of each pixel point.

[0204] In some embodiments, the second obtaining module 4553 is further configured to multiply the coordinates of each pixel point in the image to be rendered with the corresponding depth value to obtain three-dimensional coordinates of each pixel point in a coordinate system corresponding to the virtual camera; multiply the three-dimensional coordinates of each pixel point in the coordinate system corresponding to the virtual camera with an inverse matrix of the intrinsic matrix of the virtual camera to obtain three-dimensional coordinates of each pixel point in a world coordinate system; and take the three-dimensional coordinates of each pixel point in the world coordinate system as the converted three-dimensional coordinates.

[0205] In some embodiments, the device further includes a removing module configured to perform shadow removal on the shadow region in the image to be rendered based on the shadow mask to obtain a shadow-free image; the rendering module 4554 is further configured to fuse the soft shadow mask with the shadow-free image to obtain a fused image; and perform image rendering on the fused image to obtain a soft shadow image of the image to be rendered.

[0206] In some embodiments, the removing module is further configured to determine the shadow region in the image to be rendered based on the shadow mask; and perform at least one denoising process on the image to be rendered based on the shadow region to obtain the shadow-free image.

[0207] In some embodiments, the elimination module is further configured to acquire a pre-trained diffusion model; and perform at least one denoising process on the image to be rendered based on the diffusion model to obtain the shadow-free image.

[0208] In some embodiments, the generation module 4552 is further configured to predict a shadow intensity of each pixel in the image to be rendered based on the point cloud data and the shadow mask to obtain a shadow intensity value of each pixel; and obtain a soft shadow mask of the image to be rendered based on the shadow intensity value of each pixel.

[0209] In some embodiments, the apparatus further comprises a first analysis module configured to analyze the image to be rendered to obtain an illumination direction of a virtual light source corresponding to the image to be rendered; and the generation module 4552 is further configured to generate a soft shadow mask of the image to be rendered based on the point cloud data, the shadow mask, and the illumination direction of the virtual light source.

[0210] In some embodiments, the apparatus further comprises a second analysis module configured to analyze the image to be rendered to obtain a direct light component and a scattered light component corresponding to the image to be rendered; and the generation module 4552 is further configured to generate a soft shadow mask of the image to be rendered based on the point cloud data, the shadow mask, the direct light component, and the scattered light component.

[0211] In some embodiments, the generation module 4552 is further configured to determine a potential shadow area in the image to be rendered based on the direct light component and the scattered light component; predict a shadow intensity of each pixel in the image to be rendered based on the point cloud data, the shadow mask, and the potential shadow area in the image to be rendered to obtain a shadow intensity value of each pixel; and obtain a soft shadow mask of the image to be rendered based on the shadow intensity value of each pixel.

[0212] In some embodiments, the generating module 4552 is further configured to, for each pixel in the image to be rendered, perform the following processing: based on the direct light component and the scattered light component, obtain a first brightness value of the pixel in a direct light dimension corresponding to the direct light component, and a second brightness value of the pixel in a scattered light dimension corresponding to the scattered light component; based on the first brightness value and the second brightness value, select a target pixel from the plurality of pixels included in the image to be rendered; wherein the first brightness value of the target pixel is less than a first brightness value threshold, and the second brightness value of the target pixel is greater than a second brightness value threshold, the second brightness value threshold being greater than the first brightness value threshold; determine a region in which the target pixel is located in the image to be rendered as the potential shadow region in the image to be rendered.

[0213] The embodiment of the present application provides a computer program product, which comprises computer executable instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions, so that the electronic device performs the soft shadow image rendering method provided in the embodiment of the present application, for example, the soft shadow image rendering method shown in FIG. 3.

[0214] The embodiment of the present application provides a computer readable storage medium storing executable instructions, wherein the computer readable storage medium stores the computer executable instructions. When the computer executable instructions are executed by a processor, the processor will execute the soft shadow image rendering method provided in the embodiment of the present application, for example, the soft shadow image rendering method shown in FIG. 3.

[0215] In some embodiments, the computer readable storage medium can be a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), an erasable programmable read-only memory (Erasable Programmable Read-Only Memory, EPROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a CD-ROM memory. It can also be various devices including one or any combination of the above memories.

[0216] In some embodiments, the computer-executable instructions can take the form of programs, software, software modules, scripts, or code, written in any suitable programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0217] By way of example, a computer-executable instruction can, but need not, correspond to a file in a file system. A computer-executable instruction can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code.

[0218] By way of example, a computer-executable instruction can be deployed to be executed on one electronic device or on multiple electronic devices that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0219] It should be noted that, in the embodiments of the present application, the data related to the to-be-rendered image and the like need to be obtained with the permission or consent of the relevant parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0220] In summary, the embodiments of the present application have the following beneficial effects:

[0221] (1) The soft shadow image is obtained by performing image rendering on the soft shadow mask of the generated to-be-rendered image, which not only reduces the performance overhead but also improves the rendering efficiency when the soft shadow image is rendered. Since the generation of the shadow is mostly related to the stereoscopic nature of the object, the point cloud data of the to-be-rendered image is also utilized in the soft shadow image rendering process, i.e., the stereoscopic nature of the object in the to-be-rendered image in the three-dimensional environment is considered. Compared with the scheme in the related art that only utilizes the graphic geometry sample and the shadow mask to reconstruct the shadow image, the problem of the generated soft shadow image being relatively harsh is solved, and the rendering effect when the soft shadow image is rendered is improved.

[0222] (2) The shadow mask is determined based on the local minimum brightness value and the local maximum brightness value of the pixel point, which helps to distinguish the brightness change region in the to-be-rendered image, so as to facilitate the distinction between the shadow region and the non-shadow region in the to-be-rendered image, and thus improve the rendering effect of the soft shadow image in the subsequent process.

[0223] (3) Through the normalization processing of the coordinates, the influence of the size and resolution of the image on the coordinate transformation can be eliminated, so that the coordinate transformation only depends on the intrinsic matrix of the camera.

[0224] (4) In combination with the illumination direction of the virtual light emitted by the virtual light source corresponding to the image to be rendered, the soft shadow mask is generated, which not only can improve the accuracy of the generated soft shadow mask, but also can improve the rendering effect of the subsequent soft shadow image, that is, a realistic soft shadow effect can be generated, thereby enhancing the light and shadow effect and atmosphere of the soft shadow image.

[0225] (5) In the rendering process of the soft shadow image, considering the potential shadow area, not only can help simulate the lighting effect in the real world, making the rendering of the object more realistic and realistic, but also can accurately simulate the propagation path of the light in the rendering, and determine the appropriate shadow effect according to the geometric shape of the object and the position of the light source, which helps to enhance the realism and stereoscopic sense of the soft shadow rendering image, making the lighting effect more natural and realistic.

[0226] (6) By recovering the shadow-free image from the shadow image, the shadow-free image is applied in the generation process of the subsequent soft shadow image, compared with directly generating the soft shadow image based on the shadow image, not only improves the clarity and realism of the soft shadow image, but also improves the rendering effect of the soft shadow image.

[0227] The above is only an embodiment of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement and improvement made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A method for rendering a soft shadow image, the method being performed by an electronic device, the method comprising: obtaining a to-be-rendered image comprising a shadow region; generating a shadow mask of the to-be-rendered image based on the shadow region in the to-be-rendered image; obtaining point cloud data of the to-be-rendered image, and generating a soft shadow mask of the to-be-rendered image based on the point cloud data and the shadow mask; performing image rendering on the to-be-rendered image based on the soft shadow mask to obtain a soft shadow image of the to-be-rendered image. The generating of the shadow mask of the to-be-rendered image based on the shadow region in the to-be-rendered image comprises:

2. The method of claim 1, wherein, obtaining a brightness value of each pixel in the to-be-rendered image; for each pixel in the to-be-rendered image, performing the following processing to obtain a mask value of the pixel: performing dilation processing on the pixel to obtain a local maximum brightness value of the pixel, and performing erosion processing on the pixel to obtain a local minimum brightness value of the pixel; determining the mask value of the pixel based on the local maximum brightness value and the local minimum brightness value; obtaining the shadow mask of the to-be-rendered image based on the mask values of the pixels. Before the performing of the dilation processing on the pixel to obtain the local maximum brightness value of the pixel, the method further comprises:

3. The method of claim 2, wherein, obtaining a target neighborhood range of the pixel, the target neighborhood range comprising at least one adjacent pixel, the adjacent pixel being used to indicate a pixel adjacent to the pixel; The performing of the dilation processing on the pixel to obtain the local maximum brightness value of the pixel comprises: obtaining a brightness value of each adjacent pixel; selecting a maximum brightness value from the at least one brightness value, and taking the maximum brightness value as the local maximum brightness value of the pixel; The performing of the erosion processing on the pixel to obtain the local minimum brightness value of the pixel comprises: selecting a minimum brightness value from the brightness values of the adjacent pixels, and taking the minimum brightness value as the local minimum brightness value of the pixel. The determining of the mask value of the pixel based on the local maximum brightness value and the local minimum brightness value comprises:

4. The method of claim 2 or 3, wherein, obtaining a ratio of the local minimum brightness value to the local maximum brightness value of the pixel; when the ratio is less than a target ratio, determining a first constant as the mask value of the pixel; when the ratio is greater than or equal to the target ratio, determining a second constant as the mask value of the pixel; wherein the first constant is different from the second constant. The obtaining of the point cloud data of the to-be-rendered image comprises:

5. The method of any one of claims 1 to 4, wherein, obtaining a depth image of the to-be-rendered image and a coordinate of each pixel in the to-be-rendered image; based on the depth image, converting the coordinate of each pixel in the to-be-rendered image to obtain a three-dimensional coordinate of each pixel; based on the three-dimensional coordinate of each pixel in the to-be-rendered image, obtaining the point cloud data of the to-be-rendered image. The obtaining of the depth image of the to-be-rendered image comprises:

6. The method of claim 5, wherein, ​ construct a spatial coordinate system corresponding to the image to be rendered, and determine the height values of the pixel points in the image to be rendered in the spatial coordinate system; generate a depth image of the image to be rendered based on the height values.

7. The method of claim 5 or 6, wherein, The method for converting the coordinates of each pixel point in the image to be rendered based on the depth image to obtain the three-dimensional coordinates of each pixel point comprises: obtaining an intrinsic matrix of a virtual camera corresponding to the image to be rendered, the intrinsic matrix being used to indicate the focal length, principal point and distortion parameters of the virtual camera; obtaining the depth values of each pixel point in the depth image; combining the intrinsic matrix and the depth values of each pixel point to convert the coordinates of each pixel point in the image to be rendered to obtain the three-dimensional coordinates of each pixel point.

8. The method of claim 7, wherein, The method for converting the coordinates of each pixel point in the image to be rendered based on the depth image to obtain the three-dimensional coordinates of each pixel point comprises: multiplying the coordinates of each pixel point in the image to be rendered by the corresponding depth values to obtain the three-dimensional coordinates of each pixel point in the coordinate system corresponding to the virtual camera; for each pixel point, multiplying the three-dimensional coordinates of the pixel point in the coordinate system corresponding to the virtual camera by the inverse matrix of the intrinsic matrix of the virtual camera to obtain the three-dimensional coordinates of each pixel point in the world coordinate system; taking the three-dimensional coordinates of each pixel point in the world coordinate system as the converted three-dimensional coordinates.

9. The method of any one of claims 1 to 8, wherein, After the method for generating a shadow mask of the image to be rendered based on the shadow area in the image to be rendered, the method further comprises: based on the shadow mask, performing shadow removal on the shadow area in the image to be rendered to obtain a shadow-free image; The method for performing image rendering on the image to be rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered comprises: fusing the soft shadow mask and the shadow-free image to obtain a fused image; performing image rendering on the fused image to obtain a soft shadow image of the image to be rendered.

10. The method of claim 9, wherein, The method for performing shadow removal on the shadow area in the image to be rendered based on the shadow mask to obtain a shadow-free image comprises: based on the shadow mask, determining the shadow area in the image to be rendered; based on the shadow area, performing at least one denoising process on the image to be rendered to obtain the shadow-free image.

11. The method of claim 10, wherein, The method for performing at least one denoising process on the image to be rendered based on the shadow area to obtain the shadow-free image comprises: obtaining a pre-trained diffusion model; based on the diffusion model, performing at least one denoising process on the image to be rendered to obtain the shadow-free image.

12. The method of any one of claims 1 to 11, wherein, The method for generating a soft shadow mask of the image to be rendered based on the point cloud data and the shadow mask comprises: based on the point cloud data and the shadow mask, predicting the shadow intensity of each pixel point in the image to be rendered to obtain the shadow intensity values of each pixel point; Based on the shadow intensity values of each of the pixel points, a soft shadow mask of the to-be-rendered image is obtained.

13. The method of any one of claims 1 to 12, wherein, After the to-be-rendered image including the shadow region is obtained, the method further includes: parsing the to-be-rendered image to obtain an irradiation direction of a virtual light source corresponding to the to-be-rendered image; The generating of the soft shadow mask of the to-be-rendered image based on the point cloud data and the shadow mask includes: The generating of the soft shadow mask of the to-be-rendered image based on the point cloud data, the shadow mask, and the irradiation direction of the virtual light source includes:

14. The method of any one of claims 1 to 13, wherein, After the to-be-rendered image including the shadow region is obtained, the method further includes: parsing the to-be-rendered image to obtain a direct light component and a scattered light component corresponding to the to-be-rendered image; The generating of the soft shadow mask of the to-be-rendered image based on the point cloud data and the shadow mask includes: The generating of the soft shadow mask of the to-be-rendered image based on the point cloud data, the shadow mask, the direct light component, and the scattered light component includes:

15. The method of claim 14, wherein, The generating of the soft shadow mask of the to-be-rendered image based on the point cloud data, the shadow mask, the direct light component, and the scattered light component includes: Based on the direct light component and the scattered light component, a potential shadow region in the to-be-rendered image is determined; Based on the point cloud data, the shadow mask, and the potential shadow region in the to-be-rendered image, a shadow intensity of each of the pixel points in the to-be-rendered image is predicted to obtain a shadow intensity value of each of the pixel points; Based on the shadow intensity values of each of the pixel points, a soft shadow mask of the to-be-rendered image is obtained.

16. The method of claim 15, wherein, The determining of the potential shadow region in the to-be-rendered image based on the direct light component and the scattered light component includes: For each of the pixel points in the to-be-rendered image, the following processing is performed: Based on the direct light component and the scattered light component, a first brightness value of the pixel point in a direct light dimension corresponding to the direct light component and a second brightness value of the pixel point in a scattered light dimension corresponding to the scattered light component are obtained respectively; Based on the first brightness value and the second brightness value, a target pixel point is selected from a plurality of pixel points included in the to-be-rendered image; The first brightness value of the target pixel point is less than a first brightness value threshold, and the second brightness value of the target pixel point is greater than a second brightness value threshold, the second brightness value threshold being greater than the first brightness value threshold; A region in which the target pixel point is located in the to-be-rendered image is determined as the potential shadow region in the to-be-rendered image.

17. A device for rendering a soft shadow image, the device comprising: a first obtaining module configured to obtain a to-be-rendered image including a shadow region; a generating module configured to generate a shadow mask of the to-be-rendered image based on the shadow region in the to-be-rendered image; a second obtaining module configured to obtain point cloud data of the to-be-rendered image, and generate a soft shadow mask of the to-be-rendered image based on the point cloud data and the shadow mask; The rendering module is configured to perform image rendering on the image to be rendered based on the soft shadow mask to obtain a soft shadow image of the image to be rendered. 18.An electronic device comprising: a memory configured to store executable instructions; a processor configured to implement the method of rendering a soft shadow image according to any one of claims 1 to 16 when executing the executable instructions stored in the memory. 19.A computer readable storage medium storing executable instructions configured to cause a processor to implement the method of rendering a soft shadow image according to any one of claims 1 to 16 when executing the executable instructions. 20.A computer program product comprising a computer program or computer executable instructions configured to implement the method of rendering a soft shadow image according to any one of claims 1 to 16 when executed by a processor.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN113205586A

  • Method and device for generating shadow in image, electronic equipment and storage medium

    CN114626468A

  • Shadow estimation method and device, electronic equipment and readable storage medium

    CN115330926A

  • Visual three-dimensional enhancement method and device, equipment and storage medium

    CN116309121A

  • Image shadow removal method and device, storage medium and equipment

    CN117011182A

Cited By

  • Illumination light supplementing method and system for agricultural detection

    CN121510406A