Scene image display method and apparatus, and computer device and storage medium

By identifying the location of the sound source and adding disturbance effects to the special effect area, the problem of limitations in the interactive form of the special effect gameplay and the actual scene is solved, and a more vivid and dynamic special effect display is achieved.

WO2025113087A1PCT designated stage expired Publication Date: 2025-06-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/129426
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-01
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve effective interaction between special effects gameplay and actual scenes. Special effects gameplay mainly relies on visual information and cannot fully reflect the rhythm of sound and other interactive forms.

Method used

By collecting scene images of the target scene, identifying the sound source location of the sound source object, generating unit images, and adding perturbation effects to the special effect area to achieve the interaction between the special effect gameplay and the actual scene.

Benefits of technology

The spatial representation of the special effects gameplay and the actual scene is realized, which enhances the dynamicity and vividness of the special effects area, allowing users to better perceive the real scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129426_05062025_PF_FP_ABST
    Figure CN2024129426_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computers. Disclosed are a scene image display method and apparatus, and a computer device and a storage medium. The method comprises: collecting a scene image of a target scene, and identifying a sound source position of a sound source object in the scene image; generating a unit image corresponding to the scene image, and displaying the unit image; determining a special effect area in the unit image on the basis of the sound source position; and adding a disturbance effect to the image unit in the special effect area, so as to display the image unit in the special effect area according to the disturbance effect.
Need to check novelty before this filing date? Find Prior Art

Description

Scene image display method, device, computer equipment and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202311617072.6, filed on November 29, 2023, entitled “Method, device, computer equipment and storage medium for displaying scene images”. The entire contents of that application are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of computer technology, and in particular to a method, apparatus, computer equipment, and storage medium for displaying scene images. Background Art

[0004] Currently, special effects props generated by capturing scenes with the rear cameras of mobile phones, tablets, and other computer devices offer users a variety of gameplay options, such as placement, collision, and deformation. However, most special effects gameplay is based on visual information, which limits the interaction between special effects gameplay and the actual scene, making it difficult to effectively achieve this interaction.

[0005] Summary of the Invention

[0006] In view of this, the embodiments of the present disclosure provide a method, apparatus, computer device, and storage medium for displaying scene images to solve the problem of difficulty in realizing interaction between special effects gameplay and actual scenes.

[0007] In a first aspect, an embodiment of the present disclosure provides a method for displaying a scene image, comprising: acquiring a scene image of a target scene and identifying a sound source position of a sound source object in the scene image; generating a unit image corresponding to the scene image and displaying the unit image; determining a special effects area in the unit image based on the sound source position; adding a disturbance effect to the image unit in the special effects area to display the image unit in the special effects area according to the disturbance effect.

[0008] In a second aspect, an embodiment of the present disclosure provides a device for displaying a scene image, comprising: an image processing unit for capturing a scene image of a target scene and identifying a sound source position of a sound source object in the scene image; an image generation unit for generating a unit image corresponding to the scene image and displaying the unit image; a special effect area determination unit for determining a special effect area in the unit image based on the sound source position; and a disturbance display unit for adding a disturbance effect to the image unit in the special effect area, so as to display the image unit in the special effect area according to the disturbance effect.

[0009] In a third aspect, an embodiment of the present disclosure provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions being stored in the memory, and the processor executing the scene image display method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0010] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the scene image display method of the above-mentioned first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0012] FIG1 is a schematic flow chart of a method for displaying a scene image according to some embodiments of the present disclosure;

[0013] FIG2 is a schematic flow chart of another method for displaying scene images according to some embodiments of the present disclosure;

[0014] FIG3 is a schematic diagram showing a scene image according to some embodiments of the present disclosure;

[0015] FIG4 is a schematic flow chart of another method for displaying scene images according to some embodiments of the present disclosure;

[0016] FIG5 is a structural block diagram of a device for displaying scene images according to an embodiment of the present disclosure;

[0017] FIG6 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0019] Most special effects gameplay in related technologies are based on visual information, but it's difficult to represent the interaction between special effects gameplay and the actual scene through visual information, which greatly limits the form of interaction between special effects gameplay and the actual scene. For example, the special effects sound in special effects gameplay can only be displayed based on visual information, but it is impossible to reflect the rhythm of the special effects sound in the special effects gameplay.

[0020] Based on this, the disclosed technical solution processes the data collected from the actual scene based on a fusion method to project the data from the actual scene into the virtual world of the special effects scene, so as to achieve a globally consistent spatial representation of the actual scene data and the special effects scene data, and selects mutually independent image units (such as particles) as the basic units presented in the virtual space, ensuring the subsequent special effects rendering effect. By locating the sound source, the sound source position is calculated to virtually project the sound source into the virtual world of the special effects scene, realizing the visualization of the spatial sound field of the real scene, allowing users to better perceive the real scene through the special effects scene.

[0021] According to an embodiment of the present disclosure, an embodiment of a method for displaying a scene image is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0022] In this embodiment, a method for displaying a scene image is provided, which can be used on a computer device, such as a mobile phone, a tablet computer, or a computer. FIG1 is a flow chart of the method for displaying a scene image according to an embodiment of the present disclosure. As shown in FIG1 , the flow chart includes the following steps:

[0023] Step S101: Capture a scene image of a target scene and identify the sound source position of a sound source object in the scene image.

[0024] The target scene is the visual scene displayed by the video stream data, that is, the real scene that can be displayed on the display page of the computer device. The scene image is an image generated by photographing the target scene. Specifically, the scene image can be captured by the rear camera set on the computer device; the scene image can also be captured by the camera and uploaded to the computer device. Accordingly, the computer device can access the image library of the computer device according to user instructions to obtain the scene image captured by the camera. The method for capturing the scene image is not specifically limited here.

[0025] A sound source object is an object in the target scene that can emit sound, and the sound source position represents the location of the sound source object within the target scene. Specifically, after capturing an image of the target scene, the various objects contained in the image can be identified to determine the sound source position of the sound source object within the image. Alternatively, various sensors provided in the computer device can be used to capture the direction of sound within the target scene. When capturing an image of the target scene through the computer device's camera, the sound direction can be combined with the sound source position to determine the sound source object within the image.

[0026] Step S102: Generate a unit image corresponding to the scene image and display the unit image.

[0027] A unit image consists of multiple image units, which simulate and visualize the movement and behavior of sound source objects in a scene image. Specifically, the unit image can be a particle-like image, consisting of a group of tiny dots or particles. The particles that make up the unit image are pixels of a specific color, size, or shape.

[0028] Specifically, a scene image is composed of a large number of pixels with specific colors, sizes or shapes. A particle system is deployed in a computer device. According to the pixels of the scene image, the particle system can be used to simulate and generate unit images in particle form for the scene image, and the generated unit images in particle form can be displayed on the display page of the computer device. For example, dynamic effects such as fire, smoke, and explosions in the scene image can be simulated by unit images in particle form.

[0029] Step S103: determining a special effect area in the unit image based on the sound source position.

[0030] The special effects area represents the region at the sound source location where special effects related to the sound source object are generated, such as fireworks and collision effects. Specifically, a unit image is generated for a scene image, and the special effects area corresponds to the sound source location. The sound source object's location within the unit image is determined based on its location within the scene image, and the special effects area specific to the sound source object is determined based on its location within the unit image.

[0031] Specifically, the special effect area is a range area encompassing the sound source location. This special effect area can be a circular area with a radius R centered at the sound source location, or a square area with a side length L centered at the sound source location. Of course, the special effect area can also be other range areas. R and L can be determined based on actual needs and are not specifically limited here.

[0032] Step S104 : adding a disturbance effect to the image unit in the special effect area, so as to display the image unit in the special effect area according to the disturbance effect.

[0033] In order to display the disturbance effect caused by the sound source object on the image unit in the special effect area, after determining the special effect area, the disturbance effect corresponding to the sound source object is added to the image unit located in the special effect area through the corresponding interface and function provided by the graphics engine or library deployed in the computer device, so as to make the image unit in the special effect area dynamic and realize dynamic display of the image unit in the special effect area.

[0034] Taking a particle-based unit image as an example, after determining the special effect area, a disturbance effect is added to the particles in the special effect area to achieve a dynamic effect. Specifically, when adding the disturbance effect to the particles in the special effect area, a random position perturbation method can be used to add a random offset to the original position of each particle in the special effect area, causing the particle position to slightly change, thereby giving the particles in the special effect area a vivid dynamic disturbance effect.

[0035] Specifically, when adding a disturbance effect to the particles in the special effect area, a random disturbance vector can be added to the velocity vector of each particle in the special effect area in a random velocity disturbance manner to make the motion trajectory of the particle irregular, presenting a floating or trembling disturbance effect.

[0036] Specifically, when adding a disturbance effect to the particles in the special effects area, random disturbances can also be added to the color or transparency of each particle in the special effects area. By fine-tuning the color value or transparency value of the particles, the particles in the special effects area can present rich and diverse color or transparency changes, creating a dazzling or flickering disturbance effect.

[0037] Specifically, when adding a disturbance effect to the particles in the special effects area, random disturbances can also be added to the size or shape of each particle in the special effects area. By fine-tuning the size value or shape parameters of the particles, the particles can show irregular changes, such as different sizes, abnormal shapes, etc., thereby visually presenting diverse disturbance effects and variable disturbance effects.

[0038] Of course, other methods can also be used to add a disturbance effect to the particles, which are not specifically limited here. By adding a disturbance effect to the particles in the special effect area, when the particles in the special effect area are displayed according to the disturbance effect, the particles can be made more vivid, natural, and rich, increasing visual interest and dynamics.

[0039] The scene image display method provided in this embodiment realizes the unitization of the scene image by converting the scene image into a unit image, and can identify the sound source position of the sound source object in the scene image to determine the special effect area corresponding to the sound source position in the unit image, thereby projecting the sound source in the target scene into the virtual world in the special effect scene, and then adding a disturbance effect to the image unit in the special effect area to affect the spatial sound field generated by the sound source object in the target scene to the virtual world in the special effect scene, thereby realizing the visualization of the sound field in the actual environment, ensuring that the image unit display effect of the special effect area is more vivid, allowing users to better perceive the real environment in the virtual world of the special effect scene, and improving the display effect of the special effects.

[0040] In this embodiment, a method for displaying a scene image is provided, which can be used on a computer device, such as a mobile phone, a tablet computer, or a computer. FIG2 is a flow chart of the method for displaying a scene image according to an embodiment of the present disclosure. As shown in FIG2 , the flow chart includes the following steps:

[0041] Step S201 : collecting a scene image of a target scene and identifying the sound source position of a sound source object in the scene image.

[0042] Specifically, the above step S201 may include:

[0043] Step S2011: Acquire a scene image of the target scene. Detailed descriptions refer to the corresponding descriptions of the above embodiments, which will not be repeated here.

[0044] Step S2012: Collect multiple channel data of the sound source object, and determine the spatial position of the sound source object in the target scene based on the multiple channel data.

[0045] Channel data represents the data generated by separating and storing the sound signals emitted by a sound source object across different channels. Specifically, when capturing an image of the target scene, the sound emitted by the sound source object in the target scene can be collected using multiple microphones installed in a computer device, as shown in Figure 3. To ensure the stereoscopic nature of the sound, the microphones can collect data from multiple channels of the sound source object, such as dual-channel data for left and right channels.

[0046] The strength of each channel can be determined based on the data of multiple channels. The distance of the sound source from the microphone can be estimated through the channel strength, and the spatial position of the sound source object in the target scene can be determined through the sphere intersection method.

[0047] Specifically, the sound intensity received by microphones A and B is calculated using the collected channel data. Sound intensity is typically expressed as the square of the sound signal's amplitude. Assuming the amplitude of channel 1 is aA, the amplitude of channel 2 is aB, the distance between the sound source and microphone A is dA, and the distance between the sound source and microphone B is dB. The law of sound intensity attenuation indicates that sound intensity is inversely proportional to distance: aA / aB = dB / dA. This provides the ratio of the distances between the sound source and microphones A and B.

[0048] The sound source object is located at a point where the sound intensity ratio satisfies aA / aB=dB / dA. Draw a sphere with microphone A as the center and distance dA as the radius. Draw another sphere with microphone B as the center and distance dB as the radius. The intersection of these two spheres is the spatial position coordinate point of the sound source object.

[0049] Step S2013: Acquire intrinsic and extrinsic parameters of the camera, and convert the spatial position into a pixel position in the scene image based on the intrinsic and extrinsic parameters.

[0050] The pixel position represents the sound source position of the sound source object in the scene image.

[0051] The internal and external parameters of a camera refer to the internal and external parameters of the camera. The internal parameters here refer to the internal parameters of the camera on the computer device, while the external parameters refer to the external parameters of the camera on the computer device. Specifically, the internal parameters are used to describe the internal properties of the camera, and may include parameters such as focal length, pixel size, and principal point position. The external parameters are used to describe the position and orientation of the camera in the real coordinate system, and may include parameters such as the rotation matrix R and the translation vector T.

[0052] The camera's intrinsic and extrinsic parameters are used to perform a perspective projection calculation on the spatial position of the sound source object to convert the spatial position to a pixel position in the scene image. Specifically, the spatial position coordinates of the sound source object are transformed into the camera coordinate system using the rotation matrix R and the translation vector T, obtaining the position coordinates of the sound source object in the camera coordinate system. The position coordinates in the camera coordinate system are then converted into image coordinates on the scene image plane through perspective projection. These image coordinates are used to map the spatial position to the pixel position in the scene image.

[0053] Step S202: Generate a unit image corresponding to the scene image and display the unit image. Detailed descriptions refer to the corresponding descriptions of the above embodiments, which will not be repeated here.

[0054] Step S203: Determine the special effect area in the unit image based on the sound source position. Detailed descriptions refer to the corresponding descriptions of the above embodiments, which will not be repeated here.

[0055] Step S204 , adding a disturbance effect to the image unit in the special effect area, so as to display the image unit in the special effect area according to the disturbance effect.

[0056] Specifically, the above step S204 may include: identifying the current sound source intensity of the sound source object, determining a disturbance effect that matches the sound source intensity, and adding the disturbance effect to the image unit in the special effect area.

[0057] The sound source intensity is the intensity of the sound emitted by the sound source object. The sound source intensity can be determined based on the amplitude of the sound signal. The larger the amplitude, the stronger the sound source intensity. By analyzing the amplitude of the sound signal currently emitted by the sound source object, the sound source intensity corresponding to the sound source object is determined. The corresponding disturbance intensity is determined in combination with the sound source intensity. Different disturbance intensities correspond to different disturbance effects. Thus, the disturbance effect that matches the sound source intensity can be determined based on the sound source intensity. Subsequently, the disturbance intensity corresponding to the sound source intensity is added to the image unit in the special effect area to generate a disturbance effect that matches the sound source intensity, thereby dynamically displaying the image unit in the special effect area according to the disturbance effect that matches the sound source intensity.

[0058] Step S205 , receiving an editing instruction for the special effect area, and adjusting configuration parameters of the special effect area in response to the editing instruction.

[0059] Configuration parameters are pre-set adjustable parameters for special effects, such as delay time, reverberation, offset, etc. Edit commands are adjustment commands for configuration parameters triggered by the user in the special effects area.

[0060] Specifically, the user can edit the configuration parameters in the special effects area to generate the special effects effect he or she desires. Accordingly, the computer device can receive the editing instructions triggered by the user in the special effects area and respond to the editing instructions, and adjust the current configuration parameters according to the editing instructions to meet the user's needs.

[0061] Step S206: re-displaying the image unit with the disturbance effect defined by the configuration parameters added according to the adjusted configuration parameters.

[0062] The disturbance effect of the image unit in the special effects area is re-determined through the adjusted configuration parameters, and the re-determined disturbance effect is added to each image unit in the special effects area. The image unit with the disturbance effect limited by the configuration parameters is re-rendered to display the image unit with the disturbance effect limited by the configuration parameters in the special effects area.

[0063] The scene image display method provided in this embodiment uses the collected multiple sound channel data to locate the spatial position of the sound source object in the target scene, and determines the sound source position of the sound source object in the scene image based on the spatial position, so as to facilitate the virtual projection of the sound source object. The corresponding disturbance effect is added in combination with the sound source intensity emitted by the sound source object, so that the disturbance effect of the image unit displayed in the special effect area is more consistent with the real scene, ensuring that the user can better perceive the real world through the virtual world in the special effect scene. At the same time, it supports the adjustment of the configuration parameters of the special effect area, which facilitates the flexible adjustment of the disturbance effect acting on the image unit, so that the disturbance effect of the displayed image unit is more consistent with the user's expectations, thereby improving the user experience of the special effect gameplay.

[0064] In this embodiment, a method for displaying a scene image is provided, which can be used in computer devices such as mobile phones, tablet computers, and the like. FIG4 is a flow chart of the method for displaying a scene image according to an embodiment of the present disclosure. As shown in FIG4 , the flow chart includes the following steps:

[0065] Step S301: Capture a scene image of a target scene and identify the sound source position of a sound source object in the scene image. Detailed descriptions refer to the corresponding descriptions of the above embodiments, which will not be repeated here.

[0066] Step S302: Generate a unit image corresponding to the scene image and display the unit image.

[0067] Specifically, the above step S302 may include:

[0068] Step S3021: Back-project the pixel points in the scene image into spatial points in a specified coordinate system.

[0069] A scene image is composed of a large number of pixels, and its scene image has its corresponding pixel coordinate system. Each pixel point has a corresponding pixel position coordinate in the scene image's pixel coordinate system. The designated coordinate system is a pre-specified spatial coordinate system (such as the spatial coordinate system of the simultaneous localization and mapping system (SLAM) shown in Figure 3), and each position point in the designated coordinate system has a corresponding spatial position coordinate.

[0070] Combined with the conversion relationship between the pixel coordinate system and the specified coordinate system, the pixel coordinate system of the scene image is converted to the specified coordinate system, so that the spatial point corresponding to each pixel point in the scene image in the specified coordinate system can be obtained.

[0071] In some optional embodiments, the above step S3021 may include:

[0072] Step a1: obtain the current pose information of the camera and generate a depth image of the scene image in a specified coordinate system.

[0073] Step a2: Back-project the pixel points in the scene image into spatial points in a specified coordinate system based on the current pose information and the depth image.

[0074] The current pose information of the camera represents the position and orientation of the camera in three-dimensional space. Here, the current pose information represents the position and orientation of the camera of the computer device when capturing the target scene. Specifically, the current pose information of the camera can be represented by a rotation matrix, etc.

[0075] A depth image records the distance between each pixel in a scene image and the camera. It measures the distance between each object in the target scene and the camera, i.e., the distance from each pixel to the camera. Specifically, a depth image can be represented as a grayscale image, where the grayscale value of each pixel represents the distance from the camera. Brighter pixels in the depth image indicate that the pixel is closer to the camera, while darker pixels indicate that the pixel is farther away.

[0076] Combining the camera's current pose information with the target scene's 3D point cloud, the coordinates of each 3D point in the target scene in the camera coordinate system are calculated. The 3D point cloud is then projected onto the image plane to convert the coordinates of each 3D point in the camera coordinate system into image coordinates. The depth value of each pixel is then calculated based on the pose information and the projected image coordinates. By filling the calculated depth value of each pixel at the corresponding image location, a depth map corresponding to the scene image is generated.

[0077] The coordinate system of the depth map is converted to the specified coordinate system, and a depth image corresponding to the scene image is generated in the specified coordinate system. Then, using the current camera pose information and the depth image in the specified coordinate system, each pixel in the scene image is back-projected to the specified coordinate system to obtain the spatial point corresponding to each pixel in the specified coordinate system.

[0078] In the above embodiment, the pixel points in the scene image are back-projected into spatial points in the specified coordinate system in combination with the current posture information of the camera and the depth image generated in the specified coordinate system. This enables global consistency of each pixel point in the three-dimensional representation, thereby ensuring the visualization of the sound field generated by the sound source object.

[0079] In some optional embodiments, the above-mentioned step of generating a depth image of the scene image in a specified coordinate system may include: generating a relative depth map of the scene image through a depth estimation network, and aligning the depth information in the relative depth map to the specified coordinate system to obtain a depth image of the scene image in the specified coordinate system.

[0080] A depth estimation network is a pre-trained network model used to generate depth maps, such as a monocular depth estimation network. A relative depth map is an image that represents the depth information of objects in a scene image. A relative depth map uses different grayscale values ​​or colors to represent the relative depth relationships of different objects in the scene image.

[0081] Specifically, the depth estimation network is deployed in a computer device, and the computer device inputs the scene image of the target scene it collects into the depth estimation network, so that the depth estimation network outputs a relative depth map corresponding thereto, as shown in FIG3 .

[0082] Align the depth information in the relative depth map to the specified coordinate system to ensure the global consistency of the depth of each pixel. Specifically, obtain the camera parameters (intrinsic and extrinsic parameters) for generating the relative depth map, and select the corresponding reference point in the specified coordinate system; calculate the coordinate transformation matrix from the camera coordinate system to the specified coordinate system based on the camera parameters and the reference point; for each depth value in the relative depth map, convert its corresponding pixel coordinate point to a point in the camera coordinate system, and apply the coordinate transformation matrix to convert its point in the camera coordinate system to the specified coordinate system to obtain the position of each depth value in the specified coordinate system, thereby obtaining the depth image of the scene image in the specified coordinate system, as shown in Figure 3.

[0083] In the above embodiment, a relative depth map of the scene image is generated through a depth estimation network, and the depth information in the relative depth map is aligned to a specified coordinate system to obtain a depth image of the scene image in the specified coordinate system, so that the depth information has global consistency, thereby ensuring the global consistency of each pixel point in the three-dimensional representation.

[0084] Step S3022: Map the spatial points into the signed distance field to form an object surface in the signed distance field.

[0085] A signed distance field (SDF) is a data structure that represents the shape of an object in space. It distinguishes points inside an object from those outside it and records the distance from each point to the object's surface. A signed distance field can be represented as a two-dimensional or three-dimensional grid, with each voxel in the field containing a signed distance value.

[0086] For the spatial points that need to be mapped to the signed distance field, the spatial coordinates corresponding to the spatial points are converted into grid coordinates in the signed distance field, and the object surface of the sound source object in the signed distance field is formed according to the determined grid coordinates.

[0087] Step S3023 : extracting voxels located on the surface of the object from the signed distance field, and rendering the extracted voxels as image units in a unit image to generate a unit image corresponding to the scene image.

[0088] Based on the distance values ​​contained in each voxel in the signed distance field, the voxels whose signed distance values ​​represent the object surface are extracted. In other words, the voxels located on the object surface are extracted. Subsequently, the image units to be rendered are determined based on the extracted voxels located on the object surface, and the position information of each voxel located on the object surface is recorded. This information is then used for visualization rendering, forming corresponding unit images to display the shape information of the object surface.

[0089] Taking particles as image units as an example, the particles to be rendered are determined based on the extracted voxels located on the object's surface. Particle properties such as position, size, and color are also determined for each particle, along with the particle's current state in the target scene. As shown in Figure 3, in each rendered frame, particles are rendered using graphic elements such as points, lines, textures, or custom shapes, based on their properties and the particle's current state in the target scene, generating a unit image of the particles.

[0090] In some optional embodiments, the above step S3023 may include:

[0091] Step b1: Calculate the signed distance value of each voxel in the signed distance field, where the signed distance value represents the distance information between the voxel and the surface of the object.

[0092] Step b2: determining and extracting target voxels whose signed distance values ​​are specified values.

[0093] Step b3: Rendering the extracted target voxels as image units in a unit image.

[0094] A signed distance value represents the distance between each voxel in a signed distance field, representing the distance between the voxel and the surface of the object. For each voxel in the signed distance field, a matching signed distance value can be calculated using algorithms such as the flooding algorithm and the signed distance field.

[0095] The specified value is a pre-set distance value used to represent the surface of an object, for example, set to 0. The target voxel is one or more voxels located on the surface of the object. Specifically, each voxel in the signed distance field is traversed sequentially, and the signed distance value corresponding to each voxel is tested to see if it is equal to the specified value. Based on the test results, the voxels with the signed distance value equal to the specified value are extracted, and the extracted target voxels are rendered as image units in the unit image.

[0096] In the above embodiment, the target voxel located on the surface of the object is determined by the signed distance value of each voxel in the signed distance field, and the target voxel is rendered as an image unit in the unit image. In this way, the rendering effect of the image unit can be guaranteed to be consistent with the shape of the object surface, thereby improving the rendering effect of the image unit.

[0097] Step S303: Determine the special effect area in the unit image based on the sound source position. Detailed descriptions refer to the corresponding descriptions of the above embodiments, which will not be repeated here.

[0098] Step S304: Add a disturbance effect to the image unit in the special effect area to display the image unit in the special effect area according to the disturbance effect. Detailed descriptions refer to the corresponding descriptions of the above embodiments, which will not be repeated here.

[0099] The scene image display method provided in this embodiment back-projects the pixel points in the scene image into spatial points in a specified coordinate system, determines the voxels located on the surface of the object based on the signed distance field, and renders the extracted voxels as image units in the unit image, thereby applying a disturbance effect to the image units in the special effects area based on the sound source position and the rendered unit image, thereby converting the scene image captured by the camera into a unitized effect, facilitating the interaction between the sound source and the actual scene in the form of a unitized scene, thereby realizing the visualization of the spatial sound field.

[0100] In this embodiment, a device for displaying scene images is also provided. The device is used to implement the above-mentioned embodiments and preferred embodiments. The details that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0101] This embodiment provides a scene image display device, as shown in FIG5 , including:

[0102] The image processing unit 401 is configured to capture a scene image of a target scene and identify a sound source position of a sound source object in the scene image.

[0103] The image generation unit 402 is used to generate a unit image corresponding to the scene image and display the unit image.

[0104] The special effect area determination unit 403 is configured to determine a special effect area in the unit image based on the sound source position.

[0105] The disturbance display unit 404 is configured to add a disturbance effect to the image unit in the special effect area so as to display the image unit in the special effect area according to the disturbance effect.

[0106] In some optional embodiments, the image processing unit 401 may include:

[0107] The sound source position determination subunit is used to collect multiple channel data of the sound source object and determine the spatial position of the sound source object in the target scene based on the multiple channel data.

[0108] The position conversion subunit is used to obtain the intrinsic and extrinsic parameters of the camera and, based on the extrinsic and extrinsic parameters, convert the spatial position into a pixel position in the scene image.

[0109] In some optional embodiments, the disturbance display unit 404 may include:

[0110] The identification subunit is used to identify the current sound source intensity emitted by the sound source object, determine the disturbance effect that matches the sound source intensity, and add the disturbance effect to the image unit in the special effect area.

[0111] In some optional embodiments, the above device may further include:

[0112] The editing unit is used to receive an editing instruction for the special effect area, and adjust the configuration parameters of the special effect area in response to the editing instruction.

[0113] The disturbance updating unit is used to re-display the image unit with the disturbance effect defined by the configuration parameters added according to the adjusted configuration parameters.

[0114] In some optional embodiments, the image generation unit 402 may include:

[0115] The projection subunit is used to back-project the pixel points in the scene image into spatial points in a specified coordinate system.

[0116] The mapping subunit is configured to map a spatial point into a signed distance field to form an object surface in the signed distance field.

[0117] The extraction subunit is used to extract voxels located on the surface of the object in the signed distance field, and render the extracted voxels as image units in the unit image to generate a unit image corresponding to the scene image.

[0118] In some optional embodiments, the above-mentioned projection subunit is specifically used to: obtain the current posture information of the camera and generate a depth image of the scene image in a specified coordinate system; based on the current posture information and the depth image, back-project the pixel points in the scene image into spatial points in the specified coordinate system.

[0119] In some optional embodiments, the projection subunit is further configured to generate a relative depth map of the scene image through a depth estimation network, and align the depth information in the relative depth map to a specified coordinate system to obtain a depth image of the scene image in the specified coordinate system.

[0120] In some optional embodiments, the extraction subunit is specifically configured to: calculate a signed distance value of each voxel in the signed distance field, where the signed distance value represents distance information between the voxel and the object surface; and determine and extract a target voxel whose signed distance value is a specified value.

[0121] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0122] The scene image display device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0123] The scene image display device provided in this embodiment realizes the unitization of the scene image by converting the scene image into a unit image, and can identify the sound source position of the sound source object in the scene image to determine the special effect area corresponding to the sound source position in the unit image, thereby projecting the sound source in the target scene into the virtual world in the special effect scene, and then adding a disturbance effect to the image unit in the special effect area to affect the spatial sound field generated by the sound source object in the target scene to the virtual world in the special effect scene, thereby realizing the visualization of the sound field in the actual environment, ensuring that the display effect of the image unit in the special effect area is more vivid, allowing users to better perceive the real environment through the special effect scene, and improving the display effect of the special effect.

[0124] The embodiment of the present disclosure further provides a computer device having a display device for the scene image shown in FIG5 .

[0125] Please refer to Figure 6, which is a structural diagram of a computer device provided by an optional embodiment of the present disclosure. As shown in Figure 6, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 takes a processor 10 as an example.

[0126] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0127] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0128] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0129] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0130] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means, and FIG6 takes the bus connection as an example.

[0131] The input device 30 can receive input digital or character information and generate key signal input related to user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0132] The computer device further includes a communication interface for the computer device to communicate with other devices or a communication network.

[0133] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0134] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for displaying a scene image, the method comprising: Acquire a scene image of a target scene, and identify a sound source position of a sound source object in the scene image; Generate a unit image corresponding to the scene image, and display the unit image; Based on the sound source position, determining a special effect area in the unit image; A disturbance effect is added to the image unit in the special effect area to display the image unit in the special effect area according to the disturbance effect.

2. The method according to claim 1, wherein adding a disturbance effect to the image unit in the special effect area comprises: The intensity of the sound source currently emitted by the sound source object is identified, a disturbance effect matching the intensity of the sound source is determined, and the disturbance effect is added to the image unit in the special effect area.

3. The method according to claim 1 or 2, wherein after displaying the image unit in the special effect area according to the disturbance effect, the method further comprises: receiving an editing instruction for the special effect area, and adjusting configuration parameters of the special effect area in response to the editing instruction; According to the adjusted configuration parameters, the image unit with the disturbance effect defined by the configuration parameters added is displayed again.

4. The method according to claim 1, wherein the step of identifying the sound source position of the sound source object in the scene image comprises: Collecting multiple sound channel data of the sound source object, and determining the spatial position of the sound source object in the target scene according to the multiple sound channel data; Acquire intrinsic and extrinsic parameters of a camera, and based on the extrinsic and extrinsic parameters, convert the spatial position into a pixel position in the scene image; wherein the pixel position represents the sound source position of the sound source object in the scene image.

5. The method according to claim 1, wherein generating a unit image corresponding to the scene image comprises: Back-projecting the pixel points in the scene image into spatial points in a specified coordinate system; Mapping the spatial points into a signed distance field to form an object surface in the signed distance field; Voxels located on the surface of the object are extracted from the signed distance field, and the extracted voxels are rendered as image units in a unit image to generate a unit image corresponding to the scene image.

6. The method according to claim 5, wherein the back-projecting of the pixel points in the scene image into spatial points in a specified coordinate system comprises: Obtaining the current position information of the camera and generating a depth image of the scene image in a specified coordinate system; According to the current pose information and the depth image, pixel points in the scene image are back-projected into spatial points in a specified coordinate system.

7. The method according to claim 6, wherein generating a depth image of the scene image in a specified coordinate system comprises: A relative depth map of the scene image is generated through a depth estimation network, and the depth information in the relative depth map is aligned to a specified coordinate system to obtain a depth image of the scene image in the specified coordinate system.

8. The method according to claim 5, wherein extracting voxels located on the surface of the object in the signed distance field comprises: Calculating a signed distance value of each voxel in the signed distance field, wherein the signed distance value represents distance information between the voxel and the surface of the object; Identify and extract the target voxel whose signed distance value is the specified value.

9. A device for displaying a scene image, the device comprising: An image processing unit, used to collect a scene image of a target scene and identify a sound source position of a sound source object in the scene image; An image generating unit, used to generate a unit image corresponding to the scene image and display the unit image; A special effect area determination unit, configured to determine a special effect area in the unit image based on the sound source position; The disturbance display unit is used to add a disturbance effect to the image unit in the special effect area so as to display the image unit in the special effect area according to the disturbance effect.

10. A computer device comprising: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the scene image display method according to any one of claims 1 to 8 by executing the computer instructions.

11. A computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the scene image display method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Dynamic three-dimensional digital scene collection and construction system of highway transportation and working method thereof

    CN109547769A

  • Grid model reconstruction method, system and device based on two-dimensional image and medium

    CN115731365A

  • Image display method and device, equipment, storage medium and program product

    CN116304180A

  • Medical service robot

    KR1020250014579A