Scene image display method and device, computer equipment and storage medium

By identifying and processing the sound source position in the scene image, converting it into a unit image and adding disturbance effects, the limitations of the interaction between the special effect gameplay and the actual scene are solved, and the sound field visualization and special effect display effect are improved.

CN120070691APending Publication Date: 2025-05-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311617072.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

It is difficult for the prior art to achieve effective interaction between special effects gameplay and actual scenes, especially when there is insufficient visual information.

Method used

By collecting the scene image of the target scene, the sound source location of the sound source object is identified and converted into a unit image. Determine the special effect area in the cell image and add a perturbation effect to the area to display the image units in the special effect area.

Benefits of technology

It realizes the sound field visualization in the actual environment, improves the display effect of special effects, and allows users to better perceive the real environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070691A_ABST
    Figure CN120070691A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a scene image display method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a scene image of a target scene, and identifying a sound source position of a sound source object in the scene image; generating a unit image corresponding to the scene image, and displaying the unit image; determining a special effect area in the unit image based on the sound source position; and adding a disturbance effect to the image unit in the special effect area so as to display the image unit in the special effect area according to the disturbance effect. According to the technical scheme, the space sound field generated by the sound source object in the target scene can act on the virtual world in the special effect scene, and sound field visualization in the actual environment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, a computer device, and a storage medium for displaying a scene image. Background Art

[0002] Currently, special effect props generated based on the scenes captured by the rear cameras of computer devices such as mobile phones and tablets provide users with a variety of gameplay, such as special effect gameplay like placement, collision, and deformation. However, most special effect gameplay is based on visual information, making the interaction form between the special effect gameplay and the actual scene relatively limited and difficult to achieve good interaction between the special effect gameplay and the actual scene. Summary of the Invention

[0003] In view of this, embodiments of the present disclosure provide a method, an apparatus, a computer device, and a storage medium for displaying a scene image to solve the problem that it is difficult to achieve interaction between special effect gameplay and the actual scene.

[0004] In a first aspect, embodiments of the present disclosure provide a method for displaying a scene image, including: collecting a scene image of a target scene and identifying the sound source position of a sound source object in the scene image; generating a unit image corresponding to the scene image and displaying the unit image; determining a special effect area in the unit image based on the sound source position; and adding a perturbation effect to the image units in the special effect area to display the image units in the special effect area according to the perturbation effect.

[0005] The method for displaying a scene image provided by the embodiments of the present disclosure realizes the unitization of the scene image by converting the scene image into a unit image, and can identify the sound source position where the sound source object in the scene image is located to determine a special effect area corresponding to the sound source position in the unit image. Thus, the sound source in the target scene can be projected into the virtual world in the special effect scene, and then a perturbation effect is added to the image units in the special effect area to apply the spatial sound field generated by the sound source object in the target scene to the virtual world in the special effect scene, realizing the visualization of the sound field in the actual environment, ensuring that the display effect of the image units in the special effect area is more vivid, enabling the user to better perceive the real environment through the virtual world in the special effect scene, and improving the display effect of the special effects.

[0006] In a second aspect, embodiments of the present disclosure provide a device for displaying a scene image, including: an image processing unit for collecting a scene image of a target scene and identifying the sound source position of a sound source object in the scene image; an image generation unit for generating a unit image corresponding to the scene image and displaying the unit image; a special effect area determination unit for determining a special effect area in the unit image based on the sound source position; and a perturbation display unit for adding a perturbation effect to the image units in the special effect area to display the image units in the special effect area according to the perturbation effect.

[0007] In a third aspect, an embodiment of the present disclosure provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method for displaying a scene image according to the first aspect or any corresponding embodiment thereof.

[0008] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method for displaying a scene image according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 is a flowchart of a method for displaying a scene image according to some embodiments of the present disclosure;

[0011] Figure 2 is a flowchart of a method for displaying another scene image according to some embodiments of the present disclosure;

[0012] Figure 3 is a schematic diagram of displaying a scene image according to some embodiments of the present disclosure;

[0013] Figure 4 is a flowchart of a method for displaying yet another scene image according to some embodiments of the present disclosure;

[0014] Figure 5 is a block diagram of the structure of a device for displaying a scene image according to an embodiment of the present disclosure;

[0015] Figure 6 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0017] Most of the special effect gameplay in the related art is based on visual information, and it is difficult to represent the interaction between the special effect gameplay and the actual scene through visual information, resulting in great limitations in the interaction form between the special effect gameplay and the actual scene. Taking the special effect sound in the special effect gameplay as an example, the current special effect gameplay can only display the special effect sound in the special effect gameplay based on visual information, but it cannot reflect the rhythm of the special effect sound in the special effect gameplay.

[0018] Based on this, the technical solution of the present disclosure processes the data collected in the actual scene based on a fusion method to project the data in the actual scene into the virtual world in the special effect scene, so as to achieve a globally consistent spatial representation of the actual scene data and the special effect scene data, and selects independent image units (such as particles) as the basic units presented in the virtual space, ensuring the subsequent special effect rendering effect. By means of sound source localization, the sound source position is calculated to virtually project the sound source into the virtual world in the special effect scene, realizing the visualization of the spatial sound field of the real scene and enabling users to better perceive the real scene through the special effect scene.

[0019] According to an embodiment of the present disclosure, an embodiment of a method for displaying a scene image is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0020] In this embodiment, a method for displaying a scene image is provided, which can be used in computer devices such as mobile phones, tablet computers, and computers. Figure 1 is a flowchart of the method for displaying a scene image according to an embodiment of the present disclosure, as Figure 1 shown, and the process includes the following steps:

[0021] Step S101, collect a scene image of a target scene and identify the sound source position of the sound source object in the scene image.

[0022] The target scenario is the visualization scenario presented by video stream data, that is, the real scenario that can be presented on the display page of a computer device. The scenario image is an image generated by photographing the target scenario. Specifically, the scenario image can be obtained by photographing with a rear camera set on the computer device; the scenario image can also be collected by a camera and uploaded to the computer device. Correspondingly, the computer device can access the picture library of the computer device according to a user instruction to obtain the scenario image collected by the camera. The specific method of collecting the scenario image is not limited here.

[0023] The sound source object is an object that can emit sound in the target scenario, and the sound source position represents the position where the sound source object is located in the target scenario. Specifically, after the scenario image of the target scenario is collected, each object included in the scenario image can be recognized, and the sound source position of the sound source object in the scenario image can be determined; the sound direction in the target scenario can also be collected through various sensors set in the computer device, and during the process of collecting the scenario image of the target scenario through the camera of the computer device, the sound source position of the sound source object in the scenario image can be determined by combining the sound direction.

[0024] Step S102: Generate a unit image corresponding to the scenario image and display the unit image.

[0025] The unit image includes a plurality of image units, and the movement and behavior of the sound source object in the scenario image are simulated and visualized through the unit image. Specifically, the unit image can be an image in the form of particles, that is, the unit image can be an image composed of a group of tiny points or particles, and the particles forming the unit image are pixel points with specific colors, sizes or shapes.

[0026] Specifically, the scenario image is composed of a large number of pixel points with specific colors, sizes or shapes. A particle system is deployed in the computer device. According to the pixel points of the scenario image, the particle system can simulate and generate a unit image in the form of particles for the scenario image, and display the generated unit image in the form of particles on the display page of the computer device. For example, dynamic effects such as flames, smoke, and explosions in the scenario image are simulated through the unit image in the form of particles.

[0027] Step S103: Determine a special effect area in the unit image based on the sound source position.

[0028] The special effect area represents the area where special effects related to the sound source object are generated at the sound source position, such as fireworks special effects, collision special effects, etc. Specifically, the unit image is generated for the scenario image, and the special effect area corresponds to the sound source position. Therefore, the position of the sound source object in the unit image is determined according to the sound source position where the sound source object is located in the scenario image, and the special effect area for the sound source object is determined by combining the position of the sound source object in the unit image.

[0029] Specifically, the special effect area is a range area that includes the sound source position. The special effect area can be a circular area centered on the sound source position with a radius of R; it can also be a square area centered on the sound source position with a side length of L. Of course, the special effect area can also be other forms of range areas. Among them, R and L can be determined according to actual needs and are not specifically limited here.

[0030] Step S104, add a perturbation effect to the image units in the special effect area to display the image units in the special effect area according to the perturbation effect.

[0031] In order to display the perturbation effect generated by the sound source object on the image units in the special effect area, after determining the special effect area, through the corresponding interfaces and functions provided by the graphics engine or library deployed in the computer device, add a perturbation effect corresponding to the sound source object to the image units located in the special effect area to dynamicize the image units in the special effect area and achieve a dynamic display of the image units in the special effect area.

[0032] Taking the unit image in the form of particles as an example, after determining the special effect area, add a perturbation effect to the particles in the special effect area to make the particles in the special effect area achieve a dynamic effect. Specifically, when adding a perturbation effect to the particles in the special effect area, a random position perturbation method can be adopted to add a random offset to the original position of each particle in the special effect area to slightly change the position of the particle, so that the particles in the special effect area present a vivid dynamic perturbation effect.

[0033] Specifically, when adding a perturbation effect to the particles in the special effect area, a random velocity perturbation method can also be adopted to add a random perturbation vector to the velocity vector of each particle in the special effect area to make the movement trajectory of the particle irregular and present a floating or trembling perturbation effect.

[0034] Specifically, when adding a perturbation effect to the particles in the special effect area, random perturbations can also be added to the color or transparency of each particle in the special effect area. Thus, by fine-tuning the color value or transparency value of the particle, the particles in the special effect area can present rich and diverse color or transparency changes, creating a dazzling or flickering perturbation effect.

[0035] Specifically, when adding a perturbation effect to the particles in the special effect area, random perturbations can also be added to the size or shape of each particle in the special effect area. Thus, by fine-tuning the size value or shape parameter of the particle, the particle can present irregular changes, such as different sizes and abnormal shapes, etc., so as to visually present a diverse perturbation effect and a variable perturbation effect.

[0036] Of course, the perturbation effect of the particles can also be added in other ways, which is not specifically limited here. By adding the perturbation effect to the particles in the special effect area, when the particles in the special effect area are displayed according to the perturbation effect, the particles can be made more vivid, natural and rich, increasing the visual interest and dynamic sense.

[0037] The method for displaying a scene image provided in this embodiment realizes the unitization of the scene image by converting the scene image into a unit image, and can identify the sound source position where the sound source object in the scene image is located, so as to determine a special effect area corresponding to the sound source position in the unit image. Thus, the sound source in the target scene can be projected into the virtual world in the special effect scene, and then a perturbation effect is added to the image unit in the special effect area to apply the spatial sound field generated by the sound source object in the target scene to the virtual world in the special effect scene, realizing the visualization of the sound field in the actual environment, ensuring that the display effect of the image unit in the special effect area is more vivid, enabling the user to better perceive the real environment in the virtual world of the special effect scene, and improving the display effect of the special effect.

[0038] In this embodiment, a method for displaying a scene image is provided, which can be used in computer devices such as mobile phones, tablets, computers, etc. Figure 2 is a flowchart of the method for displaying a scene image according to an embodiment of the present disclosure, as Figure 2 shown, and this process includes the following steps:

[0039] Step S201, collect a scene image of the target scene and identify the sound source position of the sound source object in the scene image.

[0040] Specifically, the above step S201 may include:

[0041] Step S2011, collect a scene image of the target scene. For a detailed description, refer to the relevant description corresponding to the above embodiment, which will not be elaborated here.

[0042] Step S2012, collect multiple channel data of the sound source object, and determine the spatial position of the sound source object in the target scene according to the multiple channel data.

[0043] The channel data represents the data in which the sound signals emitted by the sound source object are separated and stored in different channels. Specifically, when the sound source object emits sound in the target scene, when collecting the scene image of the target scene, the sound emitted by the sound source object in the target scene can be collected through multiple microphone devices set in the computer device, as Figure 3 shown. To ensure the stereophonic effect of the sound, multiple channel data of the sound source object can be collected through the microphone device here, such as dual-channel data of the left channel and the right channel.

[0044] Based on the multi-channel data, the intensity of each channel can be determined. Through the channel intensity, the distance between the sound source position and the microphone can be estimated, and the spatial position of the sound source object in the target scene can be determined by the method of sphere intersection.

[0045] Specifically, the sound intensities received by microphone A and microphone B are calculated from the collected channel data. Usually, the square of the amplitude of the sound signal can be used to represent the sound intensity. Assume that the amplitude of channel 1 is aA, the amplitude of channel 2 is aB, the distance between the sound source object and microphone A is dA, and the distance between the sound source object and microphone B is dB. According to the law of sound intensity attenuation, the sound intensity is inversely proportional to the distance, that is, aA / aB = dB / dA. Thus, the distance ratio between the sound source object and microphones A and B can be obtained.

[0046] The sound source object is located at the point where the sound intensity ratio satisfies aA / aB = dB / dA. Draw a sphere with microphone A as the center and distance dA as the radius, and draw a sphere with microphone B as the center and distance dB as the radius. The intersection point of these two spheres is the spatial position coordinate point of the sound source object.

[0047] Step S2013: Obtain the internal and external parameters of the camera, and based on the internal and external parameters, convert the spatial position into the pixel position in the scene image.

[0048] Among them, the pixel position represents the sound source position of the sound source object in the scene image.

[0049] The internal and external parameters of the camera represent the internal parameters and external parameters of the camera. Here, the internal parameters represent the internal parameters of the camera lens on the computer device, and the external parameters represent the external parameters of the camera lens on the computer device. Specifically, the internal parameters are used to describe the internal attributes of the camera lens, which may include parameters such as focal length, pixel size, and principal point position. The external parameters are used to describe the position and orientation of the camera lens in the actual coordinate system, which may include parameters such as rotation matrix R and translation vector T.

[0050] Perspective projection calculation is performed on the spatial position of the sound source object using the internal and external parameters of the camera to convert the spatial position to the pixel position in the scene image. Specifically, the spatial position coordinates of the sound source object are converted to the camera coordinate system through the rotation matrix R and translation vector T to obtain the position coordinates of the sound source object in the camera coordinate system; the position coordinates in the camera coordinate system are converted to the image coordinates on the scene image plane through perspective projection, and this image coordinate is the pixel position that maps the spatial position to the scene image.

[0051] Step S202: Generate a unit image corresponding to the scene image and display the unit image. For the detailed description, please refer to the relevant description in the above embodiments, which will not be elaborated here.

[0052] Step S203: Determine the special effect area in the unit image based on the sound source position. For the detailed description, please refer to the relevant description corresponding to the above embodiments, which will not be elaborated here.

[0053] Step S204: Add a perturbation effect to the image units in the special effect area to display the image units in the special effect area according to the perturbation effect.

[0054] Specifically, the above step S204 may include: identifying the sound source intensity currently emitted by the sound source object, determining the perturbation effect matching the sound source intensity, and adding the perturbation effect to the image units in the special effect area.

[0055] The sound source intensity is the intensity of the sound emitted by the sound source object. This sound source intensity can be determined according to the amplitude of the sound signal. The larger the amplitude, the stronger the sound source intensity. By analyzing the amplitude of the sound signal currently emitted by the sound source object, the corresponding sound source intensity of the sound source object is determined. Combining the sound source intensity to determine the corresponding perturbation intensity, different perturbation intensities correspond to different perturbation effects. Thus, the perturbation effect matching the sound source intensity can be determined according to the sound source intensity. Subsequently, the perturbation intensity corresponding to the sound source intensity is added to the image units in the special effect area to generate a perturbation effect matching the sound source intensity, so as to dynamically display the image units in the special effect area according to the perturbation effect matching the sound source intensity.

[0056] Step S205: Receive an editing instruction for the special effect area, and in response to the editing instruction, adjust the configuration parameters of the special effect area.

[0057] The configuration parameters are adjustable parameters for the special effects preset to achieve various special effect effects, such as delay time, reverberation, offset, etc. The editing instruction is an adjustment instruction triggered by the user for the configuration parameters in the special effect area.

[0058] Specifically, the user can edit the configuration parameters in the special effect area to generate the special effect effect they expect. Correspondingly, the computer device can receive the editing instruction triggered by the user in the special effect area and can respond to the editing instruction to adjust the current configuration parameters to meet the user's needs.

[0059] Step S206: Redisplay the image units with the perturbation effect limited by the configuration parameters according to the adjusted configuration parameters.

[0060] Redetermine the perturbation effect of the image units in the special effect area according to the adjusted configuration parameters, and add the redetermined perturbation effect to each image unit in the special effect area, and re-render the image units with the perturbation effect limited by the configuration parameters to display the image units with the perturbation effect limited by the configuration parameters in the special effect area.

[0061] The method for displaying a scene image provided in this embodiment locates the spatial position of a sound source object in a target scene through multiple channel data collected, determines the sound source position of the sound source object in the scene image according to the spatial position, and facilitates the virtual projection of the sound source object. Combining the sound source intensity emitted by the sound source object to add a corresponding perturbation effect makes the perturbation effect of the image units displayed in the special effect area more in line with the real scene, ensuring that users can better perceive the real world through the virtual world in the special effect scene. At the same time, it supports adjusting the configuration parameters for the special effect area, facilitating flexible adjustment of the perturbation effect acting on the image units, making the perturbation effect of the displayed image units more in line with the user's expectations, and enhancing the user experience of the special effect gameplay.

[0062] In this embodiment, a method for displaying a scene image is provided, which can be used in computer devices such as mobile phones, tablets, computers, etc. Figure 4 It is a flowchart of the method for displaying a scene image according to an embodiment of the present disclosure, as Figure 4 shown, and this process includes the following steps:

[0063] Step S301, collect the scene image of the target scene and identify the sound source position of the sound source object in the scene image. For detailed description, refer to the relevant description corresponding to the above embodiment, and details will not be repeated here.

[0064] Step S302, generate a unit image corresponding to the scene image and display the unit image.

[0065] Specifically, the above step S302 may include:

[0066] Step S3021, back-project the pixel points in the scene image into spatial points in a specified coordinate system.

[0067] The scene image is composed of a large number of pixel points, and the scene image has its corresponding pixel coordinate system. Each pixel point has a corresponding pixel position coordinate in the pixel coordinate system of the scene image. The specified coordinate system is a pre-specified spatial coordinate system (such as Figure 3 the spatial coordinate system of the Simultaneous Localization and Mapping (SLAM) shown), and each position point in the specified coordinate system has a corresponding spatial position coordinate.

[0068] Combining the conversion relationship between the pixel coordinate system and the specified coordinate system, the pixel coordinate system of the scene image is converted to the specified coordinate system, so that the spatial points corresponding to each pixel point in the scene image in the specified coordinate system can be obtained.

[0069] In some optional embodiment modes, the above step S3021 may include:

[0070] Step a1, obtain the current pose information of the camera, and generate a depth image of the scene image in a specified coordinate system.

[0071] Step a2, according to the current pose information and the depth image, back-project the pixel points in the scene image into spatial points in the specified coordinate system.

[0072] The current pose information of the camera represents the position and orientation of the camera in three-dimensional space. Here, the current pose information represents the position and orientation of the camera lens of the computer device when shooting the target scene. Specifically, the current pose information of the camera can be represented by a rotation matrix, etc.

[0073] The depth image is an image that records the distance information between each pixel point in the scene image and the camera. By using the depth image, the distances of various object objects in the target scene from the camera can be measured, that is, the distance from each pixel point to the camera. Specifically, the depth image can be represented by a grayscale image. The grayscale value of each pixel point represents the distance from the corresponding position to the camera. The brighter pixels in the depth image indicate that the pixel points are closer to the camera, while the darker pixels indicate that the pixel points are farther from the camera.

[0074] Combining the current pose information of the camera and the three-dimensional point cloud of the target scene, the coordinates of each three-dimensional point in the target scene in the camera coordinate system can be calculated. Project the three-dimensional point cloud onto the image plane to convert the coordinates of each three-dimensional point in the camera coordinate system into image coordinates, and then calculate the depth value of each pixel point according to the pose information and the image coordinates obtained after projection. Fill the calculated depth value of each pixel point into the corresponding image position, and a depth map corresponding to the scene image can be generated.

[0075] Convert the coordinate system of the depth map to the specified coordinate system, and generate a depth image corresponding to the scene image in the specified coordinate system. Then, with the help of the current pose information of the camera and the depth image in the specified coordinate system, back-project each pixel point in the scene image to the specified coordinate system to obtain the spatial points corresponding to each pixel point in the specified coordinates.

[0076] In the above embodiments, by combining the current pose information of the camera and the depth image generated in the specified coordinate system, the pixel points in the scene image are back-projected into spatial points in the specified coordinate system, thereby enabling global consistency in the three-dimensional representation of each pixel point and ensuring the visualization of the sound field generated by the sound source object.

[0077] In some optional embodiments, the step of generating the depth image of the scene image in the specified coordinate system may include: generating a relative depth map of the scene image through a depth estimation network, and aligning the depth information in the relative depth map to the specified coordinate system to obtain the depth image of the scene image in the specified coordinate system.

[0078] The depth estimation network is a pre-trained network model for generating depth maps, such as a monocular depth estimation network. The relative depth map is an image representing the depth information of objects in a scene image. The relative depth map uses different grayscale values or colors to represent the relative depth relationship of different object objects in the scene image.

[0079] Specifically, the depth estimation network is deployed in a computer device. The computer device inputs the scene image of the target scene it captures into the depth estimation network, so that the depth estimation network outputs the corresponding relative depth map, as Figure 3 shown.

[0080] Align the depth information in the relative depth map to a specified coordinate system to ensure the global consistency of each pixel point in terms of depth. Specifically, obtain the camera parameters (intrinsic and extrinsic parameters) for generating the relative depth map, and select corresponding reference points in the specified coordinate system; calculate the coordinate transformation matrix from the camera coordinate system to the specified coordinate system according to the camera parameters and reference points; for each depth value in the relative depth map, convert its corresponding pixel coordinate point to a point in the camera coordinate system, and apply the coordinate transformation matrix to convert the point in the camera coordinate system to the specified coordinate system to obtain the position of each depth value in the specified coordinate system, so as to obtain the depth image of the scene image in the specified coordinate system, as Figure 3 shown.

[0081] In the above embodiments, the relative depth map of the scene image is generated through the depth estimation network, and the depth information in the relative depth map is aligned to a specified coordinate system to obtain the depth image of the scene image in the specified coordinate system, so that the depth information has global consistency, thereby ensuring the global consistency of each pixel point in the three-dimensional representation.

[0082] Step S3022, map the spatial points to the signed distance field to form the object surface in the signed distance field.

[0083] The signed distance field (SDF) is a data structure representing the shape of an object in space. It separates the points inside the object from the points outside the object through the signed distance field and records the distance of each point to the object surface. The signed distance field can be represented by a two-dimensional or three-dimensional grid, and each voxel in the signed distance field contains a signed distance value.

[0084] For the spatial points that need to be mapped to the signed distance field, convert the spatial coordinates corresponding to the spatial points to the grid coordinates in the signed distance field, and form the object surface of the sound source object in the signed distance field according to the determined grid coordinates.

[0085] Step S3023: Extract the voxels located on the object surface in the signed distance field, and render the extracted voxels as image units in the unit image to generate the unit image corresponding to the scene image.

[0086] According to the distance values contained in each voxel in the signed distance field, extract the voxels whose signed distance values represent the object surface, that is, extract the voxels located on the object surface. Subsequently, determine the image units to be rendered based on the extracted voxels located on the object surface, and record the position information of each voxel located on the object surface, so as to perform visual rendering according to the position information of each voxel to form the corresponding unit image to display the shape information of the object surface.

[0087] Taking image units as particles as an example, determine the particles to be rendered according to the extracted voxels located on the object surface, and determine the particle attributes such as the position, size, and color of each particle, and at the same time determine the current state of each particle in the target scene. As Figure 3 shown, in each rendering frame, according to the particle attributes and the current state of the particles in the target scene, use graphic elements such as points, lines, textures, or custom shapes to render the particles to generate the unit image formed by the particles.

[0088] In some optional implementation manners, the above step S3023 may include:

[0089] Step b1: Calculate the signed distance value of each voxel in the signed distance field, and the signed distance value represents the distance information between the voxel and the object surface.

[0090] Step b2: Determine and extract the target voxels with the signed distance value being a specified value.

[0091] Step b3: Render the extracted target voxels as image units in the unit image.

[0092] The signed distance value represents the distance value of each voxel in the signed distance field, and the distance between the voxel and the object surface is represented by the signed distance value. For each voxel in the signed distance field, the corresponding signed distance value can be calculated according to algorithms such as the flood algorithm and the oriented distance field.

[0093] The specified value is a pre-set distance value used to represent the object surface. For example, the specified value is set to be equal to 0. The target voxels are one or more voxels located on the object surface. Specifically, traverse each voxel in the signed distance field in turn, and detect whether the signed distance value corresponding to each voxel is equal to the specified value. Extract the voxels with the signed distance value being the specified value according to the detection result, and render the extracted target voxels as image units in the unit image.

[0094] In the above embodiments, the target voxels located on the object surface are determined by the signed distance values of the respective voxels in the signed distance field, and the target voxels are rendered as image units in the unit image. Thus, the rendering effect of the image units can be ensured to be consistent with the shape of the object surface, improving the rendering effect of the image units.

[0095] Step S303: Based on the sound source position, determine the special effect area in the unit image. For detailed description, refer to the relevant description corresponding to the above embodiments, which will not be elaborated here.

[0096] Step S304: Add a perturbation effect to the image units in the special effect area to display the image units in the special effect area according to the perturbation effect. For detailed description, refer to the relevant description corresponding to the above embodiments, which will not be elaborated here.

[0097] The method for displaying a scene image provided in this embodiment projects the pixel points in the scene image onto spatial points in a specified coordinate system, determines the voxels located on the object surface according to the signed distance field, and renders the extracted voxels as image units in the unit image. Then, according to the sound source position and the unit image obtained after rendering, a perturbation effect is applied to the image units in the special effect area, so that the scene image captured by the camera can be converted into a unitized effect, facilitating the interaction between the sound source and the actual scene in the form of a unitized scene, thereby realizing the visualization of the spatial sound field.

[0098] In this embodiment, a device for displaying a scene image is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be elaborated here. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0099] This embodiment provides a device for displaying a scene image, as Figure 5 shown, including:

[0100] An image processing unit 401, configured to collect a scene image of a target scene and identify the sound source position of a sound source object in the scene image.

[0101] An image generation unit 402, configured to generate a unit image corresponding to the scene image and display the unit image.

[0102] A special effect area determination unit 403, configured to determine a special effect area in the unit image based on the sound source position.

[0103] A perturbation display unit 404, configured to add a perturbation effect to the image units in the special effect area to display the image units in the special effect area according to the perturbation effect.

[0104] In some alternative embodiments, the above-mentioned image processing unit 401 may include:

[0105] A sound source position determination subunit, configured to collect multiple channel data of a sound source object and determine the spatial position of the sound source object in a target scene according to the multiple channel data.

[0106] A position conversion subunit, configured to obtain the internal and external parameters of a camera and convert the spatial position into a pixel position in a scene image based on the internal and external parameters.

[0107] In some alternative embodiments, the above-mentioned perturbation display unit 404 may include:

[0108] An identification subunit, configured to identify the current sound source intensity of a sound source object, determine a perturbation effect matching the sound source intensity, and add the perturbation effect to an image unit in a special effect area.

[0109] In some alternative embodiments, the above-mentioned device may further include:

[0110] An editing unit, configured to receive an editing instruction for a special effect area and adjust the configuration parameters of the special effect area in response to the editing instruction.

[0111] A perturbation update unit, configured to redisplay the image unit with the perturbation effect defined by the configuration parameters added according to the adjusted configuration parameters.

[0112] In some alternative embodiments, the above-mentioned image generation unit 402 may include:

[0113] A projection subunit, configured to back-project a pixel point in a scene image into a spatial point in a specified coordinate system.

[0114] A mapping subunit, configured to map the spatial point into a signed distance field to form an object surface in the signed distance field.

[0115] An extraction subunit, configured to extract voxels located on the object surface in the signed distance field and render the extracted voxels into image units in a unit image to generate a unit image corresponding to the scene image.

[0116] In some alternative embodiments, the above-mentioned projection subunit is specifically configured to: obtain the current pose information of a camera and generate a depth image of a scene image in a specified coordinate system; back-project a pixel point in the scene image into a spatial point in the specified coordinate system according to the current pose information and the depth image.

[0117] In some alternative embodiments, the above projection sub-unit is further specifically configured to: generate a relative depth map of the scene image through a depth estimation network, and align the depth information in the relative depth map to a specified coordinate system to obtain a depth image of the scene image in the specified coordinate system.

[0118] In some alternative embodiments, the above extraction sub-unit is specifically configured to: calculate the signed distance value of each voxel in the signed distance field, where the signed distance value represents the distance information between the voxel and the object surface; determine and extract the target voxel with the signed distance value being a specified value.

[0119] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.

[0120] The display device of the scene image in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0121] The display device of the scene image provided in this embodiment realizes the unitization of the scene image by converting the scene image into a unit image, and can identify the sound source position where the sound source object in the scene image is located, so as to determine a special effect area corresponding to the sound source position in the unit image. Thereby, the sound source in the target scene can be projected into the virtual world in the special effect scene, and then a perturbation effect is added to the image unit in the special effect area to apply the spatial sound field generated by the sound source object in the target scene to the virtual world in the special effect scene, realizing the visualization of the sound field in the actual environment, ensuring that the display effect of the image unit in the special effect area is more vivid, enabling the user to better perceive the real environment through the special effect scene, and improving the display effect of the special effect.

[0122] The present disclosure embodiment also provides a computer device having the above Figure 5 display device of the scene image shown.

[0123] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present disclosure. As Figure 6As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 Taking one processor 10 as an example in

[0124] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.

[0125] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0126] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0127] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.

[0128] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 6 Taking the connection through the bus as an example.

[0129] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The above display device includes, but is not limited to, a liquid crystal display, a light emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0130] The computer device further includes a communication interface for the computer device to communicate with other devices or communication networks.

[0131] The embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.

[0132] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for displaying a scene image, characterized in that, the method includes: Collect a scene image of a target scene and identify the sound source position of a sound source object in the scene image; Generate a unit image corresponding to the scene image and display the unit image; Based on the sound source position, determine a special effect area in the unit image; Add a perturbation effect to the image units in the special effect area to display the image units in the special effect area according to the perturbation effect.

2. The method according to claim 1, characterized in that, adding a perturbation effect to the image units in the special effect area includes: Identify the sound source intensity currently emitted by the sound source object, determine a perturbation effect matching the sound source intensity, and add the perturbation effect to the image units in the special effect area.

3. The method according to claim 1 or 2, characterized in that, After displaying the image units in the special effect area according to the perturbation effect, the method further includes: Receive an editing instruction for the special effect area, and in response to the editing instruction, adjust the configuration parameters of the special effect area; According to the adjusted configuration parameters, redisplay the image units with the perturbation effect defined by the configuration parameters added.

4. The method according to claim 1, characterized in that, identifying the sound source position of the sound source object in the scene image includes: Collect multiple channel data of the sound source object and determine the spatial position of the sound source object in the target scene according to the multiple channel data; Obtain the internal and external parameters of the camera, and based on the internal and external parameters, convert the spatial position into a pixel position in the scene image; wherein, the pixel position represents the sound source position of the sound source object in the scene image.

5. The method according to claim 1, characterized in that, generating the unit image corresponding to the scene image includes: Back-project the pixel points in the scene image into spatial points in a specified coordinate system; Map the spatial points to a signed distance field to form the object surface in the signed distance field; Extract the voxels located on the object surface in the signed distance field and render the extracted voxels as image units in the unit image to generate the unit image corresponding to the scene image.

6. The method according to claim 5, characterized in that, back-projecting the pixel points in the scene image into spatial points in a specified coordinate system includes: Obtain the current pose information of the camera and generate a depth image of the scene image in a specified coordinate system; According to the current pose information and the depth image, back-project the pixel points in the scene image into spatial points in a specified coordinate system.

7. The method according to claim 6, characterized in that, generating the depth image of the scene image in a specified coordinate system includes: Generate a relative depth map of the scene image through a depth estimation network and align the depth information in the relative depth map to a specified coordinate system to obtain the depth image of the scene image in the specified coordinate system.

8. The method according to claim 5, characterized in that, Extracting the voxels located on the surface of the object in the signed distance field includes: Calculating the signed distance value of each voxel in the signed distance field, where the signed distance value represents the distance information between the voxel and the surface of the object; Determining and extracting the target voxels with the signed distance value being a specified value.

9. A display device for a scene image, characterized in that, the device includes: An image processing unit for collecting a scene image of a target scene and identifying the sound source position of the sound source object in the scene image; An image generation unit for generating a unit image corresponding to the scene image and displaying the unit image; A special effect area determination unit for determining a special effect area in the unit image based on the sound source position; A perturbation display unit for adding a perturbation effect to the image units in the special effect area to display the image units in the special effect area according to the perturbation effect.

10. A computer device, characterized in that, it includes: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method for displaying a scene image according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer instructions for causing a computer to execute the method for displaying a scene image according to any one of claims 1 to 8.