Image shooting method and device in three-dimensional reconstruction scene, equipment and storage medium

By identifying the 3D reconstructed scene as the scene to be photographed in the in-vehicle SR environment and using a virtual camera to adjust and generate the target image, the problem of limited user screenshot function is solved, enabling free composition and creation, and improving user experience.

CN122002018APending Publication Date: 2026-05-08XG TECHNOLOGIES PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XG TECHNOLOGIES PTE LTD
Filing Date
2026-01-30
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing technologies, users can only capture fixed images in the in-vehicle SR environment through the system's built-in screenshot function, which lacks creative freedom and makes it difficult to accurately capture 3D reconstructed scene images that meet user requirements.

Method used

By determining the 3D reconstructed scene as the scene to be photographed when the photography mode is triggered, and acquiring images based on the adjusted virtual camera, users can perform six degrees of freedom of camera movement and composition in a 3D virtual space, including translation, rotation, and scaling, and generate target images by combining computer graphics rendering technology.

Benefits of technology

It enables the freezing of 3D reconstructed scenes in the in-vehicle SR environment, allowing users to freely compose and create, solving the problem of limited screenshot function, and improving creative freedom and shooting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122002018A_ABST
    Figure CN122002018A_ABST
Patent Text Reader

Abstract

The invention discloses an image shooting method and device in a three-dimensional reconstruction scene, equipment and a storage medium, and the method comprises the steps: responding to a received shooting starting instruction, and determining the three-dimensional reconstruction scene of the surrounding environment of a vehicle at the current moment as a to-be-shot three-dimensional scene; in response to the received instruction for adjusting the first virtual camera in the to-be-shot three-dimensional scene, adjusting the camera pose of the first virtual camera in the to-be-shot three-dimensional scene; and determining a target shooting image based on the adjusted camera pose and the scene data of the to-be-shot three-dimensional scene. According to the technical scheme, the three-dimensional reconstruction scene when shooting is started is determined as the to-be-shot three-dimensional scene, and image acquisition is performed on the to-be-shot three-dimensional scene based on the adjusted virtual camera; therefore, the current three-dimensional reconstruction scene of the vehicle can be frozen when the photographing mode is triggered, and a user is allowed to perform free mirror movement and composition, so that the problems that a user screenshot function is limited and creation freedom degree is lacked in a vehicle-mounted SR environment are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of intelligent driving technology, and in particular to an image capture method, apparatus, device, and storage medium for three-dimensional reconstruction scenes. Background Technology

[0002] Currently, with the development of intelligent driving technology and smart cockpits, HMI (Human-Machine Interface) is evolving from simple information display to immersive and intelligent 3D (Three-Dimensional) environmental interaction. Among these technologies, SR (Surrounding Reality) technology can visualize the real-world environment perceived by vehicle sensors (such as cameras and LiDAR) in real-time in 3D on the in-vehicle screen. However, when users want to record a moment of the SR-generated image, they can only do so through the system's built-in screenshot function, resulting in a poor user experience. Summary of the Invention

[0003] To address the aforementioned technical problems, this disclosure provides a method, apparatus, device, and storage medium for capturing images in a three-dimensional reconstructed scene.

[0004] A first aspect of this disclosure provides a method for capturing images in a vehicle 3D reconstruction scene, comprising: in response to receiving an instruction to start capturing, determining the 3D reconstruction scene of the vehicle's surrounding environment at the current moment as the 3D scene to be captured; in response to receiving an instruction to adjust a first virtual camera in the 3D scene to be captured, adjusting the camera pose of the first virtual camera in the 3D scene to be captured; and determining a target image to be captured based on the adjusted camera pose and scene data of the 3D scene to be captured.

[0005] A second aspect of this disclosure provides an image capturing device for a vehicle 3D reconstruction scene, comprising: a 3D scene determination module, configured to determine the 3D reconstruction scene of the vehicle's surrounding environment at the current moment as the 3D scene to be captured in response to receiving a command to start capturing; a camera pose adjustment module, configured to adjust the camera pose of the first virtual camera in the 3D scene to be captured in response to receiving a command to adjust the first virtual camera in the 3D scene to be captured; and a target image determination module, configured to determine a target image to be captured based on the adjusted camera pose and scene data of the 3D scene to be captured.

[0006] A third aspect of this disclosure provides a computer-readable storage medium storing a computer program for executing the image capture method in a vehicle 3D reconstruction scene provided in the first aspect embodiment.

[0007] A fourth aspect of this disclosure provides an electronic device comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the image capture method in a vehicle 3D reconstruction scene provided in the first aspect embodiment above.

[0008] A fifth aspect of this disclosure provides a computer program product that, when instructions in the computer program product are executed by a processor, performs the image capture method in a vehicle 3D reconstruction scene provided in the first aspect of the embodiment above.

[0009] The image capture method in the vehicle 3D reconstruction scene disclosed herein determines the 3D reconstruction scene at the start of photography as the 3D scene to be captured, and captures images of the 3D scene to be captured based on the adjusted virtual camera; thus, the 3D reconstruction scene of the current surrounding environment of the vehicle can be frozen when the photography mode is triggered, and allows users to freely move the camera and compose the shot, thereby solving the problem of limited screenshot function and lack of creative freedom for users in the in-vehicle SR environment. Attached Figure Description

[0010] Figure 1 A schematic diagram of a vehicle provided for an exemplary embodiment of this disclosure.

[0011] Figure 2 This is a schematic flowchart of an image capturing method provided as an exemplary embodiment of the present disclosure.

[0012] Figure 3 A schematic flowchart of an image capturing method provided for another exemplary embodiment of this disclosure.

[0013] Figure 4 A schematic flowchart of an image capturing method provided as another exemplary embodiment of this disclosure.

[0014] Figure 5 A schematic flowchart of an image capturing method provided as another exemplary embodiment of this disclosure.

[0015] Figure 6 A schematic flowchart of an image capturing method provided as another exemplary embodiment of this disclosure.

[0016] Figure 7 This is a schematic diagram of the structure of an image capturing device provided for an exemplary embodiment of the present disclosure.

[0017] Figure 8 A schematic diagram of the structure of an image capturing apparatus provided for another exemplary embodiment of this disclosure.

[0018] Figure 9 This is a schematic diagram of the structure of an image capturing device provided as another exemplary embodiment of the present disclosure.

[0019] Figure 10 A schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this disclosure. Detailed Implementation

[0020] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.

[0021] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0022] Application Overview like Figure 1 As shown, the vehicle 10 is equipped with at least one visual sensor 11 and a radar sensor 12. The visual sensor 11 and / or the radar sensor 12 can then collect data on the vehicle's surrounding environment to obtain perceived data of the environment around the vehicle 10. This perceived data can then be processed using SR (Reconstructed Environment) technology to obtain a three-dimensional reconstructed scene of the environment around the vehicle 10. SR technology is a technique that digitizes and intuitively presents the real environment through multi-sensor fusion, real-time data processing, and 3D rendering.

[0023] It should be noted that, Figure 1 For illustrative purposes only, one vision sensor 11 or multiple vision sensors 11 may be installed on the vehicle 10. This disclosure embodiment... Figure 1 The number and location of the vision sensors 11 are not limited. Of course, in addition to the vision sensors 11 and radar sensors 12, other types of sensors can be installed on the vehicle 10, and this embodiment does not limit this. In actual use, the vision sensor 11 can be a forward-looking wide-angle / narrow-angle camera of the vehicle, the vision sensor 11 can also be a side-view / surround-view / panoramic-view camera of the vehicle, or the vision sensor 11 can be an RGB-D (Red Green Blue-Depth) camera on the vehicle, and the radar sensor 12 can be a lidar, millimeter-wave radar, ultrasonic radar, etc. on the vehicle.

[0024] Furthermore, when the user triggers the command to start shooting, the image shooting method in the vehicle 3D reconstruction scene provided in this embodiment can be used to... Figure 1 The 3D reconstructed scene obtained in the image is processed to obtain the target image, which is then displayed for user use.

[0025] In related technologies, when users want to record the scene generated by SR, they can usually only use the system's built-in screenshot function. However, the system's built-in screenshot function has the following drawbacks: it can only capture a fixed scene centered on the vehicle at the current moment, and since the SR environment is changing in real time, it is difficult to accurately capture a scene that meets the user's requirements.

[0026] To address the aforementioned issues, this disclosure provides an image capture method for a vehicle 3D reconstruction scene. By determining the 3D reconstruction scene at the start of photography as the 3D scene to be captured, and acquiring images of the 3D scene to be captured based on an adjusted virtual camera, the 3D reconstruction scene of the vehicle's current surrounding environment can be frozen when the photography mode is triggered, allowing users to freely compose shots. This solves the problem of limited screenshot functionality and lack of creative freedom in the in-vehicle SR environment.

[0027] Exemplary methods Figure 2 This is a schematic flowchart illustrating an exemplary embodiment of the image capturing method provided in this disclosure. The method of this embodiment can be applied to electronic devices, such as... Figure 2 As shown, the method includes steps S21-S23.

[0028] Step S21: In response to receiving the instruction to start shooting, determine the three-dimensional reconstructed scene of the vehicle's surrounding environment at the current moment as the three-dimensional scene to be shot.

[0029] For example, in this embodiment of the disclosure, the instruction to start shooting can be a user-triggered instruction for activating the photography mode. For instance, a specific function button can be set on the vehicle's steering wheel or a specific virtual function control can be set on the vehicle's display screen, allowing the user to trigger the instruction to start shooting via the function button or virtual function control. Of course, the user can also trigger the instruction to start shooting via voice or air gestures, and this embodiment of the disclosure does not limit this. In some examples, the instruction to start shooting in step S21 can also be non-user-triggered, for example, it can be system-triggered, and this embodiment of the disclosure does not limit this either.

[0030] In some examples, when the processor of an electronic device (such as a CPU) receives a command to start recording, it can determine the 3D reconstructed scene of the vehicle's surroundings at the current moment (such as the moment the command is received) as the object to be recorded corresponding to the command to start recording. For example, as Figure 1As shown, vehicle 10 can collect real-time perception data of the vehicle's surroundings through vision sensor 11 and / or radar sensor 12, and construct a three-dimensional reconstructed scene at each moment during the driving process using the real-time collected perception data. Furthermore, if a command to start shooting is received at time t, the three-dimensional reconstructed scene corresponding to time t is determined as the three-dimensional scene to be shot.

[0031] Step S22: In response to receiving an instruction to adjust the first virtual camera in the three-dimensional scene to be photographed, adjust the camera pose of the first virtual camera in the three-dimensional scene to be photographed.

[0032] For example, a virtual camera used to observe a 3D reconstructed scene is typically bound to a vehicle and has the same default viewpoint as the vehicle, which can be a driver's viewpoint or a bird's-eye viewpoint. Therefore, the virtual camera can only capture a fixed image centered on the vehicle. Based on this, in this embodiment, the virtual camera used to observe the 3D reconstructed scene can be separated from the default viewpoint bound to the vehicle, allowing the virtual camera to move at multiple angles in 3D virtual space. The unbound virtual camera then becomes the first virtual camera in step S22. Thus, the user can perform six degrees of freedom (6-DoF) free camera movement and composition with the first virtual camera in 3D virtual space. For example, the user can drag the first virtual camera left and right on the screen with a single finger, or move it forward and backward or zoom it by pinching or opening two fingers, or rotate or orbit it by dragging it in the same direction with two fingers or by pressing and dragging a specific button with a single finger.

[0033] Furthermore, in this embodiment of the present disclosure, in response to receiving an instruction to adjust the first virtual camera in the three-dimensional scene to be captured, the camera pose of the first virtual camera in the three-dimensional scene to be captured can be adjusted. The instructions to adjust the first virtual camera in the three-dimensional scene to be captured include, but are not limited to, translation-related adjustment instructions and rotation-related adjustment instructions. In some embodiments, the first virtual camera in step S22 can also be a newly added virtual camera; this embodiment of the present disclosure does not limit this, and those skilled in the art can set it according to actual usage requirements.

[0034] In some examples, commands to adjust the first virtual camera can be triggered in multiple ways. For instance, users can drag the position or rotate the first virtual camera on the in-vehicle display screen, specify the desired viewpoint change via voice commands, or capture the user's air gestures using gesture recognition technology, thus achieving more intuitive camera pose adjustment. In some embodiments, preset camera pose templates can also be provided for users to quickly select and apply. Furthermore, during the adjustment process, the pose parameters of the first virtual camera can be updated in real time and mapped onto the 3D scene to be captured. In this way, users can preview the image effects under different camera poses in real time during the adjustment process, thereby more precisely controlling the final shooting result. The adjusted camera pose includes not only changes in position but also adjustments to directional parameters such as pitch, yaw, and roll angles to ensure that users can achieve omnidirectional viewpoint switching.

[0035] Step S23: Based on the adjusted camera pose and scene data of the 3D scene to be photographed, determine the target image to be captured.

[0036] For example, in this embodiment of the present disclosure, the scene data (i.e., resource file) of the 3D scene to be captured can be a tree-like data structure containing information on all objects, light sources, cameras, their hierarchical relationships, and transformations in the scene. For instance, the scene data of the 3D scene to be captured includes at least the geometric structure, texture information, lighting conditions, and the states of dynamic objects in the 3D reconstructed scene. Furthermore, based on the camera pose adjusted by the first virtual camera, combined with the scene data of the 3D scene to be captured, a target image can be generated using rendering techniques in computer graphics.

[0037] The image capture method provided in this embodiment determines the three-dimensional reconstructed scene at the start of photography as the three-dimensional scene to be captured, and acquires images of the three-dimensional scene to be captured based on the adjusted virtual camera. In this way, the three-dimensional reconstructed scene of the current surrounding environment of the vehicle can be frozen when the photography mode is triggered, and users can freely move the camera and compose the shot, thereby solving the problem of limited screenshot function and lack of creative freedom for users in the in-vehicle SR environment.

[0038] like Figure 3 As shown above, in the above Figure 2 Based on the illustrated embodiment, step S21 may include steps S211-S213.

[0039] Step S211: In response to receiving the instruction to start shooting, an interrupt signal is triggered.

[0040] For example, the image capture method provided in this disclosure can pause the processing and updating of external sensor data when the photography mode is activated by triggering an interrupt signal, thereby saving the three-dimensional reconstruction scene of the vehicle's surrounding environment at the current moment. The triggering of the interrupt signal ensures that the three-dimensional reconstruction scene of the vehicle's surrounding environment remains static at the current moment, avoiding real-time updates of scene data due to changes in the external environment. For instance, when the user clicks the "photography mode" button, the HMI (Human-Machine Interface) system sends a "pause update" interrupt signal or software flag to the scene reconstruction module.

[0041] Step S212: In response to the interrupt signal, stop acquiring perception data of the vehicle's surrounding environment after the current moment, so that the vehicle's 3D reconstruction scene remains in the scene at the time the interrupt signal was triggered.

[0042] For example, in this embodiment of the disclosure, in response to the triggering of an interruption signal, the acquisition of perception data of the surrounding environment collected by the vehicle's sensors after the current time can be stopped, thereby keeping the vehicle's 3D reconstruction scene in the scene at the time the interruption signal was triggered. For example, after the interruption signal is triggered, the data processing pipeline will temporarily stop pulling new data frames from upstream nodes (such as sensor drivers), making the logical time of the entire SR environment "freeze" at the current frame.

[0043] In other words, when the system receives an interrupt signal, it locks the current 3D reconstructed scene and uses it as the basis for subsequent image capture. At this time, the environmental state around the vehicle is fully preserved, including information such as geometric structure, texture details, and lighting conditions. This locking mechanism ensures that the scene will not be updated due to changes in the external environment when the user adjusts and composes the first virtual camera.

[0044] Step S213: Determine the current 3D reconstruction scene as the 3D scene to be captured.

[0045] For example, the 3D reconstructed scene at the current moment can be saved as the 3D scene to be photographed. In some embodiments, after determining that the 3D reconstructed scene at the current moment is the 3D scene to be photographed, the relevant resource files of the scene can be automatically loaded and mapped into the view range of the first virtual camera. In this way, the user can freely move the camera and compose the shot within the locked 3D reconstructed scene by adjusting the pose of the first virtual camera.

[0046] For example, when a user is driving on a highway, a uniquely designed AI-generated truck appears in the 3D reconstructed scene on the vehicle's display screen. The user immediately presses the "photograph" button on the steering wheel. At this time, the system will trigger an interrupt signal to pause the update of the SR environment and cache the complete 3D reconstructed scene data (model, textures, lighting effects, etc.) containing the truck and its surrounding environment as the 3D scene to be photographed.

[0047] The image capture method provided in this disclosure, in response to receiving a command to start capturing, triggers an interrupt signal; in response to the interrupt signal, it stops acquiring perception data of the vehicle's surrounding environment after the current moment, so that the vehicle's 3D reconstruction scene remains in the scene at the time the interrupt signal was triggered; and determines the current 3D reconstruction scene as the 3D scene to be captured. In this way, the real-time changing SR environmental data can be snapshotted to generate a static 3D scene that users can use for composition and creation.

[0048] like Figure 4 As shown above, in the above Figure 3 Based on the illustrated embodiment, step S213 may include steps S2131-S2134.

[0049] Step S2131: Determine the resource data corresponding to the scene where the vehicle is located at the current moment.

[0050] For example, the resource data corresponding to the scene where the vehicle is located at the current moment can be a tree-like data structure containing information about all objects, light sources, cameras, their hierarchical relationships, and transformations in the scene. This resource data can then be stored for later use. For instance, a rendering engine (such as Unity) can perform a deep copy of the scene graph at the current moment (i.e., the current frame) and store it in a separate memory area not accessed by the rendering thread. Here, the scene graph in the rendering engine refers to the data structure used to organize and manage all objects in the 3D reconstructed scene. This ensures the integrity and independence of the resource data, reducing interference from changes in the external environment during subsequent operations. Furthermore, by performing a deep copy of the scene graph, all detailed information about the vehicle's surrounding environment at the current moment can be preserved, including geometry, textures, lighting conditions, and the states of dynamic objects.

[0051] Step S2132: Serialize the dynamic data in the resource data to obtain the serialized dynamic data.

[0052] For example, in this embodiment of the disclosure, serialization libraries such as Protocol Buffers and Flat Buffers can be used to serialize dynamic data in resource data to obtain serialized dynamic data. Serialization of dynamic data refers to the process of converting and saving dynamic data in resource data into data that can be stored, transmitted, or reconstructed. For example, the system will traverse all dynamic objects in the scene graph, such as AI (Artificial Intelligence) generated vehicles, pedestrians, and special effects, and serialize their key state parameters. These parameters include, but are not limited to: model ID (Identity), 3D position, rotation parameters, scaling parameters, the current pose of the skeletal animation, the current state of the particle system, and material properties.

[0053] In some examples, serialized dynamic data can be stored in a specific cache area for fast subsequent access and processing. This approach not only effectively reduces data redundancy but also improves data processing efficiency. For instance, during serialization, the system compresses and encodes key state parameters of each dynamic object, thereby reducing storage space usage.

[0054] Step S2133: Store the serialized dynamic data and static data from the resource data to determine the 3D reconstruction scene at the current moment.

[0055] For example, in this embodiment of the disclosure, serialized dynamic data and static data in the resource data can be stored and integrated to determine the 3D reconstruction scene at the current moment. The static data in the resource data typically includes information such as fixed objects in the scene, lighting conditions, and environmental textures. This content does not require frequent updates, so it can be directly saved in its original format or a lightweight format. This approach not only reduces storage space usage but also improves the efficiency of subsequent loading and rendering. During the integration process, the system remaps the serialized dynamic data to the static data according to its hierarchical relationship and transformation information, thereby forming a complete resource file for the 3D reconstruction scene.

[0056] Step S2134: Determine the current 3D reconstructed scene as the 3D scene to be photographed.

[0057] For example, after storing and integrating the serialized dynamic data and the static data in the resource data, the complete resource file of the 3D reconstructed scene can be directly loaded into the view range of the first virtual camera, allowing the user to freely move the camera and compose the image.

[0058] The image capture method provided in this disclosure involves determining the resource data corresponding to the scene where the vehicle is located at the current moment; serializing the dynamic data in the resource data to obtain serialized dynamic data; storing the serialized dynamic data and the static data in the resource data to determine the 3D reconstruction scene at the current moment; and determining the 3D reconstruction scene at the current moment as the 3D scene to be captured. This effectively reduces data redundancy and improves data processing efficiency.

[0059] In some embodiments, core GPU (Graphics Processing Unit) resources such as geometry data (e.g., vertex buffers, index buffers), textures, lightmaps, and rendering targets required by the current scene can be marked as "persistent" to prevent them from being released or overwritten in subsequent regular rendering cycles, thus ensuring the integrity of the frozen scene.

[0060] like Figure 5 As shown above, in the above Figure 2 Based on the illustrated embodiment, step S24 or steps S25 and S26 are included before step S22.

[0061] Step S24: Create the first virtual camera based on the 3D reconstructed scene.

[0062] For example, in this embodiment of the disclosure, a first virtual camera can be created by adding a camera instance based on a 3D reconstructed scene.

[0063] In some examples, the creation process of the first virtual camera may include steps such as initializing camera parameters, setting the initial pose, and binding the rendering pipeline. For instance, the system can generate a first virtual camera based on the default viewpoint of the 3D reconstructed scene and assign it an independent view matrix and projection matrix to ensure that it can move flexibly without being limited by the vehicle's default viewpoint. Furthermore, the newly created first virtual camera can be loaded into the current 3D reconstructed scene for the user to adjust.

[0064] Step S25: Determine the second virtual camera that is associated with the vehicle.

[0065] For example, in this embodiment of the disclosure, a second virtual camera that is bound to the vehicle can be determined. The second virtual camera is usually used to simulate the driver's perspective or overhead perspective. Its pose is fixed to the vehicle. Therefore, the second virtual camera can usually only capture a fixed image centered on the vehicle.

[0066] Step S26: Decouple the second virtual camera from the vehicle and use the decoupled second virtual camera as the first virtual camera.

[0067] For example, in this embodiment of the disclosure, the association between the second virtual camera and the vehicle can be severed, thereby separating the second virtual camera from the fixed transformation relationship with the vehicle, and using the severed virtual camera as the first virtual camera. This breaks the forced binding between the virtual camera and the vehicle's main viewpoint, allowing the user to perform six degrees of freedom (6-DoF) free camera movement and composition in a frozen 3D reconstructed scene.

[0068] For example, the parent node of the second virtual camera can be removed from the vehicle object in the scene graph of the rendering engine, making it an independent object under the scene root node of the current 3D reconstructed scene. In this way, the movement of the second virtual camera will no longer be affected by the movement of the vehicle. Furthermore, the decoupled second virtual camera can be used as the first virtual camera.

[0069] For example, once the command to start shooting is triggered, the vehicle's display screen automatically switches to photography mode, separating the second virtual camera from the front view of the vehicle. The separated second virtual camera can then be used as the first virtual camera. Furthermore, users can slide their fingers on the touchscreen to move the first virtual camera's view from inside the vehicle to outside and then to the air, creating a bird's-eye view. They can also pinch two fingers together to bring the camera closer, making the target the main subject of the image, thus completing camera movement and composition.

[0070] It should be noted that in this embodiment, the first virtual camera can be determined by creating a new virtual camera instance or by unbinding an existing virtual camera. Those skilled in the art can choose the method for determining the first virtual camera based on actual usage, and this embodiment does not impose any limitations on this. That is, in this embodiment, step S24 can be executed first, followed by step S22. Alternatively, steps S25 and S26 can be executed first, followed by step S22.

[0071] The image capture method provided in this disclosure creates a first virtual camera based on a 3D reconstructed scene; or, determines a second virtual camera associated with the vehicle; removes the association between the second virtual camera and the vehicle, and uses the removed second virtual camera as the first virtual camera. This overcomes the fixed viewing angle limitations of in-vehicle HMIs, thereby enabling free camera movement and creative capabilities within a 3D reconstructed scene.

[0072] In some embodiments, step S26, disconnecting the second virtual camera from the vehicle and using the disconnected second virtual camera as the first virtual camera, includes: determining the parent node of the second virtual camera based on the association; removing the parent node from the vehicle node of the vehicle; and adding the removed parent node to the scene node of the 3D reconstructed scene.

[0073] For example, the parent node of the second virtual camera can be removed from the vehicle node and then remounted to the root node of the 3D reconstructed scene. This ensures that the second virtual camera can be freely adjusted independently of the vehicle. For instance, in the scene graph, the second virtual camera originally existed as a child node of the vehicle node, and its position and orientation depended on the real-time state of the vehicle; however, by moving its parent node to the scene root node, the second virtual camera gains complete independence and can move and rotate freely throughout the entire 3D reconstructed scene.

[0074] In some embodiments, an input processor dedicated to photography mode may be activated, which begins to monitor specific input events from a touchscreen, physical knob, or voice command.

[0075] like Figure 6 As shown above, in the above Figure 2 Based on the illustrated embodiment, step S23 may include steps S231-S232.

[0076] Step S231: Based on the adjusted camera pose and scene data of the 3D scene to be photographed, determine the image to be captured.

[0077] For example, in this embodiment of the present disclosure, the captured image can be generated based on the camera pose adjusted by the first virtual camera, combined with the scene data of the three-dimensional scene to be captured, and using rendering technology in computer graphics.

[0078] Step S232: In response to receiving the instruction to stylize the captured image, stylize the captured image to obtain the target captured image.

[0079] For example, the image capturing method provided in this disclosure supports further creative processing (i.e., stylization updates) of the captured image after it has been captured. For instance, if a user feels the captured image is slightly dark, they can drag the "Neon Light Intensity" slider upwards by 20% in the style adjustment bar on the display screen and select the "Rainy Night" filter. Then, in response to receiving the stylization update instruction for adjusting the "Neon Light Intensity" of the captured image and the stylization update instruction for selecting the "Rainy Night" filter, the processor performs corresponding stylization processing on the captured image to obtain the target captured image.

[0080] In some embodiments, stylizing captured images can be achieved through the following steps: First, breaking down complex AI art styles into user-understandable and adjustable parameters. For example, presenting a series of sliders or knobs on the HMI interface, such as a "Cyberpunk Degree" slider, a "Color Saturation" slider, a "Light and Shadow Contrast" slider, and a "Line Thickness" button. Alternatively, providing a set of preset style variations, such as "Cyberpunk - Rainy Night" or "Cyberpunk - Dusk," allows users to switch between different image styles with a single click. For advanced creations, allowing users to input text prompts, such as "Add more neon lights" or "Make the buildings more dilapidated," provides more precise control over the style. Second, applying the user's parameter adjustments to the 3D scene to be captured in real time. For example, adjusting post-processing shaders. Most style parameters, such as color, contrast, halo, and vignetting, can be achieved by adjusting the post-processing shader parameters at the end of the rendering pipeline in real time. Furthermore, batch modification of object material properties is possible. For example, increasing the "cyberpunk level" might uniformly increase the intensity of the self-illuminating textures of all vehicle models, or change their metallic and roughness values. Image stylization can be achieved using image generation models. For instance, when a user inputs a text prompt, the system can use the rendering result of the current viewpoint and the text prompt as input to an image-to-image generation model (such as the ControlNet model). The model will then redraw the scene and return an updated image; this process can be completed in the cloud or on edge computing units.

[0081] The image capture method provided in this disclosure determines the capture image based on the adjusted camera pose and scene data of the three-dimensional scene to be captured; in response to receiving an instruction to stylize the captured image, the captured image is stylized to obtain the target captured image. This allows for fine-tuning and re-creation of AI-generated artistic styles, thereby providing users with a deeper level of personalized experience.

[0082] In some embodiments, step S23, determining the target image based on the adjusted camera pose and scene data of the three-dimensional scene to be captured, includes: determining rendering instructions based on the adjusted camera pose and scene data of the three-dimensional scene to be captured; sending the rendering instructions to the vehicle's graphics processor, enabling the graphics processor to create an off-screen rendering target, and performing rendering operations based on the off-screen rendering target and the rendering instructions to obtain the target image; wherein the resolution of the off-screen rendering target is higher than the display resolution of the vehicle's display screen.

[0083] For example, in this embodiment of the present disclosure, after determining the rendering instructions, the rendering instructions can be sent to the vehicle's graphics processor, enabling the graphics processor to render the image at a resolution higher than the vehicle's display screen in an off-screen rendering target. Here, the off-screen rendering target refers to a specific, invisible buffer located in the GPU's video memory, used to hold and store the rendering output results. In other words, to avoid affecting the user's interactive experience and to output images at any resolution, the image rendering process in this embodiment of the present disclosure can be performed in an off-screen FBO (Frame Buffer Object). Thus, high-quality images that meet user needs can be determined through high-resolution off-screen rendering.

[0084] In some examples, the off-screen rendering target can be set to a higher resolution to ensure the output image has sufficient detail to meet the user's quality requirements for the captured image. Simultaneously, to further optimize rendering performance, the system can dynamically adjust the resolution and sampling rate of the off-screen rendering target according to actual needs. For example, a lower resolution can be selected for rendering when a quick preview is needed; while in the final output stage, a high-resolution mode is switched to obtain the best results. In some embodiments, the system also supports real-time feedback on the rendering results. For example, when the user adjusts the camera pose or stylization parameters, the system updates the image content in the rendering preview window in real time, allowing the user to intuitively see the adjustment effect. This real-time interactive method not only improves creative efficiency but also reduces operational complexity.

[0085] In some examples, the creation process for off-screen rendering targets may include allocating independent frame buffers and texture resources and binding them to the vehicle's graphics processor's rendering pipeline. In this way, the system can perform high-quality image rendering in the background, regardless of the display resolution. For example, when the user selects ultra-high definition output mode, the system automatically adjusts the resolution parameters of the off-screen rendering target to ensure that the final generated target image has higher clarity and detail.

[0086] In some examples, the system can determine which objects need their state information updated in real time based on the serialized dynamic data and embed this information into the rendering instructions. Meanwhile, for static data, the system can pre-calculate its lighting and shadow information and cache it for subsequent rendering. This layered processing approach not only reduces computational overhead during rendering but also ensures the visual consistency and stability of the captured images.

[0087] In some embodiments, the system automatically enhances rendering quality-related parameters during final rendering. For example, increasing resolution: setting the resolution of the off-screen rendering target to a size significantly higher than the screen resolution, such as 4K (3840x2160) or 8K (7680x4320). Another example is improving anti-aliasing levels: enabling or enhancing the quality level of multi-sampling anti-aliasing or temporal anti-aliasing to achieve smoother object edges. Yet another example is adjusting shadow quality and ray tracing: increasing the resolution of shadow maps, enabling more advanced soft shadow algorithms, or, where hardware supports it, performing a small amount of ray tracing to achieve more realistic reflections and ambient occlusion effects.

[0088] In some embodiments, the rendered high-resolution image can be encoded and saved to the user's storage space. Specifically, image format conversion can be performed first, such as reading the raw pixel data from the GPU memory into the CPU (Central Processing Unit) memory, where the raw pixel data is typically in RGBA format. Then, it is encoded and compressed, for example, using an image encoding library (such as the libjpeg-turbo image encoding library) to encode the pixel data into a user-selected format (such as JPEG format). Furthermore, file saving and metadata embedding are implemented, such as writing the encoded image data to the file system and selectively writing metadata from the time of capture (such as virtual geolocation, time, camera parameters, style settings, etc.) into the image's EXIF ​​(Exchangeable Image File Format) information.

[0089] In some embodiments, before receiving the instruction to start shooting, the method further includes: saving scene data of the three-dimensional reconstructed scene of the vehicle's surrounding environment at the current time and previous times to determine the three-dimensional reconstructed scene sequence; Correspondingly, step S21, in response to receiving the instruction to start shooting, determines the three-dimensional reconstructed scene of the vehicle's surrounding environment at the current moment as the three-dimensional scene to be shot, including: in response to receiving the instruction to start shooting, displaying the timeline interactive interface corresponding to the three-dimensional reconstructed scene sequence to determine the target time; and determining the three-dimensional reconstructed scene in the three-dimensional reconstructed scene sequence that corresponds to the target time as the three-dimensional scene to be shot.

[0090] For example, before triggering the shooting mode, scene data of the 3D reconstructed scene at the current moment and several previous moments can be cached, and a timeline interface can be provided to receive the target moment selected by the user. For instance, after the user triggers the command to start shooting and enters the shooting mode, they can drag the function slider on the timeline interface to go back to the current moment and any moment within N seconds before the current moment to freeze the scene, and then perform composition and stylistic updates based on this. Here, N is a natural number greater than or equal to 1. In this way, the user can be provided with the ability to go back, thereby improving the convenience of image shooting.

[0091] In some embodiments, multiple first virtual cameras can be defined to enable multiple users to enter the same 3D reconstructed scene and perform independent camera movements and compositions through their respective electronic devices (such as mobile phones, tablets, or multiple displays in a vehicle). The multiple first virtual cameras may include only multiple newly created virtual cameras, or they may include newly created virtual cameras and virtual cameras decoupled from the vehicle; this disclosure does not impose limitations on these aspects. This enhances the in-vehicle entertainment and social experience, allowing multiple users to simultaneously participate in the shooting of the 3D reconstructed scene.

[0092] In some embodiments, AI-powered composition assistance can be added. For example, after a user enters photography mode, AI analyzes the currently saved 3D reconstructed scene (i.e., the frozen scene), identifies potential subjects (such as unique vehicles and buildings), and automatically recommends multiple optimal composition angles based on photographic aesthetic principles (such as the rule of thirds and the golden ratio). The user can then choose to adopt the optimal composition angle or adjust it based on the optimal composition angle. This reduces the difficulty of shooting for the user and improves the user's shooting experience.

[0093] In some embodiments, a switch option may be provided to selectively maintain the motion state of specific dynamic elements (such as distant clouds or falling snowflakes) in the currently reconstructed 3D scene to increase the vividness of the image. In this way, the appeal of the captured image can be enhanced while keeping the subject static.

[0094] Exemplary device Figure 7 An image capturing device provided in an embodiment of this disclosure, such as Figure 7 As shown, the image capturing device 70 includes a three-dimensional scene determination module 71, a camera pose adjustment module 72, and a target image determination module 73.

[0095] The 3D scene determination module 71 is used to determine the 3D reconstructed scene of the vehicle's surrounding environment at the current moment as the 3D scene to be captured in response to receiving the instruction to start shooting. The camera pose adjustment module 72 is used to adjust the camera pose of the first virtual camera in the three-dimensional scene to be captured in response to receiving an instruction to adjust the first virtual camera in the three-dimensional scene to be captured. The target image determination module 73 is used to determine the target image based on the adjusted camera pose and scene data of the three-dimensional scene to be captured.

[0096] In some embodiments, such as Figure 8 As shown, the 3D scene determination module 71 includes an interrupt triggering unit 711, an interrupt response unit 712, and a scene determination unit 713.

[0097] Interrupt triggering unit 711 is used to trigger an interrupt signal in response to receiving a command to start shooting; Interrupt response unit 712 is used to respond to an interrupt signal and stop acquiring perception data of the vehicle's surrounding environment after the current moment, so that the vehicle's three-dimensional reconstruction scene remains in the scene at the time the interrupt signal was triggered. Scene determination unit 713 is used to determine the current 3D reconstruction scene as the 3D scene to be photographed.

[0098] In some embodiments, the scene determination unit 713 is specifically used to determine the resource data corresponding to the scene where the vehicle is located at the current moment; to serialize the dynamic data in the resource data to obtain serialized dynamic data; to store the serialized dynamic data and the static data in the resource data to determine the three-dimensional reconstruction scene at the current moment; and to determine the three-dimensional reconstruction scene at the current moment as the three-dimensional scene to be photographed.

[0099] In some embodiments, the image capturing device 70 further includes a data storage module for storing scene data of the three-dimensional reconstructed scene of the vehicle's surrounding environment at the current moment and previous moments, so as to determine the three-dimensional reconstructed scene sequence. Correspondingly, the 3D scene determination module 71 is used to respond to the received instruction to start shooting by displaying the timeline interactive interface corresponding to the 3D reconstructed scene sequence to determine the target time; and to determine the 3D reconstructed scene in the 3D reconstructed scene sequence that corresponds to the target time as the 3D scene to be shot.

[0100] In some embodiments, such as Figure 9 As shown, the image capturing device 70 also includes a virtual camera determining module 74.

[0101] The virtual camera determination module 74 is used to create a first virtual camera based on the 3D reconstructed scene; or, determine a second virtual camera that is associated with the vehicle; remove the association between the second virtual camera and the vehicle, and use the second virtual camera after removing the association as the first virtual camera.

[0102] In some embodiments, the virtual camera determination module 74 is specifically used to create a first virtual camera based on the 3D reconstructed scene; or, determine a second virtual camera that is associated with the vehicle; determine the parent node of the second virtual camera based on the association; remove the parent node from the vehicle node of the vehicle; and add the removed parent node to the scene node of the 3D reconstructed scene.

[0103] In some embodiments, the target image determination module 73 is specifically used to determine the captured image based on the adjusted camera pose and scene data of the three-dimensional scene to be captured; in response to receiving an instruction to stylize the captured image, it performs stylization processing on the captured image to obtain the target captured image.

[0104] In some embodiments, the target image determination module 73 is specifically used to determine rendering instructions based on the adjusted camera pose and scene data of the three-dimensional scene to be captured; send the rendering instructions to the vehicle's graphics processor, enabling the graphics processor to create an off-screen rendering target, and perform rendering operations based on the off-screen rendering target and the rendering instructions to obtain a target image; wherein the resolution of the off-screen rendering target is higher than the display resolution of the vehicle's display screen.

[0105] The beneficial technical effects corresponding to the exemplary embodiment of the image capturing device 70 described above can be found in the corresponding beneficial technical effects in the exemplary method section above, and will not be repeated here.

[0106] Exemplary electronic devices Figure 10 This is a structural diagram of an electronic device 10 provided in an embodiment of the present disclosure. The electronic device 10 includes at least one processor 11 and a memory 12.

[0107] The processor 11 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.

[0108] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute one or more computer program instructions to implement the image acquisition methods and / or other desired functions in the vehicle 3D reconstruction scene of the various embodiments of this disclosure described above.

[0109] In one example, the electronic device 10 may also include an input device 13 and an output device 14, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0110] The input device 13 may also include, for example, a keyboard, a mouse, etc.

[0111] The output device 14 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0112] Of course, for the sake of simplicity, Figure 10 Only some of the components of the electronic device 10 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 10 may include any other suitable components depending on the specific application.

[0113] Exemplary computer program products and computer-readable storage media In addition to the methods and devices described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions, which, when executed by a processor, cause the processor to perform the steps of the image acquisition method in a vehicle 3D reconstruction scene described in the various embodiments of this disclosure in the "Exemplary Methods" section above.

[0114] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0115] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps of the image acquisition method in a vehicle 3D reconstruction scene described in the various embodiments of this disclosure in the "Exemplary Methods" section above.

[0116] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0117] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0118] Various modifications and variations can be made to this disclosure without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.

Claims

1. A method for capturing images in a vehicle 3D reconstruction scene, comprising: In response to receiving the instruction to start shooting, the three-dimensional reconstructed scene of the vehicle's surrounding environment at the current moment is determined as the three-dimensional scene to be shot; In response to receiving an instruction to adjust the first virtual camera in the three-dimensional scene to be photographed, the camera pose of the first virtual camera in the three-dimensional scene to be photographed is adjusted; Based on the adjusted camera pose and the scene data of the three-dimensional scene to be photographed, the target image is determined.

2. The method according to claim 1, wherein, The step of responding to receiving a command to start shooting, and determining the 3D reconstructed scene of the vehicle's surrounding environment at the current moment as the 3D scene to be shot, includes: In response to receiving a command to start recording, an interrupt signal is triggered; In response to the interrupt signal, the acquisition of perception data of the vehicle's surrounding environment after the current moment is stopped, so that the three-dimensional reconstruction scene of the vehicle remains in the scene at the time when the interrupt signal was triggered. The current 3D reconstruction scene is determined to be the 3D scene to be captured.

3. The method according to claim 2, wherein, Determining that the 3D reconstructed scene at the current moment is the 3D scene to be captured includes: Determine the resource data corresponding to the current scene where the vehicle is located; The dynamic data in the resource data is serialized to obtain the serialized dynamic data; The serialized dynamic data and the static data in the resource data are stored to determine the 3D reconstruction scene at the current moment; The 3D reconstructed scene at the current moment is determined as the 3D scene to be photographed.

4. The method according to claim 1, further comprising, before responding to receiving the instruction to start shooting: Save the scene data of the three-dimensional reconstructed scene of the vehicle's surrounding environment at the current moment and previous moments to determine the three-dimensional reconstructed scene sequence; Correspondingly, the step of responding to receiving the instruction to start shooting and determining the three-dimensional reconstructed scene of the vehicle's surrounding environment at the current moment as the three-dimensional scene to be shot includes: In response to receiving a command to start shooting, the timeline interactive interface corresponding to the 3D reconstructed scene sequence is displayed to determine the target time. The three-dimensional reconstructed scene in the three-dimensional reconstructed scene sequence that corresponds to the target time is determined as the three-dimensional scene to be photographed.

5. The method according to claim 1, further comprising, before adjusting the camera pose of the first virtual camera in the three-dimensional scene to be photographed in response to receiving an instruction to adjust the first virtual camera in the three-dimensional scene to be photographed: Based on the reconstructed 3D scene, the first virtual camera is created; or, Identify a second virtual camera that is associated with the vehicle; The association between the second virtual camera and the vehicle is severed, and the second virtual camera after the association is severed is used as the first virtual camera.

6. The method according to claim 5, wherein, The step of detaching the association between the second virtual camera and the vehicle, and using the second virtual camera after detachment as the first virtual camera, includes: Based on the aforementioned relationship, the parent node of the second virtual camera is determined; Remove the parent node from the vehicle node of the vehicle; Add the removed parent node to the scene node of the 3D reconstructed scene.

7. The method according to any one of claims 1-6, wherein, The process of determining the target image based on the adjusted camera pose and the scene data of the 3D scene to be captured includes: Based on the adjusted camera pose and the scene data of the three-dimensional scene to be photographed, the image to be captured is determined; In response to receiving an instruction to stylize the captured image, the captured image is stylized to obtain the target captured image.

8. The method according to any one of claims 1-6, wherein, The process of determining the target image based on the adjusted camera pose and the scene data of the 3D scene to be captured includes: Based on the adjusted camera pose and the scene data of the 3D scene to be captured, the rendering instructions are determined; The rendering command is sent to the graphics processor of the vehicle, enabling the graphics processor to create an off-screen rendering target and perform a rendering operation based on the off-screen rendering target and the rendering command to obtain the target image. The resolution of the off-screen rendering target is higher than the display resolution of the vehicle's display screen.

9. An image capturing device for a vehicle 3D reconstruction scene, comprising: The 3D scene determination module is used to determine the 3D reconstructed scene of the vehicle's surrounding environment at the current moment as the 3D scene to be captured in response to the received instruction to start shooting; A camera pose adjustment module is used to adjust the camera pose of the first virtual camera in the three-dimensional scene to be photographed in response to receiving an instruction to adjust the first virtual camera in the three-dimensional scene to be photographed. The target image determination module is used to determine the target image based on the adjusted camera pose and the scene data of the three-dimensional scene to be captured.

10. A computer-readable storage medium storing a computer program for performing the image capture method in a vehicle three-dimensional reconstruction scene according to any one of claims 1-8.

11. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the image capture method in the vehicle three-dimensional reconstruction scene according to any one of claims 1-8.