Three-dimensional reconstruction method and electronic equipment
By performing semantic parsing and generating multi-view 2D images from the in-vehicle HMI system, and combining this with 3D reconstruction technology, a 3D model associated with the target theme is automatically generated. This solves the problem that 2D display resources in the in-vehicle HMI system cannot provide a stereoscopic experience, and achieves efficient and low-cost 3D display resource generation.
Patent Information
- Application Number
- CN202511668991.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-24
AI Technical Summary
In existing technologies, 2D display resources generated by in-vehicle HMI systems cannot provide a stereoscopic 3D display environment experience, and manually creating 3D display resources has a long creation cycle and high cost.
By performing semantic parsing based on the target theme, multi-view 2D images matching the target style are generated, and a 3D model associated with the target theme is constructed using generative models and 3D reconstruction technology, thereby realizing the automatic generation of 3D display resources.
It improves the stereoscopic and immersive feel of 3D display resources, reduces generation costs, and increases generation efficiency.
Smart Images

Figure CN121564201A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer graphics technology, and more particularly to a three-dimensional reconstruction method and electronic device. Background Technology
[0002] With the rapid development of intelligent vehicles and in-vehicle infotainment systems, users' demands for personalized experiences from in-vehicle human-machine interface (HMI) systems are increasing. In-vehicle HMI systems can display environmental perception information and vehicle status information, such as destination landmarks, overpass models, and surrounding vehicles and pedestrians identified by sensors in navigation systems. Typically, in-vehicle HMI systems have several pre-set themes, each with several display resources. When a theme is activated, the in-vehicle HMI system can replace the displayed content with the display resources corresponding to that theme. Summary of the Invention
[0003] Currently, AI technology is often used to automatically generate 2D display resources or professional designers can manually generate 3D display resources. However, automatically generated 2D display resources cannot provide a stereoscopic 3D display environment experience, and manually creating 3D display resources has a long creation cycle and high cost.
[0004] To address the aforementioned technical problems, this disclosure provides a three-dimensional reconstruction method and electronic device that can automatically create a three-dimensional model matching the theme as a 3D display resource, thus resolving the problems existing in the current technology.
[0005] The first aspect of this disclosure provides a three-dimensional reconstruction method, which includes: determining a target style, a target scene, and a target entity in the target scene based on a target theme; generating two-dimensional images of the target entity from multiple preset viewpoints based on the target style; and generating a three-dimensional model corresponding to the target entity based on the two-dimensional images from multiple preset viewpoints; wherein the three-dimensional model is associated with the target theme.
[0006] A second aspect of this disclosure provides a three-dimensional reconstruction apparatus, comprising: a theme analysis module for determining a target style, a target scene, and a target entity in the target scene based on a target theme; a two-dimensional resource generation module for generating two-dimensional images of the target entity from multiple preset viewpoints based on the target style; and a three-dimensional resource generation module for generating a three-dimensional model corresponding to the target entity based on the two-dimensional images from multiple preset viewpoints; wherein the three-dimensional model is associated with the target theme.
[0007] A third aspect of this disclosure is that embodiments of this disclosure provide a computer-readable storage medium storing a computer program for performing the three-dimensional reconstruction method provided in the first aspect.
[0008] In a fourth aspect of this disclosure, embodiments of this disclosure provide an electronic device comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to read executable instructions from the memory and execute the executable instructions to implement the three-dimensional reconstruction method provided in the first aspect.
[0009] Based on the 3D reconstruction method provided in this disclosure, semantic parsing of the target topic can be performed to determine the relevant target style, target scene, and target entity. Then, multi-view 2D images matching the target style are generated for the target entity. Finally, a 3D model associated with the target topic is constructed and used as a 3D display resource, thereby achieving automatic generation of 3D display resources. Compared to traditional 2D display resources, the 3D display resources generated by the embodiments of this disclosure can improve the sense of depth and immersion. Compared to the traditional method that relies on manual generation, the method provided by the embodiments of this disclosure can improve the generation efficiency of 3D display resources and reduce costs. Attached Figure Description
[0010] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0011] Figure 1 This is an application scenario diagram of the three-dimensional reconstruction system provided in an exemplary embodiment of this disclosure;
[0012] Figure 2 This is a schematic flowchart of a three-dimensional reconstruction method provided in an exemplary embodiment of this disclosure;
[0013] Figure 3 This is a flowchart illustrating a three-dimensional reconstruction method provided in another exemplary embodiment of this disclosure;
[0014] Figure 4 This is a flowchart illustrating a three-dimensional reconstruction method provided in yet another exemplary embodiment of this disclosure;
[0015] Figure 5 This is a flowchart illustrating the three-dimensional reconstruction method provided in the fourth exemplary embodiment of this disclosure;
[0016] Figure 6 This is a flowchart illustrating the three-dimensional reconstruction method provided in the fifth exemplary embodiment of this disclosure;
[0017] Figure 7This is a flowchart illustrating the three-dimensional reconstruction method provided in the sixth exemplary embodiment of this disclosure;
[0018] Figure 8 This is a flowchart illustrating the three-dimensional reconstruction method provided in the seventh exemplary embodiment of this disclosure;
[0019] Figure 9 This is a flowchart illustrating the three-dimensional reconstruction method provided in the eighth exemplary embodiment of this disclosure;
[0020] Figure 10 This is a flowchart illustrating the three-dimensional reconstruction method provided in the ninth exemplary embodiment of this disclosure;
[0021] Figure 11 This is a schematic diagram of the structure of a three-dimensional reconstruction apparatus provided in an exemplary embodiment of the present disclosure;
[0022] Figure 12 This is a schematic diagram of the structure of a three-dimensional reconstruction apparatus provided in another exemplary embodiment of this disclosure;
[0023] Figure 13 This is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed Implementation
[0024] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0025] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0026] Application Overview
[0027] With the rapid development of intelligent vehicles and in-vehicle infotainment systems, users' demands for personalized experiences from in-vehicle HMI systems are increasing. In-vehicle HMI systems enable interaction between users and vehicle systems. In some scenarios, in-vehicle HMI systems can display environmental perception information and vehicle status information. For example, when a vehicle approaches its destination, the in-vehicle HMI system can display landmarks from the navigation system; furthermore, when the vehicle's sensors detect surrounding vehicles and pedestrians, the in-vehicle HMI system can visually display the perceived objects.
[0028] To enrich visual expression and enhance the user's visual experience, the content displayed by the in-vehicle HMI system can typically be personalized based on different visual themes. Specifically, in-vehicle HMI systems usually have several pre-set themes, such as sports themes and holiday-themed themes. Each theme includes several display resources, such as car models and pedestrian models. When a theme is activated, the in-vehicle HMI system can utilize the corresponding display resources to display the desired content. For example, displaying pedestrian models on the interactive interface to show pedestrians detected by sensors.
[0029] Therefore, to support flexible configuration of visual themes, a large number of display resources need to be pre-generated. Currently, AI technology can be used to automatically generate 2D display resources, but the visual effect of 2D display resources is flat and cannot provide a three-dimensional 3D environment experience. In some scenarios with high requirements for visual effects, 3D display resources can be manually created by professional designers, but this is time-consuming and costly. Therefore, there is an urgent need for a method that can quickly and automatically generate 3D display resources.
[0030] This disclosure proposes a novel solution to address the aforementioned problems in current display resource generation technologies. Specifically, this disclosure provides a 3D reconstruction method that performs semantic analysis on a target subject to determine relevant target styles, scenes, and entities. It then generates multi-view 2D images matching the target style for each entity, ultimately constructing a 3D model associated with the target subject and using it as a 3D display resource, thereby achieving automatic generation of 3D display resources. Compared to traditional 2D display resources, the 3D display resources generated by this disclosure improve stereoscopic depth and immersion. Compared to traditional methods that rely on manual generation, the method provided by this disclosure improves the efficiency of 3D display resource generation and reduces costs.
[0031] Exemplary System
[0032] Figure 1 This is an application scenario diagram of a three-dimensional reconstruction system provided by an exemplary embodiment of this disclosure.
[0033] like Figure 1 As shown, in one embodiment, the 3D reconstruction system 10 and the vehicle-mounted HMI system 20 transmit data, and the vehicle-mounted HMI system 20 and the perception system 30 transmit data.
[0034] First, we will introduce the deployment methods of the 3D reconstruction system 10, the vehicle-mounted HMI system 20, and the perception system 30.
[0035] In some implementations, the 3D reconstruction system 10 can be deployed in the cloud, while the in-vehicle HMI system 20 and perception system 30 can be deployed in the vehicle. In this way, the reconstruction and storage of the 3D model can be completed using the cloud, and the vehicle only needs to load and render the 3D model. This reduces the hardware requirements of the vehicle, supports the generation of more complex 3D models, and facilitates the unified management and updating of 3D models.
[0036] In other implementations, the module for generating 3D models in the 3D reconstruction system 10 can be lightweightly deployed in an in-vehicle edge device. This allows the in-vehicle edge device to independently complete 3D model reconstruction, reducing network dependence between the vehicle and the cloud and supporting use cases where 3D models are generated offline.
[0037] In some implementations, the module in the 3D reconstruction system 10 used to store the 3D model can be deployed in the vehicle. This eliminates the need for the in-vehicle HMI system to download the 3D model from the cloud each time, improving rendering speed and adapting to different network environments.
[0038] Next, we will introduce the internal structure of the 3D reconstruction system 10, the vehicle-mounted HMI system 20, and the perception system 30.
[0039] The 3D reconstruction system 10 may include a processor 101 and a memory 102. The memory 102 can be used to store program instructions executable by the processor 101. The processor 101 can load and execute the program instructions in the memory 102 to implement the 3D reconstruction method. The memory 102 can also be used to store data, such as 2D images and 3D models generated by the processor 101.
[0040] The perception system 30 may include various sensors, such as cameras, radar, and lidar, which can be used to perceive objects around the vehicle in real time, such as surrounding vehicles, pedestrians, and road signs. The perception system 30 can transmit the perceived object information to the vehicle-mounted HMI system 20. The vehicle-mounted HMI system 20 selects the corresponding 3D model from the memory 202 according to the currently active display theme and displays it on the display screen 203, providing users with a more three-dimensional and realistic interactive experience.
[0041] The in-vehicle HMI system 20 may include a processor 201, a memory 202, and a display screen 203. The memory 202 can be used to store program instructions executable by the processor 201. The processor 201 can load and execute the program instructions in the memory 202. For example, based on the perceived object information transmitted by the perception system 30, it can determine the 3D model corresponding to the perceived object from the 3D model generated by the 3D reconstruction system 10 that corresponds to the currently active display theme, and then visualize it on the display screen 203.
[0042] In some implementations, the memory 202 can also be used to store all or part of the 3D model generated by the 3D reconstruction system 10. In this way, when the in-vehicle HMI system displays 3D content, it does not need to obtain the 3D model from the 3D reconstruction system 10, which helps reduce data transmission latency of the 3D model and improves display efficiency.
[0043] In some implementations, the aforementioned processors 101 and 201 can be single-core processors or multi-core processors; they can include general-purpose processors, such as central processing units (CPUs) and graphics processing units (GPUs), or accelerated computing units, such as neural processing units (NPUs), or dedicated processors, such as ASICs and FPGAs.
[0044] In some implementations, the aforementioned memory 102 and memory 202 may include volatile memory, such as dynamic random access memory (DRAM) and static random access memory (SRAM); and may also include non-volatile memory (NVM), such as read-only memory (ROM) and flash memory.
[0045] Exemplary methods
[0046] Figure 2 This is a schematic flowchart of a three-dimensional reconstruction method provided in an exemplary embodiment of this disclosure. Figure 2 As shown, the method includes the following steps:
[0047] Step 100: Based on the target theme, determine the target style, target scene, and target entities within the target scene.
[0048] In step 100, the 3D reconstruction device can use large language models, etc., to perform semantic analysis on the target topic, determine the target style and target scene that match the target topic, as well as the target entities that usually appear in the target scene.
[0049] In some implementations, the target style and target scene can be determined first based on the target theme. Then, by combining the target scene with a predefined in-vehicle environment ontology knowledge base, target entities within the target scene can be identified. The in-vehicle environment ontology knowledge base defines common entity types in various in-vehicle environments and can also define the relationships and probabilities of occurrence between different entity types. For example, the in-vehicle environment ontology knowledge base can define common entity types in urban environments, such as cars and buildings. For instance, semantic parsing of the target theme "cyberpunk city" can determine the target style "cyberpunk" and the target scene "urban." Then, based on the target scene and the in-vehicle environment ontology knowledge base, target entities such as "cars" and "buildings" can be identified. This ensures that the extracted entities are relevant to the target scene, avoiding the appearance of entities like "farmland" in an "urban" scene.
[0050] It should be noted that in some implementations, users can also input text containing the target theme, such as "Please generate rendering instructions based on Cyberpunk City". In this case, the large language model can directly parse the input text to determine the target theme, target style, target scene, and target entities.
[0051] In some implementations, prompt words can be generated based on the target style, target scene, and target entity extracted from the target theme, and then used to generate a 2D image in subsequent steps. For example, the target style, target scene, and target entity can be used as parameters to fill in a preset prompt word template that includes weights and grammatical structures. The template design can follow best practices of specific diffusion models (such as Stable Diffusion), such as using weighted tokens to emphasize key visual features or using negative prompt words to exclude unwanted elements (such as blurry or low-resolution). For example, using the aforementioned "cyberpunk," "urban," and "car" as parameters and inputting them into the prompt word template, the following prompt words are obtained: "Best image quality, cyberpunk style, futuristic car. Negative prompts: low resolution, blurry." Subsequently, the large language model can use the constructed prompt words as rendering instructions, and then in subsequent steps, the generative model can generate a 2D image of a cyberpunk-style car based on the rendering instructions, while avoiding the generation of low-resolution or blurry 2D images.
[0052] In some implementations, a separate prompt word can be created for each target entity, allowing for the generation of precise prompt words for each entity individually. In other implementations, multiple target entities corresponding to the same target scene can be integrated into a single prompt word.
[0053] It should be noted that the embodiments disclosed herein use Chinese prompt words as an example, which does not mean that the solution disclosed herein can only be used for Chinese prompt words. In some other scenarios, prompt words in other languages can also be generated.
[0054] In some implementations, at least one target detail can be determined based on the target style. For example, a large language model can be used to determine a set of detail descriptors based on the target style "cyberpunk". Each detail descriptor contains one aspect of the target detail, such as: material ("chrome alloy, carbon fiber"), lighting ("volume lighting, neon reflection"), and geometry ("angular, streamlined"). These detail descriptors can then be added to the aforementioned cue word template to generate the final, specific rendering instructions.
[0055] In some implementations, the 3D reconstruction device can utilize a large language model to perform the above operations. For example, it can first utilize the few-shot learning capability of the large language model, or construct a cue word chain through frameworks such as LangChain. A cue word chain is a technique using a large language model that can decompose a complex task into a series of consecutive, simple subtasks, using the output of the previous subtask as the input of the next, thereby guiding the large language model step by step to achieve the final goal. In the embodiments of this disclosure, the cue word chain can be used to first execute the first subtask, performing semantic parsing on the input text containing the target topic to extract the target style, target scene, and target entity; then, based on the target style, target scene, and target entity output from the first subtask, the second subtask is executed, filling in a preset cue word template and integrating detail descriptors to obtain the final cue word.
[0056] Step 200: Based on the target style, generate two-dimensional images of the target entity from multiple preset viewpoints.
[0057] In step 200, the 3D reconstruction device can utilize generative models, such as diffusion models or generative adversarial networks (GANs), to generate 2D images of the target entity from multiple preset viewpoints that match the target style. These preset viewpoints may include views from the front, side, rear, and 45° of the target entity.
[0058] In some implementations, a set of camera pose parameters can be pre-set for each preset viewpoint. These parameters can be vectors defined in the format of [azimuth, elevation, distance]. For example, for a frontal viewpoint, the camera pose parameters can be set to [0°, 0°, 5], indicating that the camera is directly in front of the target entity, at a horizontal line of sight, and at a distance of 5 units (e.g., meters). Multiple sets of camera pose parameters and the prompts generated in the preceding steps can then be input into a generative model (such as SyncDreamer, Zero-1-to-3, or other specific models). The generative model precisely controls the viewpoint of each 2D image, resulting in 2D images for each preset viewpoint, such as a front view, side view, rear view, and a 45-degree angle view.
[0059] In some implementations, the generative model can employ a specific model architecture, such as Zero-1-to-3 models, to ensure consistency in appearance and identity between 2D images from different preset viewpoints. Specifically, this can be achieved using shared attention mechanisms or latent spaces within the specific model architecture, ensuring that the same target entity is referenced when generating 2D images from multiple preset viewpoints. For example, the generative model can use the detailed descriptors of the aforementioned "car" entity, such as its geometric shape ("angular, streamlined"), as a shared latent vector when generating 2D images from all viewpoints. In this way, the target entity in different preset viewpoints shares the same set of visual features, ensuring that it is the same object viewed from any angle.
[0060] In some implementations, the generative model can integrate a 3D perception generator. This allows the generative model to possess basic 3D geometric perception capabilities, enabling it to understand the 3D structure of the target entity when generating 2D images. This allows it to naturally handle lighting, shadows, and occlusion relationships from different preset viewpoints, generating a set of geometrically more accurate and coherent 2D images.
[0061] Step 300: Generate a three-dimensional model corresponding to the target entity based on two-dimensional images from multiple preset perspectives.
[0062] In step 300, after generating multiple 2D images of the target entity from preset viewpoints, 3D reconstruction can be performed based on the 2D images to generate a corresponding 3D model, which is then used as a 3D display resource. It can be understood that since the target entity is determined based on a target theme, the 3D model corresponding to the target entity is associated with the target theme. For example, for the "car" entity corresponding to the "cyberpunk city" theme, a 3D model of the "car" entity can be generated based on its front view, side view, rear view, and 45-degree angle view, and this 3D model is associated with the cyberpunk theme.
[0063] As can be seen from the above technical solutions, the method provided in this disclosure can utilize a large language model to perform semantic parsing and prompt word construction on a target topic, obtaining prompt words for generating two-dimensional images corresponding to the target entity; then, it uses a generative model to generate multi-view two-dimensional images of the target entity, and finally constructs a three-dimensional model associated with the target topic and uses it as a 3D display resource, thereby realizing the automatic generation of 3D display resources. Compared with traditional 2D display resources, the 3D display resources generated by this disclosure can improve the sense of stereoscopicity and immersion; compared with the traditional method that relies on manual generation, the method provided by this disclosure can improve the generation efficiency of 3D display resources and reduce costs.
[0064] Figure 3 This is a schematic flowchart of a three-dimensional reconstruction method provided in another exemplary embodiment of this disclosure. Figure 3 As shown above, in the above Figure 2 Based on the illustrated embodiment, step 300 may include the following steps:
[0065] Step 310: Generate a triangular mesh structure corresponding to the target entity based on two-dimensional images from multiple preset viewpoints.
[0066] In step 310, a neural radiation field or Gaussian sputtering model can be used to generate the triangular mesh structure corresponding to the target entity. For example, a neural radiation field or Gaussian sputtering model can be trained or optimized using frameworks such as PyTorch or TensorFlow. Then, the neural radiation field or Gaussian sputtering model can be used to extract the triangular mesh structure corresponding to the target entity using algorithms such as Marching Cubes.
[0067] Furthermore, structured latent variable 3D generation technology can be used to generate triangular mesh structures corresponding to target entities. For example, features can first be extracted and fused from 2D images under multiple preset viewpoints and encoded as structured latent variables. Then, the decoder reconstructs the triangular mesh structure corresponding to the target entity based on the latent variables.
[0068] In addition, if there is only one viewpoint of a two-dimensional image, a single-image reconstruction technique can be used. The three-dimensional geometry of the target entity can be directly inferred by using a deep learning model based on depth map prediction (such as depth estimation models such as ZoeDepth) or shape priors, thereby obtaining the triangular mesh structure corresponding to the target entity.
[0069] In some implementations, the triangular mesh structure can be optimized after generation. For example, functions in 3D processing libraries such as PyTorch3D and Open3D can be used to perform post-processing operations such as simplification, smoothing, and UV unwrapping on the mesh to obtain a topology-optimized triangular mesh structure suitable for real-time rendering engines.
[0070] For example, for the "cyberpunk city" target theme in the aforementioned steps, after generating multiple two-dimensional images of the "car" target entity from multiple preset perspectives, the generated multiple two-dimensional images and the camera pose parameters corresponding to each two-dimensional image, such as [0°, 0°, 5] mentioned above, can be input into the Gaussian sputtering model. The Gaussian sputtering model is used to obtain the car composed of point clouds, and then the Marching Cubes algorithm is used to extract the triangular mesh structure corresponding to the "car" target entity from the point cloud.
[0071] Step 320: Generate texture maps corresponding to the triangular mesh structure based on the target style.
[0072] In step 320, after generating the triangular mesh structure, a corresponding texture map can be generated based on the target style. The style of the texture map can match the aforementioned target style. For example, based on the aforementioned generated model or by calling the automated API of professional material processing software such as Substance 3D, a complete set of physically based rendering texture maps (hereinafter referred to as the high-poly texture map set) can be generated for the triangular mesh structure. The texture map set can include texture maps from multiple aspects such as basic color texture maps, metallic texture maps, roughness texture maps, and normal texture maps.
[0073] In some implementations, considering the performance limitations of automotive hardware, the triangular mesh structure can be simplified through methods such as polygon reduction, generating a texture map set (hereinafter referred to as the low-poly texture map set) for the simplified triangular mesh structure. The low-poly texture map set can use a lower resolution than the high-poly texture map set. Then, the details of each texture map in the high-poly texture map set can be baked onto each texture map in the low-poly texture map set to obtain the final texture map corresponding to the triangular mesh structure.
[0074] In some implementations, considering the performance limitations of automotive hardware, after obtaining the texture map corresponding to the triangular mesh structure, OpenImageIO or similar libraries can be used to perform GPU-optimized compression of the texture in formats such as ASTC or BCn to reduce memory usage and bandwidth requirements, thereby ensuring real-time rendering performance on automotive hardware.
[0075] For example, for the triangular mesh structure corresponding to the aforementioned "car" target entity, the 3D reconstruction device can generate a set of texture maps for the triangular mesh structure through Substance 3D's automated API. Specifically, the body can be defined as silver using a base color texture map, as metallic paint using a metallic texture map, as smooth paint and rougher surfaces such as tires using a roughness texture map, and as detailed textures on the doors using a normal texture map.
[0076] Step 330: Map the texture map onto the triangular mesh structure to form the 3D model corresponding to the target entity.
[0077] In step 330, the generated texture map is mapped onto a triangular mesh structure. In this way, the triangular mesh structure ensures the geometry of the 3D model, and the texture map ensures the visual effect of the 3D model, thus obtaining a 3D model with information such as color, metallicity, and roughness.
[0078] For example, for the aforementioned "car" target entity, its corresponding texture map set can be mapped onto a triangular mesh to obtain a complete cyberpunk-style 3D car model.
[0079] In some implementations, after mapping texture maps to triangular mesh structures, materials can be automatically built and configured for the 3D model within the 3D engine. This allows the 3D model to possess material information, such as emissivity and transparency.
[0080] Specifically, material instances can be automatically created and configured in 3D engines such as Unity and Unreal Engine using automated scripting tools. Then, based on preset rules, the aforementioned texture maps can be automatically connected to the corresponding input nodes of the shaders, and the material parameters can be automatically adjusted to ensure that the visual style of the 3D model matches the original concept design.
[0081] For example, for the aforementioned 3D car model, an "autopainting" material instance can be created and configured using an automated scripting tool. Then, the texture map is connected to the input node corresponding to the shader of the "autopainting" material instance, and information from the texture map is read through the input node. Furthermore, according to preset rules for a "cyberpunk" style, the material parameters of the "autopainting" material instance, such as the self-illumination intensity value, are adjusted. In this way, by applying the "autopainting" material instance to the 3D car model, the shader can render the 3D car model based on the information read from the texture map and the adjusted material parameters, resulting in a visual effect where the car paint of the 3D car model matches the original concept design.
[0082] As can be seen from the above technical solutions, the method provided in this disclosure can reconstruct a three-dimensional triangular mesh from a two-dimensional image using techniques such as neural radiation fields or Gaussian sputtering, ensuring the accuracy of the geometric shape of the three-dimensional model. Subsequently, a texture map is generated and mapped onto the triangular mesh, thereby adding visual effects to the mesh to obtain a complete three-dimensional model. Furthermore, this disclosure, through texture baking and compression, ensures that the rendering of the three-dimensional model meets the performance constraints of the automotive hardware.
[0083] Figure 4 This is a flowchart illustrating a three-dimensional reconstruction method provided in yet another exemplary embodiment of this disclosure. For example... Figure 4 As shown above, in the above Figure 3 Based on the illustrated embodiment, step 310 may be followed by steps 340-350, and step 330 may include step 331:
[0084] Step 340: In response to the target entity being a biological entity, determine the skeletal structure of the target entity.
[0085] In step 340, if the target entity is a biological entity, it can be skeletally bound and skinned to add dynamic effects. The biological entity can be a pedestrian, animal, or other animal-like object.
[0086] The first step is to determine the skeleton structure of the target entity. For example, for the triangular mesh structure corresponding to the target entity, a pre-trained model (such as SMPL or a similar human parametric model) can be used to predict and match a standard skeleton structure. During this process, the Blender Python API can be used to call its built-in Auto-Rigging toolchain or similar automated binding tools to automate the determination of the skeleton structure.
[0087] For example, for the pedestrian triangular mesh structure corresponding to the "pedestrian" entity, the pedestrian triangular mesh structure can be analyzed by a pre-trained model to identify the positions of the knee joint, elbow joint, etc., and then generate a pedestrian skeleton structure containing bones.
[0088] Step 350: Based on the influence range of each bone in the skeleton structure on each vertex in the triangular mesh structure, determine the skinning weight corresponding to each vertex.
[0089] In step 350, using algorithms such as heatmap diffusion or geometric voxel binding, the skinning weight is automatically assigned to each vertex by calculating the influence range of each bone in the skeleton structure on each vertex in the triangular mesh structure. In this way, a smooth and natural joint deformation effect can be generated.
[0090] For example, for a pedestrian skeleton structure, a heatmap diffusion method can be used, treating the femur as a heat source. Its heat diffuses along the surface of the triangular mesh. Vertices 'a' closer to the femur can be assigned higher skinning weights, indicating that vertex a is more influenced by the femur; while vertices 'b' farther from the femur can be assigned lower skinning weights, indicating that vertex b is more influenced by the femur. This ensures that the thighs and calves deform naturally during pedestrian leg movements, and that the folds in the trousers change appropriately with leg movement.
[0091] Step 331: Map the texture map to the triangular mesh structure, and based on the skin weight, redirect the animation corresponding to at least one preset action to the skeleton structure to generate at least one dynamic 3D model corresponding to the target entity.
[0092] In step 331, based on skinning weights, the animation corresponding to the preset action is redirected to the skeleton structure bound to the 3D model, thus giving the 3D model a dynamic effect. The animation corresponding to the action can be an animation sequence, such as a walking animation sequence, a running animation sequence, etc.
[0093] For example, an animation library can be pre-set, containing multiple preset animation sequences corresponding to actions, such as a walking animation sequence. When it is necessary to generate a walking dynamic 3D model corresponding to a "pedestrian" entity, a standard BVH or FBX format walking animation sequence can be extracted from the animation library. Then, based on skinning weights, the walking animation sequence is mapped onto the pedestrian skeleton structure using techniques such as animation retargeting, thereby obtaining a walking dynamic 3D model corresponding to the "pedestrian" entity.
[0094] As can be seen from the above technical solutions, the method provided in this disclosure can automatically add skeletal binding to the corresponding 3D model of the biological entity; and use a skinning weight algorithm to ensure that the 3D model can deform naturally and smoothly during movement; finally, through animation redirection technology, standard animation sequences are redirected to the 3D model. In this way, the 3D model can have dynamic effects such as walking and running, providing users with a more vivid visual experience.
[0095] Figure 5 This is a flowchart illustrating the three-dimensional reconstruction method provided in the fourth exemplary embodiment of this disclosure. For example... Figure 5 As shown above, in the above Figure 2 Based on the illustrated embodiment, the following steps may be included after step 300:
[0096] Step 400: Perform quality inspection on the 3D model based on the items to be inspected.
[0097] In step 400, after generating the 3D model, the 3D model can be quality inspected for multiple dimensions of the items to be inspected, so as to ensure the practicality and aesthetics of the 3D model and make the quality of the 3D model meet the expected standards.
[0098] In some implementations, the items to be inspected in 3D model quality inspection may include geometric rationality inspection. Geometric rationality inspection may include mesh integrity analysis and physical property verification, which can be used to check the geometric structural integrity and physical laws of the 3D model, respectively.
[0099] In the mesh integrity analysis stage, libraries such as Trimesh or PyMeshFix can be used to automatically detect and attempt to repair geometric errors in the model, such as non-manifold edges, self-intersecting surfaces, internal holes, or inconsistent normal directions, thereby ensuring the stability of the 3D model in rendering and physical simulation.
[0100] In the physical property verification stage, proxy collision bodies can be automatically generated for 3D models, and their physical properties (such as mass and center of gravity) can be checked to see if they are within a reasonable range. In addition, for 3D models such as vehicles, it can also be verified whether their dimensions meet the preset vehicle category constraints (e.g., length, width, and height errors are less than 5%).
[0101] In some implementations, the items to be inspected in 3D model quality inspection may include visual quality inspection. Visual quality inspection may include perceptual metrics, style consistency scores, and texture sharpness inspection, which can be used to evaluate the visual quality of 3D models.
[0102] In the perception measurement stage, metrics such as LPIPS (Learned Perceptual Image PatchSimilarity) or SSIM (Structural Similarity Index) can be used. The 3D model is rendered from multiple preset viewpoints, and the rendered images are compared with the original 2D images generated from these preset viewpoints to quantitatively evaluate the visual similarity score between the rendered and 2D images. This visual similarity score can then be compared with a pre-set first quality control threshold. If the similarity score is lower than the first quality control threshold, the 3D model and the 2D images are considered to have significant visual differences, indicating insufficient fidelity in the 3D model, and thus the perception measurement of the 3D model is deemed unqualified. Conversely, if the similarity score is higher, the perception measurement of the 3D model is deemed qualified. In some examples, the first quality control threshold can be set to 0.1.
[0103] In the style consistency scoring stage, pre-trained visual language models, such as Contrastive Language-Image Pre-training (CLIP) models, can be used to calculate the cosine similarity between the rendered image of the 3D model and the textual description of the target topic (e.g., "cyberpunk"). This cosine similarity can then be compared to a pre-set second quality control threshold (e.g., 0.85). If the cosine similarity is higher than the second quality control threshold, the visual style of the 3D model is considered to have a high degree of matching with the target topic, and the style consistency test of the 3D model can be deemed satisfactory. In some examples, the second quality control threshold can be set to 0.85.
[0104] In the texture clarity detection stage, image processing analysis such as the Laplacian operator can be performed on the texture area of the rendered result to detect whether there are blurry, low-resolution or distorted textures, thereby ensuring the clarity of the visual details of the 3D model.
[0105] In some implementations, the items to be tested in 3D model quality inspection may include vehicle display environment adaptability testing. Vehicle display environment adaptability testing may include performance budget analysis, display environment simulation, frustum culling, and multi-level detail testing, which can be used to verify whether the resolution, contrast, and other properties of the 3D model meet the requirements of the vehicle display environment.
[0106] In the performance budget analysis phase, the number of vertices, triangles, materials, and required draw calls of the 3D model can be analyzed. This data is then compared with the performance budget of the target automotive hardware platform to determine whether the performance budget test is satisfactory. In some examples, 3D models with Vertex Count < 50k, Triangle Count < 100k, materials < 5, and required draw calls < 10 can be considered to have passed the performance budget test.
[0107] In the display environment simulation stage, the 3D model can be rendered in a simulated in-vehicle display environment to check the readability and visual comfort of the 3D model under different lighting conditions (such as strong light during the day and weak light at night), whether the contrast meets driving safety standards (such as WCAGAA level and other industry standards), and the performance of the 3D model on a high dynamic range (HDR) display, such as color performance and brightness performance.
[0108] The frustum culling and multi-level detail (LOD) testing stages verify whether the model can be effectively processed by the vehicle rendering engine's frustum culling system and check whether the corresponding Level of Detail (LOD) model of the 3D model can smoothly switch at the expected distance to ensure rendering frame rate in complex scenes. In some examples, it can be checked whether the LOD model can switch to the first level of detail at 20 meters and to the second level of detail at 50 meters.
[0109] In some implementations, the items to be inspected in 3D model quality inspection may include vehicle safety standard checks. Vehicle safety standard checks may include driver distraction detection and content compliance filtering, which can be used to ensure that the 3D model complies with vehicle safety standards.
[0110] In the driver attention distraction detection stage, image analysis algorithms can be used to detect whether there are visual elements in the 3D model that may excessively attract the driver's attention, such as flashing with a frequency higher than 2Hz or the overuse of highly saturated red warning colors, to ensure that the 3D model does not affect driving safety.
[0111] In the content compliance filtering stage, deep learning classifiers can be used to scan the texture and shape of 3D models, automatically identify and filter inappropriate (Not Safe / Suitable For Work, NSFW), offensive, or copyright-risking content, thereby ensuring that 3D models comply with local laws and regulations and brand safety requirements.
[0112] In this way, by performing tests on the aforementioned items to be tested, the quality inspection results corresponding to the 3D model can be obtained. For example, when the 3D model fails any test step among any of the aforementioned items to be tested, the 3D reconstruction device can record the failed test item, test step, as well as the specific value detected in that test step and the reason for the failure in the quality inspection report.
[0113] In some implementations, quality inspection reports can be pushed into the Continuous Integration and Continuous Deployment (CI / CD) process. This way, if a quality inspection fails, the CI / CD process or local scheduler can automatically adjust the generation parameters of the corresponding modules in the 3D reconstruction apparatus based on the quality inspection report. For example, if the geometric rationality check fails, an instruction can be issued to the module used to build the triangular mesh structure, instructing it to use higher mesh optimization parameters; if the visual quality check fails, an instruction can be issued to the cue word generator used to build cue words, instructing it to add or modify specific visual descriptors, thus forming a closed-loop, automated quality optimization process to improve the quality of the 3D model.
[0114] In some implementations, the overall workflow can be coordinated during the 3D model generation and quality inspection stages. This can involve task flow orchestration, using a Directed Acyclic Graph (DAG) to define a complete pipeline from target subject parsing to 3D asset import, precisely managing dependencies and execution order between modules. For example, the 3D reconstruction device ensures that the "quality inspection" task is only triggered after both the "2D image generation from multiple preset viewpoints" and "3D model generation" tasks have been successfully completed.
[0115] In some implementations, during the generation and quality inspection of the 3D model, GPU / CPU resources can be dynamically requested or released from container clusters such as Kubernetes based on the estimated computational load of each task (such as model complexity and image resolution) and the current load of the device where the 3D reconstruction device is located, so as to achieve efficient utilization of computing resources and parallel processing of tasks.
[0116] As can be seen from the above technical solutions, the method provided in this disclosure can automatically perform quality inspection on the generated 3D model from multiple dimensions, ensuring the quality of the 3D model. Furthermore, it can also achieve automatic detection and repair of quality problems through integration with CI / CD workflows.
[0117] Figure 6 This is a flowchart illustrating the three-dimensional reconstruction method provided in the fifth exemplary embodiment of this disclosure. For example... Figure 6 As shown above, in the above Figure 2 Based on the illustrated embodiment, the following steps may be included after step 300:
[0118] Step 500: Establish an index information database for each 3D model.
[0119] The index information repository may include a metadata information repository and a semantic vector information repository. Step 500 may include the following steps:
[0120] Step 510: Based on the metadata of each 3D model, form a metadata information database.
[0121] In step 510, a metadata file can be created for each 3D model, and a metadata information library can be formed based on the metadata files corresponding to multiple 3D models. The metadata file can be in JSON or XML format.
[0122] In some implementations, metadata includes at least one of the following: 3D model ID, target theme to which the 3D model belongs, target style corresponding to the 3D model, target entity corresponding to the 3D model, descriptive text information corresponding to the 3D model, and tag information. In addition, metadata may also include technical attribute parameters, such as the number of polygons in the 3D model and its version number.
[0123] The label information is used to describe 3D objects, such as "red sports car" or "night scene". The label information can be generated using models such as CLIP.
[0124] Step 520: Generate the semantic vector corresponding to the 3D model based on the name, description text information and label information of the 3D model.
[0125] In step 520, a text embedding model (such as Sentence-BERT) can be used to convert the name, description, and labels of the 3D model into high-dimensional semantic vectors. In this way, the high-dimensional semantic vectors can be used to achieve fuzzy retrieval based on semantic similarity.
[0126] Step 530: Based on the semantic vectors corresponding to each 3D model, a semantic vector information database is formed.
[0127] In step 530, the semantic vectors corresponding to each 3D model are integrated and stored in a vector database such as FAISS or Milvus to form a semantic vector database, so that retrieval can be performed based on semantic similarity in subsequent steps.
[0128] As can be seen from the above technical solutions, the method provided in this disclosure can build a retrieval information database for automatically generated 3D models, support precise keyword retrieval through a metadata information database, and support fuzzy retrieval through a semantic vector information database. In this way, it can accurately retrieve 3D models that meet the user's needs and preferences.
[0129] In one embodiment, asset management and deployment of the 3D model are also possible.
[0130] In some implementations, after generating the 3D model or after the 3D model passes quality inspection, the 3D model can be formatted so that the target 3D engine can recognize the 3D model and render it. The target 3D engine refers to an engine that creates and runs 3D display resources based on the 3D model, such as Unity or Unreal Engine in automotive HMI systems.
[0131] This includes converting 3D models to engine-compatible formats such as FBX and OBJ. For example, it can automatically select the best output format based on the characteristics of the target 3D engine (such as Unity or Unreal Engine) and platform (such as Android or Linux). For mobile deployments, it can prioritize converting 3D models to glTF or USDZ formats, which are more GPU memory-friendly.
[0132] In addition, during the format conversion process, data consistency and integrity checks can be performed to ensure that key information such as the geometric topology, UV mapping layout, skin weights, and materials (such as base color texture maps, metallic texture maps, and roughness texture maps) of the 3D model are transferred completely and without loss, thereby ensuring that the visual effect of the 3D model in the target 3D engine is consistent with the design draft.
[0133] In some implementations, after converting the 3D model to engine-compatible mode, the 3D model can be automatically imported into the target 3D engine and a Prefab format 3D model can be generated as a 3D display resource.
[0134] Among these features, automated engine import scripts, such as those for Unity (C# Editor Script) or Unreal Engine (Python Editor Script), can be used to batch import converted model files, texture maps, and other resources into the correct project path.
[0135] During the 3D model import process, the engine's import script can automatically create and configure material instances, connecting texture maps to the correct shader nodes. Additionally, if the model contains Levels of Detail (LOD), the generator will automatically configure LOD Groups and set the switching distance between different levels of detail to optimize rendering performance.
[0136] In some implementations, version tracking and management of different versions of 3D models can be performed, including recording version update history and maintaining the relationships between versions.
[0137] When the 3D model is updated, the version control system can compare the 3D model before and after the update, and then generate a lightweight patch package containing the changed data based on the differences. Thus, when the 3D model is updated, the vehicle HMI system only needs to download the lightweight patch package, thereby reducing data transfer and improving update speed.
[0138] Before updating to a new version based on a lightweight patch package, the vehicle-mounted HMI system can back up the current version of the 3D model. This allows the system to roll back to the previous version if an error occurs during the update, ensuring its stability and reliability.
[0139] Figure 7 This is a flowchart illustrating the three-dimensional reconstruction method provided in the sixth exemplary embodiment of this disclosure. For example... Figure 7 As shown above, in the above Figure 6 Based on the illustrated embodiment, steps 300 and thereafter may include the following steps 600-700:
[0140] Step 600: In response to the user's input of theme request information, determine the theme to be rendered and the entity to be rendered.
[0141] In step 600, the subject matter requirement information can be used to search the aforementioned search information database to determine the subject matter to be rendered and the entity to be rendered.
[0142] For example, semantic matching models, such as pre-trained large language models, can be used to parse topic demand information and classify user intent. For instance, for the topic demand information "change to a cyberpunk style", the user intent can be determined to be switching topics; for the topic demand information "what styles are available", the user intent can be determined to be querying topics; and for the topic demand information "look for something exciting", the user intent can be determined to be fuzzy search.
[0143] In this way, for the theme switching intent, a finely tuned language model (such as Fine-tuned BERT or GPT series models) can be used to extract information (slots) related to the theme, style or entity from the theme requirement information, such as [theme: cyberpunk], [entity: sports car], [scene: night], and structure it into JSON format.
[0144] Subsequently, topic features and style elements can be extracted from user input. For example, the parsed structured information or topic requirement information input by the user can be fed into a text embedding model (such as Sentence-BERT or OpenAIAda-002) to be converted into a high-dimensional semantic vector, which can then be used as the query vector. In this way, the query semantic vector can represent the deep semantic information of the structured information or topic requirement information.
[0145] Furthermore, if the user's input topic requirements contain multiple features, such as "both futuristic and glamorous", different weights can be assigned to the "futuristic" and "glamorous" features, and then multiple different features can be merged into a query vector to more accurately express the user's mixed needs.
[0146] Finally, based on the extracted information, the system can retrieve the subject and entity to be rendered that match the user's needs from the information retrieval database.
[0147] In some implementations, a two-layer semantic matching strategy of coarse-grained and fine-grained matching can be adopted. First, based on the user's input of the theme requirements, the theme category is determined, resulting in several candidate themes. Then, the specific style is determined, leading to the final theme to be rendered. This ensures that the theme to be rendered is more closely matched to the user's needs. For example, coarse-grained matching can be performed first based on the theme requirements to determine the theme category to be rendered as "retro," resulting in candidate themes "steampunk city" and "cyberpunk city." Then, fine-grained matching is performed to select "cyberpunk" as the theme to be rendered from "steampunk" and "cyberpunk."
[0148] In some implementations, the target topic and its corresponding target entities can be pre-packaged. This way, upon retrieving the topic to be rendered, the corresponding entities to be rendered can be determined; similarly, upon retrieving the entities to be rendered, the corresponding topic to be rendered can be determined. Alternatively, the search results at both the topic and entity levels can be comprehensively analyzed to obtain the final topic and entities to be rendered.
[0149] In some implementations, the ranking weight of search results can be dynamically adjusted by combining the user's input topic requirements, the user's historical behavior information, the current driving scenario (such as highway or urban area), and ambient lighting information, so as to prioritize displaying the topics and entities that best match the current context.
[0150] In some implementations, if the user's input topic is too vague or niche, a large language model can be used to expand the topic by adding relevant synonyms and hypernyms, thereby broadening the search scope and improving the accuracy of the search results. For example, if a user inputs "a bit of a cyberpunk feel," the large language model can supplement the search with keywords such as "futuristic," "neon lights," and "mechanical feel."
[0151] In some implementations, if a user inputs vague topic requirements, a large language model can be used to retrieve several candidate topics related to those requirements. Then, based on semantic similarity and other factors, the final topic to be rendered is determined from these candidate topics. For example, if a user inputs the topic requirement "I want a more interesting driving environment," the large language model can parse the requirement information to obtain features such as "interesting" and "driving environment." Based on these features, multiple candidate topics are then identified, such as "cyberpunk city," "cartoon fantasy world," and "retro steampunk." Subsequently, a text description is formed based on the topic name and tag information of each entity under the candidate topic. The semantic relevance of this description to features such as "interesting" is calculated, and "cyberpunk city," with the highest relevance, is determined as the topic to be rendered.
[0152] In some implementations, user interaction information with search results can be recorded, such as click-through rates and application durations for specific topics. This recorded interaction information can be used as optimization samples to adjust the semantic matching model. Furthermore, user preferences can be determined based on interaction information, dynamically adjusting the ranking weights of search results.
[0153] In some implementations, step 600 includes steps 610-630 as follows:
[0154] Step 610: In response to the topic requirement information input by the user, determine the first score corresponding to each target topic based on the matching relationship between the metadata of the 3D model associated with each target topic and the topic requirement information.
[0155] In step 610, sparse retrieval can be performed using metadata to quickly identify topics that are exactly matched or highly relevant to the text. For example, traditional keyword-based retrieval algorithms (such as BM25 or TF-IDF) can be used to quickly filter the metadata (such as name and tags) of each 3D model in the metadata information database, thereby obtaining the matching score corresponding to each 3D model, such as the BM25 score. In this way, the first score corresponding to the target topic can be determined by comprehensively analyzing the matching scores of all 3D models corresponding to each topic. For example, for the "car" 3D model and the "building" 3D model associated with the aforementioned "cyberpunk" topic, the matching scores between the topic requirement information and the metadata of the "car" 3D model and the "building" 3D model can be determined respectively, and then the first score corresponding to the "cyberpunk" topic can be obtained through comprehensive analysis.
[0156] In one example, based on the user's input theme request "I want a cyberpunk style," the metadata of all 3D models is analyzed. For instance, 3D models tagged with "cyberpunk," "car," and "red" have a higher match score with the theme request; 3D models tagged with "cyberpunk" and "building" have a higher match score; while 3D models tagged with "classical" and "building" have a lower match score. Thus, based on the match scores of all 3D models, a first score for the "cyberpunk" theme and a first score for the "classical" theme can be determined.
[0157] Step 620: Based on the semantic similarity between the semantic vectors corresponding to the 3D models associated with each target topic and the topic requirement information, determine the second score corresponding to each target topic.
[0158] In step 620, a dense search can be performed using semantic vectors to determine the semantically most similar topics. For example, a nearest neighbor (ANN) search can be performed in a vector database using a query vector generated based on topic demand information, and the similarity score between the query vector and the semantic vector corresponding to the 3D model can be calculated using cosine similarity or dot product. In this way, a second score corresponding to each target topic can be determined based on the similarity scores of all 3D models corresponding to each target topic.
[0159] In one example, a query vector can be generated based on the topic requirement "I want cyberpunk style," and then the cosine similarity between the query vector and the semantic vectors of all 3D models can be calculated. For example, the semantic vector of a car 3D model corresponding to the cyberpunk theme has a high cosine similarity to the query vector. Thus, based on the matching scores of all 3D models, a second score can be determined for the "cyberpunk" theme and a second score for the "classical" theme.
[0160] Step 630: The first score and the second score are weighted and summed. Based on the weighted summation result, the subject to be rendered is determined in each target subject, and at least one target entity associated with the subject to be rendered is determined as the entity to be rendered.
[0161] In step 630, the first score and the second score are weighted and summed according to pre-set weight parameters to obtain the final comprehensive score. This comprehensive score allows the determination of the subject to be rendered, as well as the entities associated with that subject.
[0162] In one example, based on the overall score, the "Cyberpunk" theme is selected as the theme to be rendered between the "Cyberpunk" and "Classical" themes. This allows target entities associated with the "Cyberpunk" theme, such as "cars," "buildings," and "streets" under the "Cyberpunk" theme, to be identified as entities to be rendered.
[0163] In addition, strategies such as Reciprocal Rank Fusion (RRF) can be used to rank target topics or entities based on a comprehensive score, and the ranking results can be fed back to the user so that the user can select the topic or entity to be rendered.
[0164] Step 700: When any entity to be rendered is detected, display the 3D model corresponding to that entity on the display screen.
[0165] In step 700, the in-vehicle HMI system activates the theme to be rendered. Thus, if an entity to be rendered is detected, its corresponding 3D model can be displayed on the screen. For example, if the sensing system detects a car, it can call up the 3D model corresponding to the "Car" entity under the "Cyberpunk" theme and display it on the screen. In this way, the user can see a cyberpunk-style 3D car on the in-vehicle HMI system's screen.
[0166] In some implementations, the in-vehicle HMI system activates the theme to be rendered, which can replace the 3D model currently displayed on the screen. For example, if the screen is currently displaying a 3D model of a cartoon-themed building, it can be replaced with a 3D model corresponding to a cyberpunk theme.
[0167] In some implementations, asynchronous loading and streaming can be used to quickly load the 3D model corresponding to the entity to be rendered into the in-vehicle HMI system, facilitating rapid display when the entity is detected. For example, an asynchronous API interface can be pre-configured, allowing the in-vehicle HMI system to request and load 3D models from local storage or the cloud without blocking the main thread. Furthermore, for large-scale 3D models, such as environmental 3D models, a chunked streaming loading method can be used for gradual rendering, improving the perceived response speed for the user.
[0168] In some implementations, a dynamic replacement interface can be pre-configured. When the target 3D engine of the in-vehicle HMI system is running, the dynamic replacement interface can be used to dynamically replace entities to be rendered in the scene without reloading the entire scene. This enables smooth switching of display themes.
[0169] As can be seen from the above technical solutions, the method provided in this disclosure can accurately determine the user's intended rendering theme and the 3D model to be rendered based on the user's input theme requirements, and then perform rendering through the in-vehicle HMI system. This can provide users with a personalized visual interaction experience.
[0170] Figure 8 This is a flowchart illustrating the three-dimensional reconstruction method provided in the seventh exemplary embodiment of this disclosure. For example... Figure 8 As shown above, in the above Figure 7 Based on the illustrated embodiment, step 700 may include the following steps:
[0171] Step 710: In response to the perception that any entity to be rendered is a biological entity, a target action is determined in at least one preset action based on the speed of any entity to be rendered.
[0172] In step 710, if the sensing system detects a movable biological entity such as a pedestrian, it can select a target action that matches the movement speed of the biological entity from multiple preset actions. For example, for a "pedestrian" entity, if the perceived speed of the pedestrian is greater than 5 m / s, "running" can be identified as the target action; if the perceived speed of the pedestrian is less than 0.2 m / s, "standing" can be identified as the target action; and if the perceived speed of the pedestrian is between 0.2 m / s and 5 m / s, "walking" can be identified as the target action.
[0173] In some implementations, a mapping relationship between movement speed and target action can be set separately for different types of biological entities to ensure the accuracy and rationality of target action matching. For example, unlike the mapping relationship between the movement speed (greater than 5m / s) of the "pedestrian" entity and the "running" action, for the "cat" entity, if the perceived movement speed of the cat is greater than 1.5m / s, then "running" can be determined as the target action.
[0174] Step 720: In at least one dynamic 3D model corresponding to any entity to be rendered, determine the dynamic 3D model to be rendered corresponding to the target action.
[0175] In step 720, after determining the target action for the entity to be rendered, a dynamic 3D model to be rendered can be selected from the pre-established dynamic 3D model corresponding to the entity. For example, for the aforementioned "pedestrian" entity, if "walking" is determined as the target action, a walking dynamic 3D model can be selected as the dynamic 3D model to be rendered from the dynamic 3D model corresponding to the "pedestrian" entity. It should be noted that during the determination of the dynamic 3D model to be rendered, the target theme can also be considered to ensure that the selected dynamic 3D model matches the currently active theme. For example, if the "cyberpunk" theme is currently active, a walking dynamic 3D model associated with the cyberpunk theme can be selected, rather than a walking dynamic 3D model associated with the classical theme.
[0176] Step 730: Display the dynamic 3D model to be rendered on the screen.
[0177] In step 730, after determining the dynamic 3D model to be rendered, the in-vehicle HMI system can load the dynamic 3D model onto the display screen. This allows users to not only see 3D pedestrians on the display screen but also observe the dynamic changes of the pedestrians, providing a more vivid visual experience.
[0178] As can be seen from the above technical solutions, the method provided in this disclosure can select a matching dynamic 3D model to be rendered based on the real-time movement speed of the biological entity. In this way, the movement state of the biological entity can be displayed through the dynamic 3D model to be rendered, which is beneficial to enhancing the fun of the in-vehicle HMI display and the user's immersion.
[0179] Figure 9 This is a flowchart illustrating the three-dimensional reconstruction method provided in the eighth exemplary embodiment of this disclosure. For example... Figure 9 As shown above, in the above Figure 7 Based on the illustrated embodiment, step 300 is followed by step 800, and step 600 is followed by step 900:
[0180] Step 800: Based on user preference information and historical trend data, determine the cache path corresponding to each target topic; cache the 3D model associated with the target topic through the cache path corresponding to the target topic.
[0181] In step 800, in order to speed up the loading of the 3D model and improve the smoothness of the dynamic 3D model display, the vehicle-mounted HMI system may adopt cache management technology.
[0182] In some implementations, user preference information, individual user historical trend data, and group historical trend data can be analyzed to predict whether users are interested in various target topics. Based on the prediction results, the target topic caching path can be determined and cached. For example, if it is predicted that users have a high level of interest in the "cyberpunk" topic, the 3D model associated with the "cyberpunk" topic can be cached on the vehicle during vehicle startup or idle periods.
[0183] In some implementations, a multi-level caching architecture can be adopted. Based on factors such as user interest in each target topic and the frequency of use of each target topic, the corresponding caching level for each target topic is determined, and then the 3D models associated with the target topics are cached in the corresponding level. For example, a three-level caching architecture of "cloud-vehicle-memory" can be established. The cloud can be a remote server cluster that transmits data with the in-vehicle HMI system; the vehicle can be the vehicle's built-in storage device; and the memory can be the vehicle's built-in system memory. In this way, 3D models associated with target topics with relatively high usage frequency can be downloaded from the cloud to the vehicle, while 3D models associated with target topics with even higher usage frequency can be further loaded into memory.
[0184] Step 900: Based on the cache path corresponding to the theme to be rendered, obtain the 3D model associated with the theme to be rendered.
[0185] In step 900, for a cached theme to be rendered, the 3D model can be directly retrieved from the cache path corresponding to the theme. This shortens the data reading path for the 3D model and improves its loading speed. For example, after caching the 3D model associated with the "Cyberpunk" theme on the vehicle, if a 3D model of a car with the "Cyberpunk" theme needs to be displayed, the in-vehicle HMI system can directly read the 3D model locally on the vehicle without loading it from the cloud. This avoids display lag in the in-vehicle HMI system due to network latency.
[0186] As can be seen from the above technical solutions, the method provided in this disclosure can predict and cache the 3D models that users may need by analyzing user behavior and trends. This shortens the data acquisition path for 3D models, reduces the loading latency of 3D models, and ensures that the in-vehicle HMI system can smoothly render 3D models on the display screen.
[0187] Figure 10 This is a flowchart illustrating the three-dimensional reconstruction method provided in the ninth exemplary embodiment of this disclosure. (As shown...) Figure 10 As shown, the method includes a 3D model generation stage and a user interaction stage.
[0188] like Figure 10 As shown, the 3D model generation stage may include the following steps:
[0189] Step 11, Theme Definition. First, the designer determines the target theme to be generated and inputs the corresponding target theme information. For example, the designer determines that the target theme is "Cyberpunk City".
[0190] Step 12, Prompt Generation. The large language model analyzes the target theme information input by the designer and determines the prompts used to generate the corresponding 2D images of the target entities. For example, the large language model analyzes the target theme information, identifies target entities such as the "car" entity and the "building" entity, and generates the prompt "a floating car with a flowing light effect" for the "car" entity and the prompt "rectangular neon billboard, futuristic skyscraper" for the "building" entity.
[0191] Step 13, Multi-view Image Generation. Based on the prompt words, a generative model is used to generate two-dimensional images of the target entity from multiple preset viewpoints. For example, based on the prompt words "a hovering car with a flowing light effect," two-dimensional images of the "car" entity from multiple preset viewpoints can be generated.
[0192] Step 14, 3D Reconstruction. Based on 2D images from multiple preset viewpoints, reconstruct the 3D model corresponding to the target entity. For example, 2D images of a "car" entity from multiple preset viewpoints can be converted into a cyberpunk-style 3D model of a car, such as a 3D model of a hovercar.
[0193] Step 15: Add motion effects to the biological entity. If the target entity is a biological entity, motion effects are added to the 3D model through skeletal rigging and skinning to obtain a dynamic 3D model corresponding to the target entity. For example, if the target entity is a "pedestrian" entity, a dynamic 3D model of a pedestrian walking and a dynamic 3D model of a pedestrian running can be generated.
[0194] Step 16, Quality Inspection. A quality inspection is performed from multiple dimensions. If the quality inspection fails, the process returns to Step 13 and repeats the multi-view image generation and 3D reconstruction steps until the quality inspection is passed. If the quality inspection is passed, the process proceeds to Step 17. For example, the topological structure of the 3D model, visual style consistency, and performance in an in-vehicle rendering environment can be checked. Step 17 is then initiated after these checks are passed.
[0195] Step 17, Format Conversion. If the quality check passes, convert the 3D model to a format such as FBX or OBJ. For example, first convert the 3D model to FBX format, then import the FBX format 3D model into the target 3D engine (such as Unity), and use the target 3D engine to generate a Prefab format 3D model as a 3D display resource.
[0196] Step 18, Asset Storage. Store the quality-checked and format-converted 3D models in the asset library, and establish a retrieval information database for each 3D model. For example, the 3D model of a hovercar can be stored in the asset library, and corresponding semantic tags can be created for it, including "cyberpunk city," etc. Furthermore, the object type corresponding to the 3D model can be associated, such as associating the 3D model of the hovercar with the sedan type.
[0197] like Figure 10 As shown, the user interaction phase may include the following steps:
[0198] Step 21, User Input. The user inputs their topic requirements via voice or text. For example, the user inputs "Make the world cyberpunk" via voice.
[0199] Step 22, Semantic Parsing. The large language model can parse the topic information input by the user to determine the user's intent. For example, the large language model can parse "make the world cyberpunk," extract the feature "cyberpunk," and then determine that the user's intent is to switch to a topic related to "cyberpunk."
[0200] Step 23: Match 3D Models. The system retrieves the entities associated with the theme to be rendered from the index database, and then searches the asset library for the corresponding 3D models of those entities. If a 3D model is found, proceed to Step 24; otherwise, proceed to Step 26. For example, the retrieval module might match the "Cyberpunk City" theme asset package in the asset library, which includes 3D models of various target entities under the "Cyberpunk City" theme, such as the 3D model of the hovercar corresponding to the "Sedan" entity.
[0201] Step 24: Load the 3D model. If a 3D model corresponding to the entity to be rendered is found, it can be loaded. For example, an in-vehicle HMI system can load 3D models from the asset package corresponding to the "Cyberpunk City" theme, such as the 3D model of the hovercar corresponding to the "Sedan" entity.
[0202] In addition, in some examples, recommended information can be generated based on the topic to be rendered and fed back to the user. The 3D model corresponding to the entity to be rendered is then loaded after the user confirms it.
[0203] Step 25, Real-time Rendering. If the vehicle's perception system detects an entity to be rendered, it can render the corresponding 3D model. For example, when the onboard sensors detect a real car, the onboard HMI's target 3D engine can render a 3D model of a hovercraft on the display screen. In this way, the onboard HMI system can render a cyberpunk driving perspective on the display screen.
[0204] Step 26: Generate feedback information. If no 3D model corresponding to the entity to be rendered is found, feedback information indicating no corresponding solution will be generated and sent to the user. This feedback information may include alternative suggestions, which provide existing theme options that are semantically similar to the theme requirements. Furthermore, based on the user's theme requirements, the user's theme requirements can be recorded in a list of themes to be developed for later generation of themes based on those requirements.
[0205] For example, if a user enters "I want the Ancient Egyptian Pyramid style", and no related themes or asset packs are found, the system can generate the feedback message "Sorry, there are currently no Ancient Egyptian style themes. We recommend you try the Mystic Oriental or Classical European styles."
[0206] Exemplary device
[0207] The three-dimensional reconstruction method provided by the embodiments of this disclosure has been described above. It is understood that, in order to realize the various functions of the three-dimensional reconstruction method, the three-dimensional reconstruction device may include corresponding hardware and software for implementing the hardware functions.
[0208] Those skilled in the art will readily recognize that the steps of the three-dimensional reconstruction method described in conjunction with the embodiments of this disclosure can be implemented in hardware or in a combination of software-driven hardware. Whether a function is implemented in hardware or software-driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0209] Figure 11 This is a schematic diagram of the structure of a three-dimensional reconstruction apparatus provided in an exemplary embodiment of this disclosure. Figure 11 As shown, in one embodiment, the 3D reconstruction device 1100 includes: a topic parsing module 1110, a 2D resource generation module 1120, and a 3D resource generation module 1130.
[0210] The topic analysis module 1110 is used to determine the target style, target scene, and target entities in the target scene based on the target topic;
[0211] The two-dimensional resource generation module 1120 is used to generate two-dimensional images of the target entity from multiple preset viewpoints based on the target style;
[0212] The 3D resource generation module 1130 is used to generate a 3D model corresponding to the target entity based on 2D images from multiple preset viewpoints; wherein the 3D model is associated with the target theme.
[0213] In an exemplary embodiment, the 3D resource generation module 1130 is used to: generate a triangular mesh structure corresponding to a target entity based on 2D images from multiple preset viewpoints; generate a texture map corresponding to the triangular mesh structure based on a target style; and map the texture map onto the triangular mesh structure to form a 3D model corresponding to the target entity.
[0214] In an exemplary embodiment, the 3D resource generation module 1130 is configured to: determine the skeleton structure of the target entity in response to the target entity being a biological entity; determine the skinning weight corresponding to each vertex based on the influence range of each bone in the skeleton structure on each vertex in the triangular mesh structure; and map the texture map to the triangular mesh structure, and based on the skinning weight, redirect the animation corresponding to at least one preset action to the skeleton structure to generate at least one dynamic 3D model corresponding to the target entity.
[0215] Figure 12 This is a schematic diagram of the structure of a three-dimensional reconstruction apparatus provided in another exemplary embodiment of this disclosure.
[0216] like Figure 12As shown, in an exemplary embodiment, the 3D reconstruction device 1100 further includes a quality inspection module 1140, which is used to: perform quality inspection on the 3D model based on the items to be inspected; wherein the items to be inspected include at least one of geometric rationality inspection, visual quality inspection, vehicle display environment adaptability inspection, and vehicle safety standard inspection.
[0217] like Figure 12 As shown, in an exemplary embodiment, the 3D reconstruction device 1100 further includes an index module 1150, configured to: form a metadata information library based on the metadata of each 3D model; the metadata includes at least one of the following: 3D model ID, target theme to which the 3D model belongs, target style corresponding to the 3D model, target entity corresponding to the 3D model, descriptive text information corresponding to the 3D model, and tag information; generate a semantic vector corresponding to the 3D model based on the name, descriptive text information, and tag information of the 3D model; and form a semantic vector information library based on the semantic vectors corresponding to each 3D model.
[0218] like Figure 12 As shown, in an exemplary embodiment, the 3D reconstruction apparatus 1100 further includes a rendering module 1160, configured to: in response to user-inputted topic requirement information, determine a first score corresponding to each target topic based on the matching relationship between the metadata of the 3D model associated with each target topic and the topic requirement information; determine a second score corresponding to each target topic based on the semantic similarity between the semantic vector corresponding to the 3D model associated with each target topic and the topic requirement information; perform a weighted summation of the first score and the second score, determine the topic to be rendered in each target topic based on the weighted summation result, and determine at least one target entity associated with the topic to be rendered as the entity to be rendered; and display the 3D model corresponding to any entity to be rendered on the display screen when any entity to be rendered is detected.
[0219] like Figure 12 As shown, in an exemplary embodiment, the rendering module 1160 is configured to: in response to the perception that any entity to be rendered is a biological entity, determine a target action in at least one preset action based on the speed of any entity to be rendered; determine the dynamic three-dimensional model to be rendered corresponding to the target action in at least one dynamic three-dimensional model corresponding to any entity to be rendered; and display the dynamic three-dimensional model to be rendered on the display screen.
[0220] like Figure 12 As shown, in one exemplary embodiment, the 3D reconstruction apparatus further includes an access module 1170, configured to: determine the cache path corresponding to each target topic based on user preference information and historical trend data; cache the 3D model associated with the target topic through the cache path corresponding to the target topic; and obtain the 3D model associated with the topic to be rendered based on the cache path corresponding to the topic to be rendered.
[0221] Exemplary electronic devices
[0222] Figure 13 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Figure 13 As shown, the electronic device 1300 includes at least one processor 1310 and a memory 1320.
[0223] The processor 1310 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1300 to perform desired functions.
[0224] The memory 1320 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1310 may execute one or more computer program instructions to implement the three-dimensional reconstruction methods and / or other desired functions of the various embodiments of this disclosure described above.
[0225] In one example, the electronic device 1300 may also include an input device 1330 and an output device 1340, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0226] The input device 1330 may also include, for example, a keyboard, a mouse, etc.
[0227] The output device 1340 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0228] Of course, for the sake of simplicity, Figure 13 Only some of the components of the electronic device 1300 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 1300 may include any other suitable components depending on the specific application.
[0229] Exemplary computer program products and computer-readable storage media
[0230] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps of the bandwidth control methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0231] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0232] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the bandwidth control methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.
[0233] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0234] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0235] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context explicitly states otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0236] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0237] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0238] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A three-dimensional reconstruction method, the method comprising: Based on the target theme, determine the target style, target scene, and target entities within the target scene; Based on the target style, generate two-dimensional images of the target entity from multiple preset viewpoints; Based on the two-dimensional images from the multiple preset perspectives, a three-dimensional model corresponding to the target entity is generated; wherein, the three-dimensional model is associated with the target subject.
2. The three-dimensional reconstruction method according to claim 1, wherein, The step of generating a 3D model corresponding to the target entity based on the 2D images from the multiple preset viewpoints includes: Based on the two-dimensional images from the multiple preset perspectives, a triangular mesh structure corresponding to the target entity is generated. Based on the target style, generate the texture map corresponding to the triangular mesh structure; The texture map is mapped onto the triangular mesh structure to form the three-dimensional model corresponding to the target entity.
3. The three-dimensional reconstruction method according to claim 2, wherein, After generating the triangular mesh structure corresponding to the target entity, the process further includes: In response to the fact that the target entity is a biological entity, the skeletal structure of the target entity is determined; Based on the influence range of each bone in the skeleton structure on each vertex in the triangular mesh structure, the skinning weight corresponding to each vertex is determined; The process of forming the three-dimensional model corresponding to the target entity includes: Based on the skinning weights, the animation corresponding to at least one preset action is redirected to the skeleton structure to generate at least one dynamic 3D model corresponding to the target entity.
4. The three-dimensional reconstruction method according to any one of claims 1 to 3, wherein, After generating the 3D model corresponding to the target entity, the process further includes: The quality of the 3D model is inspected based on the items to be inspected; wherein the items to be inspected include at least one of geometric rationality inspection, visual quality inspection, vehicle display environment adaptability inspection, and vehicle safety standard inspection.
5. The three-dimensional reconstruction method according to claim 1, wherein, After generating the 3D model corresponding to the target entity, the process further includes: Based on the metadata of each of the three-dimensional models, a metadata information database is formed; the metadata includes at least one of the following: three-dimensional model ID, target theme to which the three-dimensional model belongs, target style corresponding to the three-dimensional model, target entity corresponding to the three-dimensional model, descriptive text information corresponding to the three-dimensional model, and tag information; Based on the name of the 3D model, the descriptive text information, and the label information, a semantic vector corresponding to the 3D model is generated. A semantic vector information database is formed based on the semantic vectors corresponding to each of the three-dimensional models.
6. The three-dimensional reconstruction method according to claim 5, wherein, After generating the 3D model corresponding to the target entity, the process also includes: In response to the topic requirement information input by the user, a first score is determined for each of the target topics based on the matching relationship between the metadata of the 3D model associated with each target topic and the topic requirement information; Based on the semantic similarity between the semantic vectors corresponding to the three-dimensional models associated with each target topic and the topic requirement information, a second score is determined for each target topic. The first score and the second score are weighted and summed. Based on the weighted summation result, the subject to be rendered is determined in each of the target subjects, and at least one target entity associated with the subject to be rendered is determined as the entity to be rendered. When any entity to be rendered is detected, the 3D model corresponding to that entity is displayed on the screen.
7. The three-dimensional reconstruction method according to claim 6, wherein, The step of displaying the 3D model corresponding to any entity to be rendered on the display screen when any entity to be rendered is detected includes: In response to the perception that any of the entities to be rendered is a biological entity, a target action is determined in at least one preset action based on the speed of any of the entities to be rendered. In at least one dynamic 3D model corresponding to any entity to be rendered, determine the dynamic 3D model to be rendered corresponding to the target action; The dynamic 3D model to be rendered is displayed on the screen.
8. The three-dimensional reconstruction method according to claim 5, wherein, After generating the 3D model corresponding to the target entity, the process further includes: Based on user preference information and historical trend data, determine the cache path corresponding to each target topic; The 3D model associated with the target topic is cached through the cache path corresponding to the target topic; After determining the subject to be rendered among the target subjects, the process further includes: Based on the cache path corresponding to the theme to be rendered, obtain the 3D model associated with the theme to be rendered.
9. An electronic device, comprising: One or more processors, and a memory; the memory storing computer instructions; the computer instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1 to 8.