Three-dimensional model generation method, device and equipment

By acquiring the 3D base model and basic layout rules corresponding to the vehicle environment, and combining them with user style requirements to generate textures and arrange and render them, the problem of low efficiency in generating personalized 3D architectural models in traditional technologies is solved, and efficient and automated stylized model generation is achieved.

CN121544795APending Publication Date: 2026-02-17XG TECHNOLOGIES PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511653495.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Traditional technologies are inefficient at generating 3D architectural models that meet users' personalized style requirements. Manual modeling is costly and time-consuming, making it difficult to quickly respond to users' personalized style needs.

Method used

By acquiring a pre-generated 3D base model and basic layout rules corresponding to the vehicle's environment, determining texture files based on the user's style requirements, applying them to the 3D base model, arranging them based on the basic layout rules, and rendering them in real time, a 3D architectural model that meets the user's style requirements is generated.

Benefits of technology

It enables the automated generation of 3D architectural models that meet users' personalized style requirements, improves generation efficiency, and ensures seamless and dense arrangement between architectural models to construct a visually rich and dynamic virtual world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121544795A_ABST
    Figure CN121544795A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional model generation method, device and equipment. The method comprises the steps that at least one three-dimensional basic model corresponding to the current environment of a vehicle and a basic layout rule of the at least one three-dimensional basic model are acquired; in response to received style description content input by a user, determining a target style of at least one three-dimensional fundamental model; based on the target style, determining at least one group of mapping files matched with the target style, and respectively applying the at least one group of mapping files to the corresponding three-dimensional basic models in the at least one three-dimensional basic model to obtain at least one target three-dimensional model of the target style; arranging the at least one target three-dimensional model based on a basic layout rule; and rendering and displaying the arranged at least one target three-dimensional model in real time. Compared with a manual modeling mode in the traditional technology, the scheme can automatically generate the 3D building model meeting the personalized style requirement of the user, so that the generation efficiency of the stylized three-dimensional model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer graphics technology, and in particular to a method, apparatus and device for generating three-dimensional models. Background Technology

[0002] With the development of intelligent vehicle technology and the growth of users' demand for personalized experiences, rendering a virtual environment in real time in the vehicle HMI that corresponds to the external world but is artistically processed has become a very attractive development direction.

[0003] In traditional technology, creating three-dimensional (3D) architectural models with a unified artistic style for large-scale urban environments involves a professional team of 3D artists investing a lot of time in manual modeling, texture painting, and layout. This process is costly and time-consuming, making it difficult to quickly respond to users' personalized style requirements, resulting in poor efficiency in generating 3D models that meet users' personalized style needs.

[0004] Therefore, how to automatically generate 3D architectural models that meet users' personalized style needs has become an urgent problem to be solved. Summary of the Invention

[0005] To address the aforementioned technical problems, this disclosure provides a method, apparatus, and device for generating 3D models, which can automatically generate 3D architectural models that meet users' personalized style requirements, thereby solving the problem of poor efficiency in generating 3D models that meet users' personalized style requirements.

[0006] One aspect is the provision of a method for generating three-dimensional models, including:

[0007] Obtain at least one 3D base model corresponding to the current environment of the vehicle and the basic layout rules of the at least one 3D base model;

[0008] In response to the style description content received from user input, the target style of the at least one 3D primitive model is determined;

[0009] Based on the target style, at least one set of texture files matching the target style is determined, and the at least one set of texture files is applied to the corresponding 3D base models in the at least one 3D base model to obtain at least one target 3D model of the target style; wherein, the set of texture files includes textures of the 3D model from different viewpoints;

[0010] The at least one target 3D model is arranged based on the basic layout rules;

[0011] The at least one target 3D model after being arranged is rendered and displayed in real time.

[0012] In another aspect, a three-dimensional model generation apparatus is provided, comprising:

[0013] The first acquisition module is used to acquire at least one three-dimensional base model corresponding to the current environment of the vehicle and the basic layout rules of the at least one three-dimensional base model;

[0014] The first determining module is used to determine the target style of the at least one three-dimensional primitive model in response to the style description content received by the user input.

[0015] The second determining module is used to determine at least one set of texture files that match the target style based on the target style;

[0016] The first processing module is used to apply the at least one set of texture files to the corresponding three-dimensional base models in the at least one three-dimensional base model to obtain the at least one target three-dimensional model of the target style;

[0017] A model arrangement module is used to arrange the at least one target 3D model based on the basic layout rules;

[0018] The rendering and display module is used to render and display the at least one target 3D model after it has been arranged in real time.

[0019] In another aspect, embodiments of this disclosure provide a computer program product that, when executed by an instruction processor, performs the three-dimensional model generation method proposed in the first aspect of this disclosure.

[0020] In another aspect, an electronic device is proposed, comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the three-dimensional model generation method described in the first aspect above.

[0021] The disclosed solution can obtain a pre-generated 3D base model and basic layout rules corresponding to the vehicle's environment, and obtain textures matching the target style according to the user's style requirements, applying the textures to the 3D base model. Therefore, it can generate a target 3D model that meets the user's style requirements. Multiple stylized 3D models are then arranged based on the basic layout rules, and the arranged stylized 3D models are rendered and displayed in real time, allowing the user to see a stylized virtual world corresponding to the real world. Thus, compared to the manual modeling methods of traditional technologies, the disclosed solution can automatically generate 3D architectural models that meet the user's personalized style requirements, thereby improving the efficiency of stylized 3D model generation. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of an automatic urban environment generation system provided in an exemplary embodiment of this disclosure.

[0023] Figure 2 This is a flowchart illustrating a three-dimensional model generation method provided in an exemplary embodiment of this disclosure.

[0024] Figure 3 This is a flowchart illustrating a three-dimensional model generation method provided in another exemplary embodiment of this disclosure.

[0025] Figure 4 This is a flowchart illustrating a three-dimensional model generation method provided in yet another exemplary embodiment of this disclosure.

[0026] Figure 5 This is a flowchart illustrating a three-dimensional model generation method provided in another exemplary embodiment of this disclosure.

[0027] Figure 6 This is a flowchart illustrating a three-dimensional model generation method provided in yet another exemplary embodiment of this disclosure.

[0028] Figure 7 This is a schematic diagram of the structure of a three-dimensional model generation apparatus provided in yet another exemplary embodiment of this disclosure.

[0029] Figure 8 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0030] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.

[0031] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0032] Application Overview

[0033] In traditional techniques, creating 3D architectural models with a unified artistic style for large-scale urban environments involves a significant investment of time by professional 3D artist teams to manually model, texture, and lay out the models, then manually place and adjust them within the engine to construct the 3D model of the urban scene. This process is costly and time-consuming. If a user changes their style requirements, professional technicians must manually remodel the model, making it difficult to quickly respond to users' personalized style needs. Consequently, the efficiency of generating 3D models that meet users' individual style requirements is relatively low.

[0034] In some related technologies, photogrammetry and other techniques are used to reconstruct 3D models of cities from aerial images of the real world, or to post-process existing 3D models or scenes by applying global style filters, such as changing the tone and contrast. However, the former usually contains a lot of redundant data and has a chaotic topological structure, making it difficult to modify the artistic style and optimize real-time rendering. The latter, through global filters, can only change the final rendered image and cannot delve into the material and texture level of the 3D model, thus failing to achieve a deep and consistent artistic style transformation.

[0035] Therefore, how to automatically generate 3D architectural models that meet users' personalized style needs has become an urgent problem to be solved.

[0036] Based on the aforementioned technical problems, the solution disclosed herein can obtain a pre-generated 3D base model and basic layout rules corresponding to the vehicle's environment, and obtain textures matching the target style according to the user's style requirements, applying the textures to the 3D base model. Therefore, it can generate a target 3D model that meets the user's style requirements. Multiple stylized 3D models are then arranged based on the basic layout rules, and the arranged stylized 3D models are rendered and displayed in real time, allowing the user to see a stylized virtual world corresponding to the real world. Thus, compared to the manual modeling methods of traditional technologies, the solution disclosed herein can automatically generate 3D architectural models that meet the user's personalized style requirements, thereby improving the efficiency of stylized 3D model generation.

[0037] Exemplary System

[0038] Figure 1 This is a schematic diagram of an automated urban environment generation system provided as an exemplary embodiment of this disclosure. The system achieves an efficient and flexible automated content production process by decoupling building layout, style generation, and scene construction. For example, Figure 1 As shown, the system mainly includes three core stages: the design of the 3D base model, the generation of stylized materials, and the construction and rendering of dense scenes.

[0039] In some examples, the first stage above, the design of the 3D model, may include: defining basic layout rules between 3D models, pre-designing 3D models with various geometric structures, and creating a model database. After professional technicians pre-design 3D models with various geometric structures and define the layout rules between them, the 3D models with various geometric structures and the basic layout rules can be stored in the model database so that the corresponding 3D models can be directly retrieved from the model database later.

[0040] In some examples, stage two above, generating stylized materials, may include the following process:

[0041] (1) After the user inputs the style description content, in response to the style description content, the style description content is parsed based on the large model to extract style cue words (i.e. visual descriptors), and the target style is determined based on the style cue words.

[0042] (2) Determine the current environment of the vehicle based on the vehicle's environmental perception technology, and obtain at least one three-dimensional basic model corresponding to the current environment of the vehicle from the basic model database; perform multi-view rasterization on each of the at least one three-dimensional basic models to generate multiple two-dimensional reference images.

[0043] (3) Using the target style, guide the style transfer model to process multiple two-dimensional reference images to generate multiple stylized two-dimensional architectural images; and process multiple two-dimensional architectural images based on the material generation model to generate a set of texture files (i.e. materials), and apply the set of texture files to the corresponding three-dimensional base model to obtain a stylized three-dimensional model.

[0044] In some examples, the above-mentioned stage three: dense scene construction and rendering may include the following process: first, obtain the target layout algorithm and basic layout rules, and arrange at least one stylized 3D model based on the target layout algorithm and basic layout rules to generate a dense city layout; then, the dense city layout can be optimized (e.g., removing occluded and invisible 3D models), and then the rendering engine is used to render the optimized dense city layout in real time, and the rendered image is displayed on the screen.

[0045] For a detailed description of each of the above stages, please refer to the detailed description in the following embodiments. The embodiments disclosed herein will not be repeated here.

[0046] The solution disclosed herein can obtain a pre-generated 3D base model and basic layout rules corresponding to the vehicle's environment, and obtain textures that match the target style according to the user's style requirements and apply the textures to the 3D base model. Therefore, it can generate a target 3D model that meets the user's style requirements. Compared with the manual modeling method in traditional technology, the solution disclosed herein can automatically generate 3D architectural models that meet the user's personalized style requirements, thus improving the efficiency of stylized 3D model generation.

[0047] Moreover, since multiple stylized 3D models can be arranged based on basic layout rules, and the arranged stylized 3D models can be rendered and displayed in real time, it ensures that the building models can be arranged seamlessly and densely to create a visually rich and empty urban environment. As a result, users can see a dynamic and stylized virtual world that corresponds to the real world.

[0048] Figure 2This is a schematic flowchart illustrating a three-dimensional model generation method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices (e.g., vehicles), meaning that the executing entity of this embodiment can be a vehicle or a chip within a vehicle.

[0049] For example, such as Figure 2 As shown, the above method may include the following steps:

[0050] Step 201: Obtain at least one 3D base model and at least one basic layout rule of the 3D base model corresponding to the current environment of the vehicle.

[0051] In some embodiments, when the vehicle is in an autonomous driving or navigation scenario, the vehicle can determine its current environment based on positioning technology (e.g., LiDAR, Global Navigation Satellite System, or visual positioning) and high-precision maps, and obtain at least one 3D fundamental model and at least one basic layout rule of the 3D fundamental model corresponding to the vehicle's current environment from a fundamental model database. The fundamental model database can be a local fundamental model database or a server-side fundamental model database.

[0052] In some embodiments, the above-mentioned basic layout rules define the connection and combination rules between different models, which provide a basis for the subsequent arrangement of different 3D models.

[0053] For example, the above-mentioned basic layout rules may include, but are not limited to: a standardized grid system (e.g., a 10x10 meter basic plot), and standardized connection points or sockets set at the edges and specific locations of the three-dimensional base model, and the connection points of the three-dimensional base model are classified (e.g., street interfaces, roof interfaces), etc.

[0054] In some examples, a series of 3D base models can be pre-created by professional designers using 3D modeling software such as Blender and Maya. Alternatively, a series of 3D base models can be generated with the assistance of a procedural algorithm, the L-System. This series of 3D base models can include 3D base models of different objects (such as roads, rivers, buildings, or trees) with different contours, heights, and structural features. The 3D base models can be basic 3D geometries, and all 3D base models use low-polygon-count 3D geometries. For example, a 3D base model of a building could include towers, slab buildings, L-shaped buildings, and buildings with roofs of different slopes. This series of generated 3D base models can then be stored uniformly in a database, either locally or on a server.

[0055] For example, suppose the vehicle is currently in a downtown intersection, which includes a hospital, roads, and a shopping mall. After determining that the vehicle is currently in a downtown intersection, the 3D base models corresponding to the hospital, roads, and shopping mall can be obtained from the local base model database, along with the attribute information of these three 3D base models. Then, based on the attribute information of these three 3D base models, the basic layout rules between these three 3D base models can be determined.

[0056] Step 202: In response to the style description content received from the user input, determine the target style of at least one three-dimensional primitive.

[0057] In some examples, the style description content mentioned above can be text-based style description content entered by the user on the human-computer interaction screen, voice-based style description content entered by the user based on the voice function, an identifier selected by the user in the menu interface to indicate the style description content, or image-based style description content used as a style reference. The target style mentioned above can be a keyword used to describe the target style.

[0058] For example, taking style description content as a text type, when a user enters the style description content "Chinese ink painting style" on the vehicle's central control screen, the style description content is analyzed to determine that the target style of at least one three-dimensional model can be represented by the following style prompt word: "black and white gray tone, Xuan paper texture, brush strokes, ink stains".

[0059] Step 203: Based on the target style, determine at least one set of texture files that match the target style, and apply the at least one set of texture files to the corresponding 3D base model in at least one 3D base model to obtain at least one target 3D model of the target style.

[0060] One set of texture files includes textures of the 3D model from different perspectives.

[0061] In some examples, at least one set of texture files mentioned above refers to physically based rendering texture files (PBR); where textures of a 3D model from different perspectives may include: albedo maps, roughness maps, metallicity maps, and normal maps, etc.

[0062] In some examples, at least one set of texture files mentioned above may be pre-generated or generated in real time.

[0063] In some embodiments, when at least one set of texture files is pre-generated, in response to multiple sets of texture files including a target style that are pre-stored in the cloud, at least one set of texture files matching the target style can be obtained from the multiple sets of texture files, and the at least one set of texture files can be applied to the corresponding three-dimensional base model in at least one three-dimensional base model to obtain at least one target three-dimensional model.

[0064] For example, suppose multiple sets of texture files are pre-stored in the cloud, including: a Chinese ink painting style texture file A and a steampunk style texture file B. If the target style is Chinese ink painting style, a request can be sent to the cloud to retrieve at least one set of texture files from the cloud-stored Chinese ink painting style texture files that matches the vehicle's current environment.

[0065] In other embodiments, when the above-mentioned at least one set of texture files is generated in real time, at least one three-dimensional primitive model can be rasterized to obtain a multi-view image, and the multi-view image can be converted into a target style architectural image. Then, at least one set of texture files in the target style can be generated based on the target style architectural image. For details, please refer to the detailed description in the following embodiments. The embodiments disclosed herein will not be repeated here.

[0066] In some embodiments, when at least one target 3D model of a stylized nature is pre-generated, in response to multiple 3D models with different geographical locations and applied styles being pre-stored in the cloud, at least one target 3D model matching the current environment and target style is obtained from the multiple 3D models.

[0067] For example, suppose the cloud pre-stores multiple 3D models with different geographical locations and application styles, including: a traditional Chinese ink painting style 3D model of a city center location, and a traditional Chinese ink painting style 3D model of a residential area. If the target style is traditional Chinese ink painting and the vehicle is currently located in the city center, a request can be sent to the cloud to retrieve the traditional Chinese ink painting style 3D model of the city center location from the multiple 3D models stored in the cloud.

[0068] In some examples, the above-described application of at least one set of texture files to corresponding 3D base models in at least one 3D base model to obtain at least one target 3D model with a target style refers to binding physical property textures from different viewpoints corresponding to each base model to the corresponding 3D base model, so as to convert the 3D base model with "blank geometry" into a target 3D model with "realistic physical material appearance". That is, the target 3D model is a materialized 3D model compared to the textureless 3D base model. It should be noted that for the process of applying at least one set of texture files to corresponding 3D base models in at least one 3D base model, please refer to the detailed description in related technologies, which will not be repeated here.

[0069] Step 204: Arrange at least one target 3D model based on the basic layout rules.

[0070] In some examples, since the basic layout rules define the connection rules between 3D models, a procedural layout engine can be used to arrange at least one target 3D model based on the basic layout rules to ensure that the models can be accurately and seamlessly combined in a procedural manner, thereby obtaining at least one target 3D model after arrangement.

[0071] For example, suppose at least one target 3D model includes 3D model A, 3D model B, and 3D model C. If the basic layout rules define 3D model A and 3D model C as adjacent, and 3D model C as adjacent to 3D model B, based on the type and connection points of each 3D model, then based on the basic layout rules, the above 3D models A to C can be arranged as follows: 3D model A is connected to 3D model C, and 3D model C is connected to 3D model B, that is, 3D model A -- 3D model C -- 3D model B.

[0072] Step 205: Render and display at least one of the arranged target 3D models in real time.

[0073] In some examples, at least one arranged target 3D model can be loaded into the real-time 3D rendering engine (e.g., Unreal Engine, Unity, etc.) of the target platform (e.g., an in-vehicle human-machine interface). Using the real-time 3D rendering engine, based on the position and orientation of a preset virtual camera, the at least one arranged target 3D model is rendered in real-time at a preset frame rate (e.g., 60 FPS) to generate an image sequence, which is then displayed on a display device. In this way, an immersive, dynamic, and uniformly stylized virtual world corresponding to the external world is presented to the user through the display device.

[0074] The disclosed solution can obtain a pre-generated 3D base model and basic layout rules corresponding to the vehicle's environment, and obtain textures matching the target style according to the user's style requirements, applying the textures to the 3D base model. Therefore, it can generate a target 3D model that meets the user's style requirements. Multiple stylized 3D models are then arranged based on the basic layout rules, and the arranged stylized 3D models are rendered and displayed in real time, allowing the user to see a stylized virtual world corresponding to the real world. Thus, compared to the manual modeling methods of traditional technologies, the disclosed solution can automatically generate 3D architectural models that meet the user's personalized style requirements, thereby improving the efficiency of stylized 3D model generation.

[0075] like Figure 3 As shown above, in the above Figure 2 Based on the illustrated embodiment, step 203 above may include the following steps:

[0076] Step 2031: Perform multi-view rendering on the target 3D model to obtain multiple 2D reference images.

[0077] The target 3D model is any one of at least one 3D model, and the multiple 2D reference images correspond to different viewpoints.

[0078] In some embodiments, for a target 3D model in at least one 3D model, a set of standardized multi-view virtual cameras is set in a virtual 3D space to render the target 3D model (i.e., a textureless model) to generate a 2D reference image (also known as a "line drawing") that can accurately depict the outline, structure, and perspective relationship of the target 3D model from different viewpoints; wherein, the set of multi-view virtual cameras can cover six viewpoints respectively (e.g., covering front, side, top viewpoints and key viewpoints such as 45-degree angle), or may also include other possible key viewpoints. This disclosure embodiment does not limit this and is determined according to the specific use case.

[0079] Step 2032: Using the target style, guide the preset image generation model to process multiple two-dimensional reference images to generate multiple two-dimensional architectural images in the target style.

[0080] In some examples, the image generation model described above can be a conditional image generation model, such as Stable Diffusion with ControlNet, or Stable Diffusion powered by IP-Adapter. The model takes style cues representing the target style and multiple 2D reference images as input. Based on the conditional image generation model, it generates geometric structures that follow the individual 2D reference images, while simultaneously applying the target style to the image content. In other words, it fuses the target style and the structures of the multiple 2D reference images to obtain multiple 2D architectural images corresponding to the multiple 2D reference images. These multiple 2D architectural images are visually unified with the target style and geometrically consistent with the multiple 2D reference images (i.e., the original 3D primitives corresponding to the multiple 2D reference images).

[0081] Step 2033: Process multiple two-dimensional architectural images based on a preset material generation model to obtain the first set of texture files that match the target style.

[0082] The first set of texture files can be any set of texture files from at least one set of texture files.

[0083] In some examples, the aforementioned material generation model can be the multimodal diffusion model Material-Anything, the text2Material diffusion model which generates material textures based on text descriptions, or other possible material generation models. Assuming the material generation model is the Material-Anything model, multiple 2D architectural images are input into the Material-Anything model. Based on the Material-Anything model, visual features of the multiple 2D architectural images are extracted. Based on the structural features and semantic features related to the target style included in the visual features, a set of texture files matching the size and style of the multiple 2D architectural images is generated; this is the first set of texture files. For a detailed explanation of how the Material-Anything model generates texture files, please refer to the descriptions in related technologies; these will not be repeated here.

[0084] In some embodiments, step 2033 may specifically include: rendering each target 3D model in at least one target 3D model according to a preset multiple perspectives to generate multiple 2D rendered images; determining the perceptual loss based on the multiple 2D rendered images and the corresponding multiple 2D reference images; and adjusting the first set of texture files using the perceptual loss so that the perceptual loss corresponding to the adjusted first set of texture files is minimized.

[0085] In some examples, the preset multiple viewpoints mentioned above are the same as the viewpoints used when generating two-dimensional reference images in the above embodiments. For example, the preset multiple viewpoints may include six viewpoints covering key viewpoints such as front, side, top view and 45-degree angle.

[0086] In some embodiments, the aforementioned preset multiple viewpoints refer to the orientation and position of the virtual camera. Based on the orientation and position of the virtual camera, a rendering engine is used to render each target 3D model in at least one target 3D model, generating a 2D rendered image corresponding to each viewpoint. Image features are extracted from the 2D rendered image and the corresponding 2D reference image under each viewpoint, and a pre-trained deep neural network can be used to calculate the perceptual loss between the rendered image features and the reference image features. The first set of texture files is then adjusted using a backpropagation algorithm. Based on the adjusted first set of texture files, multiple 2D rendered images are re-rendered, and the perceptual loss is calculated. This process is iteratively executed until the maximum number of iterations is reached or the perceptual loss corresponding to the first set of texture files is less than a preset threshold. This minimizes the perceptual loss corresponding to the adjusted first set of texture files, ultimately resulting in an optimized first set of texture files. This improves the visual quality of the texture files and their adaptability to the 3D model.

[0087] The technical solution provided in this disclosure involves rendering the target 3D base model from multiple perspectives to obtain multiple 2D reference images. Then, the target style-guided image generation model is used to process the multiple 2D reference images to generate multiple 2D architectural images in the target style. Finally, the material generation model is used to process the multiple 2D architectural images to obtain a set of texture files that match the target style. Therefore, this solution can generate a set of texture files that are visually unified and geometrically consistent with the original base model structure in real time.

[0088] like Figure 4 As shown above, in the above Figure 2 Based on the illustrated embodiment, step 202 above can specifically include the following steps 2021 and 2022:

[0089] Step 2021: In response to the style description content input by the user, process the style description content based on the preset large language model to obtain the visual descriptor.

[0090] In some examples, the large language model described above can be a multimodal language model. For instance, a large language model can be a generative pre-trained Transformer 4withVision (GPT-4V) model with visual capabilities, which is a large model that supports the analysis of both image and text information; another example is a finely tuned contrastive language-image pre-training (CLIP) model.

[0091] For example, taking GPT-4V as the large language model. When a user inputs the style description "ink painting style" on the vehicle's infotainment system, in response to this style description, GPT-4V performs semantic analysis and parsing on the style description to obtain a series of specific visual descriptors such as "black and white gray tones, Xuan paper texture, brushstrokes, and ink smudges," which can also be called style prompt words.

[0092] Step 2022: Determine the target style based on the visual descriptor.

[0093] In some embodiments, after obtaining the visual descriptor, the visual descriptor can be identified as a style cue word for the target style, thereby determining the style of the 3D model to be generated as the target style based on the style cue word.

[0094] The technical solution provided in this disclosure can respond to the style description content input by the user, process the style description content based on a preset large language model to obtain a visual descriptor, and determine the target style based on the visual descriptor. Therefore, it can accurately convert the fuzzy natural language input by the user into a structured visual descriptor through the semantic understanding of the large model, thereby realizing both interaction with the user and accurately parsing the user's style requirements.

[0095] like Figure 5 As shown above, in the above Figure 2 Based on the illustrated embodiment, step 204 above can specifically include the following steps 2041 and 2042:

[0096] Step 2041: Obtain the preset target layout algorithm.

[0097] In some examples, the target layout algorithm described above can be a Wave Function Collapse (WFC) algorithm, or other generation algorithms based on example samples. The target layout algorithm defines the position, rotation, scaling, and type of each building instance (e.g., a stylized target 3D model) in 3D space, as well as the local adjacency rules and constraints between models.

[0098] For example, the WFC algorithm is used as the target layout algorithm. The WFC algorithm generates a building layout scheme that is natural at a macro level and conflict-free at a micro level by analyzing and learning the local adjacency rules (such as connection point matching) between building models and using an iterative constraint propagation method. In addition, macro-level planning can be superimposed, such as using Berlin noise to generate a height map to control the building density and skyline of the virtual city area to be generated, or pre-setting main roads, parks and other structures to guide the layout.

[0099] In some embodiments, the layout rules between various architectural models (e.g., stylized target 3D models) can be determined in advance by learning and analyzing example samples of different styles and city locations, thus obtaining a target layout algorithm. This target layout algorithm can be deployed locally or in the cloud. When it is necessary to arrange at least one stylized target 3D model, the target layout algorithm can be invoked from the cloud or locally.

[0100] Step 2042: Arrange at least one target 3D model based on the target layout algorithm and basic layout rules.

[0101] In some examples, a programmatic layout engine can be used to identify at least one target 3D model that can be adjacent and one that cannot be adjacent, based on the WFC algorithm. Then, combined with basic layout rules, a layout scheme for at least one target 3D model can be determined. Based on the layout scheme of the at least one target 3D model, the at least one target 3D model is arranged to generate a digital layout blueprint that meets the constraints of the WFC algorithm and the basic layout rules. That is, the digital layout blueprint can simultaneously consider physical collision-free, road system continuity and urban planning aesthetic principles, thus ensuring that the target 3D models are arranged seamlessly and densely to create a visually rich and empty urban environment.

[0102] For example, taking the layout rules of the Jiangnan water town as the target layout algorithm, the programmatic layout engine, based on the layout rules of the Jiangnan water town and the basic layout rules between the stylized target 3D models, densely arranges at least one stylized target 3D model (such as rivers, trees, ancient buildings, etc.) along the virtual waterway to form a well-arranged ancient town.

[0103] For example, consider the layout rules of an industrialized urban area using a target layout algorithm. Utilizing a programmatic layout engine, based on the layout rules of the industrialized urban area and the fundamental layout rules between stylized target 3D models, at least one stylized target 3D model (e.g., factories, clock towers, and residences) is stacked together at varying heights to form a vast and complex industrialized urban area.

[0104] The technical solution provided in this disclosure can arrange at least one target 3D model based on the acquired target layout algorithm and basic layout rules, so that the generated digital layout blueprint conforms to the constraints of the target layout algorithm and basic layout rules. As a result, the generated digital layout blueprint is more reasonable and aesthetically pleasing, which helps to improve the user's visual experience.

[0105] like Figure 6 As shown above, in the above Figure 2 Based on the illustrated embodiment, step 205 above can specifically include the following steps 2051 and 2052:

[0106] Step 2051: Perform target processing on at least one of the arranged target 3D models.

[0107] The target processing includes at least one of the following: removing at least one occluded 3D model from the target 3D model, and combining 3D models with the same material and structure from at least one target 3D model.

[0108] For example, consider the process of removing at least one occluded 3D model from a target 3D model. After arranging at least one target 3D model, it can be pre-calculated whether there are any 3D models that are occluded and invisible from the viewpoint to be rendered. If there are occluded 3D models, they can be removed. If there are no occluded 3D models, the number of at least one target 3D model remains unchanged.

[0109] For example, consider the target processing as combining at least one target 3D model with the same material and structure. After arranging at least one target 3D model, all target 3D models are traversed to identify whether there are target 3D models using the same structure and material. After identifying multiple target 3D models using the same structure and material, all multiple target 3D models using the same structure and material are combined. That is, all multiple target 3D models using the same structure and material are rendered as a group of 3D models in a single Draw Call, such as 100 streetlights in a city or 1000 trees in a forest.

[0110] In other embodiments, because multiple geometric models with different levels of complexity (i.e., different levels of detail) can be pre-created for the same object based on a Level of Detail (LOD) algorithm, each geometric model can include multiple geometric models with different levels of complexity, such as high-complexity models or low-complexity models, for at least one target 3D model. This allows for dynamic switching of the 3D model corresponding to each object based on the rendering perspective (i.e., distance from the virtual camera), ensuring visual realism for near-field and distant objects, and avoiding the subsequent rendering of a large number of highly detailed 3D models.

[0111] It should be noted that when the target processing simultaneously includes removing at least one occluded 3D model from the target 3D model and combining 3D models with the same material and structure from the at least one target 3D model, the step of combining 3D models with the same material and structure from the at least one target 3D model can be performed first, followed by the step of removing the occluded 3D model from the at least one target 3D model; alternatively, the step of removing the occluded 3D model from the at least one target 3D model can be performed first, followed by the step of combining 3D models with the same material and structure from the at least one target 3D model. The specific method can be determined based on actual usage, and this embodiment does not limit this.

[0112] Step 2052: Render and display at least one processed target 3D model in real time.

[0113] In some embodiments, after removing at least one occluded 3D model from the target 3D model, for the combined target 3D models with the same structure and material, the CPU sends a single DrawCall to the GPU. This allows the real-time 3D rendering engine to render the combined target 3D models with the same structure and material in parallel and in real-time at a preset viewpoint (i.e., the position and orientation of a preset virtual camera) and a preset frame rate. This also enables real-time rendering of other target 3D models to generate a 2D image, which is then displayed in real-time on the display device (i.e., the in-vehicle terminal). Thus, while the vehicle is in motion, an immersive, dynamic, and stylized virtual world corresponding to the external world can be presented on the in-vehicle screen as the vehicle's environment changes. For a detailed description of the specific implementation of the rendering technology, please refer to the relevant technical specifications; this disclosure will not elaborate further.

[0114] In addition, some special effects can be added during the rendering process. For example, when rendering at least one target 3D model of "steampunk style" in real time, steam particle animation effects can be added to further enhance the atmosphere of the steampunk world.

[0115] The technical solution provided in this disclosure has two advantages. First, it effectively reduces the number of models that need to be rendered by eliminating occluded 3D models in at least one target 3D model after arrangement, thus avoiding unnecessary subsequent rendering calculations. Second, it allows for the combination of 3D models with the same material and structure from at least one target 3D model, thus compressing multiple Draw Calls into a single Draw Call during rendering. This enables parallel processing of 3D models with the same material and structure, significantly reducing CPU overhead. Therefore, even with limited resources on an in-vehicle platform, this reduces rendering overhead, ensuring smooth real-time rendering.

[0116] Exemplary device

[0117] Figure 7 This is a schematic diagram of a three-dimensional model generation apparatus provided as an exemplary embodiment of the present disclosure. The apparatus can be installed in electronic devices such as terminal devices and servers, or on objects such as vehicles, to execute the three-dimensional model generation method of any of the embodiments described above.

[0118] like Figure 7As shown, the aforementioned device 300 may include: a first acquisition module 301, which can be used to acquire at least one 3D base model corresponding to the current environment of the vehicle and the basic layout rules of the at least one 3D base model; a first determination module 302, which can be used to determine the target style of the at least one 3D base model in response to the style description content received by the user; a second determination module 303, which can be used to determine at least one set of texture files matching the target style based on the target style; a first processing module 304, which can be used to apply the at least one set of texture files to the corresponding 3D base models in the at least one 3D base model respectively to obtain the at least one target 3D model of the target style; a model arrangement module 305, which can be used to arrange the at least one target 3D model based on the basic layout rules; and a rendering and display module 306, which can be used to render and display the arranged at least one target 3D model in real time.

[0119] In one possible implementation, the second determining module 303 described above can be specifically used to: perform multi-view rendering on the target 3D base model to obtain multiple 2D reference images; wherein the target 3D base model is any one of at least one 3D base model, and the multiple 2D reference images correspond to different viewpoints; using the target style, guide a preset image generation model to process the multiple 2D reference images to generate multiple 2D architectural images of the target style; process the multiple 2D architectural images based on a preset material generation model to obtain a first set of texture files matching the target style; wherein the first set of texture files is any one of the at least one set of texture files.

[0120] In one possible implementation, the second determining module 303 described above can be specifically used to: render each target 3D model in the at least one target 3D model according to a preset multiple perspectives to generate multiple 2D rendered images; determine the perceptual loss based on the multiple 2D rendered images and the corresponding multiple 2D reference images; and adjust the first set of texture files using the perceptual loss so that the perceptual loss corresponding to the adjusted first set of texture files is minimized.

[0121] In one possible implementation, the first processing module 304 may be specifically configured to: in response to multiple sets of texture files, including the target style, pre-stored in the cloud, obtain at least one set of texture files matching the target style from the multiple sets of texture files, and apply the at least one set of texture files to the corresponding 3D base models in the at least one 3D base model respectively, to obtain the at least one target 3D model of the target style; or, in response to multiple 3D models with different address locations and different styles pre-stored in the cloud, obtain at least one target 3D model matching the current environment and the target style from the multiple 3D models.

[0122] In one possible implementation, the first determining module 302 may be specifically used to: respond to the style description content input by the user, process the style description content based on a preset large language model to obtain a visual descriptor; and determine the target style based on the visual descriptor.

[0123] In one possible implementation, the model layout module 305 is specifically used to: obtain a preset target layout algorithm; and arrange the at least one target 3D model based on the target layout algorithm and the basic layout rules.

[0124] In one possible implementation, the rendering and display module 306 is specifically used for: performing target processing on the at least one target 3D model after arrangement; wherein, the target processing includes at least one of the following: removing occluded 3D models in the at least one target 3D model, combining 3D models with the same material and the same structure in the at least one target 3D model; and rendering and displaying the processed at least one target 3D model in real time.

[0125] The beneficial technical effects corresponding to the exemplary embodiments of this device can be found in the corresponding beneficial technical effects of the exemplary method section above, and will not be repeated here.

[0126] Exemplary electronic devices

[0127] Figure 8 A structural diagram of an electronic device provided in an embodiment of this disclosure includes at least one processor 111 and a memory 112.

[0128] The processor 111 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 11 to perform desired functions.

[0129] The memory 112 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 111 may execute one or more computer program instructions to implement the masking sound generation methods and / or other desired functions of the various embodiments of this disclosure described above.

[0130] In one example, the electronic device 11 may also include an input device 113 and an output device 114, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0131] The input device 113 may include various sensors, including but not limited to: a distance sensor for detecting the distance between a target object and the vehicle; an image sensor for acquiring information about the vehicle's surrounding environment. In some examples, the input device may also include a pressure sensor for detecting seat pressure to determine the presence and location of passengers; a temperature sensor for monitoring the temperature inside the cabin; a humidity sensor for monitoring the humidity inside the cabin to assist in regulating the in-vehicle environment; an air quality sensor for monitoring in-vehicle air quality, such as carbon dioxide and volatile organic compounds (VOCs); a light sensor for detecting the intensity of light inside and outside the vehicle; an acceleration sensor for detecting changes in the vehicle's acceleration; a distance sensor for detecting the distance between the vehicle and other objects; a touchscreen sensor for interaction with the vehicle's infotainment system; biometric sensors, such as fingerprint recognition and facial recognition; a heart rate monitor for monitoring the driver's heart rate; a sound sensor for voice recognition and interaction to enable voice control; a seat sensor for monitoring seat usage, such as whether the seat is occupied and the passenger's body size; and wireless communication sensors, such as Bluetooth and Wi-Fi, for connecting to smart devices to achieve data transmission and remote control. In addition to the examples given above, the input device may include more or fewer sensors, which will not be elaborated here.

[0132] The output device 114 can output various information or signals to other hardware or devices, which may include displays, car audio systems, seats, windows, steering wheels, communication networks, and their connected remote output devices. The displays may include multiple different displays such as a driver's side display, a passenger side display, and a rear-seat display. The car audio system may include multiple speakers located in different positions within the vehicle cabin, and each display or speaker can operate independently.

[0133] Of course, for the sake of simplicity, Figure 8 Only some of the components of the electronic device 8 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 8 may include any other suitable components depending on the specific application.

[0134] Exemplary computer program products and computer-readable storage media

[0135] In addition to the methods and apparatus described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions that, when executed by a processor, cause the processor to perform the steps in the three-dimensional model generation methods of the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0136] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0137] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the three-dimensional model generation methods of the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0138] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0139] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0140] Various modifications and variations can be made to this disclosure without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.

Claims

1. A three-dimensional model generation method, comprising: obtaining at least one three-dimensional base model corresponding to an environment in which a vehicle is currently located and a basic layout rule of the at least one three-dimensional base model; determining a target style of the at least one three-dimensional base model in response to a received style description content of a user input; determining at least one set of map files matching the target style based on the target style, and applying the at least one set of map files to corresponding three-dimensional base models in the at least one three-dimensional base model respectively to obtain at least one target three-dimensional model of the target style; wherein the set of map files includes maps of a three-dimensional model at different viewing angles; arranging the at least one target three-dimensional model based on the basic layout rule; rendering and displaying the arranged at least one target three-dimensional model in real time.

2. The method of claim 1, wherein, The determination of the at least one set of map files matching the target style based on the target style comprises: performing multi-view rendering on a target three-dimensional base model to obtain a plurality of two-dimensional reference images; wherein the target three-dimensional base model is any one of the at least one three-dimensional base model, and the plurality of two-dimensional reference images correspond to different viewing angles respectively; using the target style to guide a preset image generation model to process the plurality of two-dimensional reference images to generate a plurality of two-dimensional architectural images of the target style; processing the plurality of two-dimensional architectural images based on a preset material generation model to obtain a first set of map files matching the target style; wherein the first set of map files is any one of the at least one set of map files.

3. The method of claim 2, wherein, The processing of the plurality of two-dimensional architectural images based on the preset material generation model to obtain the first set of map files matching the target style comprises: rendering each target three-dimensional model in the at least one target three-dimensional model according to a plurality of preset viewing angles to generate a plurality of two-dimensional rendering images; determining a perception loss based on the plurality of two-dimensional rendering images and the corresponding plurality of two-dimensional reference images; adjusting the first set of map files using the perception loss to minimize the perception loss corresponding to the adjusted first set of map files.

4. The method of claim 1, wherein, The determination of the at least one set of map files matching the target style based on the target style and the application of the at least one set of map files to the corresponding three-dimensional base models in the at least one three-dimensional base model to obtain the at least one target three-dimensional model of the target style comprises: in response to the cloud side pre-storing a plurality of sets of map files of a plurality of styles including the target style, obtaining the at least one set of map files matching the target style from the plurality of sets of map files, and applying the at least one set of map files to the corresponding three-dimensional base models in the at least one three-dimensional base model to obtain the at least one target three-dimensional model of the target style; or in response to the cloud side pre-storing a plurality of three-dimensional models of different address locations and applying different styles, obtaining the at least one target three-dimensional model matching the current environment and the target style from the plurality of three-dimensional models.

5. The method of claim 1, wherein, The target style of the at least one three-dimensional base model is determined in response to the received style description content, including: In response to the style description content input by the user, the style description content is processed based on a preset large language model to obtain a visual descriptor; The target style is determined based on the visual descriptor.

6. The method of claim 1, wherein, The arrangement of the at least one target three-dimensional model based on the basic layout rule includes: A preset target layout algorithm is obtained; The at least one target three-dimensional model is arranged based on the target layout algorithm and the basic layout rule.

7. The method of claim 1, wherein, The arranged at least one target three-dimensional model is rendered and displayed in real time, including: The arranged at least one target three-dimensional model is processed; wherein the target processing includes at least one of the following: removing occluded three-dimensional models existing in the at least one target three-dimensional model, combining three-dimensional models with the same material and the same structure in the at least one target three-dimensional model; The processed at least one target three-dimensional model is rendered and displayed in real time.

8. A three-dimensional model generation device, comprising: A first acquisition module for acquiring at least one three-dimensional base model corresponding to the current environment of the vehicle and the basic layout rule of the at least one three-dimensional base model; A first determination module for determining the target style of the at least one three-dimensional base model in response to the received style description content input by the user; A second determination module for determining at least one set of map files matched with the target style based on the target style; A first processing module for applying the at least one set of map files to the corresponding three-dimensional base model in the at least one three-dimensional base model respectively to obtain the at least one target three-dimensional model of the target style; A model arrangement module for arranging the at least one target three-dimensional model based on the basic layout rule; A rendering and display module for rendering and displaying the arranged at least one target three-dimensional model in real time.

9. A computer readable storage medium, the storage medium storing a computer program, the computer program being used to execute the three-dimensional model generation method of any one of claims 1-7.

10. An electronic device, comprising: A processor; A memory for storing executable instructions of the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the three-dimensional model generation method of any one of claims 1-7.