A sparse voxel-based high-fidelity three-dimensional asset local editing method and device

By embedding local regions into a 3D model using a sparse voxel-based method and merging them with non-editable regions, the problem of unclear boundaries between editable and non-editable regions is solved, achieving high-fidelity local editing while preserving the geometric structure and material properties of non-editable regions.

CN122265607APending Publication Date: 2026-06-23BEIHANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2026-04-21
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

In the process of local editing of high-fidelity 3D assets, existing technologies make it difficult to precisely constrain the boundary between the editing area and the non-editing area. After editing the target area, spatial drift is likely to occur, and the geometric structure, texture details and physical material properties of the non-editing area are prone to degradation.

Method used

A sparse voxel-based approach is adopted to embed local regions into the voxel representation of a 3D model. The editing outline is delineated and editing labels are assigned through a visual recognition algorithm. Evolutionary driving forces are used to gradually deform and merge with the non-editable region. Morphological deviations are monitored to generate correction constraints, ensuring the consistency between the local structure and the non-editable region.

Benefits of technology

It achieves a balance between local topology modification and global structure preservation, improving the accuracy, stability and fidelity of editing, and reducing geometric deformation and material degradation in non-editing areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265607A_ABST
    Figure CN122265607A_ABST
Patent Text Reader

Abstract

The application discloses a kind of high-fidelity three-dimensional asset local editing method and device based on sparse voxel.It includes embedding the local area to be edited into the voxel representation of three-dimensional model;And according to editing instruction, the evolution driving force of local area is determined to drive the gradual evolution of the morphology of local area, form new local structure;Further, the uneditable area in the three-dimensional model to be edited is used as position constraint reference, the new local structure is fused with the uneditable area;During the fusion process, the morphology deviation of the uneditable area is monitored, and the deviation correction constraint is generated based on the morphology deviation to inhibit the unexpected change of the uneditable area;Finally, the target three-dimensional model after local editing is output.The application can realize local topological modification of editing area while better maintaining the geometric structure, texture details and physical material properties of non-editing area, thereby improving the accuracy, stability and fidelity of high-fidelity three-dimensional asset local editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D modeling technology, and more specifically to a method and apparatus for local editing of high-fidelity 3D assets based on sparse voxels. Background Technology

[0002] With the development of 3D generation and 3D digital content creation technologies, the demand for local editing of existing 3D assets is constantly increasing. In application scenarios such as game development, digital content creation, and virtual display, it is often necessary to replace, delete, or adjust the shape of local components in 3D assets, while requiring that the geometric structure, texture details, and physical material properties of the unedited areas remain as intact as possible. For high-fidelity 3D assets with high-frequency geometric details and native physically rendered materials, how to maintain the consistency of non-edited areas while achieving local topological changes has become a pressing technical problem that needs to be solved in this field.

[0003] Currently, the mainstream methods for local editing of 3D assets include:

[0004] 1. 3D editing methods based on 2D diffusion priors or sample-by-sample inversion optimization. These methods typically first render the 3D asset as a 2D image from multiple perspectives, then edit it using a 2D diffusion model or inversion process, and finally project the edited result back into 3D space. While this method can achieve a certain degree of zero-sample editing, it usually requires a long period of iterative optimization, incurs significant computational overhead, and struggles to maintain global high-frequency details during geometric deformation, easily leading to oversmoothing and inconsistencies across multiple views.

[0005] 2. Multi-view feedforward editing-based solutions. This type of solution reconstructs the 3D result by locally editing 2D images from multiple perspectives. Although it improves inference speed, it is prone to problems such as spatial artifacts, geometric collapse, and material degradation during multi-view fusion due to the lack of strict 3D geometric constraints at the underlying level. This is especially true in scenes with complex physical materials, which can easily lead to color distortion, shadow solidification, and loss of high-frequency details.

[0006] 3. Solutions based on direct editing in 3D space. This type of method is closer to the native 3D editing paradigm than the aforementioned 2D paths, and can improve response speed. However, when editing locally, it usually relies on coarse-grained methods such as bounding boxes, spheres, or ordinary masks to define the editing area. When the boundary of the area to be edited is complex and closely fits the surrounding structure, coarse-grained area definition methods are difficult to accurately constrain the editing range, which can easily cause editing disturbances to cross the boundary, thereby destroying the original topology, texture features, or material properties of the unedited area.

[0007] In addition, the above schemes all lack a physical alignment mechanism based on the underlying discrete 3D representation and an effective constraint mechanism for the trajectory generated in the continuous latent space, making it difficult to effectively balance "local precise editing" and "global structure preservation". When the target area undergoes a large topological change, the feature perturbation during the editing process is prone to leak to or drift to the non-editing area, ultimately causing geometric deformation, texture distortion and appearance attribute degradation in the non-editing area.

[0008] Therefore, how to overcome the above-mentioned shortcomings remains a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0009] In view of the above problems, this invention proposes a high-fidelity 3D asset local editing method and apparatus based on sparse voxels. It aims to solve the problems in the existing technology of high-fidelity 3D asset local editing process, such as the difficulty in accurately constraining the boundary between the editing area and the non-editing area, the easy spatial drift after the target area is edited, and the easy degradation of the geometric structure, texture details and physical material properties of the non-editing area in the subsequent generation process. In this way, it can realize local topological reconstruction of the editing area, while maintaining the structural consistency and appearance consistency of the non-editing area, and improve the accuracy, stability and fidelity of 3D asset local editing under complex boundary conditions.

[0010] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, embodiments of the present invention provide a high-fidelity local editing method for 3D assets based on sparse voxels, comprising: Receive editing instructions for the 3D model to be edited, the editing instructions specifying the local area to be edited and the editing target; The local region is embedded into the voxel representation of the 3D model to be edited; In response to the editing command, the evolutionary driving force of the local region is determined, and the morphology of the local region is driven to undergo gradual evolution according to the evolutionary driving force to form a new local structure; Using the uneditable region in the 3D model to be edited as the positional constraint reference, the new local structure is fused with the uneditable region; during the fusion process, the morphological deviation of the uneditable region is monitored, and a correction constraint is generated based on the morphological deviation to suppress the unexpected changes of the uneditable region; Output the target 3D model after partial editing.

[0011] Preferably, embedding the local region into the voxel representation of the 3D model to be edited includes: The 3D model to be edited is converted into a discretized voxel representation; Based on a visual recognition algorithm, the target editing outline is automatically drawn on the surface of the 3D model to be edited according to the editing instructions; Edit labels are assigned to each voxel in the voxel representation according to the target edit contour; the edit labels include at least a first type of label indicating that modification is allowed and a second type of label indicating that modification is prohibited.

[0012] This step provides a consistent and clear foundation for subsequent editing, alignment, and protection.

[0013] Preferably, in response to the editing instruction, determining the evolutionary driving force of the local region includes: Within one or more time steps, a first evolution trend and a second evolution trend of the local region are predicted in parallel according to the editing instructions; the first evolution trend represents the evolution direction of maintaining the current form, and the second evolution trend represents the evolution direction of approaching the editing target; Calculate the difference between the first evolutionary trend and the second evolutionary trend, and use the difference as the evolutionary driving force.

[0014] Preferably, calculating the difference between the first evolutionary trend and the second evolutionary trend, and using the difference as the evolutionary driving force, specifically includes: Within each micro-time step, the velocity difference between the evolution direction representing the current morphology and the evolution direction representing the editing target is dynamically calculated, and this velocity difference is used as a numerical driving force to progressively drive the local region to deform.

[0015] Based on the original model, this application only pushes the editing area to change in the target direction, which can realize local structural modification and reduce interference with the main body.

[0016] Preferably, merging the new local structure with the non-editable region includes: Using the non-editable region in the 3D model to be edited as a fixed reference, spatial transformation is used to find the matching position with the highest degree of overlap between the new local structure and the fixed reference. The new local structure is fused with the 3D model to be edited according to the matching position.

[0017] This application uses the preserved area as a reference to correct the position of the new results and completes the fusion under a unified coordinate system, which can solve the problems of misalignment between the old and new structures, unnatural splicing, and local drift.

[0018] Preferably, merging the new local structure with the non-editable region further includes: During the process of spatially matching the new local structure with the fixed reference object, the connectivity state of the new local structure is automatically evaluated; Discrete units that exist in isolation in space, have no topological connection with the main structure, and have a geometric size smaller than a preset threshold are identified as invalid noise and removed from the fusion result.

[0019] Preferably, monitoring the morphological deviation of the non-editable region and generating correction constraints based on the morphological deviation includes: Using the uneditable region as a reference, calculate the deviation energy of the uneditable region in the current editing result relative to the uneditable region in the original 3D model to be edited; The deviation energy is converted into a correction constraint to suppress unexpected deformation of the non-editable region.

[0020] Preferably, converting the deviation energy into a correction constraint includes: In each iteration of material and detail generation cycle, the algorithm model is used to predict the final edited state by using the intermediate state data that is not yet fully edited. High-fidelity data of the non-editable regions in the original 3D model to be edited are used as a reference standard; The predicted state and the reference standard are compared one by one within the uneditable region to generate a quantified error energy value, which is used to characterize the degree of deviation of the uneditable region from the intended editing. The error energy value is converted into a gradient traction force to correct the deviation, and the average intensity of the gradient traction force in the editing area is calculated. Then, dynamic truncation and attenuation are performed on the traction force that exceeds the preset force.

[0021] In the subsequent geometric refinement and material generation process, this application continuously uses the preserved areas in the original model as a reference to flexibly protect these areas, thus preserving their original geometric details and material properties.

[0022] Secondly, this application also provides a high-fidelity 3D asset local editing apparatus based on sparse voxels, the apparatus including a module for performing the high-fidelity 3D asset local editing method based on sparse voxels as described in any of the preceding claims. Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the high-fidelity 3D asset local editing method based on sparse voxels as described in any of the preceding claims.

[0023] This invention provides a high-fidelity 3D asset local editing method and apparatus based on sparse voxels. Addressing issues such as unclear editing boundaries, positional drift after local structure generation, and damage to details and materials in non-edited areas during existing high-fidelity 3D asset local editing processes, this invention proposes a complete technical solution from region division, local structure generation, position correction to subsequent protection. This solution ensures that the edited area can complete structural changes while suppressing the spread of editing disturbances to surrounding areas, ultimately improving the accuracy, stability, and fidelity of the local editing results.

[0024] Through the aforementioned technical means, this invention can better preserve the geometric structure, texture details, and physical material properties of non-editing areas while realizing local topological modifications in the editing area, thereby improving the accuracy, stability, and fidelity of local editing of high-fidelity 3D assets.

[0025] Compared with the prior art, the beneficial effects of the above-mentioned technical solutions provided by the embodiments of the present invention include at least the following: 1. It can achieve both local editing and overall preservation, that is, while performing local replacement, deletion or shape adjustment, it can better maintain the original state of the unedited area; 2. By directly incorporating the information of the editing area into the underlying 3D representation of the 3D model, it is possible to more accurately distinguish between the editable parts and the retained parts, thereby improving the accuracy and stability of editing under complex boundary conditions; 3. After the local structure is generated, the position of the newly generated local structure is corrected and then merged using the non-edited area as a reference. This can reduce the problems of misalignment, gaps, overlap and local drift between the edited area and the original subject, and improve the structural integrity and naturalness of the editing result. 4. By continuously referencing the information of the non-editable areas in the original 3D model and providing flexible protection for these areas, problems such as surface detail smoothing, texture distortion, color deviation, and degradation of material properties such as roughness and metallicity can be reduced, thereby better preserving the high-frequency details and material effects of the unedited areas.

[0026] Combining spatial positioning correction and generation process protection: the former is used to solve the problem of the structure not connecting after editing, and the latter is used to solve the problem of the details and materials of the unedited area being damaged. The two work together to achieve both local editability and overall high fidelity. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0028] Figure 1 This is an example flowchart of a high-fidelity 3D asset local editing method based on sparse voxels provided in an embodiment of the present invention. Detailed Implementation

[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] First, let's explain the terminology used in this embodiment: High-fidelity 3D assets can be understood as 3D models that have been completed and are rich in detail. They not only have shape information, but also material information such as color, surface roughness, and metallic feel. Partial editing refers to modifying only a part of the model, such as replacing a part, deleting an auxiliary structure, or changing a local shape, while the remaining areas that are not specified for modification should remain unchanged as much as possible.

[0031] While related local editing schemes based on sparse voxels or native 3D latent space have improved editing efficiency, they still have the following drawbacks: First, the lack of accurate and physically alignable underlying boundary representations between the edited and non-edited regions makes it difficult to achieve stable editing at complex boundaries; second, after the target region completes topological evolution, the generated result is prone to spatial drift relative to the original asset, leading to misalignment in subsequent fusion; third, in the subsequent shape and texture generation process, the non-edited region lacks continuous and flexible latent space constraints, making it difficult to effectively preserve high-frequency geometric details and native physical material features.

[0032] To address this, this invention discloses a high-fidelity 3D asset local editing method based on sparse voxels. The core idea involves clearly distinguishing between "editable regions" and "non-editable regions," then allowing changes only to the editable regions while simultaneously using two layers of protection to constrain the non-editable regions. The first layer of protection operates in 3D space to address issues such as positional offsets and misalignments between the edited structure and the original model. The second layer of protection operates during the generation process to prevent areas not originally intended for modification from being damaged during subsequent generation, thereby preserving as much of the original model's geometric details and material effects as possible.

[0033] This invention does not rely on complex inversion optimization for each object to be edited, and can achieve high-fidelity local editing on the basis of existing 3D generation and editing frameworks. Therefore, it is suitable for application scenarios such as game model editing, digital content production, virtual display, and 3D asset reuse.

[0034] In one specific embodiment, this application provides a high-fidelity 3D asset local editing method based on sparse voxels, comprising: Receive editing instructions for the 3D model to be edited, the editing instructions specifying the local area to be edited and the editing target; The local region is embedded into the voxel representation of the 3D model to be edited; In response to the editing command, the evolutionary driving force of the local region is determined, and the morphology of the local region is driven to undergo gradual evolution according to the evolutionary driving force to form a new local structure; Using the uneditable region in the 3D model to be edited as the positional constraint reference, the new local structure is fused with the uneditable region; during the fusion process, the morphological deviation of the uneditable region is monitored, and a correction constraint is generated based on the morphological deviation to suppress the unexpected changes of the uneditable region; Finally, the target 3D model with partial editing is output.

[0035] In one alternative implementation, the 3D model to be edited and the editing instructions required by the user are first obtained, wherein the editing instructions specify the local area to be edited and the editing target. For example, the editing instructions may be to replace a component, delete a local structure, or change a part to another form.

[0036] In this embodiment, the editing command divides the 3D model to be edited into two parts: the area to be modified and the area not to be changed. Optionally, the former is defined as the editing area and the latter as the non-editing area.

[0037] Furthermore, the editing region is embedded into the voxel representation of the 3D model to be edited, the steps of which include: Based on a visual recognition algorithm, the target editing outline is automatically and accurately drawn on the surface of the 3D model to be edited according to the editing instructions; in this embodiment, the visual recognition algorithm adopts the existing P3-SAM algorithm. In order to completely incorporate the boundaries of the editing area into the 3D model, the 3D model to be edited is converted into a discretized voxel representation; that is, the 3D model to be edited is decomposed and converted into countless tiny micro cubes. Traditional micro cubes only store geometric shapes and colors and materials. However, this embodiment further converts the outlined editing area information into explicit "editing labels". The editing labels include at least a first type of label indicating that modification is allowed and a second type of label indicating that modification is prohibited.

[0038] In some implementations, the number "1" represents that the voxel is in a region where modification is allowed, and "0" represents that the voxel belongs to a region that must be preserved unchanged; at the same time, the edit tag is directly concatenated and appended to the end of the underlying data of each corresponding voxel.

[0039] This application, through a data fusion operation that directly "forces labeling" at the underlying microstructure, allows for absolute boundary awareness when performing any subsequent reconstruction calculations. It only needs to read the numerical label at the end of each square to accurately lock onto and freely replace squares with "1" while fixing squares with "0" as inviolable rigid references. This completely solves the industry problem of accidental changes and damage to surrounding components due to unclear boundary divisions during previous editing processes, thus completely resolving the issue from the underlying data source.

[0040] In one alternative implementation, after the region is divided, local editing begins. This embodiment does not rewrite the entire model; instead, it uses the original model as a foundation and only moves the edited area in the desired direction. For example, a local area might have originally been a wheel, which might be replaced with another component; or it might have had a decoration, which might be removed. The task at this stage in this embodiment is to gradually grow the specified area into a new shape while minimizing disruption to the original overall structure.

[0041] Understandingly, a 3D model is generated in a high-dimensional latent space. Initially, it starts as a bunch of random noise points or random numbers. Then, based on given conditions such as an image, it gradually evolves from noise with no information into a feature vector in the latent space with some information. Decoding and reconstruction then yields the specific 3D model result. Therefore, the modification process actually modifies the changes made during the gradual denoising process to obtain a feature vector from the initial noise.

[0042] Accordingly, in this embodiment, after determining the local modification region, shape changes are guided through velocity differential within the implicit data space. Specifically, in response to editing instructions, the evolutionary driving force of the local region is determined, and based on the evolutionary driving force, the morphology of the local region is driven to undergo gradual evolution to form a new local structure.

[0043] In some specific implementation schemes, the steps include: In each tiny time step, a first evolution trend and a second evolution trend of the local region are predicted in parallel according to the editing instructions; the first evolution trend represents the evolution direction of maintaining the current form, and the second evolution trend represents the evolution direction of approaching the editing target; Then, the velocity difference between the evolution direction representing the current form and the evolution direction representing the editing target is dynamically calculated, and this velocity difference is used as a numerical driving force to progressively drive the local region to deform.

[0044] This embodiment achieves efficient and in-depth local structure replacement and elimination by continuously comparing and superimposing the difference values, without going through a tedious and time-consuming reverse deduction process, while keeping the main structure unchanged.

[0045] Meanwhile, the actual physical shape change in this embodiment is not just a change in surface color, but a modification at the local topology level, such as adding, replacing, or removing a local structure.

[0046] In one alternative implementation, the newly generated local results are positionally corrected to avoid misalignment with the original model. When a new shape grows in a specified area, it cannot be directly stitched with the original 3D model because the newly generated local structure often exhibits an overall positional shift due to the lack of absolute coordinate anchor points in the implicit space.

[0047] To address this challenge, this application proposes using the uneditable region in the 3D model to be edited as a positional constraint reference, and merging the new local structure with the uneditable region; that is, using the uneditable region in the 3D model to be edited as an absolutely static fixed reference, and finding the matching position with the highest degree of overlap between the new local structure and the fixed reference through spatial transformation; The new local structure is then merged with the 3D model to be edited at the matching location.

[0048] Optionally, after finding the optimal matching position, the merging is not performed directly and crudely, because the generation process often generates some unwanted scattered fragments or floating noise. In order to effectively identify and remove these, some implementations automatically evaluate the connectivity of the new local structure; discrete units that exist in isolation in space, have no topological connection with the main structure, and have a geometric size smaller than a preset threshold are judged as invalid noise and removed from the fusion result; finally, only the purified and completely aligned new main structure is retained, and it is adaptively fused with the preserved skeleton of the original 3D model, thereby completely eliminating splicing gaps and misalignments, and ensuring the structural integrity and naturalness of the edited 3D model.

[0049] This embodiment improves the integrity and naturalness of the structure after local editing by using the preserved area as a reference to correct the position of the editing result and then merging it.

[0050] In one alternative implementation, during subsequent generation, morphological deviations of the uneditable region are monitored, and corrective constraints are generated based on the morphological deviations to suppress unexpected changes in the uneditable region, continuously protect the reserved region, and prevent details and materials from being damaged.

[0051] Understandably, generating a 3D model is not merely a process of "initial shape formation." It also requires completing the model with finer surface undulations and gradually generating complex physical material information such as color, roughness, and metallic luster. During this global generation phase, without intervention, areas not intended for modification will be automatically "redrawn." This manifests as the original clear micro-textures of the reserved areas being smoothed out, and material colors turning black or reflecting light incorrectly, significantly degrading the overall visual appeal. To completely resolve this issue, this embodiment introduces a second core protection mechanism.

[0052] In a preferred embodiment, the mechanism includes: In each cycle of refining materials and details, the deviation energy of the uneditable area in the current editing result relative to the uneditable area in the original 3D model to be edited is continuously calculated, using the uneditable area as a reference benchmark. For example, based on the current incomplete data state, the algorithm model pre-calculates and "predicts" the final generated image result, retrieves the "reserved area" range, and automatically compares the differences between the "predicted result" and the "reference standard" within this reserved range to obtain a quantified error value, i.e., error energy. The larger the error energy value, the more severely the reserved area is deviated.

[0053] After confirming the error energy, this embodiment does not adopt a simple and crude "direct copy and paste" approach to forcibly overwrite the data, because forcibly injecting the original data would result in extremely abrupt fault lines or band-like visual defects at the boundary between the old and new data. Instead, this application converts the deviation energy into a gradient traction force used to correct the deviation. To ensure that the pull-back action is gentle enough and does not disrupt the overall generation pattern, the average intensity of the gradient traction force within the editing area is further calculated. When it is found that the traction force for a certain data point is too strong, a dynamic limiting mechanism is automatically triggered to proportionally weaken and cut off the excessive force, thereby suppressing unexpected deformation of the non-editable area.

[0054] The intelligent flow-limiting traction provided in this embodiment can be applied to the current generated data extremely smoothly and continuously, gently pulling the deviated generation trajectory back to the original model state in each small step. Through this flexible and coordinated calculation, it can accurately lock and restore the original fine structure and physical material properties in the preserved area, while ensuring that the newly generated edited parts and the preserved main body are visually integrated, truly achieving a high-fidelity editing effect of "only modifying the specified parts, without ever damaging the whole".

[0055] Through the aforementioned continuous reference protection, the present invention can better preserve the original fine structure, color distribution, and physical material effects of the unedited area; at the same time, since the entire change process always uses the original model as a reference, this application is more likely to preserve the main body parts of the original model that have not been specified for modification compared to directly regenerating the entire model.

[0056] In one optional implementation, the final output is the partially edited target 3D model. In this embodiment, the output meets two requirements: firstly, the edited area has undergone local replacement, deletion, or morphological changes according to the user's requirements; secondly, the unedited areas retain as much of the original model's geometry, texture details, and material representation as possible. From an overall perspective, this invention does not simply pursue "creating something new," but emphasizes "only modifying what needs to be modified and preserving what shouldn't be modified." Therefore, it is more suitable for local editing scenarios of high-fidelity 3D assets.

[0057] In related technologies, the high-fidelity 3D asset local editing method based on sparse voxels in this application can be applied to a device or computer-readable storage medium. Exemplarily, the device includes a server and a terminal device. The server can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, in-vehicle device, super mobile personal computer, netbook, or cellular phone, etc.; the terminal device can be a virtual reality display device, mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, in-vehicle device, super mobile personal computer, netbook, or cellular phone, etc. This application embodiment does not impose any special limitations on the specific form of the device.

[0058] Furthermore, in a specific embodiment, the implementation process of the present invention will be described using the editing of a local component of a high-fidelity 3D vehicle body asset as an example. This 3D asset includes the vehicle body, wheels, supporting structures, and surface material information, possessing both a relatively complete geometric structure and material attributes such as color, surface roughness, and metallic texture. Now, it is necessary to replace a local wheel component in this 3D asset, while the vehicle body, supporting components, and other unspecified modification areas should remain as unchanged as possible. The specific editing process is described below. Figure 1 .

[0059] First, the system inputs the original 3D asset and receives the corresponding editing requirements. Based on these requirements, the system divides the original 3D asset into regions, designating the area to be replaced (the wheel area) as the editing region and the main body of the vehicle, supporting structures, and other areas as non-editing regions. Subsequently, the system writes the region information, along with the geometric and material information of the original 3D asset, into a unified underlying 3D representation. This ensures that subsequent processing can always distinguish which areas are allowed to be modified and which should be retained for reference.

[0060] After the region is divided, the system performs local structure generation based on the original 3D asset. According to the editing requirements, the edited area is gradually formed into new local component shapes, thus obtaining the replaced local structure. During this process, the non-editable area also participates in the generation as part of the input; therefore, the non-editable area will undergo slight changes during this process.

[0061] After generating a new local structure, the system further corrects the position of this local result. Since the newly generated local structure may shift relative to the original 3D asset during the generation process, this embodiment uses the non-editable region as a fixed reference to search for the spatial position of the newly generated part, finding a position that better matches the original main body. After determining a suitable position, the newly generated local structure is merged with the original non-editable region, and isolated noise blocks or unreasonable fragments generated during the merging process are removed, ensuring that the merged overall 3D structure remains continuous and natural. In this way, inconsistencies in the non-editable region caused by the previous generation result are avoided.

[0062] After completing position correction and fusion, the system proceeds to the subsequent geometry and material generation stage. To prevent detail loss, texture distortion, or material degradation in unedited areas during subsequent processing, this embodiment continuously uses the unedited areas of the original 3D asset as a reference, applying flexible protection to these areas to preserve as much of the original high-frequency geometric details, color information, and material properties such as surface roughness and metallicity as possible during subsequent generation. Unlike directly forcibly replacing the original unedited areas back into the result, this embodiment uses a continuous and coordinated approach to protect the unedited areas, thus maintaining the overall consistency of the editing result while reducing the possibility of collateral damage to the unedited areas.

[0063] After the above processing, the system outputs the edited 3D asset. In the output, the originally designated wheel area has been partially replaced, while the main body of the vehicle and the remaining unedited areas still retain their original structural shape and material effects. Therefore, this invention can reduce interference with surrounding areas during the editing process while achieving local topology modification, minimize misalignment between old and new structures, and improve the accuracy, stability, and fidelity of local editing of high-fidelity 3D assets.

[0064] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0065] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A high-fidelity 3D asset local editing method based on sparse voxels, characterized in that, include: Receive editing instructions for the 3D model to be edited, the editing instructions specifying the local area to be edited and the editing target; The local region is embedded into the voxel representation of the 3D model to be edited; In response to the editing command, the evolutionary driving force of the local region is determined, and the morphology of the local region is driven to undergo gradual evolution according to the evolutionary driving force to form a new local structure; Using the uneditable region in the 3D model to be edited as the positional constraint reference, the new local structure is fused with the uneditable region; during the fusion process, the morphological deviation of the uneditable region is monitored, and a correction constraint is generated based on the morphological deviation to suppress the unexpected changes of the uneditable region; Output the target 3D model after partial editing.

2. The high-fidelity 3D asset local editing method based on sparse voxels as described in claim 1, characterized in that, Embedding the local region into the voxel representation of the 3D model to be edited includes: The 3D model to be edited is converted into a discretized voxel representation; Based on a visual recognition algorithm, the target editing outline is automatically drawn on the surface of the 3D model to be edited according to the editing instructions; Edit labels are assigned to each voxel in the voxel representation according to the target edit contour; the edit labels include at least a first type of label indicating that modification is allowed and a second type of label indicating that modification is prohibited.

3. The high-fidelity 3D asset local editing method based on sparse voxels as described in claim 1, characterized in that, In response to the editing instruction, determine the evolutionary driving force of the local region; including: Within one or more time steps, a first evolution trend and a second evolution trend of the local region are predicted in parallel according to the editing instructions; the first evolution trend represents the evolution direction of maintaining the current form, and the second evolution trend represents the evolution direction of approaching the editing target; Calculate the difference between the first evolutionary trend and the second evolutionary trend, and use the difference as the evolutionary driving force.

4. The high-fidelity 3D asset local editing method based on sparse voxels as described in claim 3, characterized in that, The step of calculating the difference between the first evolutionary trend and the second evolutionary trend, and using the difference as the evolutionary driving force, specifically includes: Within each micro-time step, the velocity difference between the evolution direction representing the current morphology and the evolution direction representing the editing target is dynamically calculated, and this velocity difference is used as a numerical driving force to progressively drive the local region to deform.

5. The high-fidelity 3D asset local editing method based on sparse voxels as described in claim 1, characterized in that, The process of merging the new local structure with the non-editable region includes: Using the non-editable region in the 3D model to be edited as a fixed reference, spatial transformation is used to find the matching position with the highest degree of overlap between the new local structure and the fixed reference. The new local structure is fused with the 3D model to be edited according to the matching position.

6. The high-fidelity 3D asset local editing method based on sparse voxels as described in claim 5, characterized in that, The fusion of the new local structure with the non-editable region further includes: During the process of spatially matching the new local structure with the fixed reference object, the connectivity state of the new local structure is automatically evaluated; Discrete units that exist in isolation in space, have no topological connection with the main structure, and have a geometric size smaller than a preset threshold are identified as invalid noise and removed from the fusion result.

7. The high-fidelity 3D asset local editing method based on sparse voxels as described in claim 1, characterized in that, Monitoring the shape deviation of the non-editable region and generating correction constraints based on the shape deviation, including: Using the uneditable region as a reference, calculate the deviation energy of the uneditable region in the current editing result relative to the uneditable region in the original 3D model to be edited; The deviation energy is converted into a correction constraint to suppress unexpected deformation of the non-editable region.

8. The high-fidelity 3D asset local editing method based on sparse voxels as described in claim 7, characterized in that, Converting the deviation energy into correction constraints includes: In each iteration of material and detail generation cycle, the predicted state after the final editing is inferred by using the intermediate state data that is not yet fully edited; High-fidelity data of the non-editable regions in the original 3D model to be edited are used as a reference standard; The predicted state and the reference standard are compared one by one within the uneditable region to generate a quantified error energy value, which is used to characterize the degree of deviation of the uneditable region from the intended editing. The error energy value is converted into a gradient traction force to correct the deviation, and the average intensity of the gradient traction force in the editing area is calculated. Then, dynamic truncation and attenuation are performed on the traction force that exceeds the preset force.

9. A high-fidelity 3D asset local editing device based on sparse voxels, characterized in that, Includes a module for performing the sparse voxel-based high-fidelity 3D asset local editing method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the high-fidelity 3D asset local editing method based on sparse voxels as described in any one of claims 1 to 8.