Three-dimensional model generation method and apparatus, computer device, and storage medium

By generating multi-view image sets and combining them with a low-rank adaptation module and a 3D reconstruction model, the style and details of the 3D model are optimized, solving the problem of generating high-fidelity 3D models that conform to a specific style in existing technologies, and realizing efficient and personalized 3D model generation.

CN120747356BActive Publication Date: 2026-05-29PING AN TECH (BEIJING) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (BEIJING) CO LTD
Filing Date
2025-06-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently generate 3D models that meet specific style requirements and possess high-fidelity 3D details from a single 2D image input, especially in areas like finance and digital healthcare where customized needs are difficult to meet.

Method used

By generating a multi-view image set with a consistent style, using a low-rank adaptation module for consistent style training, and combining it with a pre-set 3D reconstruction model, an initial 3D model is generated. Then, by rendering and refining the 2D reference image, the geometry and texture of the initial 3D model are optimized to obtain the target 3D model.

Benefits of technology

It enables the efficient generation of 3D models that meet specific style requirements and have high-fidelity 3D details with a small amount of input information, solving the problems of difficult style customization and low efficiency, and meeting the needs of the financial industry and digital healthcare for high-quality and personalized virtual avatars.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747356B_ABST
    Figure CN120747356B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional model generation method and device, computer equipment and a storage medium. The method generates a multi-view image set with consistent style based on input images through computer vision and machine learning technology; the multi-view image set is used for consistent style training on a preset low-rank adaptive module to obtain a trained low-rank adaptive module; an initial three-dimensional model with a target style is generated by combining the trained low-rank adaptive module and a preset first three-dimensional reconstruction model; an intermediate three-dimensional model is generated based on the multi-view image set, and the intermediate three-dimensional model is rendered according to a plurality of preset views to obtain a plurality of two-dimensional reference images. The application can be applied to financial technology, digital medical and other business scenarios. The application combines the injection of a target style with high-fidelity detail optimization to ensure that the finally generated three-dimensional model not only accurately maintains the expected style, but also presents rich geometric structures and clear textures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of 3D reconstruction technology and financial technology, and particularly to a 3D model generation method, apparatus, computer equipment, and storage medium. Background Technology

[0002] Currently, generating 3D virtual humans from single 2D images is an important task in the field of computer vision, with broad application prospects in video games, metaverse, and virtual reality. Among existing technologies, fractional distillation sampling (SDS) is one of the mainstream methods. This type of method typically uses a pre-trained 2D diffusion model to generate high-quality 2D images, and combines projection relationships with the SDS loss function to guide the optimization of an initial 3D model (e.g., a model based on 3D Gaussian Splatting (3DGS) or Neural Radiance Field (NeRF), thereby gradually generating richly detailed 3D virtual images, such as 3D virtual humans.

[0003] However, existing methods still face challenges in generating 3D virtual avatars that meet specific business needs, such as generating 3D avatars with a specific brand style or marketing campaign theme in the financial industry. Similarly, in the field of digital healthcare, such as customizing a personalized virtual therapist avatar for an online consultation platform or rehabilitation guidance application, there is also the challenge of efficiently achieving style customization while ensuring the realism of the model's 3D details. Standard diffusion models generate results with diverse styles, making it difficult to directly meet precise customization needs, often requiring extensive trial-and-error adjustments or complex model retraining. Furthermore, how to efficiently combine the style consistency of 2D images with the high-quality geometric and textural details of 3D models, especially when only a single input image is available, remains a technical problem that needs to be solved. In other words, existing technologies lack a method that can efficiently generate 3D models (such as virtual avatars) that both meet specific style requirements and possess high-fidelity 3D details based on limited input information. Summary of the Invention

[0004] This invention provides a method, apparatus, computer device, and storage medium for generating three-dimensional models, aiming to efficiently generate three-dimensional models that meet specific style requirements and possess high-fidelity three-dimensional details based on a small amount of input information.

[0005] In a first aspect, embodiments of the present invention provide a method for generating a three-dimensional model. The method is applied to a three-dimensional model generation system and includes the following steps: generating a multi-view image set with a consistent style based on an input image; training a preset low-rank adaptation module with the consistent style using the multi-view image set to obtain a trained low-rank adaptation module; combining the trained low-rank adaptation module with a preset first three-dimensional reconstruction model to generate an initial three-dimensional model with a target style; generating an intermediate three-dimensional model based on the multi-view image set; rendering the intermediate three-dimensional model according to multiple preset perspectives to obtain multiple two-dimensional reference images; refining each of the multiple two-dimensional reference images to generate multiple two-dimensional supervision images; and optimizing the geometry and texture of the initial three-dimensional model using the multiple two-dimensional supervision images to obtain a target three-dimensional model.

[0006] Secondly, embodiments of the present invention also provide a three-dimensional model generation apparatus, comprising: a multi-view generation unit for generating a multi-view image set with a consistent style based on an input image; a style adaptation training unit for training a preset low-rank adaptation module with a consistent style using the multi-view image set to obtain a trained low-rank adaptation module; an initial model generation unit for combining the trained low-rank adaptation module with a preset first three-dimensional reconstruction model to generate an initial three-dimensional model with a target style; an intermediate model processing unit for generating an intermediate three-dimensional model based on the multi-view image set, and rendering the intermediate three-dimensional model according to preset multiple views to obtain multiple two-dimensional reference images; an image refining unit for refining the multiple two-dimensional reference images to generate multiple two-dimensional supervision images; and a model optimization unit for optimizing the geometry and texture of the initial three-dimensional model using the multiple two-dimensional supervision images to obtain a target three-dimensional model.

[0007] Thirdly, embodiments of the present invention also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0008] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the method described in the first aspect.

[0009] This invention provides a method, apparatus, computer device, and storage medium for generating a 3D model. The method includes: generating a multi-view image set with a consistent style based on an input image; training a preset low-rank adaptation module with the consistent style using the multi-view image set to obtain a trained low-rank adaptation module; combining the trained low-rank adaptation module with a preset first 3D reconstruction model to generate an initial 3D model with a target style; generating an intermediate 3D model based on the multi-view image set; rendering the intermediate 3D model according to multiple preset viewpoints to obtain multiple 2D reference images; refining each of the multiple 2D reference images to generate multiple 2D supervision images; and optimizing the geometry and texture of the initial 3D model using the multiple 2D supervision images to obtain a target 3D model.

[0010] This invention first generates a multi-view image set with a consistent style based on the input image, overcoming the limitation of insufficient information from a single input image and laying the foundation for accurate style extraction and reliable reference for 3D details. Next, a low-rank adaptation module of the multi-view image set is used for consistent style training. This trained low-rank adaptation module is then combined with a first 3D reconstruction model to generate an initial 3D model with the target style. This series of operations achieves efficient capture and injection of a specific artistic style, avoiding the complex and time-consuming retraining of the entire model for different styles. This significantly improves the efficiency and flexibility of style customization, solving the problems of difficulty and low efficiency in style customization in the prior art.

[0011] Meanwhile, to ensure high detail quality in the final model, this embodiment generates an intermediate 3D model in parallel based on a multi-view image set. The intermediate 3D model is then rendered to obtain a 2D reference image, which is further refined to generate a 2D supervisory image. This approach focuses on generating a 2D visual reference rich in high-quality geometric and textural details, providing strong guidance for subsequent 3D model optimization.

[0012] Finally, by utilizing 2D supervised images, the geometry and texture of the initial 3D model are optimized to obtain the target 3D model. This method fuses the initial model, which already possesses the target style, with high-quality detail information. This strategy of separating style injection and detail enhancement to a certain extent and then performing targeted optimization ensures that the final generated target 3D model not only accurately maintains the specific artistic style desired by the input image but also presents rich geometric structures and clear texture details. Therefore, this embodiment of the invention can efficiently generate high-quality 3D models that meet specific style requirements and possess high-fidelity 3D details with only a small amount of input information, effectively solving the contradiction of style and detail being difficult to achieve simultaneously in the prior art. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A flowchart illustrating the three-dimensional model generation method provided in an embodiment of the present invention;

[0015] Figure 2 A schematic diagram of a sub-process of the three-dimensional model generation method provided in an embodiment of the present invention;

[0016] Figure 3 A schematic block diagram of a three-dimensional model generation device provided in an embodiment of the present invention;

[0017] Figure 4 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0020] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0021] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0022] Please see Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating the three-dimensional model generation method provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of a sub-process of the 3D model generation method provided in an embodiment of the present invention. The 3D model generation method is applied to 3D model generation systems, such as in scenarios like virtual customer service generation in the financial industry or digital healthcare, and virtual avatar customization in metaverse marketing activities.

[0023] In existing technologies, generating 3D virtual humans or objects from single-view images faces a core challenge: how to precisely control the specific artistic style of the final 3D model, such as color, material, and clothing features, while ensuring the model possesses rich geometric details and high-definition textures. Traditional approaches, such as training large-scale 3D generative models from scratch for a specific style, consume enormous computational resources and are time-consuming, making them unsuitable for rapid customization needs. Relying solely on simple style transfer techniques often results in the 3D model losing important geometric and textural details while maintaining the style, making the model appear lifeless and failing to meet the needs of financial institutions for high-quality, personalized virtual avatars in digital marketing, for example.

[0024] The 3D model generation method proposed in this embodiment aims to effectively solve the aforementioned problems through a series of collaborative steps, namely, to generate high-quality 3D models with specific styles and rich details from a single or a small number of input images at a lower cost and with higher efficiency.

[0025] Figure 1 This is a flowchart illustrating the three-dimensional model generation method provided in an embodiment of the present invention. As shown in the figure, the method includes the following steps S100-600.

[0026] Step S100: Generate a multi-view image set with a consistent style based on the input image.

[0027] In this embodiment, a bank needs to customize a virtual financial advisor avatar for its wealth management business. The designer provides a pre-designed frontal image of the virtual character as input. This image reflects the bank's desired professional and approachable brand image, including specific clothing style, facial features, and overall color scheme.

[0028] After receiving the input image, the 3D model generation system (hereinafter referred to as the system) feeds it into a pre-trained multi-view diffusion model. This multi-view diffusion model can understand the 3D structural information of the input image and generate images of the virtual character from different perspectives, such as four different perspectives: left 45-degree angle, right 45-degree angle, and back view. These generated multi-view images maintain consistent style features with the input image, including clothing style, color saturation, and lighting conditions, forming a style-consistent multi-view image set.

[0029] Step S200: Use the multi-view image set to perform consistent style training on the preset low-rank adaptation module to obtain the trained low-rank adaptation module.

[0030] In this embodiment, the system uses the multi-view image set generated in step S100 as training data to train the low-rank adaptation module. The low-rank adaptation module learns and encodes specific style information of the virtual financial advisor by analyzing common features in these images, including brand color preferences, clothing texture features, and facial expression styles. After training, the module can parameterize these style features to form reusable style codes. Here, the predefined low-rank adaptation module refers to one whose structure is already defined, whose parameters are in an initial state, and which is ready to learn and encode specific styles before specific style training.

[0031] Step S300: Combine the trained low-rank adaptation module with the preset first 3D reconstruction model to generate an initial 3D model with the target style.

[0032] In this embodiment, the system combines the trained low-rank adaptation module with the pre-trained first 3D reconstruction model. Specifically, the parameters of the low-rank adaptation module are injected into the network structure of the first 3D reconstruction model, enabling the originally general 3D reconstruction model to generate 3D models with specific styles.

[0033] Using this fused model to process multi-view image sets, the system generated an initial 3D virtual financial advisor model. This initial 3D model successfully maintained the brand style required by the bank, but details such as facial texture and clothing folds still need improvement. The pre-trained first 3D reconstruction model, also known as the preset first 3D reconstruction model, refers to a pre-selected and prepared, pre-trained deep learning model with general 3D reconstruction capabilities.

[0034] Step S400: Generate an intermediate 3D model based on the multi-view image set, and render the intermediate 3D model according to multiple preset viewpoints to obtain multiple 2D reference images.

[0035] In this embodiment, to obtain reference information for detail enhancement, the system uses another 3D reconstruction pipeline focused on geometric and textural details in parallel. The same multi-view image set is input into the second 3D reconstruction model to generate an intermediate 3D model. While this intermediate 3D model may not be as consistent in style as the initial 3D model, it performs better in terms of geometric accuracy and textural detail.

[0036] The system renders the central 3D model from multiple preset standard perspectives, such as front, left, right, and back, to obtain a set of 2D reference images. These reference images contain rich detail information, providing important references for subsequent optimization.

[0037] Step S500: Refine the multiple two-dimensional reference images to generate multiple two-dimensional supervision images.

[0038] In this embodiment, to further improve detail quality, the system refines the two-dimensional reference image obtained in step S400. This process enhances visual quality indicators such as texture detail and edge sharpness in the reference image through image enhancement. The refined image, as a two-dimensional supervisory image, contains high-quality detail information, such as clear facial features and realistic clothing textures.

[0039] Step S600: Using multiple two-dimensional supervised images, optimize the geometry and texture of the initial three-dimensional model to obtain the target three-dimensional model.

[0040] In this embodiment, the system utilizes the high-quality two-dimensional supervisory image generated in step S500 to optimize the initial three-dimensional model generated in step S300. The specific process is as follows:

[0041] First, the system renders the initial 3D model from the same perspective as the 2D supervised image, obtaining the corresponding rendered image. Then, the difference between the rendered image and the 2D supervised image is calculated, reflecting the deficiencies in detail representation of the initial 3D model.

[0042] Based on the calculated differences, the system adjusts the geometric and texture parameters of the initial 3D model using an optimization algorithm. This optimization process is iterative; each iteration reduces the difference between the rendered image and the supervised image, gradually improving the detail representation of the 3D model.

[0043] After multiple rounds of optimization, the system finally generated the target 3D model. This model not only perfectly maintained the brand style characteristics required by the bank, but also had rich details, such as realistic facial expressions, fine clothing textures, and natural lighting effects.

[0044] In summary, this embodiment, through the coordinated operation of the above steps, firstly efficiently injects the target style into the initial 3D model using a low-rank adaptation module, then independently generates high-quality 2D detail references, and finally uses these references to supervise and optimize the initial model, thereby effectively solving the problem of single-view... Figure 3 When generating 3D models, it is difficult to balance the contradiction between customizing specific styles and presenting high details. This provides an effective way for industries such as finance to quickly generate high-quality, personalized 3D virtual images.

[0045] In another embodiment, besides generating virtual customer service or virtual financial advisors in the financial industry as described above, this method can also be applied to digital healthcare scenarios. For example, a professional virtual doctor assistant avatar can be customized for the online consultation platform of a large general hospital. This virtual doctor assistant can be used in scenarios such as the hospital's triage system, initial patient consultation, health knowledge dissemination, medication guidance, or postoperative rehabilitation follow-up. By providing friendly and professional interaction, it improves the patient service experience and the efficiency of medical information transmission, and alleviates the consultation pressure on medical staff. In this embodiment, the frontal concept design drawing of the virtual doctor assistant provided by the hospital is used as the input image. The system first generates a set of images containing multiple perspectives of the doctor assistant with a consistent style based on the input image. Then, it uses this image set to train a low-rank adaptation module to capture its unique professional style, and combines the trained module with the first 3D reconstruction model to generate an initial 3D model that initially possesses the doctor's style. At the same time, the system also generates a more detailed intermediate 3D model based on the multi-view image set through a second 3D reconstruction model, and renders a 2D baseline image. These 2D baseline images are refined to enhance details, forming multiple 2D supervision images. Finally, using these high-quality 2D supervision images, the geometry and texture of the initial 3D model are meticulously optimized, resulting in the final target virtual doctor assistant model that maintains the professional doctor's style while presenting rich 3D details.

[0046] In one embodiment, step S100, generating a multi-view image set with a consistent style based on the input image, includes:

[0047] The input image is input into a multi-view diffusion model, and the multi-view diffusion model generates multiple images of the input image from different viewpoints that are consistent with the style of the input image, thus obtaining the multi-view image set.

[0048] In this embodiment, the 3D model generation system first receives a frontal image of a virtual financial advisor provided by the designer. This input image forms the basis for all subsequent generation work, accurately defining the core visual elements of the virtual financial advisor and the brand perception to be conveyed.

[0049] The system takes this single frontal image as input and feeds it into a pre-trained multi-view diffusion model. This model has been trained on a large number of images with different people, professions, and styles, along with their corresponding multi-view data, enabling it to infer the 3D structure of a single image and generate new perspective images. In this example, the model can understand the professional characteristics of a virtual financial advisor, their expected professional image, and their approximate geometric shape.

[0050] After receiving a frontal image of a virtual financial advisor, the multi-view diffusion model begins the generation process. The model's internal mechanism captures defined stylistic features from the input image, such as the style of the business suit, facial expressions, and overall color tone, and strictly maintains these features when generating new perspectives.

[0051] The advantage of multi-view diffusion models lies in their ability to maintain a high degree of style consistency when generating images from different perspectives. Specifically, the model continuously references the style features of the input image during the generation process. For example, when generating an image from a 45-degree angle on the left, the model not only needs to correctly represent the side profile of the person but also ensure that the color and gloss of the clothing are consistent with the frontal image. Similarly, when generating a back view, although the back details cannot be directly observed from the input image, the model reasonably infers and generates a back image that conforms to the overall style based on learned prior knowledge.

[0052] In this embodiment, the multi-view diffusion model generates four images from different perspectives for the virtual financial advisor: front, left 45-degree angle, right 45-degree angle, and back. These four images together constitute the multi-view image set. Each generated image strictly maintains the stylistic features consistent with the input image to reproduce the core visual elements of the virtual financial advisor defined in the original input image and the brand perception to be conveyed.

[0053] The multi-view image set generated in this way provides a rich and consistent source of information for subsequent 3D reconstruction. Compared with traditional single-view reconstruction, multi-view information can significantly reduce ambiguity in 3D reconstruction and improve reconstruction quality. At the same time, since the images from all perspectives maintain a high degree of stylistic consistency, the final 3D model can present the desired brand image when viewed from any angle, ensuring visual consistency of the virtual wealth advisor in various application scenarios.

[0054] In this embodiment, the multi-view diffusion model employs a Transformer-based architecture to capture and preserve style features. Upon receiving a frontal image of a virtual financial advisor, the model first converts the image into a high-dimensional feature representation using an encoder network. During this process, a Transformer-based attention mechanism plays a crucial role: it automatically identifies and focuses on important regions in the image, such as the financial advisor's face, key parts of the clothing (e.g., suit lapels, buttons), and the overall color scheme.

[0055] In this specific application scenario, when the model analyzes the input image of a financial advisor, the attention mechanism assigns different weights to different image regions. For example, for the key style element of a dark blue suit, the model forms a strong representation in the feature space, recording information such as the suit's specific hue (RGB values), texture (inferred from lighting and shadow), and tailoring style (identified through contour lines). Simultaneously, for facial features, the model captures professional yet approachable facial expression details, such as the slightly upturned corners of the mouth and confident eyes—micro-expression features.

[0056] When generating new perspective images, the model ensures stylistic consistency through a conditional generation process. When generating a side view, instead of simply rotating the image, the model understands the suit's structure in three-dimensional space—it knows how the shoulder line should appear from the side, how the lapel's three-dimensionality should be expressed, and how to maintain the same deep blue hue under different lighting angles. This understanding ensures that the generated side image conforms to realistic perspective while perfectly preserving the original design's stylistic intent.

[0057] In one embodiment, step S200, which involves using the multi-view image set to perform consistent style training on a preset low-rank adaptation module to obtain the trained low-rank adaptation module, includes:

[0058] The multi-view image set is used as training data, and the low-rank adaptation module is trained by fine-tuning, so that the low-rank adaptation module learns and encodes the style features in the multi-view image set, thus obtaining the trained low-rank adaptation module.

[0059] In this embodiment, the low-rank adaptation module used is specifically a low-rank adaptive (LoRA) module. Continuing from the multi-view image set generated in the previous step, which includes images of the virtual financial advisor from multiple different perspectives, the system uses this as dedicated training data to fine-tune an initial low-rank adaptation module. Here, the low-rank adaptation module is attached to a specific layer of a pre-trained deep learning model with image generation capabilities. This deep learning model can be referred to as the pre-trained model, which is also the first 3D reconstruction model described in subsequent step S300.

[0060] During fine-tuning, the original weight parameters of the pre-trained model remain frozen, and only the low-rank parameters within the low-rank adaptation module are trained and optimized. The low-rank adaptation module is configured to learn and capture consistent style features common to all images in the multi-view image set. These style features are key elements defining the unique visual identity of the virtual financial advisor, such as clothing style, color scheme, material representation, facial features and expressions, lighting, and atmosphere.

[0061] By iteratively training on these multi-view images rich in style information, the low-rank adaptation module can effectively learn and encode (or compress or parameterize) these complex style features into its internal low-rank parameters, namely the low-rank decomposition of the weight matrix. This process does not require training a massive model from scratch; instead, it specifically learns new styles by injecting and optimizing the low-rank adaptation matrix into specific layers of a pre-trained model, such as attention or feedforward layers in a Transformer architecture, or relevant layers in a diffusion model like the U-Net structure.

[0062] After training, a trained low-rank adaptation module is obtained. This module has internalized the style-specific information about the virtual financial advisor extracted from the multi-view image set, forming an efficient and reusable style encoding. This trained low-rank adaptation module can be considered a plugin or adapter for this specific style. It can be combined with the first 3D reconstruction model in subsequent steps to give the model the ability to generate an initial 3D model with the target style, namely the style specific to the virtual financial advisor, without requiring a complete retraining of the massive first 3D reconstruction model. This significantly reduces the cost and time of style customization.

[0063] In one embodiment, S300, the step of combining the trained low-rank adaptation module with a preset first 3D reconstruction model to generate an initial 3D model with the target style, includes:

[0064] S310. The weights of the trained low-rank adaptation module are fused with the weights of the pre-trained first 3D reconstruction model to obtain a fusion model with style encoding capability.

[0065] In this embodiment, a low-rank adaptation module, which has been trained and has learned and encoded the specific style information of the virtual financial advisor, is included. Simultaneously, the system pre-configures a first 3D reconstruction model. This first 3D reconstruction model is a pre-trained model, i.e., the pre-trained model mentioned above.

[0066] The weight fusion process specifically refers to applying the parameters, i.e., the weights, of the low-rank adaptation matrix learned from the trained low-rank adaptation module to the corresponding layer of the first 3D reconstruction model. For example:

[0067] If the low-rank adaptation module is stored as a separable weight file, then when loading the first 3D reconstruction model, this trained low-rank adaptation module is loaded at the same time, and during the forward propagation of the model, the weight changes introduced by the low-rank adaptation module are added to the original weights of the corresponding layer of the pre-trained model.

[0068] In this way, the style-specific information learned by the low-rank adaptation module is effectively injected or grafted into the first 3D reconstruction model. This enables the first 3D reconstruction model, which originally generated a general style, to now specifically generate the target style of a virtual financial advisor. This model, after weighted fusion, can be called a fusion model. This fusion model retains the original 3D reconstruction capabilities and image feature understanding of the first 3D reconstruction model, and with the support of the low-rank adaptation module, it possesses the ability to accurately encode and reproduce a specific target style.

[0069] S320. The multi-view image set is reconstructed using the fusion model to generate an initial three-dimensional model with the target style.

[0070] In this embodiment, after obtaining a fusion model with style encoding capabilities, the system uses the multi-view image set generated in step S100 as input to feed the fusion model.

[0071] When processing these multi-view images, the fusion model integrates stylistic guidance from a low-rank adaptation module. This means that when parsing the input image and generating the geometry and texture of the 3D model, the model strictly follows the style characteristics of the virtual financial advisor encoded by the trained low-rank adaptation module. For example, when generating texture maps for the model, it ensures that the clothing color is the specific deep blue learned by the low-rank adaptation module; when shaping facial geometry, it tends to generate friendly and professional facial expressions that conform to the low-rank adaptation module's learning.

[0072] Ultimately, the fusion model outputs an initial 3D model. This initial 3D model reflects the target style shared with the original input image and the multi-view image set in terms of 3D geometry and surface texture—that is, the specific brand style of the virtual financial advisor. For example, when viewed from different angles, the clothing, hairstyle, skin tone, and expression of the generated initial 3D model should maintain a high degree of consistency with the style of the input image.

[0073] It's important to emphasize that while this initial 3D model already possesses the correct target style, its geometric details and texture refinement may still have room for improvement. For example, subtle wrinkles on the face and the details of the fabric texture in the clothing may not be perfect. This is precisely what needs further optimization in subsequent steps. However, at this stage, ensuring the accuracy and consistency of the style is the primary goal, and this is precisely achieved through the effective combination of the low-rank adaptation module and the first 3D reconstruction model.

[0074] In one embodiment, step S400 involves generating an intermediate 3D model based on the multi-view image set, and rendering the intermediate 3D model according to multiple preset viewpoints to obtain multiple 2D reference images, including:

[0075] S410. Input the multi-view image set into a preset second three-dimensional reconstruction model to generate the intermediate three-dimensional model.

[0076] In this embodiment, in parallel or independently of step S300, which generates an initial 3D model with the target style, the system also executes a process for acquiring high-quality geometric and textural details. In this process, the system also uses the multi-view image set generated in step S100, namely, multi-view images of the virtual financial advisor from the front, left 45 degrees, right 45 degrees, and back, as input.

[0077] However, the input this time is a second 3D reconstruction model. This second 3D reconstruction model may differ from the first 3D reconstruction model mentioned in step S300. Its main design goal and optimization direction focus more on accurately recovering the 3D geometric structure and generating high-fidelity texture details from multi-view images, rather than prioritizing the injection of a specific artistic style.

[0078] By inputting the multi-view image set into this second 3D reconstruction model, the model processes these images and reconstructs an intermediate 3D model. The core advantage of this intermediate 3D model lies in the accuracy of its geometry and the richness of its texture details. Although this intermediate 3D model may not be strongly style-guided by the low-rank adaptation module during generation like the initial 3D model, and therefore its overall style may lean towards realism or the model's default style, potentially differing somewhat from the specific brand style of the virtual financial advisor, it provides a high-quality geometric and texture blueprint for subsequent optimization.

[0079] S420. Render the intermediate three-dimensional model from multiple preset perspectives to obtain multiple two-dimensional reference images.

[0080] In this embodiment, after the intermediate 3D model is generated through the second 3D reconstruction model, the system will render the intermediate 3D model from a set of pre-set or selected virtual camera perspectives.

[0081] These preset viewpoints can be consistent with the viewpoints used when generating the multi-view image set, such as front, left 45 degrees, right 45 degrees, and back, or other specific viewpoints that help to showcase the details of the model. The rendering process projects the three-dimensional geometric and texture information of the intermediate 3D model onto a two-dimensional plane, generating a series of two-dimensional images.

[0082] These two-dimensional images rendered from the intermediate 3D model are called 2D reference images. Because the intermediate 3D model itself has high geometric accuracy and rich texture details, these 2D reference images also inherit these advantages. They contain clearer and more detailed visual information than images directly rendered from the initial 3D model, such as more defined contour edges and richer surface details.

[0083] This set of multiple 2D baseline images will serve as the basis for the refining process in subsequent step S500, and will ultimately be used to generate 2D supervisory images to guide the optimization of the initial 3D model. Their core value lies in providing a high-quality, richly detailed 2D visual reference.

[0084] In one embodiment, step S500, refining the multiple two-dimensional reference images to generate multiple two-dimensional supervision images, includes:

[0085] S510. Add preset noise to each of the multiple two-dimensional reference images to obtain multiple noisy images.

[0086] In this embodiment, the system takes the detailed two-dimensional reference images generated in step S420, which are multiple two-dimensional images rendered from the intermediate three-dimensional model. To further improve the quality of these images, especially to enhance their detail and fix some minor imperfections that may occur during the rendering process, the system first performs a preprocessing step on these two-dimensional reference images: adding noise.

[0087] For each 2D reference image, the system applies a pre-defined amount of random noise. The type (e.g., Gaussian noise) and intensity (noise level) of this noise are controllable parameters. The purpose of adding noise is to better stimulate the diffusion model to generate richer and more natural details in subsequent denoising steps. Image quality is improved through a process of destruction followed by reconstruction.

[0088] After this step, each 2D reference image is transformed into a corresponding noisy image. These noisy images are visually slightly blurry or grainy than the original 2D reference image, but they prepare for the subsequent refining steps.

[0089] S520. Denoise the multiple noisy images to generate multiple two-dimensional supervised images with enhanced details.

[0090] After obtaining the noisy images, the system inputs them one by one into a pre-trained diffusion model. This diffusion model is used for image denoising and image super-resolution or detail enhancement tasks. Trained on a large number of clean images and their corresponding noisy images, it has learned how to recover clean, clear, and detailed images from noisy images.

[0091] When a noisy 2D reference image is input into the diffusion model, the model performs a denoising process. In this process, the diffusion model does more than simply remove noise; it also understands and reconstructs the content of the image based on its learned prior knowledge of the image. This allows it to generate more refined details than the original input, i.e., the 2D reference image before noise addition, and to repair some potential flaws, thus improving the overall visual quality of the image. For example, it can make clothing textures more clearly visible, facial features more vivid, and edges sharper.

[0092] The images output after the diffusion model processing are the two-dimensional supervised images. Compared with the previous two-dimensional baseline images, these two-dimensional supervised images have significantly reduced noise levels; richer and clearer details, such as enhanced visual elements like textures and edges; and higher overall visual quality, potentially appearing more natural and realistic.

[0093] This set of refined 2D supervisory images, generated from multiple perspectives, will serve as powerful guiding signals for optimizing the geometry and texture of the initial 3D model in subsequent step S600 due to their high quality and rich detail. They provide crucial visual supervision for the final generation of a high-quality target 3D model.

[0094] In one embodiment, step S600, using multiple two-dimensional supervised images, optimizes the geometry and texture of the initial three-dimensional model to obtain a target three-dimensional model, including:

[0095] S610. Render the initial three-dimensional model from multiple preset perspectives to obtain multiple rendered images.

[0096] In this embodiment, the system first obtains the initial 3D model generated in step S300. This initial 3D model already possesses the target style guided by the low-rank adaptation module, such as the specific brand style of a virtual financial advisor, but may still be insufficient in terms of geometric details and texture refinement.

[0097] To optimize this initial 3D model, the system renders it from a set of preset virtual camera perspectives. These preset perspectives should be strictly consistent with the perspectives used when generating the 2D supervision image in step S520, i.e., the perspectives used to render the intermediate 3D model in step S420. This ensures that the rendered image can be directly compared with the 2D supervision image from the corresponding perspective when calculating the loss later.

[0098] Through this rendering process, the system obtains a set of 2D images generated from the initial 3D model from multiple different viewpoints, which can be called rendered images. Each rendered image corresponds to a specific viewpoint and reflects the geometric appearance and texture of the initial 3D model under that viewpoint.

[0099] S620. Calculate the loss function between the multiple rendered images and the two-dimensional supervised image from the corresponding viewpoint.

[0100] In this embodiment, the system compares each rendered image generated in step S610 with the corresponding two-dimensional supervisory image generated in step S520 from the same viewpoint.

[0101] The comparison is performed by calculating a loss function between the rendered image and the 2D supervised image. This loss function quantifies the difference between the rendered image and the 2D supervised image. For multiple images from different viewpoints, the loss values ​​calculated for each viewpoint can be weighted and averaged or summed to obtain a total loss value. This total loss value comprehensively reflects the overall difference between the current initial 3D model and the high-quality 2D supervised image across all supervised viewpoints. The greater the difference, the higher the loss value.

[0102] S630. Based on the loss function, the geometric and texture parameters of the initial 3D model are optimized by gradient descent to obtain the target 3D model.

[0103] In this embodiment, after calculating the value of the loss function, the system will use a gradient-based optimization algorithm to adjust the parameters of the initial 3D model in order to minimize the loss function.

[0104] The parameters of the initial 3D model include:

[0105] Geometric parameters: For example, if the model uses a representation such as point cloud, mesh, neural radiation field (NeRF) or 3D Gaussian Splatting, the geometric parameters may include the coordinates of the points, the positions of the mesh vertices, the network weights of the neural field, the center position / rotation / scaling of the Gaussian sphere, etc.

[0106] Texture parameters: such as pixel values ​​of the texture map, vertex colors, network weights that control color in the neural field, color / transparency of the Gaussian sphere, etc.

[0107] The optimization process is iterative. In each iteration: based on the parameters of the current initial 3D model, S610 is executed to render the rendered image; S620 is executed to calculate the loss between the rendered image and the 2D supervised image; the gradient of the loss function with respect to the geometric and texture parameters of the initial 3D model is calculated; based on the calculated gradient, the parameters of the initial 3D model are updated according to the rules of the optimization algorithm.

[0108] Through continuous iteration of this process, the geometry of the model will gradually approach the accurate shape reflected by the two-dimensional supervised image, and the texture will become clearer and richer, thereby gradually reducing the difference between the rendered image and the two-dimensional supervised image.

[0109] The optimization process stops when the loss function converges to a certain extent or after reaching a preset number of iterations. The resulting 3D model, optimized for geometry and texture, is the target 3D model to be generated in this embodiment of the invention. This target 3D model not only maintains the correct target style initially injected by the low-rank adaptation module, but also, guided by 2D supervised images, significantly improves its geometric details, such as subtle changes in facial expressions and the naturalness of clothing folds, as well as its texture quality, such as the realism of skin and the clarity of materials, thereby achieving high-quality 3D model generation.

[0110] Figure 3 This is a schematic block diagram of a three-dimensional model generation device provided in an embodiment of the present invention. Figure 3 As shown, corresponding to the above-described 3D model generation method, the present invention also provides a 3D model generation device 100. This device can be configured in terminals such as desktop computers, tablet computers, and laptops. For details, please refer to... Figure 3 The device includes a multi-view generation unit 110, a style adaptation training unit 120, an initial model generation unit 130, an intermediate model processing unit 140, an image refinement unit 150, and a model optimization unit 160.

[0111] A bank needed to create a virtual financial advisor avatar for its wealth management business. Designers provided a pre-designed frontal image of the virtual character as input, reflecting the bank's desired professional and approachable brand image, including specific clothing style, facial features, and overall color scheme.

[0112] The multi-view generation unit 110 is used to generate a set of multi-view images with a consistent style based on the input image. Specifically, the multi-view generation unit 110 is configured to input the input image into a pre-trained multi-view diffusion unit after receiving it. This multi-view diffusion unit can understand the three-dimensional structural information of the input image and generate images of the virtual character from different perspectives, such as generating images from four different perspectives: left 45 degrees, right 45 degrees, and back. These generated multi-view images maintain consistent style features with the input image, including clothing style, color saturation, lighting conditions, etc., forming a set of multi-view images with a consistent style.

[0113] The style adaptation training unit 120 is used to train a pre-defined low-rank adaptation module with consistent style using the multi-view image set, resulting in a trained low-rank adaptation module. Specifically, the style adaptation training unit 120 is configured to use the multi-view image set generated by the multi-view generation unit 110 as training data to train the low-rank adaptation module. The low-rank adaptation module learns and encodes specific style information of the virtual bank financial advisor by analyzing common features in these images, including brand color preferences, clothing texture features, and facial expression styles. After training, the module can parameterize these style features to form a reusable style code.

[0114] The initial model generation unit 130 is used to combine the trained low-rank adaptation module with a pre-set first 3D reconstruction model to generate an initial 3D model with a target style. The initial model generation unit 130 is configured to combine the trained low-rank adaptation module with the pre-trained first 3D reconstruction unit. Specifically, the parameters of the low-rank adaptation module are injected into the network structure of the first 3D reconstruction unit, enabling the originally general 3D reconstruction model to acquire the ability to generate 3D models with a specific style.

[0115] Using this fused unit to process the multi-view image set, the initial model generation unit 130 generated an initial 3D virtual financial advisor model. This initial 3D virtual financial advisor model successfully maintained the brand style required by the bank, but details such as facial texture and clothing folds still need improvement.

[0116] The intermediate model processing unit 140 is used to generate an intermediate 3D model based on the multi-view image set, and to render the intermediate 3D model according to multiple preset viewpoints to obtain multiple 2D reference images. The intermediate model processing unit 140 is configured to input the same multi-view image set into the second 3D reconstruction unit to generate an intermediate 3D model. Although this intermediate 3D model may not be as consistent in style as the initial 3D model, it performs better in terms of geometric accuracy and texture detail.

[0117] The intermediate model processing unit 140 renders the intermediate 3D model from multiple preset standard perspectives, such as front, left, right, and back, to obtain a set of 2D reference images. These reference images contain rich detail information, providing important references for subsequent optimization.

[0118] The image refining unit 150 is used to refine multiple two-dimensional reference images to generate multiple two-dimensional supervision images. The image refining unit 150 is configured to refine the two-dimensional reference images obtained by the intermediate model processing unit 140. This process improves visual quality indicators such as texture details and edge sharpness in the reference images through image enhancement. The refined images, as two-dimensional supervision images, contain high-quality detail information, such as clear facial features and realistic clothing textures.

[0119] The model optimization unit 160 is used to optimize the geometry and texture of the initial 3D model using multiple 2D supervised images to obtain the target 3D model. The model optimization unit 160 also uses the high-quality 2D supervised images generated by the image refining unit 150 to optimize the initial 3D model generated by the initial model generation unit 130. The specific process is as follows:

[0120] First, the model optimization unit 160 renders the initial 3D model from the same perspective as the 2D supervised image, obtaining the corresponding rendered image. Then, the difference between the rendered image and the 2D supervised image is calculated, reflecting the deficiencies in detail representation of the initial 3D model.

[0121] Based on the calculated differences, the model optimization unit 160 adjusts the geometric and texture parameters of the initial 3D model using an optimization algorithm. This optimization process is iterative; each iteration reduces the difference between the rendered image and the supervised image, gradually improving the detail representation of the 3D model.

[0122] After multiple rounds of optimization, the final target 3D model was obtained. This model not only perfectly maintains the brand style characteristics required by the bank, but also has rich details, such as realistic facial expressions, fine clothing textures, and natural lighting effects.

[0123] In summary, this embodiment, through the collaborative work of the aforementioned units, firstly efficiently injects the target style into the initial 3D model via a low-rank adaptation module, then independently generates high-quality 2D detail references, and finally utilizes these references to supervise and optimize the initial model, thereby effectively solving the problem of single-view... Figure 3 When generating 3D models, it is difficult to balance the contradiction between customizing specific styles and presenting high details. This provides an effective way for industries such as finance to quickly generate high-quality, personalized 3D virtual images.

[0124] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned device and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.

[0125] The above-described device can be implemented as a computer program, which can be used in, for example... Figure 4 It runs on the computer device shown.

[0126] Please see Figure 4 , Figure 4 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.

[0127] See Figure 4 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.

[0128] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a method.

[0129] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.

[0130] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a method.

[0131] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0132] The processor 502 is used to run a computer program 5032 stored in a memory to implement the three-dimensional model generation method of the present invention.

[0133] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0134] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0135] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions implement the three-dimensional model generation method of the embodiments of the present invention.

[0136] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0137] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0138] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0139] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0140] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0141] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating a three-dimensional model, the method being applied to a three-dimensional model generation system, characterized in that, The method includes the following steps: Generate a multi-view image set with a consistent style based on the input image; The pre-defined low-rank adaptation module is trained with a consistent style using the multi-view image set to obtain the trained low-rank adaptation module. By combining the trained low-rank adaptation module with the preset first 3D reconstruction model, an initial 3D model with the target style is generated, including: The weights of the trained low-rank adaptation module are fused with the weights of the pre-trained first 3D reconstruction model to obtain a fusion model with style encoding capability. The fusion model is used to perform 3D reconstruction processing on the multi-view image set to generate an initial 3D model with the target style. An intermediate 3D model is generated based on the multi-view image set, and the intermediate 3D model is rendered according to multiple preset viewpoints to obtain multiple 2D reference images. The multiple two-dimensional reference images are refined to generate multiple two-dimensional supervision images, including: Preset noise is added to each of the multiple two-dimensional reference images to obtain multiple noisy images; The multiple noisy images are denoised to generate multiple two-dimensional supervised images with enhanced details; Using multiple 2D supervised images, the geometry and texture of the initial 3D model are optimized to obtain the target 3D model, including: The initial 3D model is rendered from multiple preset perspectives to obtain multiple rendered images; Calculate the loss function between the multiple rendered images and the corresponding two-dimensional supervised images; Based on the loss function, the geometric and texture parameters of the initial 3D model are optimized by gradient descent to obtain the target 3D model.

2. The method according to claim 1, characterized in that, The generation of a multi-view image set with a consistent style based on the input image includes: The input image is input into a multi-view diffusion model, and the multi-view diffusion model generates multiple images of the input image from different viewpoints that are consistent with the style of the input image, thus obtaining the multi-view image set.

3. The method according to claim 2, characterized in that, The step of using the multi-view image set to perform consistent style training on a preset low-rank adaptation module to obtain the trained low-rank adaptation module includes: The multi-view image set is used as training data, and the low-rank adaptation module is trained by fine-tuning, so that the low-rank adaptation module learns and encodes the style features in the multi-view image set, thus obtaining the trained low-rank adaptation module.

4. The method according to claim 1, characterized in that, The intermediate 3D model is generated based on the multi-view image set, and the intermediate 3D model is rendered according to multiple preset viewpoints to obtain multiple 2D reference images, including: The multi-view image set is input into a preset second three-dimensional reconstruction model to generate the intermediate three-dimensional model; The intermediate 3D model is rendered from multiple preset perspectives to obtain multiple 2D reference images.

5. A three-dimensional model generation device, characterized in that, include: A multi-view generation unit is used to generate a set of multi-view images with a consistent style based on the input image; The style adaptation training unit is used to perform consistent style training on the preset low-rank adaptation module using the multi-view image set to obtain the trained low-rank adaptation module. An initial model generation unit, used to combine the trained low-rank adaptation module with a preset first 3D reconstruction model, generates an initial 3D model with the target style, including: The weights of the trained low-rank adaptation module are fused with the weights of the pre-trained first 3D reconstruction model to obtain a fusion model with style encoding capability. The fusion model is used to perform 3D reconstruction processing on the multi-view image set to generate an initial 3D model with the target style. An intermediate model processing unit is used to generate an intermediate 3D model based on the multi-view image set, and to render the intermediate 3D model according to multiple preset viewpoints to obtain multiple 2D reference images. An image refining unit is used to refine multiple two-dimensional reference images respectively to generate multiple two-dimensional supervision images, including: Preset noise is added to each of the multiple two-dimensional reference images to obtain multiple noisy images; The multiple noisy images are denoised to generate multiple two-dimensional supervised images with enhanced details; The model optimization unit is used to optimize the geometry and texture of the initial 3D model using multiple 2D supervised images to obtain the target 3D model, including: The initial 3D model is rendered from multiple preset perspectives to obtain multiple rendered images; Calculate the loss function between the multiple rendered images and the corresponding two-dimensional supervised images; Based on the loss function, the geometric and texture parameters of the initial 3D model are optimized by gradient descent to obtain the target 3D model.

6. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-4.

7. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, implement the method as described in any one of claims 1-4.