Rendering method, device, equipment and storage medium for fusion of model and scene
By displaying the skeleton model in a three-dimensional scene and using a large language model for background fusion, the target fusion image is generated, and the existing three-dimensional modeling technology is solved, with a more realistic three-dimensional image and a lower threshold for use.
Patent Information
- Application Number
- CN202510082345.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing three-dimensional modeling technology has high operating threshold and high cost, and the rendering effect is unreal, which reduces the user experience.
By displaying the skeleton model in a three-dimensional scene, the large language model is used to fuse the skeleton model with the rendering renderings of the three-dimensional scene to generate a target fusion image.
It improves the authenticity and user experience of three-dimensional images, reduces the difficulty and cost of modeling, and lowers the threshold for users to use.
Smart Images

Figure CN119516078B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a rendering method, device, equipment and storage medium for fusion of model and scene. Background Art
[0002] In the field of 3D modeling, 3D modeling is usually used to add 3D models. For example, after adding some highly complex 3D models such as people and animals to the 3D scene, the 3D model is integrated with the 3D scene using 3D rendering technology. However, the above process has high operational thresholds and high cost inputs. For example, a large amount of professional technical reserves are required, and the modeling cost is high. At the same time, the modeling difficulty of the above process is high, which further increases the modeling cost. Moreover, the rendering effect of the above process is not realistic. For example, for rendering human models, it is highly complex and users are highly sensitive to it. Most of the existing rendering methods have unrealistic rendering effects, which reduces the user experience. Summary of the invention
[0003] The present disclosure provides a rendering method, device, equipment and storage medium for fusion of model and scene to solve or alleviate one or more technical problems in the prior art.
[0004] In a first aspect, the present disclosure provides a rendering method for fusing a three-dimensional model with a three-dimensional scene, comprising:
[0005] Display three-dimensional scenes;
[0006] A skeleton model representing a skeleton structure of a three-dimensional model is displayed in a three-dimensional scene;
[0007] In response to the adjustment confirmation operation on the skeleton model, a fusion image generation request is sent to the server to obtain a target fusion image, wherein the target fusion image is obtained by fusing the model rendering with the rendering rendering of the three-dimensional scene; the model rendering is obtained by using a large language model to perform background fusion between the three-dimensional model corresponding to the skeleton model and the rendering rendering of the three-dimensional scene.
[0008] In a second aspect, the present disclosure provides a rendering device for fusing a three-dimensional model with a three-dimensional scene, comprising:
[0009] The client is used to display a three-dimensional scene, display a skeleton model representing the skeleton structure of the three-dimensional model in the three-dimensional scene; and in response to an adjustment confirmation operation on the skeleton model, send a fusion image generation request to the server to obtain a target fusion image;
[0010] The server is used to obtain the fused image generation request, obtain and send the target fused image;
[0011] Among them, the target fusion image is obtained by fusing the model rendering and the rendering rendering of the three-dimensional scene; the model rendering is obtained by using a large language model to perform background fusion of the three-dimensional model corresponding to the skeleton model and the rendering rendering of the three-dimensional scene.
[0012] In a third aspect, an electronic device is provided, including:
[0013] at least one processor; and
[0014] a memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.
[0016] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute any method according to the embodiments of the present disclosure.
[0017] In a fifth aspect, a computer program product is provided, comprising a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.
[0018] The beneficial effects of the technical solution provided by the present disclosure include at least:
[0019] In this way, since the target fusion image obtained by the disclosed solution is obtained by fusing the model rendering and the rendering of the three-dimensional scene, and the model rendering is obtained by using the large language model, the effect of the target fusion image is more realistic, which greatly improves the user experience. Moreover, since the disclosed solution effectively uses the image processing capabilities of the large language model, compared with the existing 3D rendering technology, the disclosed solution effectively reduces the difficulty of modeling, thereby providing strong support for effectively reducing the cost of modeling. In addition, the disclosed solution does not require users to have professional knowledge, which greatly reduces the user's usage threshold and operation difficulty, thereby improving the user experience.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments provided according to the present disclosure and should not be regarded as limiting the scope of the present disclosure.
[0022] Figure 1 This is a schematic flow chart of a rendering method for fusing a three-dimensional model with a three-dimensional scene according to an embodiment of the present application. Figure 1 ;
[0023] Figure 2 is a schematic diagram of a skeleton model for describing the bone structure of a three-dimensional character model according to an embodiment of the present application;
[0024] Figure 3 This is a schematic flow chart of a rendering method for fusing a three-dimensional model with a three-dimensional scene according to an embodiment of the present application. Figure 2 ;
[0025] Figure 4 is a schematic diagram of a character model setting interface according to an embodiment of the present application;
[0026] Figure 5 is a schematic diagram of a scene after an adjustment confirmation operation on a skeleton model according to an embodiment of the present application;
[0027] Figure 6 A target mask image for indicating the area where the skeleton model is located in the rendering effect image according to an embodiment of the present application;
[0028] Figure 7 is a target posture graph of a three-dimensional model corresponding to a skeleton model according to an embodiment of the present application;
[0029] Figure 8 is a target depth map for representing depth information of a three-dimensional model corresponding to a skeleton model in a three-dimensional scene according to an embodiment of the present application;
[0030] Fig. 9 It is a schematic diagram of an image generation process in a rendering method for fusing a three-dimensional model with a three-dimensional scene according to an embodiment of the present application;
[0031] Fig.10 The present invention is a flowchart of a method for rendering a 3D model and a 3D scene according to an embodiment of the present invention in a specific example. Figure 1 ;
[0032] Fig.11 The present invention is a flowchart of a method for rendering a 3D model and a 3D scene according to an embodiment of the present invention in a specific example. Figure 2 ;
[0033] Fig.12 is a 3D rendering of a three-dimensional scene according to an embodiment of the present application;
[0034] Fig.13 is a schematic diagram of the structure of a rendering device for fusing a three-dimensional model with a three-dimensional scene according to an embodiment of the present application;
[0035] Fig.14 The present invention is a block diagram of an electronic device for implementing the rendering method for fusing a three-dimensional model with a three-dimensional scene according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0036] The present disclosure will be further described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0037] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure.
[0038] The disclosed solution provides a rendering method for the fusion of a three-dimensional model and a three-dimensional scene, which can add some highly complex three-dimensional models (such as people, animals, etc.) to a specified three-dimensional scene, thereby obtaining a three-dimensional image with a more realistic and natural image effect at a low cost; moreover, the disclosed solution does not require users to have professional knowledge, so the usage threshold is low, thereby improving the user experience.
[0039] Specifically, Figure 1 This is a schematic flow chart of a rendering method for fusing a three-dimensional model with a three-dimensional scene according to an embodiment of the present application. Figure 1 The method may be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0040] Further, the method includes at least part of the following contents. Figure 1 As shown, including:
[0041] Step S101: displaying a three-dimensional scene.
[0042] Step S102: displaying a skeleton model representing the skeleton structure of the three-dimensional model in the three-dimensional scene.
[0043] For instance, in one example, the three-dimensional model may be a three-dimensional character model, an animal model, or a three-dimensional building model, etc., and the present disclosure does not impose any specific limitation on this.
[0044] Furthermore, it should be noted that the skeleton model may refer to a basic model used to describe the skeleton structure of a three-dimensional model to simulate the overall structure of the three-dimensional model. For example, in one example, the skeleton model may be specifically Figure 2 The basic model model used to describe the skeleton structure of the three-dimensional character model shown in the figure can mainly show the basic torso features of the character model; or, in another example, the skeleton model can also be specifically a model used to describe the architectural structure of the three-dimensional building model to simulate the architectural structure of the building model. In this way, it is convenient for subsequent users to adjust the posture of the skeleton model in the three-dimensional scene.
[0045] Furthermore, in a specific example, after displaying the skeleton model, the skeleton model in the three-dimensional scene may be adjusted, for example, the position and posture of the model may be adjusted, so that the three-dimensional model required by the user may be obtained.
[0046] Step S103: In response to the adjustment confirmation operation on the skeleton model, a fused image generation request is sent to the server to obtain a target fused image.
[0047] Here, the target fused image is obtained by fusing the model rendering and the rendering of the three-dimensional scene; in other words, the target fused image is an image fused with the model rendering and the rendering.
[0048] Furthermore, in one example, the model rendering is obtained by using a large language model to perform background fusion between the three-dimensional model corresponding to the skeleton model (that is, the skeleton model after the adjustment confirmation operation) and the rendering rendering of the three-dimensional scene.
[0049] Here, it should be noted that, in addition to being a large language model, the model used to generate the model rendering can also be an image generation model, such as a stable diffusion (SD) model, etc., and the present disclosure does not impose any specific restrictions on this.
[0050] In this way, since the target fusion image obtained by the disclosed solution is obtained by fusing the model rendering and the rendering rendering of the three-dimensional scene, and the model rendering is obtained using a large language model, the effect of the target fusion image is more realistic, greatly improving the user experience.
[0051] Moreover, since the disclosed solution effectively makes use of the image processing capabilities of the large language model, compared with the existing 3D rendering technology, the disclosed solution effectively reduces the difficulty of modeling, thereby providing strong support for effectively reducing the cost of modeling.
[0052] In addition, the disclosed solution does not require users to have professional knowledge, which greatly reduces the user's usage threshold and operation difficulty, thereby improving the user experience.
[0053] Figure 3 This is a schematic flow chart of a rendering method for fusing a three-dimensional model with a three-dimensional scene according to an embodiment of the present application. Figure 2 The method can optionally be applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices. It is understandable that the above Figure 1 and Figure 2 The relevant contents of the method shown can also be applied to this example, and the relevant contents will not be described in detail in this example.
[0054] Further, the method includes at least part of the following contents. Specifically, Figure 3 As shown, including:
[0055] Step S301: The client displays a three-dimensional scene.
[0056] Step S302: The client displays a skeleton model representing the skeleton structure of the three-dimensional model in the three-dimensional scene.
[0057] Here, for the relevant contents of the three-dimensional model and the skeleton model, please refer to the above examples, which will not be repeated here.
[0058] Step S303: the client sends a fusion image generation request to the server in response to the adjustment confirmation operation on the skeleton model.
[0059] Step S304: the server executes the rendering process, the image generation process and the fusion process in response to the fusion image generation request to obtain the target fusion image.
[0060] Here, the rendering process (which may be simply referred to as process a) is used to render a three-dimensional scene to obtain a rendering effect graph. For example, the three-dimensional scene is rendered using 3D rendering technology to obtain a rendering effect graph.
[0061] Furthermore, the image generation process (which may be abbreviated as process b) is used to utilize the large language model to perform background fusion of the three-dimensional model corresponding to the skeleton model and the rendering effect image of the three-dimensional scene to obtain the model rendering image.
[0062] Furthermore, the fusion process (which may be abbreviated as process c) is used to superimpose the model effect image and the rendering effect image to obtain a target fused image.
[0063] Step S305: The client displays the target fused image.
[0064] That is to say, after the client sends a fused image generation request to the server, the server responds to the fused image generation request and executes the rendering process, image generation process and fusion process to obtain the target fused image; further, the server sends the target fused image to the client for display in the client.
[0065] In this way, the disclosed solution can utilize multiple processing flows to obtain the target fused image. Compared with the existing solution of utilizing 3D rendering technology to obtain a three-dimensional image with a three-dimensional model, the target fused image obtained by the disclosed solution is more realistic, greatly improving the user experience. Moreover, it effectively reduces the difficulty of modeling, thereby providing strong support for effectively reducing the modeling cost.
[0066] In a specific example, the model effect diagram can be obtained in the following manner; specifically, the image generation process described above (i.e., process b described above) specifically includes:
[0067] Step b01: Determine the position information of the skeleton model after the adjustment confirmation operation in the three-dimensional scene and the target area where it is located.
[0068] Here, in this example, the posture information of the skeleton model may include the posture and position of the skeleton model in the three-dimensional scene.
[0069] Step b02: Call the large language model and generate a model rendering that is integrated with the background of the rendering based on at least the rendering rendering, the posture information of the skeleton model in the three-dimensional scene and the target area where it is located (that is, the posture information and the target area where the skeleton model is located in the three-dimensional scene after the adjustment confirmation operation), and the model description information of the three-dimensional model corresponding to the skeleton model.
[0070] For instance, in one example, the model description information of the three-dimensional model corresponding to the skeleton model may be the appearance features of the three-dimensional model. For instance, in one example, for a character model, the model description information may specifically be the appearance features of the character model. This provides support for subsequently obtaining the model renderings required by the user.
[0071] In this way, the disclosed solution can utilize the large language model in the image generation process to obtain a model rendering that is integrated with the background of the rendering rendering. In this way, there is no need to go through a complex three-dimensional modeling and rendering process, which improves design efficiency and flexibility, while also lowering the usage threshold and the required computing resources. This enables the model rendering to be generated quickly and accurately, laying the foundation for subsequent high-quality fused images (i.e., target fused images) with added three-dimensional models.
[0072] Furthermore, in a specific example, the model description information of the three-dimensional model corresponding to the skeleton model may be obtained in the following manner, specifically including:
[0073] Display a parameter setting interface for setting parameters of the three-dimensional model corresponding to the skeleton model;
[0074] In response to a setting confirmation operation on the parameter setting interface, model description information for constraining appearance features of the three-dimensional model corresponding to the skeleton model is obtained.
[0075] For example, adding a three-dimensional character model to a three-dimensional scene, such as Figure 4 As shown, the user can set the appearance features of the character model on the character model setting interface corresponding to the skeleton model (corresponding to the parameter setting interface above), such as body shape, eye color, facial expression, face shape, hairstyle, hair color and clothing, and then use the above setting items to constrain the appearance features of the character model.
[0076] In other words, the disclosed solution can utilize the parameter setting interface to constrain the appearance features of the three-dimensional model corresponding to the skeleton model, thereby meeting the user's customized needs for the appearance features of the three-dimensional model, laying the foundation for subsequently obtaining a target fusion image that meets user needs and thereby improving the user experience.
[0077] Further, in a specific example, the model rendering may be obtained in the following manner; specifically, the above-mentioned calling of the large language model and generating the model rendering integrated with the background of the rendering rendering based on at least the rendering rendering, the pose information of the skeleton model in the three-dimensional scene and the target area where the skeleton model is located, and the model description information of the three-dimensional model corresponding to the skeleton model (for example, step b02) may specifically include:
[0078] Step b02 - 1 : Based on the target area where the skeleton model is located in the three-dimensional scene after the adjustment confirmation operation, a target mask image capable of indicating the area where the skeleton model is located in the rendering effect image is obtained.
[0079] For example, continue to add a three-dimensional character model to the three-dimensional scene. After adjusting and confirming the operation, the following is obtained: Figure 5 The scene graph shown is based on Figure 5 The target area where the skeleton model is located in the scene graph shown in the figure is obtained Figure 6 The target mask diagram shown in Figure 1 provides a basis for subsequently guiding the large language model to generate model effect diagrams.
[0080] Step b02-2: Based on the posture information of the skeleton model in the three-dimensional scene after the adjustment confirmation operation, at least generate posture control information for the three-dimensional model corresponding to the skeleton model.
[0081] Step b02-3: Call the large language model, and generate a model rendering that matches the image size of the target mask image and is integrated with the background of the rendering rendering based at least on the target mask image, posture control information, rendering rendering, and model description information of the three-dimensional model corresponding to the skeleton model.
[0082] That is to say, in the process of calling the large language model to generate a model rendering, the parameters input to the large language model include not only the rendering rendering and the model description information, but also the posture control information and the target mask map obtained above; here, the target mask map is used to guide the large language model to generate a model rendering that matches the specified area in the rendering rendering; the posture control information is used to guide the large language model to generate a model rendering of a three-dimensional model with a specified posture.
[0083] Further, in an example, the above-mentioned step of generating at least the pose control information of the three-dimensional model corresponding to the skeleton model based on the pose information of the skeleton model in the three-dimensional scene after the adjustment confirmation operation (for example, step b02-2) may specifically include:
[0084] The first preprocessing network is used to identify the pose information of the skeleton model after the adjustment confirmation operation to obtain a target pose graph of the three-dimensional model corresponding to the skeleton model.
[0085] Here, the target posture graph can represent the posture control information of the three-dimensional model corresponding to the skeleton model after the adjustment confirmation operation.
[0086] For example, let's continue to add a 3D character model to a 3D scene. At this time, we can use the control network (Control Net) as a preprocessor to Figure 5 The pose information of the skeleton model in the scene graph is identified to obtain Figure 7 The target pose graph of the three-dimensional model corresponding to the skeleton model shown. That is, in one example, the input of the large language model may specifically include: a target mask graph, a target pose graph, a rendering effect graph, and model description information of the three-dimensional model corresponding to the skeleton model.
[0087] Furthermore, in order to further improve the real effect of the model rendering, the disclosed solution can also obtain image depth information, and then use the image depth information as the input of the large language model, so as to further improve the image effect. For example, in one example, the image depth information can be obtained in the following manner, specifically including:
[0088] Based on the position information of the skeleton model in the three-dimensional scene after the adjustment confirmation operation and the target area in which it is located, image depth information of the three-dimensional model corresponding to the skeleton model in the three-dimensional scene is generated.
[0089] At this time, the above-mentioned calling of the large language model and generating a model rendering that matches the image size of the target mask image and is integrated with the background of the rendering rendering image based on at least the target mask image, the posture control information, the rendering rendering image, and the model description information of the three-dimensional model corresponding to the skeleton model (for example, step b02-3) may also specifically include:
[0090] The large language model is called, and based on the image depth information, target mask map, pose control information, rendering effect map and model description information of the three-dimensional model corresponding to the skeleton model, a model rendering map that matches the image size of the target mask map and is integrated with the background of the rendering effect map is generated.
[0091] That is to say, the input of the large language model includes the relevant parameters shown in the above step b02-3, and further includes the image depth information. In this way, the model rendering generated by the large language model is more realistic, thereby laying the foundation for subsequently obtaining the target fusion image that meets user needs and improving user experience.
[0092] Here, in an example, the above-mentioned generating the image depth information of the three-dimensional model corresponding to the skeleton model in the three-dimensional scene based on the posture information of the skeleton model in the three-dimensional scene after the adjustment confirmation operation and the target area in which the skeleton model is located may specifically include:
[0093] The second preprocessing network is used to perform posture and depth recognition on the skeleton model located in the three-dimensional scene after the adjustment confirmation operation to obtain a target depth map.
[0094] Here, the target depth map can represent the depth information of the three-dimensional model corresponding to the skeleton model after the adjustment confirmation operation in the three-dimensional scene.
[0095] It should be noted that the second pre-processing network and the first pre-processing network described above can be the same processing network or two different processing networks, which is not specifically limited in the present disclosure.
[0096] For example, let's continue to add a 3D character model to a 3D scene. At this time, we can use the control network (Control Net) as a preprocessor to Figure 5 The skeleton model in the three-dimensional scene shown in the figure is used for pose and depth recognition to obtain Figure 8 The target depth map shown is capable of representing the depth information of the three-dimensional model corresponding to the skeleton model in the three-dimensional scene.
[0097] That is to say, in the process of calling the large language model to generate the model rendering, the parameters input to the large language model include not only the rendering rendering, the model description information, the target pose map and the target mask map, but also the target depth map mentioned above.
[0098] For example, Fig. 9 Schematic diagram of the image generation process in the disclosed solution; Fig. 9 As shown, the steps of the image generation process may specifically include:
[0099] Step S901: determining the position and posture information of the skeleton model after the adjustment confirmation operation in the three-dimensional scene and the target area where the skeleton model is located.
[0100] Step S902: based on the target area where the skeleton model is located in the three-dimensional scene after the adjustment confirmation operation, a target mask image capable of indicating the area where the skeleton model is located in the rendering effect image is obtained.
[0101] Step S903: using the first preprocessing network, identifying the pose information of the skeleton model after the adjustment confirmation operation, so as to obtain a target pose graph of the three-dimensional model corresponding to the skeleton model.
[0102] Step S904: using the second preprocessing network, performing posture and depth recognition on the skeleton model located in the three-dimensional scene after the adjustment confirmation operation to obtain a target depth map.
[0103] Step S905: Call the large language model, and generate a model rendering that matches the image size of the target mask image and is integrated with the background of the rendering rendering based on at least the target mask image, the target posture image, the target depth image, the rendering rendering, and the model description information of the three-dimensional model corresponding to the skeleton model.
[0104] In this way, the disclosed solution can obtain the model rendering without going through a complex 3D modeling and rendering process, thereby improving design efficiency and flexibility, while also lowering the usage threshold, laying the foundation for subsequently obtaining the target fusion image that meets user needs and improving user experience.
[0105] In a specific example of the disclosed solution, before sending a fused image generation request to a server in response to an adjustment confirmation operation on a skeleton model to obtain a target fused image, the following may also be included:
[0106] In response to the image input operation, a target reference image is obtained.
[0107] Here, the target reference image is used to constrain the appearance features (such as facial features or clothing features, etc.) of the three-dimensional model corresponding to the skeleton model during the generation process of the model rendering (such as in the image generation process); the appearance features of the three-dimensional model in the model rendering are associated with the appearance features of the target object in the target reference image.
[0108] That is to say, in one example, in addition to inputting model description information for constraining the appearance features of the three-dimensional model corresponding to the skeleton model, the user can also input a specified target reference image, and then further constrain the appearance features of the three-dimensional model corresponding to the skeleton model through the appearance features of the target object in the target reference image, thereby meeting the user's customized needs for image generation and further improving the user experience.
[0109] Furthermore, in a specific example, after obtaining the target reference image uploaded by the user, the server can also obtain the final model rendering in the following manner:
[0110] Mode 1: that is, calling the large language model as described above, and generating a model rendering that is integrated with the background of the rendering rendering based on at least the rendering rendering, the pose information of the skeleton model in the three-dimensional scene and the target area where the skeleton model is located, and the model description information of the three-dimensional model corresponding to the skeleton model (for example, step b02), which may specifically include:
[0111] Recognize the appearance features of the target object in the target reference image;
[0112] The large language model is called, and based on the appearance features of the target object in the target reference image, the rendering effect diagram, the posture information of the skeleton model in the three-dimensional scene and the target area where it is located, and the model description information of the three-dimensional model corresponding to the skeleton model, a model rendering diagram fused with the background of the rendering effect diagram is generated.
[0113] That is to say, in this example, the parameters input to the large language model may include, in addition to the relevant parameters in the above examples, the appearance features of the target object in the target reference image. In this way, the appearance features of the target object in the target reference image are used to guide the large language model to generate a model rendering whose appearance features meet user needs, thereby further meeting the user's customized needs for image generation and further improving the user experience.
[0114] For instance, in a specific example, first, the above-described method can be used to obtain a target mask map that can indicate the area where the skeleton model is located in the rendering effect image, a target posture map that represents the posture control information, and the appearance features of the target object in the target reference image can be identified. Secondly, a large language model is called, and based on the appearance features of the target object in the target reference image, the target mask map, the target posture map, the rendering effect image, and the model description information, a model rendering image that is integrated with the background of the rendering effect image is generated.
[0115] Alternatively, in another example, first, the above-described method can be used to obtain a target mask map that can indicate the area where the skeleton model is located in the rendering effect image, a target pose map that represents the pose control information, and a target depth map, as well as identify the appearance features of the target object in the target reference image. Secondly, a large language model is called, and based on the appearance features of the target object in the target reference image, the target mask map, the target pose map, the target depth map, the rendering effect image, and the model description information, a model rendering image that is integrated with the background of the rendering effect image is generated.
[0116] Method 2: After the server obtains the model rendering, it can also adjust the model rendering according to the appearance characteristics of the target reference image; specifically, it includes:
[0117] Recognize the appearance features of the target object in the target reference image;
[0118] Based on the appearance features of the target object in the target reference image, the appearance features of the three-dimensional model in the generated model rendering are adjusted to update the model rendering.
[0119] That is to say, in this example, the method described in the above example can be used to first obtain a model rendering, and then adjust the appearance features of the three-dimensional model in the obtained model rendering based on the appearance features of the target object in the target reference image, such as facial features or clothing features.
[0120] For example, in one example, the appearance features of the target object in the target reference image and the model rendering are input into the feature transfer model (e.g., the reactor (ReActor) model, etc.), so as to use the feature transfer model to replace the appearance features of the 3D model in the model rendering with the appearance features of the target object in the target reference image. In this way, the user's customized needs for image generation are further met, and the user experience is further improved.
[0121] For example, adding a three-dimensional character model to a three-dimensional scene, such as Fig.10As shown, on the client side, a skeleton model corresponding to a three-dimensional character model is added to a created three-dimensional scene, and the client responds to the user's adjustment operations, such as adjusting the position and posture of the skeleton model in the three-dimensional scene, and responds to the setting operation of the character model corresponding to the skeleton model, so as to obtain the appearance features of the character model; further, after completing the adjustment confirmation operation on the skeleton model, a fusion image generation request is sent to the server; accordingly, after responding to the fusion image generation request, the server renders the three-dimensional scene to obtain a rendering effect diagram, and uses a large language model to generate a model effect diagram of the character model, and then fuses the rendering effect diagram and the model effect diagram to obtain a target fusion image.
[0122] Here, it should be noted that the target fused image can be obtained in the following two ways:
[0123] Method 1: After obtaining the model image of the character model uploaded by the user (corresponding to the above target reference image), and determining that the clothing of the character model in the model image needs to be saved, identify the clothing features of the character model in the model image; call the large language model, and generate a model rendering image based on at least the clothing features of the character model in the model image, the target mask image, the posture control information, the rendering effect image, and the model appearance information (corresponding to the above model description information). Here, the clothing features of the character model in the model rendering image match the clothing features of the character model in the model image, for example, the clothing of the character model in the model rendering image is the clothing of the character model in the model image. Finally, the generated model rendering image is fused with the rendering effect image to obtain the target fused image.
[0124] Method 2: When a model image of a character model uploaded by a user (corresponding to the target reference image above) is obtained and it is determined that there is no need to save the clothing of the character model in the model image, the large language model is called to generate a model rendering of the character model, the facial features of the character model in the model image are identified, and the facial features of the character model in the model image are migrated to the model rendering to replace the facial features of the character model in the model rendering. For example, the facial features of the character model in the model image and the model rendering are input into a feature migration model to utilize the feature migration model to migrate the facial features of the character model in the model image to the model rendering; finally, the updated model rendering is fused with the rendering rendering to obtain a target fused image.
[0125] It can be understood that the above two feature migration methods are only exemplary. In practical applications, other processing methods can also be selected. The present disclosure does not specifically limit the details of the feature migration method.
[0126] It should be noted that, in one example, after obtaining the target fusion image, if the three-dimensional scene needs to be adjusted, such as adding a target object (such as a pillow, etc.), you can return to the scene layout interface to rearrange the three-dimensional scene and re-execute the above method. Or, in another example, determine the area where the target object to be added needs to be located in the target fusion image, call the large language model, and obtain the target fusion image containing the target object based on the area where the target object needs to be located in the target fusion image, the prompt word of the target object, and the target fusion image. In this way, the optimization of the target fusion image is achieved, thereby further improving the user experience.
[0127] The disclosed solution is further explained below with reference to specific examples; the disclosed solution proposes a rendering method for fusing a three-dimensional model with a three-dimensional scene. Specifically, a pre-designed three-dimensional scene is first rendered using 3D rendering technology to obtain a 3D rendering image (corresponding to the rendering effect image described above), and then an artificial intelligence (AI) image generation technology is used to generate a model effect image of a three-dimensional model (such as a high-complexity model, such as a person, an animal, a liquid, a flame, a smoke, an explosion, etc.) that matches a specified area in the 3D rendering image. Finally, the 3D rendering image is fused with the model effect image to obtain a final fused effect image (corresponding to the target fused image described above); in this way, a highly complex three-dimensional model is added to a three-dimensional scene in a low-cost and low-threshold manner, and the image effect of the final image is more realistic and natural, thereby improving the user experience.
[0128] Specifically, taking adding a three-dimensional character model to a three-dimensional scene as an example, Fig.11 As shown, the main steps of the rendering method include:
[0129] Step 1101: Create a three-dimensional scene.
[0130] For example, in 3D software, 3D modeling is used to build scenes for content other than three-dimensional models to obtain the desired three-dimensional scene.
[0131] Step 1102: Add a skeleton model capable of representing the skeleton structure of a three-dimensional character model to the created three-dimensional scene.
[0132] Here, the skeleton model is used to describe the virtual skeleton structure of the three-dimensional character model, such as Figure 2 As shown, the skeleton model can show the torso features of the three-dimensional character model, so that it is convenient to quickly adjust the position, posture, etc. of the skeleton model in the three-dimensional scene later.
[0133] Step 1103: adjust the posture of the skeleton model in the three-dimensional scene, set the appearance features of the skeleton model, and initiate a fused image generation request after the adjustment is confirmed.
[0134] Here, we can use Figure 5 The position and posture of the skeleton model in the three-dimensional scene can be adjusted in the manner shown. Figure 4 The character model setting interface shown is used to set the body shape, eye color, facial expression, face shape, hairstyle, hair color, clothing, etc. of the skeleton model.
[0135] Step 1104: After receiving the fused image generation request, the server executes a rendering process to render the three-dimensional scene using 3D rendering technology to obtain a 3D rendering image.
[0136] It should be noted that, in one example, before rendering the three-dimensional scene, the skeleton model (or the adjusted skeleton model) in the three-dimensional scene can be hidden first, and then the three-dimensional scene without the skeleton model can be rendered using 3D rendering technology to obtain the following: Fig.12 The 3D rendering shown does not contain a skeleton model, which can effectively improve the rendering effect.
[0137] Step 1105: The server executes an image generation process to generate a model rendering of a three-dimensional character model corresponding to the adjusted skeleton model using AI image generation technology.
[0138] Step 1105 - 1 : obtaining the position information of the adjusted skeleton model in the three-dimensional scene, the target area where the skeleton model is located, and the appearance features of the character model corresponding to the adjusted skeleton model.
[0139] Step 1105-2: According to the target area where the adjusted skeleton model is located in the three-dimensional scene, find the area matching the target area from the 3D rendering image, and perform mask processing on the area to obtain Figure 6 The mask image shown (corresponding to the target mask image described above).
[0140] Step 1105 - 3 : Obtain appearance prompt information according to the appearance features of the character model corresponding to the adjusted skeleton model.
[0141] Step 1105-4: Use the Control Net preprocessor to identify the posture information of the adjusted skeleton model to obtain Figure 7 The target posture diagram of the three-dimensional character model corresponding to the skeleton model shown in the figure, and the posture and depth recognition of the adjusted skeleton model are performed to obtain Figure 8 The target depth map is shown.
[0142] It should be pointed out that in practical applications, the control network preprocessor can also be built with an open source posture library (OpenPose), which can use the trained convolutional neural network (CNN) to detect the key points of the human body on the adjusted skeleton model, and recognize the detected key points of the human body to obtain the human posture map of the adjusted skeleton model (corresponding to the target depth map mentioned above). In this way, accurate recognition of the human posture of the skeleton model is achieved.
[0143] Here, the key points of the human body may specifically include the head, shoulders, elbows, wrists, hips, knees, etc.
[0144] Step 1105-6: Input the mask image, target posture image, target depth image, 3D rendering image and appearance prompt information into a stable diffusion (SD) model to generate a model effect image that matches the image size of the mask image and is integrated with the background of the 3D rendering image.
[0145] Step 1106: The server executes a fusion process to superimpose the 3D rendering image and the model rendering image to obtain a fusion rendering image with a three-dimensional character model added thereto.
[0146] Here, the disclosed solution can also perform secondary adjustments on the following contents in the fusion effect diagram: three-dimensional scene, character position, rendering perspective, etc. The disclosed solution does not impose specific restrictions on the contents of the secondary adjustments.
[0147] In summary, compared with the prior art of using 3D rendering technology to achieve the integration of three-dimensional models and three-dimensional scenes, the disclosed solution has the following advantages:
[0148] First, the image effect is better. Compared with the existing technology, the disclosed solution first uses 3D rendering technology to render the three-dimensional scene to obtain a rendering effect picture, then calls the large language model to generate a model effect picture of the required three-dimensional model, and then superimposes the rendering effect picture with the model effect picture, so as to obtain a high-quality and more realistic effect picture.
[0149] Second, the threshold for use is lower. The disclosed solution does not require complex 3D modeling and rendering processes to obtain the required renderings. Moreover, it does not require users to have professional rendering knowledge, which improves design efficiency and flexibility and greatly reduces the user's threshold for use and investment costs.
[0150] Third, the user experience is better. The disclosed solution abandons the method of using 3D rendering technology to achieve the fusion of different images. Instead, it uses 3D rendering technology to obtain the rendering effect of the three-dimensional scene, and uses AI image generation technology to obtain the model effect of the three-dimensional model, and then obtains the fusion image of the two. Moreover, users can also freely adjust the input of AI generation technology to obtain the final effect that meets user needs. In this way, the user's customized needs for three-dimensional image design are met, thereby improving the user experience.
[0151] The disclosed solution also provides a rendering device for integrating a three-dimensional model with a three-dimensional scene, such as Fig.13 As shown, including:
[0152] The client 1301 is used to display a three-dimensional scene, display a skeleton model representing the skeleton structure of the three-dimensional model in the three-dimensional scene; and in response to an adjustment confirmation operation on the skeleton model, send a fusion image generation request to the server to obtain a target fusion image;
[0153] The server 1302 is used to obtain the fused image generation request, obtain and send the target fused image;
[0154] Among them, the target fusion image is obtained by fusing the model rendering image with the rendering rendering image of the three-dimensional scene; the model rendering image is obtained by using a large language model to perform background fusion of the three-dimensional model corresponding to the skeleton model and the rendering rendering image of the three-dimensional scene.
[0155] In a specific example of the disclosed solution, the server is further used to execute a rendering process, an image generation process, and a fusion process in response to a fusion image generation request;
[0156] The rendering process is used to render the three-dimensional scene to obtain a rendering effect diagram;
[0157] The image generation process is used to use the large language model to perform background fusion between the three-dimensional model corresponding to the skeleton model and the rendering effect image of the three-dimensional scene to obtain the model rendering image;
[0158] The fusion process is used to superimpose the model effect image and the rendering effect image to obtain a target fused image.
[0159] In a specific example of the disclosed solution, the server is specifically used to execute the following image generation process:
[0160] Determine the position information of the skeleton model after the adjustment confirmation operation in the three-dimensional scene and the target area where it is located;
[0161] A large language model is called, and a model rendering that is integrated with the background of the rendering rendering is generated based on at least the rendering rendering, the pose information of the skeleton model in the three-dimensional scene and the target area where the skeleton model is located, and the model description information of the three-dimensional model corresponding to the skeleton model.
[0162] In a specific example of the disclosed solution, the server is specifically used to execute the following image generation process:
[0163] Based on the target area where the skeleton model is located in the three-dimensional scene after the adjustment confirmation operation, a target mask image capable of indicating the area where the skeleton model is located in the rendering effect image is obtained;
[0164] Based on the posture information of the skeleton model in the three-dimensional scene after the adjustment confirmation operation, at least generating posture control information for the three-dimensional model corresponding to the skeleton model;
[0165] The large language model is called, and based on at least the target mask map, the posture control information, the rendering effect map and the model description information of the three-dimensional model corresponding to the skeleton model, a model rendering map that matches the image size of the target mask map and is integrated with the background of the rendering effect map is generated.
[0166] In a specific example of the disclosed solution, the server is specifically used to execute the following image generation process:
[0167] The first preprocessing network is used to identify the posture information of the skeleton model after the adjustment confirmation operation to obtain a target posture graph of the three-dimensional model corresponding to the skeleton model, wherein the target posture graph can characterize the posture control information of the three-dimensional model corresponding to the skeleton model after the adjustment confirmation operation.
[0168] In a specific example of the disclosed solution, the server is specifically used to execute the following image generation process:
[0169] Based on the position information of the skeleton model in the three-dimensional scene and the target area after the adjustment confirmation operation, generate image depth information of the three-dimensional model corresponding to the skeleton model in the three-dimensional scene;
[0170] The large language model is called, and based on the image depth information, target mask map, pose control information, rendering effect map and model description information of the three-dimensional model corresponding to the skeleton model, a model rendering map that matches the image size of the target mask map and is integrated with the background of the rendering effect map is generated.
[0171] In a specific example of the disclosed solution, the server is specifically used to execute the following image generation process:
[0172] The second preprocessing network is used to perform posture and depth recognition on the skeleton model located in the three-dimensional scene after the adjustment confirmation operation to obtain a target depth map, wherein the target depth map can characterize the depth information of the three-dimensional model corresponding to the skeleton model after the adjustment confirmation operation in the three-dimensional scene.
[0173] In a specific example of the disclosed solution, the client is also used for:
[0174] Display a parameter setting interface for setting parameters of the three-dimensional model corresponding to the skeleton model;
[0175] In response to a setting confirmation operation on the parameter setting interface, model description information for constraining appearance features of the three-dimensional model corresponding to the skeleton model is obtained.
[0176] In a specific example of the disclosed solution, the client is also used for:
[0177] In response to an image input operation, a target reference image is obtained, wherein the target reference image is used to constrain the appearance features of the three-dimensional model corresponding to the skeleton model during the generation of the model rendering; wherein the appearance features of the three-dimensional model in the model rendering are associated with the appearance features of the target object in the target reference image.
[0178] In a specific example of the disclosed solution, the server is specifically used to execute the following image generation process:
[0179] Recognize the appearance features of the target object in the target reference image;
[0180] The large language model is called, and based on the appearance features of the target object in the target reference image, the rendering effect diagram, the posture information of the skeleton model in the three-dimensional scene and the target area where it is located, and the model description information of the three-dimensional model corresponding to the skeleton model, a model rendering diagram fused with the background of the rendering effect diagram is generated.
[0181] In a specific example of the disclosed solution, the server is specifically used to execute the following image generation process:
[0182] After generating a model rendering that is integrated with the background of the rendering rendering, the appearance features of the target object in the target reference image are identified;
[0183] Based on the appearance features of the target object in the target reference image, the appearance features of the three-dimensional model in the generated model rendering are adjusted to update the model rendering.
[0184] For the description of the specific functions and examples of each unit of the device in the embodiment of the present disclosure, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0185] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0186] Fig.14 FIG. 1 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Fig.14 As shown, the electronic device includes: a memory 1410 and a processor 1420, and the memory 1410 stores a computer program that can be run on the processor 1420. The number of memories 1410 and processors 1420 can be one or more. The memory 1410 can store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device executes the method provided by the above method embodiment. The electronic device may also include: a communication interface 1430, which is used to communicate with external devices and perform data exchange transmission.
[0187] If the memory 1410, the processor 1420 and the communication interface 1430 are implemented independently, the memory 1410, the processor 1420 and the communication interface 1430 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.14 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0188] Optionally, in a specific implementation, if the memory 1410, the processor 1420 and the communication interface 1430 are integrated on a chip, the memory 1410, the processor 1420 and the communication interface 1430 can communicate with each other through an internal interface.
[0189] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor supporting the Advanced RISC Machines (ARM) architecture.
[0190] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may also include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM).
[0191] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present disclosure is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (for example: coaxial cable, optical fiber, data subscriber line (Digital Subscriber Line, DSL)) or wireless (for example: infrared, Bluetooth, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more available media integrations. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)), etc. It is worth noting that the computer-readable storage medium mentioned in the present disclosure may be a non-volatile storage medium, in other words, a non-transient storage medium.
[0192] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0193] In the description of the embodiments of the present disclosure, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.
[0194] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means or, for example, A / B can mean A or B. "And / or" in this article is only a way to describe the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0195] In the description of the embodiments of the present disclosure, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "plurality" means two or more.
[0196] The above description is only an exemplary embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A rendering method for integrating a three-dimensional model with a three-dimensional scene, comprising: Display three-dimensional scenes; A skeleton model representing a skeleton structure of a three-dimensional model is displayed in a three-dimensional scene; In response to the adjustment confirmation operation on the skeleton model, a fusion image generation request is sent to the server to obtain a target fusion image, wherein the target fusion image is obtained by fusing the model effect image with the rendering effect image of the three-dimensional scene; the model effect image is obtained by using a large language model to perform background fusion of the three-dimensional model corresponding to the skeleton model and the rendering effect image of the three-dimensional scene; The target fused image is obtained after the server responds to the fused image generation request and executes the rendering process, the image generation process and the fusion process; The rendering process is used to render the three-dimensional scene to obtain a rendering effect diagram; The image generation process is used to use the large language model to perform background fusion between the three-dimensional model corresponding to the skeleton model and the rendering effect image of the three-dimensional scene to obtain the model rendering image; The fusion process is used to superimpose the model effect image and the rendering effect image to obtain a target fused image.
2. The method according to claim 1, wherein: The image generation process includes: Determine the position information of the skeleton model after the adjustment confirmation operation in the three-dimensional scene and the target area where it is located; A large language model is called, and a model rendering that is integrated with the background of the rendering rendering is generated based on at least the rendering rendering, the pose information of the skeleton model in the three-dimensional scene and the target area where the skeleton model is located, and the model description information of the three-dimensional model corresponding to the skeleton model.
3. The method according to claim 2, wherein: The calling of the large language model and generating a model rendering that is integrated with the background of the rendering rendering based on at least the rendering rendering, the pose information of the skeleton model in the three-dimensional scene and the target area where the skeleton model is located, and the model description information of the three-dimensional model corresponding to the skeleton model, includes: Based on the target area where the skeleton model is located in the three-dimensional scene after the adjustment confirmation operation, a target mask image capable of indicating the area where the skeleton model is located in the rendering effect image is obtained; Based on the posture information of the skeleton model in the three-dimensional scene after the adjustment confirmation operation, at least generating posture control information for the three-dimensional model corresponding to the skeleton model; The large language model is called, and based on at least the target mask map, the posture control information, the rendering effect map and the model description information of the three-dimensional model corresponding to the skeleton model, a model rendering map that matches the image size of the target mask map and is integrated with the background of the rendering effect map is generated.
4. The method according to claim 3, wherein: The step of generating at least posture control information for the three-dimensional model corresponding to the skeleton model based on the posture information of the skeleton model in the three-dimensional scene after the adjustment confirmation operation comprises: The first preprocessing network is used to identify the posture information of the skeleton model after the adjustment confirmation operation to obtain a target posture graph of the three-dimensional model corresponding to the skeleton model, wherein the target posture graph can characterize the posture control information of the three-dimensional model corresponding to the skeleton model after the adjustment confirmation operation.
5. The method according to claim 3, further comprising: Based on the position information of the skeleton model in the three-dimensional scene and the target area after the adjustment confirmation operation, generate image depth information of the three-dimensional model corresponding to the skeleton model in the three-dimensional scene; The calling of the large language model and generating a model rendering that matches the image size of the target mask image and is integrated with the background of the rendering rendering based at least on the target mask image, the posture control information, the rendering rendering, and the model description information of the three-dimensional model corresponding to the skeleton model, includes: The large language model is called, and based on the image depth information, target mask map, pose control information, rendering effect map and model description information of the three-dimensional model corresponding to the skeleton model, a model rendering map that matches the image size of the target mask map and is integrated with the background of the rendering effect map is generated.
6. The method according to claim 5, wherein: The step of generating image depth information of a three-dimensional model corresponding to the skeleton model in the three-dimensional scene based on the posture information of the skeleton model in the three-dimensional scene after the adjustment confirmation operation and the target area in which the skeleton model is located comprises: The second preprocessing network is used to perform posture and depth recognition on the skeleton model located in the three-dimensional scene after the adjustment confirmation operation to obtain a target depth map, wherein the target depth map can characterize the depth information of the three-dimensional model corresponding to the skeleton model after the adjustment confirmation operation in the three-dimensional scene.
7. The method according to any one of claims 2 to 6, further comprising: Display a parameter setting interface for setting parameters of the three-dimensional model corresponding to the skeleton model; In response to a setting confirmation operation on the parameter setting interface, model description information for constraining appearance features of the three-dimensional model corresponding to the skeleton model is obtained.
8. The method according to any one of claims 2 to 6, further comprising: In response to an image input operation, a target reference image is obtained, wherein the target reference image is used to constrain the appearance features of the three-dimensional model corresponding to the skeleton model during the generation of the model rendering; wherein the appearance features of the three-dimensional model in the model rendering are associated with the appearance features of the target object in the target reference image.
9. The method according to claim 8, wherein: The calling of the large language model and generating a model rendering that is integrated with the background of the rendering rendering based on at least the rendering rendering, the pose information of the skeleton model in the three-dimensional scene and the target area where the skeleton model is located, and the model description information of the three-dimensional model corresponding to the skeleton model, includes: Recognize the appearance features of the target object in the target reference image; The large language model is called, and based on the appearance features of the target object in the target reference image, the rendering effect diagram, the pose information of the skeleton model in the three-dimensional scene and the target area where it is located, and the model description information of the three-dimensional model corresponding to the skeleton model, a model rendering diagram fused with the background of the rendering effect diagram is generated.
10. The method according to claim 8, after generating the model effect image integrated with the background of the rendering effect image, the method further comprises: Recognize the appearance features of the target object in the target reference image; Based on the appearance features of the target object in the target reference image, the appearance features of the three-dimensional model in the generated model rendering are adjusted to update the model rendering.
11. A rendering device for integrating a three-dimensional model with a three-dimensional scene, comprising: A client, used to display a three-dimensional scene, and to display a skeleton model representing a skeleton structure of a three-dimensional model in the three-dimensional scene; and in response to the adjustment confirmation operation on the skeleton model, sending a fused image generation request to the server to obtain a target fused image; The server is used to obtain the fused image generation request, obtain and send the target fused image; The target fused image is obtained by fusing the model rendering with the rendering rendering of the three-dimensional scene; the model rendering is obtained by fusing the three-dimensional model corresponding to the skeleton model with the rendering rendering of the three-dimensional scene using a large language model; The server is further configured to execute a rendering process, an image generation process, and a fusion process in response to a fusion image generation request; The rendering process is used to render the three-dimensional scene to obtain a rendering effect diagram; The image generation process is used to use the large language model to perform background fusion between the three-dimensional model corresponding to the skeleton model and the rendering effect image of the three-dimensional scene to obtain the model rendering image; The fusion process is used to superimpose the model effect image and the rendering effect image to obtain a target fused image.
12. The device according to claim 11, wherein The server is specifically used to execute the following image generation process: Determine the position information of the skeleton model after the adjustment confirmation operation in the three-dimensional scene and the target area where it is located; A large language model is called, and a model rendering that is integrated with the background of the rendering rendering is generated based on at least the rendering rendering, the pose information of the skeleton model in the three-dimensional scene and the target area where the skeleton model is located, and the model description information of the three-dimensional model corresponding to the skeleton model.
13. The device according to claim 12, wherein: The server is specifically used to execute the following image generation process: Based on the target area where the skeleton model is located in the three-dimensional scene after the adjustment confirmation operation, a target mask image capable of indicating the area where the skeleton model is located in the rendering effect image is obtained; Based on the posture information of the skeleton model in the three-dimensional scene after the adjustment confirmation operation, at least generating posture control information for the three-dimensional model corresponding to the skeleton model; The large language model is called, and based on at least the target mask map, the posture control information, the rendering effect map and the model description information of the three-dimensional model corresponding to the skeleton model, a model rendering map that matches the image size of the target mask map and is integrated with the background of the rendering effect map is generated.
14. The device according to claim 13, wherein: The server is specifically used to execute the following image generation process: The first preprocessing network is used to identify the posture information of the skeleton model after the adjustment confirmation operation to obtain a target posture graph of the three-dimensional model corresponding to the skeleton model, wherein the target posture graph can characterize the posture control information of the three-dimensional model corresponding to the skeleton model after the adjustment confirmation operation.
15. The device according to claim 13, wherein: The server is specifically used to execute the following image generation process: Based on the position information of the skeleton model in the three-dimensional scene and the target area after the adjustment confirmation operation, generate image depth information of the three-dimensional model corresponding to the skeleton model in the three-dimensional scene; The large language model is called, and based on the image depth information, target mask map, pose control information, rendering effect map and model description information of the three-dimensional model corresponding to the skeleton model, a model rendering map that matches the image size of the target mask map and is integrated with the background of the rendering effect map is generated.
16. The device according to claim 15, wherein: The server is specifically used to execute the following image generation process: The second preprocessing network is used to perform posture and depth recognition on the skeleton model located in the three-dimensional scene after the adjustment confirmation operation to obtain a target depth map, wherein the target depth map can characterize the depth information of the three-dimensional model corresponding to the skeleton model after the adjustment confirmation operation in the three-dimensional scene.
17. The device according to any one of claims 12 to 16, wherein: The client is also used to: Display a parameter setting interface for setting parameters of the three-dimensional model corresponding to the skeleton model; In response to a setting confirmation operation on the parameter setting interface, model description information for constraining appearance features of the three-dimensional model corresponding to the skeleton model is obtained.
18. The device according to any one of claims 12 to 16, wherein: The client is also used to: In response to an image input operation, a target reference image is obtained, wherein the target reference image is used to constrain the appearance features of the three-dimensional model corresponding to the skeleton model during the generation of the model rendering; wherein the appearance features of the three-dimensional model in the model rendering are associated with the appearance features of the target object in the target reference image.
19. The device according to claim 18, wherein: The server is specifically used to execute the following image generation process: Recognize the appearance features of the target object in the target reference image; The large language model is called, and based on the appearance features of the target object in the target reference image, the rendering effect diagram, the posture information of the skeleton model in the three-dimensional scene and the target area where it is located, and the model description information of the three-dimensional model corresponding to the skeleton model, a model rendering diagram fused with the background of the rendering effect diagram is generated.
20. The device according to claim 18, wherein The server is specifically used to execute the following image generation process: After generating a model rendering that is integrated with the background of the rendering rendering, the appearance features of the target object in the target reference image are identified; Based on the appearance features of the target object in the target reference image, the appearance features of the three-dimensional model in the generated model rendering are adjusted to update the model rendering.
21. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
23. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Mechanical arm control method, device and equipment based on large visual model and storage medium
CN118143940A
Sparse visual angle three-dimensional reconstruction method based on depth prior information
CN118657888A