Controllable generation method and device of three-dimensional effect picture, equipment and storage medium

By displaying a 2D main image in a 3D scene design interface and combining scene adaptation and 3D rendering, controllable generation from 2D images to 3D renderings is achieved, solving the problems of high rendering cost of 3D engines and poor controllability of AI rendering, and improving generation efficiency and realism.

CN121353552AActive Publication Date: 2026-01-16HANGZHOU QUNHE INFORMATION TECHNOLOGIES CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202511903144.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-01-16
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

Among existing 3D rendering technologies, 3D engine rendering is costly and difficult for non-professional users to use, while AI rendering has poor controllability, resulting in insufficient commercial applicability.

Method used

By displaying a two-dimensional main image in the three-dimensional scene design interface, the system generates an initial three-dimensional model and candidate scenes in response to user operations. Combined with scene adaptation and a 3D rendering engine, it achieves controllable generation of two-dimensional images into three-dimensional renderings.

Benefits of technology

It lowers the barrier to entry, enhances the intuitiveness and realism of the interaction, supports flexible scene selection, avoids repeated debugging, and ensures the preservation of the original characteristics of the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353552A_ABST
    Figure CN121353552A_ABST
Patent Text Reader

Abstract

The invention provides a controllable generation method and device of a three-dimensional effect picture, equipment and a storage medium. The method comprises the following steps: displaying a two-dimensional main body graph of a target object in a three-dimensional scene design interface; in response to an intelligent task operation for a two-dimensional main body graph, displaying an initial three-dimensional model of the two-dimensional main body graph in a first display area of the three-dimensional scene design interface, and displaying a plurality of candidate three-dimensional scenes selected by a user in a second display area of the three-dimensional scene design interface; in response to a scene adaptation operation, displaying a target three-dimensional scene in a main display area of the three-dimensional scene design interface; and in response to a rendering operation for the target three-dimensional scene, displaying a target rendering effect picture in the main display area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and in particular to a controllable method, apparatus, device, and storage medium for generating three-dimensional renderings. Background Technology

[0002] In fields such as architectural design and product display, 3D renderings are a core medium for meeting visualization needs. The practicality, efficiency, and quality of their generation technology directly impact industry application efficiency and commercial value. Currently, mainstream 3D rendering generation technologies fall into two categories: one is 3D engine rendering, which can accurately simulate physical photography and is suitable for scenarios with high precision and realism requirements. However, 3D modeling is costly, making it difficult for non-professional clients with only image materials to use, thus limiting its application scope. The other is Artificial Intelligence (AI) rendering, which can generate renderings using only text or images. It has a low barrier to entry and high efficiency, but suffers from poor controllability and physical inaccuracies, reducing its overall commercial viability. Therefore, there is an urgent need for a controllable generation method that balances autonomy and realism to obtain 3D renderings from 2D images. Summary of the Invention

[0003] This disclosure provides a method, apparatus, device, and storage medium for the controllable generation of three-dimensional renderings, in order to solve or alleviate one or more technical problems in the prior art.

[0004] In a first aspect, this disclosure provides a controllable method for generating three-dimensional renderings based on two-dimensional images, including: Display the two-dimensional main image of the target object in the 3D scene design interface; In response to intelligent task operations on the two-dimensional main image, an initial three-dimensional model of the two-dimensional main image is displayed in the first display area of ​​the three-dimensional scene design interface, and multiple candidate three-dimensional scenes are displayed in the second display area of ​​the three-dimensional scene design interface for the user to select. In response to the scene adaptation operation, the target 3D scene is displayed in the main display area of ​​the 3D scene design interface. The target 3D scene is the 3D scene obtained by adapting the initial 3D model to the selected 3D scene to be processed. The 3D scene to be processed is one of the multiple candidate 3D scenes. In response to the rendering operation for the target 3D scene, the target rendering effect image is displayed in the main display area; wherein, the target rendering effect image is the rendering effect image after transferring the main detail features of the target object in the 2D main image to the 3D rendering image using a large model; the 3D rendering image is obtained after rendering the target 3D scene using a 3D rendering engine.

[0005] Secondly, this disclosure provides a controllable generation device for obtaining a three-dimensional rendering based on a two-dimensional image, comprising: The task triggering unit is used to display a two-dimensional main image of the target object in the three-dimensional scene design interface, and in response to intelligent task operation on the two-dimensional main image, to display an initial three-dimensional model of the two-dimensional main image in the first display area of ​​the three-dimensional scene design interface, and to display multiple candidate three-dimensional scenes for the user to select in the second display area of ​​the three-dimensional scene design interface. A scene adaptation unit is used to respond to a scene adaptation operation and display a target 3D scene in the main display area of ​​the 3D scene design interface. The target 3D scene is a 3D scene obtained by adapting the initial 3D model to the selected 3D scene to be processed. The 3D scene to be processed is one of the multiple candidate 3D scenes. The AI ​​rendering unit is used to respond to the rendering operation for the target 3D scene and display the target rendering effect in the main display area; wherein, the target rendering effect is the rendering effect after transferring the main detail features of the target object in the 2D main image to the 3D rendering image using a large model; the 3D rendering image is obtained by rendering the target 3D scene using a 3D rendering engine.

[0006] Thirdly, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.

[0007] Fourthly, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of the present disclosure.

[0008] Fifthly, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of the present disclosure.

[0009] The beneficial effects of the technical solution provided in this disclosure include at least the following: This disclosed solution displays a 2D main image of the target object in a 3D scene design interface. In response to intelligent task operations, it displays the initial 3D model of the 2D main image and multiple candidate 3D scenes in different areas. Then, in response to scene adaptation operations, it displays the target 3D scene after adaptation of the initial 3D model and the 3D scene. Finally, through rendering operations, it displays the target rendered image after merging the initial 3D model of the 2D main image into the 3D scene. This achieves a unified interface integration from 2D design to 3D effect, lowering the operational threshold and improving interactive intuitiveness. Furthermore, it supports flexible scene selection and real-time preview of adaptation effects, avoiding the inefficiency of repeated debugging. Simultaneously, combining the spatial realism of 3D rendering with the detail transfer capability of large models ensures scene fusion effects while highly preserving the original features of the target object, significantly improving the realism and detail accuracy of the final rendering.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments provided according to this disclosure and should not be construed as limiting the scope of this disclosure.

[0012] Figure 1 This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 1 ; Figure 2 This is an illustrative schematic diagram of a three-dimensional scene design interface according to an embodiment of this application; Figure 3 This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 2 ; Figure 4 This is an illustrative schematic diagram showing the original two-dimensional interface according to an embodiment of this application; Figure 5 This is an illustrative schematic diagram showing the interface of the original two-dimensional drawing and the two-dimensional main body drawing according to an embodiment of this application; Figure 6 This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 3 ; Figure 7This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 4 ; Figure 8 This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 5 ; Figure 9 This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 6 ; Figure 10 This is a schematic diagram of the structure of a controllable generation device for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application; Figure 11 This is a block diagram of an electronic device used to implement the controllable generation method of obtaining a three-dimensional rendering based on a two-dimensional image according to the embodiments of this disclosure. Detailed Implementation

[0013] The present disclosure will now be described in further detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0014] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0015] In fields such as architectural design and product display, 3D renderings are a key medium for visually presenting target objects. The realism, controllability, and technical barriers of the generated effects directly impact industry application efficiency and commercial value. Currently, mainstream 3D rendering technologies fall into two categories: one is 3D engine rendering, which can accurately simulate physical photography and is suitable for scenarios with high precision and realism requirements. However, 3D modeling is costly, making it difficult for non-professional clients with only image materials to use, thus limiting its application scope. The other is AI rendering, which can generate renderings using only text or images. It has a low barrier to entry and high efficiency, but suffers from poor controllability and physical inaccuracies, reducing its overall commercial viability.

[0016] Based on this, the disclosed solution provides a controllable method for generating 3D renderings from 2D images. This method uses a 2D main image of the target object as a foundation, combining scene adaptation operations to achieve precise matching between the target object and any scene. Furthermore, the scene adaptation process can be adjusted according to user needs (e.g., selecting a 3D scene and adjusting the target object's pose within the 3D scene), offering strong controllability and meeting personalized user requirements. Finally, a 3D rendering engine and AI technology are used in collaboration to complete the rendering, generating the target rendering image. This achieves the conversion from 2D images to 3D renderings, effectively solving the problems of high modeling barriers in traditional 3D rendering engines and poor controllability in AI rendering. It reduces operational difficulty while maintaining physical realism, meeting users' needs for 3D renderings.

[0017] Specifically, Figure 1 This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 1 This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0018] Furthermore, the method includes at least a portion of the following: For example... Figure 1 As shown, it includes: Step S101: Display the two-dimensional main image of the target object in the 3D scene design interface.

[0019] Here, the two-dimensional subject image can specifically be an image in which the background has been removed and only the target object (or subject object, such as a person, product, pet, etc.) is retained.

[0020] Step S102: In response to the intelligent task operation for the two-dimensional main image, the initial three-dimensional model of the two-dimensional main image is displayed in the first display area of ​​the three-dimensional scene design interface, and multiple candidate three-dimensional scenes for the user to select are displayed in the second display area of ​​the three-dimensional scene design interface.

[0021] Here, the intelligent task operation refers to the active interactive operation initiated by the user in the 3D scene design interface for the two-dimensional main image of the displayed target object, which triggers the system to automatically complete the initial 3D model generation and the display of candidate 3D scenes. It can be a form of interaction that the user can intuitively perceive, such as clicking interface buttons or selecting menus.

[0022] Furthermore, it can be understood that the initial three-dimensional model of the two-dimensional subject image displayed in the first display area can specifically be the initial three-dimensional model of the target object corresponding to the two-dimensional subject image.

[0023] Furthermore, the first display area and the second display area are different display areas in the 3D scene design interface, thus providing users with richer display content and further enhancing the user's interactive experience.

[0024] Furthermore, in one example, the multiple candidate 3D scenes displayed in the second display area can be specifically 3D scenes that match the main category of the target object. This can greatly improve interaction efficiency and user experience, while also ensuring the rationality and realism of the renderings, thus achieving "intelligent assistance".

[0025] Step S103: In response to the scene adaptation operation, display the target 3D scene in the main display area of ​​the 3D scene design interface.

[0026] Here, the target 3D scene is the 3D scene obtained by scene adaptation (or scene fusion) of the initial 3D model and the selected 3D scene to be processed.

[0027] Furthermore, the 3D scene to be processed is one of the multiple candidate 3D scenes. For example, the 3D scene to be processed is a 3D scene selected by the user from multiple candidate 3D scenes; or, it is an automatically adapted 3D scene.

[0028] In other words, the target 3D scene is a scene fusion map obtained by fusing the initial 3D model of the 2D main image with the selected or adapted 3D scene to be processed.

[0029] Step S104: In response to the rendering operation for the target 3D scene, display the target rendering effect image in the main display area.

[0030] Here, the target rendering image is the rendering image after transferring the main details of the target object in the two-dimensional subject image to the three-dimensional rendering image using a large model; the three-dimensional rendering image is obtained by rendering the target three-dimensional scene using a 3D rendering engine.

[0031] In this way, the disclosed solution displays the two-dimensional main image of the target object in the 3D scene design interface, and in response to intelligent task operations, displays the initial 3D model of the two-dimensional main image and multiple candidate 3D scenes in different areas. Then, in response to scene adaptation operations, it displays the target 3D scene after the initial 3D model and the 3D scene are adapted. Finally, through rendering operations, it displays the target rendered image after the initial 3D model of the two-dimensional main image is merged into the 3D scene. In this way, on the one hand, it realizes the unified interface integration of the entire process from 2D design to 3D effect, reduces the operation threshold and improves the intuitiveness of interaction. On the other hand, it supports flexible scene selection and real-time preview of adaptation effect, avoiding the inefficiency of repeated debugging. At the same time, combined with the spatial realism of 3D rendering and the detail transfer capability of large models, it not only ensures the scene fusion effect, but also highly preserves the original features of the target object, significantly improving the realism and detail accuracy of the final image.

[0032] For example, such as Figure 2 As shown in one example, the 3D scene design interface includes a first display area and a second display area; wherein, the first display area is used to display an initial 3D model, and the second display area is used to display multiple candidate 3D scenes for the user to select, such as candidate 3D scene 1, candidate 3D scene 2, and candidate 3, etc.

[0033] Furthermore, in one example, the 3D scene design interface also includes a main display area, which is used to display the target 3D scene after the initial 3D model and the selected 3D scene to be processed are adapted, and to display the target rendering effect after the initial 3D model of the 2D main image is merged into the 3D scene.

[0034] It should be noted that the main display area is the core display area in the 3D scene design interface, showcasing the most essential and important content, such as... Figure 2 As shown, the main display area is located in the middle of the 3D scene design interface, and its display area is larger than that of other display interfaces.

[0035] It is understood that this disclosed solution does not impose specific limitations on the display area and display position of the first display area, the second display area, and the main display area, and can be configured according to the actual situation.

[0036] Furthermore, in one example, to further enhance the interactive experience and meet the user's customization needs, the 3D scene design interface may also include parameter setting areas, other tool areas, etc. This disclosure does not impose specific restrictions on the display areas required in the 3D scene design interface, and can be configured according to the actual situation.

[0037] Figure 3This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 2 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figure 1 and Figure 2 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.

[0038] Furthermore, the method includes at least a portion of the following: For example... Figure 3 As shown, it includes: Step S301: In response to the input operation, display the original two-dimensional image of the target object in the three-dimensional scene design interface.

[0039] For example, such as Figure 4 As shown, the user can input the original 2D image of any target object (such as the camera shown in the figure). Furthermore, in response to the user's input operation, the original 2D image of the target object is displayed in the 3D scene design interface.

[0040] It is understandable that in practical applications, other function buttons can be set in the 3D scene design interface, such as the "Re-upload" button for re-transmission operations, the "Cutout" button, etc. This public solution does not impose specific restrictions on this.

[0041] Step S302: In response to the image cutout operation, display the two-dimensional main image of the target object in the three-dimensional scene design interface.

[0042] Here, the two-dimensional subject image is obtained by cutting out the area where the target object is located in the original two-dimensional image.

[0043] For example, such as Figure 4 As shown, users can trigger the cutout operation by clicking the "Cutout" button and be redirected to the following page. Figure 5 The displayed interface shows the cutout result (i.e., the 2D main image) to the user. Further, in one example, such as... Figure 5 As shown, while displaying the cutout results in the 3D scene design interface, the original image before cutout processing, i.e., the original 2D image, can also be displayed, allowing users to intuitively experience the comparison effect before and after cutout. Furthermore, in this scene, function buttons such as "re-upload" can also be set in the 3D scene design interface, so that users can make adjustments based on their personalized needs and further improve the interactive effect.

[0044] Step S303: Responding to intelligent task operations on the two-dimensional subject graph (e.g., responding to...) Figure 5The "Submit Task" button shown is triggered to display the initial 3D model of the 2D main image in the first display area of ​​the 3D scene design interface, and multiple candidate 3D scenes for the user to select are displayed in the second display area of ​​the 3D scene design interface.

[0045] Step S304: In response to the scene adaptation operation, display the target 3D scene in the main display area of ​​the 3D scene design interface.

[0046] Here, the target 3D scene is the 3D scene obtained by adapting the initial 3D model to the selected 3D scene to be processed; the 3D scene to be processed is one of the multiple candidate 3D scenes.

[0047] Step S305: In response to the rendering operation for the target 3D scene, display the target rendering effect image in the main display area.

[0048] Here, the target rendering image is the rendering image after transferring the main details of the target object in the two-dimensional subject image to the three-dimensional rendering image using a large model; the three-dimensional rendering image is obtained by rendering the target three-dimensional scene using a 3D rendering engine.

[0049] For details regarding the first display area, the second display area, the main display area, the target 3D scene, and the target rendering effect, please refer to the above descriptions; they will not be repeated here.

[0050] In this way, the disclosed solution first responds to the input operation by displaying the original 2D image of the target object, and then responds to the cutout operation by displaying the 2D main image of the target object. On the one hand, it provides a flexible operation path, first displaying the original 2D image to allow the user to guide the initial form of the target object, and then using automated cutout operation to accurately extract the area where the target object is located to generate the 2D main image. This reduces the operation difficulty for non-professional users, automates and visualizes the extraction process of the target object, and avoids the problems of "extracting irrelevant areas" or "missing the target part", ensuring the clarity and target orientation of the 2D main image. On the other hand, by removing redundant background information from the original 2D image through cutout operation, the subsequent steps of 3D model generation and scene adaptation focus on the target object itself, avoiding the interference of background information on model construction and scene fusion, improving the accuracy of 3D effect generation, and laying the foundation for the subsequent generation of high-quality 3D renderings.

[0051] Furthermore, in a specific example, the method further includes: The object features of the target object in the original two-dimensional image are identified; wherein the object features of the target object include at least one of the following: primitive information, subject category, and spatial perspective information.

[0052] For example, in one instance, after responding to the cutout operation, the object features of the target object in the original 2D image can also be identified, thus providing data support for the subsequent generation of high-quality 3D renderings.

[0053] Furthermore, in one example, the aforementioned primitive information refers to the geometric information of the target object in the original two-dimensional image, used to describe the main outline, size, position, and deformation effect of the target object, such as, but not limited to, at least one of the following: the original image size, the anchor point information of the main body of the target object (e.g., the key reference points of the main body, etc.).

[0054] Furthermore, in one example, the aforementioned subject category can be used to characterize the object type to which the target object belongs in the original two-dimensional image. This facilitates the determination of the candidate three-dimensional scene that matches the subject type, providing strong support for the rationality and realism of the subsequently generated rendering.

[0055] In one example, the aforementioned spatial perspective information refers to data describing the shooting perspective of the target object in the original two-dimensional image, such as, but not limited to, at least one of the following: shooting pitch angle, field of view angle, etc. In this way, it can be ensured that the subsequently generated three-dimensional model can accurately restore the posture of the target object and avoid the problem of perspective misalignment.

[0056] In this way, the disclosed solution identifies the object features (including primitive information, subject category, spatial perspective information, etc.) of the target object in the original 2D image, providing structured data support for subsequent 3D model generation and scene adaptation. Primitive information ensures the accuracy of the basic structure of the 3D model; the subject category clarifies the type of target object, allowing for targeted matching to the 3D scene; and spatial perspective information helps calibrate the orientation and posture of the target object in 3D space (or 3D scene), avoiding scene fusion misalignment caused by perspective deviations. This improves the processing efficiency of subsequent steps, enhances the consistency between the generated 3D model and the target object in the original 2D image, and provides a precise data foundation for generating high-quality final renderings.

[0057] Furthermore, in another specific example, the method further includes: Based on the two-dimensional subject image and the pixel features of the original two-dimensional image, a contour mask image is obtained that can describe the contour of the target object in the original two-dimensional image.

[0058] For example, in one instance, after responding to the cutout operation, a two-dimensional main image of the target object is obtained; further, based on the two-dimensional main image of the target object and the pixel features of the original two-dimensional image, a contour mask image that can describe the contour of the target object in the original two-dimensional image is obtained, thus providing data support for the subsequent generation of high-quality three-dimensional renderings.

[0059] Here, in a specific example, the aforementioned contour mask image refers to an image that accurately describes the contour range of the target object in the original two-dimensional image. For example, based on the pixel features of the two-dimensional main image and the original two-dimensional image (such as the red, green, and blue values ​​(i.e., RGB values)), a black and white mask image that can accurately capture the edge details of the target object can be generated, avoiding the error of manually outlining the contour, and thus providing a boundary benchmark and accuracy guarantee for the subsequent generation of three-dimensional effect images.

[0060] Thus, based on the two-dimensional main image and the pixel features of the original two-dimensional image, this disclosed solution obtains a contour mask image that describes the outline of the target object in the original two-dimensional image. This avoids subjective errors and tedious operations when manually outlining the outline, preserves richer edge details (such as subtle protrusions / indentations), and ensures that the outline range is consistent with the actual shape of the target object in the original two-dimensional image, providing a precise outline reference for subsequent operations. On the other hand, since this contour mask image provides key boundary constraint data for subsequent three-dimensional model generation steps, it can effectively avoid model edge deformation or misalignment with the scene caused by outline blurring, enhancing the efficiency of the process and the reliability of the results.

[0061] Figure 6 This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 3 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figures 1-5 The relevant content of the method shown in any of the attached figures can also be applied to this example, and the relevant content will not be described again in this example.

[0062] Furthermore, the method includes at least a portion of the following: For example... Figure 6 As shown, it includes: Step S601: In response to the input operation, display the original two-dimensional image of the target object in the three-dimensional scene design interface.

[0063] Step S602: In response to the image cutout operation, display the two-dimensional main image of the target object in the three-dimensional scene design interface.

[0064] Here, the two-dimensional subject image is obtained by cutting out the area where the target object is located in the original two-dimensional image.

[0065] For example, in one instance, after the image cutout operation, a contour mask image that describes the outline of the target object in the original 2D image can also be obtained. For details regarding contour mask images, please refer to the above description; they will not be repeated here.

[0066] Furthermore, in one example, after the image matting operation, the main category of the target object can also be obtained, thus providing a basis for filtering multiple candidate 3D scenes to be identified subsequently.

[0067] Step S603: In response to the intelligent task operation for the two-dimensional subject image, obtain a contour mask image that can describe the outline of the target object in the original two-dimensional image, for example, obtain the black and white mask image obtained above that can accurately capture the edge details of the target object.

[0068] Step S604: Call the large model and obtain the initial three-dimensional model of the two-dimensional main body image based on the two-dimensional main body image and the contour mask image.

[0069] In this example, the large model may specifically be a large language model or an image generative model, etc. This disclosure does not limit the specific model.

[0070] Further, step S604 can be specifically described as: inputting the two-dimensional main image and the contour mask image into the large model to obtain the initial three-dimensional model of the two-dimensional main image output by the large model.

[0071] Furthermore, in one example, while generating the initial 3D model using a large model, the model pose information of this initial 3D model can also be obtained using the large model. Here, the model pose information of the initial 3D model can be used to describe the spatial transformation information of the initial 3D model when it is restored to the 2D main image. For example, it may include, but is not limited to, at least one of the following: rotation parameters, orientation parameters, etc. of the initial 3D model. In this way, data support is provided for the subsequent adaptation of the initial 3D model to the 3D scene, making the shape of the target object in the target 3D scene more consistent with the core visual features of the original 2D image, thereby improving the realism and coordination of the scene adaptation.

[0072] Step S605: In the first display area of ​​the three-dimensional scene design interface, the initial three-dimensional model of the two-dimensional main image is displayed, and in the second display area of ​​the three-dimensional scene design interface, multiple candidate three-dimensional scenes for the user to select are displayed.

[0073] Step S606: In response to the scene adaptation operation, display the target 3D scene in the main display area of ​​the 3D scene design interface.

[0074] Here, the target 3D scene is the 3D scene obtained by adapting the initial 3D model to the selected 3D scene to be processed; the 3D scene to be processed is one of the multiple candidate 3D scenes.

[0075] Step S607: In response to the rendering operation for the target 3D scene, display the target rendering effect image in the main display area.

[0076] Here, the target rendering image is the rendering image after transferring the main details of the target object in the two-dimensional subject image to the three-dimensional rendering image using a large model; the three-dimensional rendering image is obtained by rendering the target three-dimensional scene using a 3D rendering engine.

[0077] For details regarding the first display area, the second display area, the main display area, the target 3D scene, and the target rendering effect, please refer to the above descriptions; they will not be repeated here.

[0078] Thus, this disclosed solution provides a detailed approach to obtaining an initial 3D model. This approach first obtains a 2D main image and a contour mask image that describes the outline of the target object in the original 2D image. Then, a large model is called, and the initial 3D model of the 2D main image is obtained by combining the 2D main image and the contour mask image. In this way, the boundary outline of the target object is accurately determined by the contour mask image, avoiding the problem of model edge deformation caused by outline blurring. At the same time, by using the large model and combining the 2D main image and the contour mask image, the 3D form of the target object (such as volume, proportion, structure, etc.) is more accurately restored, thereby improving the structural rationality and realism of the initial 3D model. This provides strong data support for the subsequent generation of the target 3D scene and the final generation of the target rendering effect image.

[0079] Figure 7 This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 4 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figures 1-6 The relevant content of the method shown in any of the attached figures can also be applied to this example, and the relevant content will not be described again in this example.

[0080] Furthermore, the method includes at least a portion of the following: For example... Figure 7 As shown, it includes: Step S701: In response to the input operation, display the original two-dimensional image of the target object in the three-dimensional scene design interface.

[0081] Step S702: In response to the image cutout operation, display the two-dimensional main image of the target object in the three-dimensional scene design interface.

[0082] Here, the two-dimensional subject image is obtained by cutting out the area where the target object is located in the original two-dimensional image.

[0083] Furthermore, in one example, after the image matting operation, the object features of the target object, such as primitive information and spatial perspective information, can also be obtained.

[0084] Step S703: In response to the intelligent task operation for the two-dimensional subject image, display the initial three-dimensional model of the two-dimensional subject image in the first display area of ​​the three-dimensional scene design interface, and display multiple candidate three-dimensional scenes for the user to select in the second display area of ​​the three-dimensional scene design interface.

[0085] The process for generating the initial 3D model can be found in the above description and will not be repeated here.

[0086] Step S704: In response to the scene selection operation, display the selected 3D scene to be processed in the main display area.

[0087] Here, the 3D scene to be processed is one of the multiple candidate 3D scenes.

[0088] Step S705: In response to a model selection operation on the original model in the 3D scene to be processed displayed in the main display area, the original model in the selected state is obtained.

[0089] Step S706: In response to the replacement operation, display the target 3D scene in the main display area.

[0090] Here, the target 3D scene is the 3D scene obtained by replacing the original model in the selected state in the 3D scene to be processed with the initial 3D model.

[0091] In other words, this publicly available solution can complete the scene adaptation process through multiple steps, such as scene selection and replacement. This process is simple and straightforward, improving the user's control experience and providing strong support for meeting users' personalized needs.

[0092] Step S707: In response to the rendering operation for the target 3D scene, display the target rendering effect image in the main display area.

[0093] Here, the target rendering image is the rendering image after transferring the main details of the target object in the two-dimensional subject image to the three-dimensional rendering image using a large model; the three-dimensional rendering image is obtained by rendering the target three-dimensional scene using a 3D rendering engine.

[0094] In this way, the disclosed solution first responds to the scene selection operation, displaying the selected 3D scene to be processed in the main display area; then responds to the model selection operation, obtaining the original model in the selected state; and finally responds to the replacement operation, replacing the original model in the 3D scene to be processed with the initial 3D model, and displaying the target 3D scene after replacement in the main display area. In this way, users can flexibly adapt the initial 3D model to the 3D scene to be processed through clear step-by-step operations, greatly reducing the difficulty of scene adaptation. At the same time, users can independently select the 3D scene to be processed and the original model to be replaced, accurately matching user needs and improving the flexibility of scene adaptation. In addition, the entire process is triggered by user-initiated operation, and the replacement result is displayed in real time in the main display area, allowing users to intuitively confirm the adaptation effect, reducing subsequent adjustment costs, significantly improving the generation efficiency of the target 3D scene, and quickly providing a scene foundation that meets the requirements for subsequent rendering stages.

[0095] Furthermore, in a specific example, the target 3D scene can be displayed in the following manner; specifically, the above-described response to the replacement operation, displaying the target 3D scene in the main display area (step S706 above), can specifically include: Step S706-1: In response to the replacement operation, the original model in the selected state in the three-dimensional scene to be processed is replaced with the initial three-dimensional model to obtain the initial three-dimensional scene.

[0096] Step S706-2: Obtain the primitive information and / or spatial perspective information of the target object in the original two-dimensional image.

[0097] Step S706-3: Adjust the initial 3D model in the initial 3D scene according to the primitive information and / or spatial perspective information of the target object in the original 2D image to obtain the target 3D scene.

[0098] For example, in one instance, the primitive information of the target object in the original two-dimensional image can be obtained, and the three-dimensional model in the initial three-dimensional scene can be adjusted according to the primitive information to obtain the target three-dimensional scene. In this way, the accuracy of the basic structure of the three-dimensional model is guaranteed.

[0099] Alternatively, in another example, the spatial perspective information of the target object in the original 2D image can be obtained, and the 3D model in the initial 3D scene can be adjusted based on the spatial perspective information to obtain the target 3D scene. In this way, the spatial shape and perspective characteristics of the target object can be restored by relying on the spatial perspective information, avoiding problems such as scene fusion misalignment and scale imbalance caused by perspective deviation.

[0100] Alternatively, in another example, the primitive information and spatial perspective information of the target object in the original 2D image can be obtained, and the 3D model in the initial 3D scene can be adjusted according to the primitive information and spatial perspective information to obtain the target 3D scene. This improves the generation efficiency and accuracy of the target 3D scene, and further ensures that the rendered image can not only fit the characteristics of the original material, but also have excellent 3D scene presentation effect to meet user needs.

[0101] Step S706-4: Display the target 3D scene in the main display area.

[0102] In this way, the disclosed solution, in response to the replacement operation, replaces the original model in the selected state in the 3D scene to be processed with the initial 3D model to obtain the initial 3D scene. Furthermore, the initial 3D model in the initial 3D scene is adjusted to obtain the target 3D scene, which is then displayed in the main display area. On the one hand, it can accurately constrain the shape of the 3D model based on primitive information, avoiding structural distortion after replacement. At the same time, it combines spatial perspective information to calibrate the perspective of the 3D model, restoring the perspective relationship of the 2D main image, thereby reducing problems such as size deviation, positional offset, and perspective misalignment, and ensuring the consistency and coordination between the target object and the original material in the target 3D scene. On the other hand, by automatically completing information acquisition and model adjustment by the system, the workload of manual model adjustment by users is reduced, the operation threshold for non-professional users is significantly lowered, human error and time consumption are reduced, the generation efficiency of the target 3D scene is significantly improved, and the production cycle is shortened.

[0103] Here, in a specific example, after obtaining the initial 3D scene, the initial 3D model in the initial 3D scene can be adjusted to meet the user's personalized design needs. Specifically, the target 3D scene can be obtained in the following way: Specifically, the adjustment of the initial 3D model in the initial 3D scene based on the primitive information and / or spatial perspective information of the target object in the original 2D image to obtain the target 3D scene (step S706-3 above) can specifically include: Step S706-3-1: Obtain the model pose information of the initial three-dimensional model.

[0104] Here, in one example, the model pose information of the initial 3D model can be used to describe the spatial transformation information of the initial 3D model to the 2D main image. For example, it may include, but is not limited to, at least one of the following: rotation parameters, orientation parameters, etc. of the initial 3D model.

[0105] For example, model pose information can specifically include rotation parameters, such as rotation angle values ​​based on the x-axis, y-axis, and z-axis. Specifically, if the target object in the original 2D image is tilted at 45°, the model pose information will include a parameter for rotating 45° around the y-axis, ensuring that after this rotation, the projection of the 3D model in the main view is consistent with the appearance and pose of the 2D image. This ensures that the model in the target 3D scene and the original 2D image design are highly compatible in spatial features, improving the realism and harmony of scene adaptation.

[0106] Step S706-3-2: Based on the model pose information of the initial 3D model and the primitive information and / or spatial perspective information of the target object in the original 2D image, adjust the initial 3D model in the initial 3D scene to obtain the target 3D scene.

[0107] For example, in one instance, the initial 3D model in the initial 3D scene can be adjusted based on the model pose information of the initial 3D model and the primitive information of the target object in the original 2D image to obtain the target 3D scene. In this way, the morphological matching degree between the model in the target 3D scene and the original 2D image can be improved.

[0108] Alternatively, in another example, the initial 3D model in the initial 3D scene can be adjusted based on the model pose information of the initial 3D model and the spatial perspective information of the target object in the original 2D image to obtain the target 3D scene. In this way, the target 3D scene not only conforms to the 3D spatial logic, but also fits the shooting perspective characteristics of the original 2D image, thus enhancing the realism and coordination of the scene.

[0109] Alternatively, in another example, the initial 3D model in the initial 3D scene can be adjusted based on the model pose information of the initial 3D model, as well as the primitive information and spatial perspective information of the target object in the original 2D image, to obtain the target 3D scene. In this way, the model in the target 3D scene is highly consistent with the original 2D image in terms of shape and perspective, laying the foundation for rendering a high-quality 3D effect image.

[0110] It should be noted that after obtaining the target 3D scene, a pose adjustment tool can also be displayed on the initial 3D model in the target 3D scene displayed in the main display area. In this way, the pose of the initial 3D model can be further adjusted by using the pose adjustment tool (for example, translation or rotation).

[0111] In this way, the present solution utilizes the model pose information of the acquired initial 3D model and the primitive information and / or spatial perspective information of the target object in the original 2D image to adjust the initial 3D model and obtain the target 3D scene. In this way, the reference parameters for spatial transformation of the initial 3D model can be clearly defined with the help of the model pose information, and the size ratio of the initial 3D model can be accurately calibrated with the primitive information. The rotation posture and viewpoint orientation of the model can be optimized with the spatial perspective information, avoiding problems such as size deviation, position offset, and viewpoint misalignment of the initial 3D model. This ensures the automation and traceability of the adjustment process, accurately eliminates the deviations in posture, proportion, and structure between the initial 3D model and the original 2D image, and makes the initial 3D model in the target 3D scene retain the basic form of the 3D model and be highly consistent with the core visual features of the original 2D image design, significantly improving the realism and coordination of scene adaptation.

[0112] Figure 8 This is an illustrative flowchart of a controllable generation method for obtaining a three-dimensional rendering based on a two-dimensional image according to an embodiment of this application. Figure 5 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figures 1-7 The relevant content of the method shown in any of the attached figures can also be applied to this example, and the relevant content will not be described again in this example.

[0113] Furthermore, the method includes at least a portion of the following: For example... Figure 8 As shown, it includes: Step S801: In response to the input operation, display the original two-dimensional image of the target object in the three-dimensional scene design interface.

[0114] Step S802: In response to the image cutout operation, display the two-dimensional main image of the target object in the three-dimensional scene design interface.

[0115] Here, the two-dimensional subject image is obtained by cutting out the area where the target object is located in the original two-dimensional image.

[0116] Step S803: In response to the intelligent task operation for the two-dimensional subject image, display the initial three-dimensional model of the two-dimensional subject image in the first display area of ​​the three-dimensional scene design interface, and display multiple candidate three-dimensional scenes for the user to select in the second display area of ​​the three-dimensional scene design interface.

[0117] Step S804: In response to the scene adaptation operation, display the target 3D scene in the main display area of ​​the 3D scene design interface.

[0118] Here, the target 3D scene is the 3D scene obtained by adapting the initial 3D model to the selected 3D scene to be processed; the 3D scene to be processed is one of the multiple candidate 3D scenes.

[0119] Step S805: In response to the rendering operation for the target 3D scene, an initial rendering effect image is obtained.

[0120] Here, the initial rendering effect is obtained by rendering the target three-dimensional scene using a 3D rendering engine.

[0121] For example, in one example, the rendering operation refers to the active interactive operation initiated by the user in the 3D scene design interface for the generated target 3D scene, which triggers the system to generate the initial rendering effect map.

[0122] Furthermore, in practical applications, users can freely and precisely perform 3D design processes before rendering. For example, after completing independent 3D design such as model adjustment, material settings, and lighting settings, users can trigger rendering operations to render the target 3D scene and obtain the initial rendering effect.

[0123] Step S806: Call the large model, and based on the initial rendering effect image, the two-dimensional main image, and the contour mask image that can describe the contour of the target object in the original two-dimensional image, obtain the target rendering effect image, so as to display the target rendering effect image in the main display area.

[0124] Here, the main detail features of the target object in the target rendering effect image are matched with the main detail features of the target object in the two-dimensional main image.

[0125] Here, the large model can be specifically a large language model or an image generative model, etc. This disclosed solution does not limit the specific model.

[0126] Further, step S806 can be specifically as follows: inputting the initial rendering effect image, the two-dimensional main image, and the outline mask image (e.g., a black and white mask image) into the large model to obtain the target rendering effect image output by the large model.

[0127] In other words, after the rendering operation is triggered in this disclosed solution, firstly, the 3D rendering engine needs to be called to render the target 3D scene that incorporates the initial 3D model. Then, AI technology is used to carry out detail transfer (e.g., accurately transfer the main detail features of the target object in the 2D main image to the corresponding position of the main body in the target rendered image) and style transfer (e.g., transfer the specific style of the target object in the 2D main image to the target object in the target rendered image to unify the visual style) to further correct the initial rendered image and ensure that the main body in the final target rendered image is consistent with the 2D main image.

[0128] In this way, the disclosed solution can, in response to rendering operations on the target 3D scene, first use a 3D rendering engine to render the target 3D scene to obtain an initial rendering effect image, then call a large model and combine it with the initial rendering effect image, the 2D main image, and the contour mask image to obtain and display the target rendering effect image. In this way, the 3D rendering engine ensures the physical realism of the 3D scene in the initial rendering effect image (such as lighting, materials, texture, etc.), and then accurately transfers the main details of the target object through the large model and the 2D main image and contour mask image, avoiding the detail deviation and physical unrealism problems caused by AI rendering relying solely on text / images; at the same time, it enables the target rendering effect image to have both the scene realism of 3D engine rendering and a high degree of matching with the details of the target object in the 2D main image, meeting the user's dual needs for realistic scenes and accurate details.

[0129] The following combination Figure 9 This disclosure will be described in detail, and specifically, this disclosure may include the following steps: Step S901: In the 3D scene design interface, users can upload any 2D image (i.e., the original 2D image mentioned above). The 2D image can be the original image with a background, or it can be an image with the background removed and only the main body (i.e., the target object mentioned above) retained.

[0130] Step S902: In response to the image cutout operation, automatically process the uploaded 2D image and generate a series of data; further, the following data can be generated: Perform a cutout operation on the 2D image to generate the main image (i.e., the two-dimensional main image mentioned above). Based on the main image, a main mask image (corresponding to the outline mask image described above) is generated according to the RGB values ​​to describe the outline of the main image. The primitive information for generating 2D images may include, for example, image size, anchor point information of the subject, etc. Algorithm calibration technology is used to identify the subject category of the 2D image and the perspective relationship of the subject in the 2D image (that is, the spatial perspective information mentioned above, such as the pitch angle and field of view of the original image).

[0131] It should be noted that this disclosed solution does not impose specific restrictions on the order of the obtained main image, main mask image, primitive information, main category, and perspective relationship.

[0132] Step S903: Based on the main image and main mask image obtained in step S902, generate a 3D model (corresponding to the initial 3D model mentioned above) and model pose information (used to describe the spatial transformation information of the 2D model corresponding to the 2D image restored from the 3D model) using AI large model technology.

[0133] Here, the above-mentioned use of AI large model technology to generate a complete 3D model can specifically include: automatically implementing the AI ​​material translation step, so that the geometric vertices and materials of the generated 3D model are close to the 2D model corresponding to the 2D image, so as to restore the 2D image to the greatest extent. In this way, the subsequently generated rendering effect image is highly matched with the detailed features of the main subject in the main image.

[0134] Here, it is understandable that the 3D model obtained at this point does not aim to completely reproduce the geometric vertices and materials of the 2D model in the 2D image. The final rendered image can be obtained by subsequent correction using AI technology.

[0135] Step S904: Based on the perspective relationship and primitive information obtained in step S902 and the 3D model obtained in step S903, adapt the 3D scene selected by the user and make automatic adjustments to obtain the final target 3D scene adapted to the 3D model.

[0136] Here, on the one hand, since step S902 identifies the subject category of the target object, a more suitable 3D scene can be recommended; on the other hand, the perspective relationship obtained by the algorithm calibration can automatically adjust the pose of the 3D model and its subject in the 3D scene, thereby restoring the perspective relationship of the 2D model in the 2D image.

[0137] Step S905: In the 3D scene design interface, users can freely and accurately design the models, materials, lighting, etc. in the target 3D scene; and after the design is completed, the initial rendering image is obtained by rendering through the 3D rendering engine (i.e. the initial rendering effect image mentioned above).

[0138] Step S906: Based on the main image, the main mask image, and the initial rendering image obtained in step S905, and using AI large model technology, style and detail transfer correction are achieved to obtain the final target rendering image.

[0139] In summary, this solution has the following advantages: First, this disclosed solution replaces the complex manual operations of traditional 3D design with automated processing. That is, users do not need professional modeling skills to automatically generate an initial 3D model through intelligent task operations; and 2D main images and outline masking images can be automatically generated through image cutout operations without the need for manual outlining. In other words, in this disclosed solution, users only need to go through a simple "upload-select scene-trigger operation" process to advance the workflow, reducing cumbersome operations and lowering the technical threshold for generating 3D renderings.

[0140] Second, this disclosed solution ensures that the features of the two-dimensional material are not lost through multi-stage design. That is, the initial three-dimensional model is generated based on the two-dimensional main image and the outline mask image. At the same time, during scene adaptation, the model is adjusted by combining the primitive information and / or spatial perspective information of the original image to avoid problems such as size deviation and perspective misalignment. In the rendering stage, the detailed features of the two-dimensional main image are transferred to the three-dimensional rendering image through the large model. The 3D rendering engine is used to ensure the realism of the scene's lighting and materials, thereby ensuring the physical reality of the three-dimensional scene of the target rendering effect image.

[0141] Third, users can independently select the appropriate 3D scene to be processed from the candidate scenes, and can also preview the adaptation effect by dragging the model to adjust its position and rotating the view in the main display area. In other words, users can adjust different display effects according to their needs to meet the personalized design preferences in different scenarios, improve the user experience, and at the same time improve the autonomy and controllability of the generated effect.

[0142] This disclosure provides a controllable device for generating three-dimensional renderings based on two-dimensional images, such as... Figure 10 As shown, the device includes: The task triggering unit 1001 is used to display a two-dimensional main image of the target object in the three-dimensional scene design interface, and in response to the intelligent task operation on the two-dimensional main image, to display an initial three-dimensional model of the two-dimensional main image in the first display area of ​​the three-dimensional scene design interface, and to display multiple candidate three-dimensional scenes for the user to select in the second display area of ​​the three-dimensional scene design interface. The scene adaptation unit 1002 is used to display a target 3D scene in the main display area of ​​the 3D scene design interface in response to the scene adaptation operation. The target 3D scene is a 3D scene obtained by adapting the initial 3D model to the selected 3D scene to be processed. The 3D scene to be processed is one of the multiple candidate 3D scenes. AI rendering unit 1003 is used to display a target rendering effect image in the main display area in response to a rendering operation for the target 3D scene; wherein, the target rendering effect image is a rendering effect image after transferring the main detail features of the target object in the 2D main image to the 3D rendering image using a large model; the 3D rendering image is obtained after rendering the target 3D scene using a 3D rendering engine.

[0143] In a specific example of the disclosed solution, the task triggering unit is specifically used for: In response to an input operation, the original two-dimensional image of the target object is displayed in the three-dimensional scene design interface; In response to the image cutout operation, a two-dimensional main image of the target object is displayed in the 3D scene design interface; wherein, the two-dimensional main image is obtained by cutting out the area where the target object is located in the original two-dimensional image.

[0144] In a specific example of the scheme disclosed herein, the task triggering unit is further configured to: The object features of the target object in the original two-dimensional image are identified; wherein the object features of the target object include at least one of the following: primitive information, subject category, and spatial perspective information.

[0145] In a specific example of the scheme disclosed herein, the task triggering unit is further configured to: Based on the two-dimensional subject image and the pixel features of the original two-dimensional image, a contour mask image is obtained that can describe the contour of the target object in the original two-dimensional image.

[0146] In a specific example of the disclosed solution, the scene adaptation unit is specifically used for: In response to a scene selection operation, the selected 3D scene to be processed is displayed in the main display area; In response to a model selection operation on the original model in the 3D scene to be processed displayed in the main display area, the original model in the selected state is obtained. In response to the replacement operation, the target 3D scene is displayed in the main display area; wherein, the target 3D scene is the 3D scene obtained by replacing the original model in the selected state of the 3D scene to be processed with the initial 3D model.

[0147] In a specific example of the disclosed solution, the scene adaptation unit is specifically used for: In response to the replacement operation, the original model in the selected state in the three-dimensional scene to be processed is replaced with the initial three-dimensional model to obtain the initial three-dimensional scene; Obtain the primitive information and / or spatial perspective information of the target object in the original two-dimensional image; Based on the primitive information and / or spatial perspective information of the target object in the original two-dimensional image, the initial three-dimensional model in the initial three-dimensional scene is adjusted to obtain the target three-dimensional scene; The target 3D scene is displayed in the main display area.

[0148] In a specific example of the disclosed solution, the scene adaptation unit is specifically used for: Obtain the model pose information of the initial 3D model; Based on the model pose information of the initial 3D model, and the primitive information and / or spatial perspective information of the target object in the original 2D image, the initial 3D model in the initial 3D scene is adjusted to obtain the target 3D scene.

[0149] In a specific example of the disclosed solution, the task triggering unit is specifically used for: In response to intelligent task operations on a two-dimensional subject image, a contour mask image that can describe the contour of the target object in the original two-dimensional image is obtained. The large model is invoked, and an initial three-dimensional model of the two-dimensional main body image is obtained based on the two-dimensional main body image and the contour mask image; The initial 3D model of the 2D main image is displayed in the first display area of ​​the 3D scene design interface.

[0150] In a specific example of the disclosed solution, the AI ​​rendering unit is specifically used for: In response to a rendering operation on the target 3D scene, an initial rendering effect image is obtained, wherein the initial rendering effect image is obtained by rendering the target 3D scene using a 3D rendering engine; The large model is invoked, and based on the initial rendering effect image, the two-dimensional main image, and the contour mask image that can describe the outline of the target object in the original two-dimensional image, the target rendering effect image is obtained, and the target rendering effect image is displayed in the main display area; wherein, the main detail features of the target object in the target rendering effect image match the main detail features of the target object in the two-dimensional main image.

[0151] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.

[0152] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0153] Figure 11This is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 11 As shown, the electronic device includes a memory 1110 and a processor 1120. The memory 1110 stores a computer program that can run on the processor 1120. The number of memories 1110 and processors 1120 can be one or more. The memory 1110 can store one or more computer programs, which, when executed by the electronic device, cause the electronic device to perform the method provided in the above-described method embodiments. The electronic device may also include a communication interface 1130 for communicating with external devices and performing data exchange and transmission.

[0154] If the memory 1110, processor 1120, and communication interface 1130 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0155] Optionally, in a specific implementation, if the memory 1110, processor 1120 and communication interface 1130 are integrated on a single chip, the memory 1110, processor 1120 and communication interface 1130 can communicate with each other through an internal interface.

[0156] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.

[0157] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct RAMBUS RAM (DR RAM).

[0158] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, Bluetooth, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)). It is worth noting that the computer-readable storage media mentioned in this disclosure can be non-volatile storage media; in other words, it can be non-transient storage media.

[0159] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0160] In the description of the embodiments of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0161] In the description of the embodiments disclosed herein, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0162] In the description of embodiments of this disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.

[0163] The above description is merely an exemplary embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.

Claims

1. A controllable generation method for obtaining a three-dimensional effect drawing based on a two-dimensional image, comprising: displaying a two-dimensional subject drawing of a target object in a three-dimensional scene design interface; in response to an intelligent task operation on the two-dimensional subject drawing, displaying an initial three-dimensional model of the two-dimensional subject drawing in a first display area of the three-dimensional scene design interface, and displaying a plurality of candidate three-dimensional scenes selected by a user in a second display area of the three-dimensional scene design interface; in response to a scene adaptation operation, displaying a target three-dimensional scene in a main display area of the three-dimensional scene design interface, wherein the target three-dimensional scene is a three-dimensional scene obtained by adapting the initial three-dimensional model to a selected three-dimensional scene to be processed; and the three-dimensional scene to be processed is one of the plurality of candidate three-dimensional scenes; in response to a rendering operation on the target three-dimensional scene, displaying a target rendering effect drawing in the main display area; wherein the target rendering effect drawing is a rendering effect drawing obtained by migrating subject detail features of a target object in the two-dimensional subject drawing to a three-dimensional rendering drawing using a large model; and the three-dimensional rendering drawing is obtained by rendering the target three-dimensional scene using a 3D rendering engine.

2. The method of claim 1, wherein, The displaying of the two-dimensional subject drawing of the target object in the three-dimensional scene design interface comprises: in response to an input operation, displaying an original two-dimensional drawing of the target object in the three-dimensional scene design interface; in response to a matting operation, displaying a two-dimensional subject drawing of the target object in the three-dimensional scene design interface; wherein the two-dimensional subject drawing is obtained by performing a matting operation on a region in which the target object is located in the original two-dimensional drawing.

3. The method of claim 2, further comprising: identifying object features of the target object in the original two-dimensional drawing; wherein the object features of the target object include at least one of the following: primitive information, subject category, and spatial perspective information.

4. The method of claim 2, further comprising: based on the two-dimensional subject drawing and pixel features of the original two-dimensional drawing, obtaining a contour mask drawing capable of describing a contour of the target object in the original two-dimensional drawing.

5. The method according to any one of claims 2-4, wherein, The displaying of the target three-dimensional scene in the main display area of the three-dimensional scene design interface in response to the scene adaptation operation comprises: in response to a scene selection operation, displaying the selected three-dimensional scene to be processed in the main display area; in response to a model selection operation on an original model in the three-dimensional scene to be processed displayed in the main display area, obtaining the original model in a selected state; in response to a replacement operation, displaying the target three-dimensional scene in the main display area; wherein the target three-dimensional scene is a three-dimensional scene obtained by replacing the original model in the selected state in the three-dimensional scene to be processed with the initial three-dimensional model.

6. The method of claim 5, wherein, The displaying of the target three-dimensional scene in the main display area in response to the replacement operation comprises: in response to a replacement operation, replacing the original model in the selected state in the three-dimensional scene to be processed with the initial three-dimensional model to obtain an initial three-dimensional scene; obtaining primitive information and / or spatial perspective information of the target object in the original two-dimensional drawing; According to the primitive two-dimensional graph in the target object information and / or spatial perspective information, the initial three-dimensional model in the initial three-dimensional scene is adjusted to obtain the target three-dimensional scene; In the main display area, the target three-dimensional scene is displayed.

7. The method of claim 6, wherein, According to the primitive two-dimensional graph in the target object information and / or spatial perspective information, the initial three-dimensional model in the initial three-dimensional scene is adjusted to obtain the target three-dimensional scene, comprising: Obtain the model pose information of the initial three-dimensional model; According to the initial three-dimensional model of the model pose information, and the target object in the original two-dimensional graph in the primitive information and / or spatial perspective information, the initial three-dimensional model in the initial three-dimensional scene is adjusted to obtain the target three-dimensional scene.

8. The method of any one of claims 2-4, wherein, The initial three-dimensional model of the two-dimensional subject graph is displayed in the first display area of the three-dimensional scene design interface in response to the intelligent task operation of the two-dimensional subject graph, comprising: In response to the intelligent task operation of the two-dimensional subject graph, a contour mask graph capable of describing the target object in the contour of the original two-dimensional graph is obtained; Call a large model, and based on the two-dimensional subject graph and the contour mask graph, obtain the initial three-dimensional model of the two-dimensional subject graph; In the first display area of the three-dimensional scene design interface, the initial three-dimensional model of the two-dimensional subject graph is displayed.

9. The method of any one of claims 2-4, wherein, In response to the rendering operation of the target three-dimensional scene, the target rendering effect graph is displayed in the main display area, comprising: In response to the rendering operation of the target three-dimensional scene, an initial rendering effect graph is obtained, wherein the initial rendering effect graph is obtained by rendering the target three-dimensional scene using a 3D rendering engine; Call a large model, and based on the initial rendering effect graph, the two-dimensional subject graph, and the contour mask graph capable of describing the target object in the contour of the original two-dimensional graph, obtain the target rendering effect graph to display the target rendering effect graph in the main display area; wherein the main body detail features of the target object in the target rendering effect graph match the main body detail features of the target object in the two-dimensional subject graph.

10. A controllable generation device for obtaining a three-dimensional effect graph based on a two-dimensional image, comprising: A task triggering unit is configured to display a two-dimensional subject graph of a target object in a three-dimensional scene design interface, display an initial three-dimensional model of the two-dimensional subject graph in a first display area of the three-dimensional scene design interface in response to an intelligent task operation of the two-dimensional subject graph, and display a plurality of candidate three-dimensional scenes selected by a user in a second display area of the three-dimensional scene design interface; A scene adaptation unit is configured to display a target three-dimensional scene in a main display area of the three-dimensional scene design interface in response to a scene adaptation operation, wherein the target three-dimensional scene is a three-dimensional scene obtained by scene adaptation of the initial three-dimensional model and a selected three-dimensional scene to be processed; the three-dimensional scene to be processed is one of the plurality of candidate three-dimensional scenes. an AI rendering unit, configured to, in response to a rendering operation on the target three-dimensional scene, display a target rendering effect picture in the main display area; wherein the target rendering effect picture is a rendering effect picture obtained by using a large model to migrate subject detail features of a target object in a two-dimensional subject picture to a three-dimensional rendering picture; and the three-dimensional rendering picture is obtained by using a 3D rendering engine to render the target three-dimensional scene.

11. The apparatus of claim 10, wherein, The task triggering unit is specifically configured to: in response to an input operation, display an original two-dimensional picture of the target object in the three-dimensional scene design interface; in response to a cutout operation, display a two-dimensional subject picture of the target object in the three-dimensional scene design interface; wherein the two-dimensional subject picture is obtained by performing cutout processing on a region in which the target object is located in the original two-dimensional picture.

12. The apparatus of claim 11, wherein, The scene adaptation unit is specifically configured to: in response to a scene selection operation, display a selected three-dimensional scene to be processed in the main display area; in response to a model selection operation on an original model in the three-dimensional scene to be processed displayed in the main display area, obtain the original model in a selected state; in response to a replacement operation, display the target three-dimensional scene in the main display area; wherein the target three-dimensional scene is a three-dimensional scene obtained by replacing the original model in the selected state in the three-dimensional scene to be processed with the initial three-dimensional model.

13. The apparatus of claim 12, wherein, The scene adaptation unit is specifically configured to: in response to a replacement operation, replace the original model in the selected state in the three-dimensional scene to be processed with the initial three-dimensional model to obtain an initial three-dimensional scene; obtain graph element information and / or spatial perspective information of the target object in the original two-dimensional picture; adjust the initial three-dimensional model in the initial three-dimensional scene according to the graph element information and / or spatial perspective information of the target object in the original two-dimensional picture to obtain the target three-dimensional scene; display the target three-dimensional scene in the main display area.

14. The apparatus of any one of claims 11-13, wherein, The task triggering unit is specifically configured to: in response to an intelligent task operation on the two-dimensional subject picture, obtain a contour mask picture capable of describing a contour of the original two-dimensional picture in which the target object is located; invoke a large model, and based on the two-dimensional subject picture and the contour mask picture, obtain an initial three-dimensional model of the two-dimensional subject picture; display the initial three-dimensional model of the two-dimensional subject picture in a first display area of the three-dimensional scene design interface.

15. The apparatus of any one of claims 11-13, wherein, The AI rendering unit is specifically configured to: in response to a rendering operation on the target three-dimensional scene, obtain an initial rendering effect picture, wherein the initial rendering effect picture is obtained by using a 3D rendering engine to render the target three-dimensional scene; invoke a large model, and based on the initial rendering effect picture, the two-dimensional subject picture, and a contour mask picture capable of describing a contour of the original two-dimensional picture in which the target object is located, obtain the target rendering effect picture to display the target rendering effect picture in the main display area; wherein subject detail features of the target object in the target rendering effect picture match subject detail features of the target object in the two-dimensional subject picture.

16. An electronic device, comprising: at least one processor; and ​ a memory in communication with the at least one processor; wherein the memory has stored instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.

17. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, the computer instructions are for causing the computer to perform the method of any one of claims 1-9.

18. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Image display method, image display device, electronic equipment and storage medium

    CN114092675A

  • Display graph generation method and device, equipment and storage medium

    CN117788304A

  • Effect picture generation method and device, electronic equipment and storage medium

    CN118097083A

  • Image processing method and device, equipment and storage medium

    CN118397202A

  • Model and scene fusion rendering method and device, equipment and storage medium

    CN119516078A