3D relief photo frame automatic generation method capable of customizing relief shapes and characters

By using deep learning and Boolean operation techniques, an end-to-end automated process from image input to embossed frame generation has been achieved, solving the problems of low efficiency and lack of personalization in the design and generation of 3D embossed frames, and realizing fast and high-quality personalized customization.

CN121810933APending Publication Date: 2026-04-07XIAMEN SUXIANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing 3D relief photo frame design and generation suffers from low customization efficiency, insufficient personalization, and fragmented processes, making it difficult to quickly personalize user needs.

Method used

By employing deep learning image segmentation and Boolean operation techniques, an end-to-end automated process is achieved from image input to embossed frame generation, including image super-resolution enhancement, instance segmentation, Boolean operation fusion of embossing and frame, and generation and positioning of text grids.

Benefits of technology

It enables rapid and automated generation of personalized 3D relief photo frames, reducing manual intervention costs, improving production efficiency, and ensuring high-quality and personalized customization of reliefs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121810933A_ABST
    Figure CN121810933A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic generation method of a 3D embossment photo frame capable of customizing embossment shapes and characters, and belongs to the technical field of computer vision, three-dimensional reconstruction and computer graphics. A high resolution image and a pixel level mask are generated to precisely define a relief core region. Secondly, Boolean operation fusion of the embossment and the photo frame is carried out, a 3D photo frame grid is generated through loading or programming, and an embossment main body grid and the photo frame grid are fused through Boolean union set operation; and finally, generating and finally fusing character grids, generating 3D character grids according to custom characters input by a user, positioning the 3D character grids to a designated area of a photo frame through a rigid body transformation matrix, executing Boolean union set operation again, and outputting a complete 3D model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of computer vision, 3D reconstruction and computer graphics technology, and specifically relates to a method for automatically generating 3D relief frames with custom relief shapes and text. Background Technology

[0002] Currently, in the field of personalized 3D relief photo frame design and generation, existing technologies mainly suffer from the following pain points: 1. Low customization efficiency: The shape design and text embedding of traditional 3D relief photo frames rely on manual modeling. From creative conception to model output, it often takes several hours or even days, which is difficult to meet users' needs for rapid customization.

[0003] 2. Lack of personalization: Existing automation solutions are mostly limited to fixed templates, and users cannot freely define the shape and outline of the relief (such as irregular shapes and artistic outlines) and the style of the text (such as font, size, and relief depth), resulting in serious product homogenization.

[0004] 3. Fragmented process: From image input, relief shape design, text embedding to frame generation, each step is an independent operation, lacking an end-to-end automated collaborative process. This leads to the continuous accumulation of errors in each step, ultimately affecting the overall quality of the 3D relief frame.

[0005] In view of this, this solution was developed. Summary of the Invention

[0006] In view of the shortcomings of the prior art, the technical problem to be solved by the present invention is to provide an automatic generation method for 3D relief photo frames with custom relief shapes and text. It can be applied to the automated design and manufacturing of personalized 3D relief crafts, and solves the problems of shape and text design relying on manual labor, low generation efficiency and insufficient personalization in the traditional 3D relief photo frame customization process.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a method for automatically generating 3D relief photo frames with custom relief shapes and text, comprising the following steps: S1. Deep learning-based image segmentation: Super-resolution enhancement and instance segmentation are performed on a single image uploaded by the user to generate a high-resolution image and a pixel-level mask to accurately define the core region of the relief. S2. Boolean operation fusion of relief and frame: Load or programmatically generate 3D frame mesh, and fuse the relief main mesh with the frame mesh through Boolean union operation; S3. Generation and final fusion of text mesh: Based on the user-input custom text, a 3D text mesh is generated using a font library and extrusion algorithm. The text mesh is then located to the specified area of ​​the frame using a rigid body transformation matrix. A Boolean union operation is then performed again to fuse the text mesh with the relief-frame combined mesh, outputting a complete 3D model that integrates relief, frame, and text.

[0008] Furthermore, the image segmentation in step S1 uses a Mask R-CNN or U-Net model to achieve pixel-level separation between the core subject and the background, and marks the text region and key image feature points.

[0009] Furthermore, the super-resolution enhancement in step S1 employs an ESRGAN or RCAN deep learning model, with a focus on enhancing high-frequency textures.

[0010] Furthermore, the Boolean union operation in step S2 is implemented through a 3D modeling software library to ensure the watertightness of the merged mesh.

[0011] Furthermore, in step S3, the text grid positioning is controlled by a rigid body transformation matrix to determine its spatial position, orientation, and scaling. The positioning area includes the lower half of the frame or other user-specified areas.

[0012] Furthermore, an automatic generation module for 3D embossed photo frames with custom embossed shapes and text is applied, the automatic generation module including an input layer, a data preprocessing layer, a core production layer, and a fusion output layer.

[0013] Furthermore, the input layer includes uploading images, entering custom text, and selecting templates.

[0014] Furthermore, the data preprocessing layer includes a first preprocessing pipeline, a second preprocessing pipeline, and a third preprocessing pipeline; The first preprocessing pipeline is used for image processing to generate high-resolution images and binarized masks as input to the core production layer; The second preprocessing pipeline is used to process custom text, producing a 3D text mesh that is input into the fusion output layer; The third preprocessing pipeline is used for template processing, producing a 3D frame mesh which is then input into the fusion output layer.

[0015] Furthermore, the core production layer is started immediately after the first preprocessing pipeline input, and through single-view depth prediction and predicted depth reprocessing, a high-precision core relief mesh is finally output.

[0016] Furthermore, the fusion output layer performs a Boolean union on the contents output from the first preprocessing pipeline, the second preprocessing pipeline, the third preprocessing pipeline, and the core production layer to output a 3D relief model.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. End-to-end automation: This invention constructs an end-to-end automated workflow from multi-view 2D image input to final integrated (including frame and text) 3D relief mesh output. Multiple preprocessing pipelines eliminate the need for tedious manual operations such as 3D modeling, digital sculpting, pose alignment, or geometric repair.

[0018] 2. Zero human intervention cost: Compared to the traditional method of creating reliefs by relying on 3D artists to manually design and carve them, this invention automatically completes the geometric integration of the frame and text through Boolean operations in steps (2) and (3). This automation completely replaces high-cost manual labor, greatly reducing the technical threshold and labor costs of relief production.

[0019] 3. High production efficiency: By eliminating the time-consuming manual modeling and modification process, this invention can shorten the production cycle of reliefs from several days or weeks to several hours or even minutes (depending on computing resources), realizing rapid customization and efficient production of personalized 3D reliefs. Attached Figure Description

[0020] Figure 1 This is a flowchart of a method for automatically generating 3D relief photo frames with custom relief shapes and text, according to the present invention. Detailed Implementation

[0021] To make the above features and advantages of the present invention more apparent and understandable, specific embodiments are described below in conjunction with the accompanying drawings for detailed explanation.

[0022] like Figure 1 As shown, this embodiment provides a method for automatically generating 3D embossed photo frames with custom embossed shapes and text, including the following steps: S1. Deep learning-based image segmentation: Super-resolution enhancement and instance segmentation are performed on a single image uploaded by the user to generate a high-resolution image and a pixel-level mask to accurately define the core region of the relief. Step S1: The core objective of this stage is to improve the quality of the input image and accurately define the core custom area (ROI) of the embossed frame, so as to provide high-quality data support for subsequent embossing design, text embedding and 3D modeling.

[0023] To address the common issues of low resolution and blurred details in user-uploaded photos (life photos, commemorative photos, etc.) due to device limitations, distance, and lighting, this invention employs deep learning super-resolution models such as ESRGAN and RCAN to enhance high-frequency textures related to embossing effects—such as hair strands in portraits, clothing folds, and the layers of vegetation in landscapes. This not only improves image realism but also provides reliable features for embossing depth design, texture restoration, and text integration, preventing the finished product from appearing flat and lifeless.

[0024] Focusing on core design elements, this invention generates pixel-level masks using high-precision instance segmentation models such as Mask R-CNN and U-Net to accurately separate core subjects like portraits and pets from the background, while simultaneously marking text areas and key image feature points. This process not only eliminates background interference but also provides a basis for layered design and depth-differentiated settings of the "subject + text," ensuring that the relief core is prominent, the layers are distinct, and personalized needs are met.

[0025] S2. Boolean operation fusion of relief and frame: Load or programmatically generate 3D frame mesh, and fuse the relief main mesh with the frame mesh through Boolean union operation; Step S2: In this stage, the geometric merging of the relief body and the custom frame is performed.

[0026] Boolean Operation: First, load or programmatically generate a custom 3D frame mesh. Then, perform a Boolean union operation in Constructive Solid Geometry (CSG) using a 3D modeling software library (such as CGAL or libigl).

[0027] S3. Generation and final fusion of text mesh: Based on the user-input custom text, a 3D text mesh is generated using a font library and extrusion algorithm. The text mesh is then located to the specified area of ​​the frame using a rigid body transformation matrix. A Boolean union operation is then performed again to fuse the text mesh with the relief-frame combined mesh, outputting a complete 3D model that integrates relief, frame, and text.

[0028] Step S3: This stage integrates personalized text with embossed frames.

[0029] Text Mesh Generation and Positioning: Based on user-inputted custom text (string), a corresponding 3D text mesh is generated using a font library and an extrusion algorithm. Then, by applying a rigid transformation matrix, the spatial position, orientation, and scaling of the text mesh are precisely controlled, placing it within a specified area (e.g., the lower half) of the 3D embossed frame.

[0030] Boolean Union: Perform the Boolean union operation again to merge the positioned text mesh with the "emboss-frame" combined mesh generated in step 7. Finally, output a complete and unified 3D mesh model integrating the embossing, frame, and custom text.

[0031] This embodiment also provides an automatic generation module for 3D embossed photo frames with custom embossed shapes and text. It applies an automatic generation method for 3D embossed photo frames with custom embossed shapes and text, including an input layer, a data preprocessing layer, a core production layer, and a fusion output layer.

[0032] The input layer includes uploading images, entering custom text, and selecting templates.

[0033] The operation process is initiated by the user through the front-end interface: 1. Upload Image: The user uploads a 2D image (single image) of the target subject.

[0034] 2. Input text: The user enters a custom string (custom text) in the specified text box.

[0035] 3. Select a template: The user selects one from the UI interface provided by the system (such as a list of thumbnails).

[0036] The data preprocessing layer includes a first preprocessing pipeline, a second preprocessing pipeline, and a third preprocessing pipeline; After receiving user input (images, text, and templates), the system immediately starts three parallel preprocessing pipelines to prepare the data required for subsequent algorithms: The first preprocessing pipeline performs image processing: First, the system sends a single image uploaded by the user into this pipeline; second, step S1 is executed: first, "image super-resolution" is performed to enhance details, followed by "subject segmentation" to extract the target mask; finally, the output is a high-resolution image and a binarized mask. This output will serve as the input to the core generation layer.

[0037] The second preprocessing pipeline is used to process custom text: First, the system sends the user-inputted custom text into this pipeline; second, step S3 is executed: the "3D Text Meshization" module converts the 2D string into 3D geometry using an extrusion algorithm; finally, the output is a 3D text mesh. This output awaits delivery to the fusion output layer.

[0038] The third preprocessing pipeline is used for template processing: First, the system sends the template ID selected by the user into this pipeline; Step S2 is executed: The system uses this ID to search in the internal "Template Library"; Then, after a successful search, the corresponding 3D frame model is loaded and normalized, finally producing a 3D frame mesh. This output will wait to be sent to the fusion output layer.

[0039] The core production layer starts immediately after the first preprocessing pipeline input. Through single-view depth prediction and predicted depth reprocessing, it finally outputs a high-precision core relief mesh, as follows: First, single-view depth prediction is performed: the "high-resolution image + mask" produced by pipeline A is fed into the "single-view depth prediction" deep learning model to predict a dense depth map. Second, predicted depth reprocessing is performed: using this depth map, the predicted depth is reprocessed, and finally, a high-precision core relief mesh is output: after depth processing, a high-precision core relief mesh with rich geometric details and accurate shape is output.

[0040] The fusion output layer performs a Boolean union on the outputs of the first preprocessing pipeline, the second preprocessing pipeline, the third preprocessing pipeline, and the core production layer to output a 3D relief model, as follows: 1. Data Collection: The “CSG Fusion Engine” module in steps S2 and S3 is started and simultaneously receives data from the following three sources: the core relief mesh from the core generation layer, the 3D text mesh from the data preprocessing layer - second preprocessing pipeline, and the 3D frame mesh from the data preprocessing layer - third preprocessing pipeline.

[0041] 2. Boolean operation: The fusion engine performs two consecutive "Boolean union" operations: First operation: core relief mesh + 3D frame mesh, second operation: result of the first operation + 3D text mesh.

[0042] 3. Final output: The system generates a single, watertight 3D relief model that integrates embossing, frames, and text, and presents it to the user or makes it available for download.

[0043] The foregoing has shown and described the basic principles and main features of the present invention, as well as its advantages. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the present invention. Various changes and modifications can be made to the present invention without departing from its spirit and scope. All such changes and modifications fall within the scope of the present invention as claimed, which is defined by the appended claims and their equivalents.

Claims

1. A method for automatically generating 3D relief photo frames with customizable relief shapes and text, characterized in that: Includes the following steps: S1. Deep learning-based image segmentation: Super-resolution enhancement and instance segmentation are performed on a single image uploaded by the user to generate a high-resolution image and a pixel-level mask to accurately define the core region of the relief. S2. Boolean operation fusion of relief and frame: Load or programmatically generate 3D frame mesh, and fuse the relief main mesh with the frame mesh through Boolean union operation; S3. Generation and final fusion of text mesh: Based on the user-input custom text, a 3D text mesh is generated using a font library and extrusion algorithm. The text mesh is then located to the specified area of ​​the frame using a rigid body transformation matrix. A Boolean union operation is then performed again to fuse the text mesh with the relief-frame combined mesh, outputting a complete 3D model that integrates relief, frame, and text.

2. The method for automatically generating a 3D relief photo frame with custom relief shapes and text according to claim 1, characterized in that: The image segmentation in step S1 uses Mask R-CNN or U-Net models to achieve pixel-level separation between the core subject and the background, and marks text regions and key image feature points.

3. The method for automatically generating a 3D relief photo frame with custom relief shapes and text according to claim 1, characterized in that: The super-resolution enhancement in step S1 uses an ESRGAN or RCAN deep learning model, with a focus on enhancing high-frequency textures.

4. The method for automatically generating a 3D relief photo frame with custom relief shapes and text according to claim 1, characterized in that: The Boolean union operation in step S2 is implemented through a 3D modeling software library to ensure the watertightness of the merged mesh.

5. The method for automatically generating a 3D embossed photo frame with custom embossed shapes and text according to claim 1, characterized in that: In step S3, the text grid positioning is controlled by a rigid body transformation matrix to determine its spatial position, orientation, and scaling. The positioning area includes the lower half of the frame or other user-specified areas.

6. The method for automatically generating a 3D relief photo frame with custom relief shapes and text according to claim 1, characterized in that: An automatic generation module for 3D embossed photo frames with custom embossed shapes and text is provided. The automatic generation module includes an input layer, a data preprocessing layer, a core production layer, and a fusion output layer.

7. The method for automatically generating a 3D relief photo frame with custom relief shapes and text according to claim 6, characterized in that: The input layer includes uploading images, entering custom text, and selecting templates.

8. The method for automatically generating a 3D relief photo frame with custom relief shapes and text according to claim 6, characterized in that: The data preprocessing layer includes a first preprocessing pipeline, a second preprocessing pipeline, and a third preprocessing pipeline; The first preprocessing pipeline is used for image processing to generate high-resolution images and binarized masks as input to the core production layer; The second preprocessing pipeline is used to process custom text, producing a 3D text mesh that is input into the fusion output layer; The third preprocessing pipeline is used for template processing, producing a 3D frame mesh which is then input into the fusion output layer.

9. The method for automatically generating a 3D relief photo frame with custom relief shapes and text according to claim 8, characterized in that: The core production layer is started immediately after the first preprocessing pipeline input, and through single-view depth prediction and predicted depth reprocessing, it finally outputs a high-precision core relief mesh.

10. The method for automatically generating a 3D relief photo frame with custom relief shapes and text according to claim 9, characterized in that: The fusion output layer performs a Boolean union on the contents output from the first preprocessing pipeline, the second preprocessing pipeline, the third preprocessing pipeline, and the core production layer to output a 3D relief model.