Hierarchical Constraint Modeling Method and System for Scene Assets

CN122550831APending Publication Date: 2026-08-11JIHUA LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]本申请的主要目的在于提供一种场景资产的分层约束建模方法及系统,旨在解决现有技术中在对象高精度、环境低精度的前提下,难以实现对象模型与环境模型之间空间关系的一致性的技术问题

Benefits of technology

[0018]本申请提供了一种场景资产的分层约束建模方法,基于目标对象的原始场景环绕图像与原始局部图像,构建满足预设精度要求的对象层数据与环境层数据;基于对象层数据,构建对象模型,并基于所述环境层数据,构建环境模型;将对象模型和环境模型映射至同一参考坐标系中,基于对象模型与所述环境模型在参考坐标系中的空间关系约束,求解对象模型相对于环境模型的目标位姿变换矩阵;基于目标位姿变换矩阵,将对象模型嵌入环境模型,生成缺陷仿真场景下的双层场景资产。本申请将对象层数据和环境层数据从同一真实采集场景中分离出来,构建高精度的对象模型和低精度的环境模型,通过尺度一致性约束、姿态一致性约束、接触一致性约束、遮挡一致性约束和穿插惩罚约束求解对象模型嵌入环境模型的最优位姿变换,能够保证对象模型与环境模型之间的空间关系一致,提升缺陷仿真的真实性和训练数据质量,输出结果能够直接服务于缺陷注入、仿真渲染和训练数据构建,解决了在对象高精度、环境低精度的前提下,难以实现对象模型与环境模型之间空间关系的一致性的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550831A_ABST
    Figure CN122550831A_ABST
Patent Text Reader

Abstract

This application discloses a hierarchical constraint modeling method and system for scene assets, relating to the field of 3D reconstruction technology. The method includes: constructing object layer data and environment layer data that meet preset accuracy requirements based on the original scene surround image and original local image of the target object; constructing an object model based on the object layer data, and constructing an environment model based on the environment layer data; mapping the object model and environment model to the same reference coordinate system, and solving the target pose transformation matrix of the object model relative to the environment model based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system; and embedding the object model into the environment model based on the target pose transformation matrix to generate a two-layer scene asset under a defect simulation scene. Through the above method, the spatial relationship between the object model and the environment model is ensured to be consistent, and the output results can directly serve defect injection, simulation rendering, and training data construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of 3D reconstruction technology, and in particular to a hierarchical constraint modeling method and system for scene assets. Background Technology

[0002] In industrial inspection, defect generation, simulation rendering, and synthetic training data construction scenarios, it is typically necessary to build a high-precision 3D model of the target object to facilitate subsequent defect injection, local geometric editing, texture modification, and rendering on the object's surface. Simultaneously, the target object is usually situated in a real-world environment, such as on a workbench, tooling fixture, conveyor line, ground, or other supporting structure. For subsequent simulation and data generation, the environment does not need to achieve the same high-precision detail as the object, but it must provide scene information such as the object's location, background, support relationships, occlusion relationships, and shooting perspective. However, most existing solutions process the scene and object separately. While high-precision modeling of only the target object yields relatively detailed geometric and texture information, it cannot reconstruct the object's spatial position, contact relationships, and occlusion relationships in the real environment. Conversely, uniformly reconstructing the entire scene incurs unnecessary high-precision environment modeling overhead and limits subsequent defect injection and local editing of the object. Even if both the object model and the environment model are obtained simultaneously, problems such as inconsistent scale, inconsistent pose, offset placement, floating or intersecting contact surfaces, and incorrect occlusion relationships may still occur when there is a lack of a unified coordinate reference and embedding constraints. These problems will directly affect the realism of the defect simulation and the quality of the training data. The final generated image will have a greater difference from the actual detection scene, which will affect defect recognition.

[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a hierarchical constraint modeling method and system for scene assets, which aims to solve the technical problem in the prior art that it is difficult to achieve consistency in the spatial relationship between the object model and the environment model under the premise of high object accuracy and low environment accuracy.

[0005] To achieve the above objectives, this application provides a hierarchical constraint modeling method for scene assets, the method comprising: Based on the original scene surround image and original local image of the target object, construct object layer data and environment layer data that meet the preset accuracy requirements; Based on the object layer data, an object model is constructed, and based on the environment layer data, an environment model is constructed. The object model and the environment model are mapped to the same reference coordinate system. Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, the target pose transformation matrix of the object model relative to the environment model is solved. Based on the target pose transformation matrix, the object model is embedded into the environment model to generate a two-layer scene asset in the defect simulation scenario.

[0006] In one embodiment, the step of mapping the object model and the environment model to the same reference coordinate system, and solving the pose transformation matrix of the object model relative to the environment model based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, includes: The object model and the environment model are mapped to the same reference coordinate system to determine the initial pose of the object model; Obtain the initial pose transformation matrix of the object model relative to the environment model, and adjust the initial pose of the object model based on the initial pose transformation matrix to obtain the transformed pose of the object model. Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, the comprehensive constraint error corresponding to the transformed pose of the object model is calculated. Based on the comprehensive constraint error, the initial pose transformation matrix is ​​iteratively optimized to determine the target pose transformation matrix under the spatial relationship constraint.

[0007] In one embodiment, the step of calculating the comprehensive constraint error corresponding to the transformed pose of the object model based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system includes: Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, the target quantization errors corresponding to the transformed pose of the object model are determined to be scale error, posture error, contact error, occlusion error, and interpenetration error. The spatial relationship constraints include scale consistency constraints, posture consistency constraints, contact consistency constraints, occlusion consistency constraints, and interpenetration penalty constraints. Based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error, the comprehensive constraint error is determined.

[0008] In one embodiment, the step of determining the comprehensive constraint error based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error further includes: Obtain the position coordinates of the feature points of the calibration board in the original scene surround image and the original local image in the reference coordinate system, and obtain the position coordinates of the feature points of the calibration board in the reconstructed coordinate system; The scaling error is determined based on the position coordinates of the feature points of the calibration plate in the reference coordinate system and the position coordinates of the feature points of the calibration plate in the reconstructed coordinate system.

[0009] In one embodiment, the step of determining the comprehensive constraint error based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error further includes: Based on the bottom surface region of the object in the object model, determine the object's bottom surface normal vector; Perform plane fitting on the vertices of the supporting surfaces in the environmental model to determine the normal vectors of the environmental supporting surfaces; Based on the normal vector of the object's bottom surface and the normal vector of the environmental support surface, determine the normal alignment error; Based on the object model, the principal axis direction of the target object within the environmental support surface is extracted; A reference direction is determined based on the coded orientation of the calibration board in the original scene surround image and the original local image; Based on the spindle direction and the reference direction, determine the spindle direction alignment error; The attitude error is determined based on the normal alignment error and the principal axis alignment error.

[0010] In one embodiment, the step of determining the comprehensive constraint error based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error further includes: Based on the bottom candidate contact points in the object model and the environmental support surface in the environment model, calculate the signed distance between the bottom candidate contact points and the environmental support surface; The contact error is determined based on the signed distance between the bottom candidate contact point and the environmental support surface, as well as the contact tolerance.

[0011] In one embodiment, the step of determining the comprehensive constraint error based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error further includes: The object model is projected onto the camera view corresponding to the original scene surround image to generate the object projection outline and projection depth map; The object mask corresponding to the original scene surrounding image is used as the real observation contour; Based on the object's projected contour and the actual observed contour, the contour matching error is determined; Based on the projected depth map, the depth order error is determined; The occlusion error is determined based on the contour matching error and the depth order error.

[0012] In one embodiment, the step of constructing object layer data and environment layer data that meet preset accuracy requirements based on the original scene surround image and the original local image of the target object includes: Distortion correction is performed on the original scene surround image and the original local image of the target object to obtain the corrected scene image and the corrected local image; Generate the object mask of the corrected scene image and the object mask of the corrected local image; Based on the object mask of the corrected scene image, the environment region is extracted from the corrected scene image to obtain environment layer data that meets the preset accuracy requirements; Based on the object mask of the corrected local image, the object region is extracted from the corrected local image to obtain object layer data that meets the preset accuracy requirements.

[0013] In one embodiment, the step of generating the object mask of the corrected scene image and the object mask of the corrected local image includes: A seed frame image is determined from the corrected scene image and the corrected local image, and a prompt message is generated based on the object foreground and background points of the seed frame image; The seed frame image and the prompt information of the seed frame image are input into the object segmentation model to obtain the initial object mask of the seed frame image; The initial object mask of the seed frame image is modified to obtain the object mask of the seed frame image; Based on the object mask of the seed frame image and the object segmentation model, the object mask of the corrected scene image and the object mask of the corrected local image are determined frame by frame.

[0014] Furthermore, to achieve the above objectives, this application also proposes a hierarchical constraint modeling system for scene assets, which includes: The dual-layer data construction module is used to construct object-layer data and environment-layer data that meet preset accuracy requirements based on the original scene surround image and the original local image of the target object. A two-layer model building module is used to build an object model based on the object layer data and an environment model based on the environment layer data. The constraint embedding module is used to map the object model and the environment model to the same reference coordinate system, and solve the target pose transformation matrix of the object model relative to the environment model based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system. The scene asset output module is used to embed the object model into the environment model based on the target pose transformation matrix to generate a two-layer scene asset in the defect simulation scene.

[0015] In addition, to achieve the above objectives, this application also proposes a hierarchical constraint modeling device for scene assets. The hierarchical constraint modeling device for scene assets includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the hierarchical constraint modeling method for scene assets as described above.

[0016] In addition, to achieve the above objectives, the present invention also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the hierarchical constraint modeling method for scene assets as described above.

[0017] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the hierarchical constraint modeling method for scene assets as described above.

[0018] This application provides a hierarchical constraint modeling method for scene assets. Based on the original scene surround image and original local image of the target object, object layer data and environment layer data that meet preset accuracy requirements are constructed. Based on the object layer data, an object model is constructed, and based on the environment layer data, an environment model is constructed. The object model and the environment model are mapped to the same reference coordinate system. Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, the target pose transformation matrix of the object model relative to the environment model is solved. Based on the target pose transformation matrix, the object model is embedded into the environment model to generate a two-layer scene asset under the defect simulation scene. This application separates object-layer data and environment-layer data from the same real-world acquisition scene, constructing a high-precision object model and a low-precision environment model. By using scale consistency constraints, pose consistency constraints, contact consistency constraints, occlusion consistency constraints, and interpenetration penalty constraints, the optimal pose transformation of the object model embedded in the environment model is solved. This ensures the consistency of the spatial relationship between the object model and the environment model, improves the realism of defect simulation and the quality of training data, and the output results can directly serve defect injection, simulation rendering, and training data construction. It solves the technical problem of difficulty in achieving consistency of the spatial relationship between the object model and the environment model under the premise of high object precision and low environment precision. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating an embodiment of the hierarchical constraint modeling method for scenario assets in this application. Figure 2 This is a flowchart illustrating Embodiment 2 of the hierarchical constraint modeling method for scenario assets in this application. Figure 3 A simplified flowchart illustrating the hierarchical constraint modeling method for scene assets provided in Embodiment 2 of this application; Figure 4 This is a schematic diagram of the module structure of the hierarchical constraint modeling system for scene assets in an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the hierarchical constraint modeling method for scene assets in this application embodiment.

[0022] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0023] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0024] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0025] The main solution of this application embodiment is as follows: Based on the original scene surround image and original local image of the target object, construct object layer data and environment layer data that meet the preset accuracy requirements; based on the object layer data, construct an object model, and based on the environment layer data, construct an environment model; map the object model and the environment model to the same reference coordinate system, and based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, solve the target pose transformation matrix of the object model relative to the environment model; based on the target pose transformation matrix, embed the object model into the environment model to generate a two-layer scene asset under the defect simulation scene.

[0026] This application provides a solution that separates object-layer data and environment-layer data from the same real-world acquisition scene, constructs a high-precision object model and a low-precision environment model, and solves the optimal pose transformation of the object model embedded in the environment model through scale consistency constraints, pose consistency constraints, contact consistency constraints, occlusion consistency constraints, and interpenetration penalty constraints. This ensures the consistency of the spatial relationship between the object model and the environment model, improves the realism of defect simulation and the quality of training data, and the output results can directly serve defect injection, simulation rendering, and training data construction.

[0027] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, a hierarchical constraint modeling device for scene assets, etc. This embodiment does not specifically limit it. The following uses a hierarchical constraint modeling device for scene assets as an example to describe this embodiment and the following embodiments.

[0028] This application provides a hierarchical constraint modeling method for scene assets, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the hierarchical constraint modeling method for scenario assets in this application.

[0029] In this embodiment, the hierarchical constraint modeling method for scene assets includes steps S10~S40: Step S10: Based on the original scene surround image and the original local image of the target object, construct object layer data and environment layer data that meet the preset accuracy requirements; It should be noted that the target object is the object whose 3D model needs to be built for defect injection on its surface. Since the target object is typically in a real-world environment—for example, placed on a workbench, in a tooling fixture, on the ground, or on other supporting structures—if the target object cannot be injected and rendered within a realistic environmental context for defect simulation scenarios, the difference between the generated image and the actual detection scene will increase, thus affecting the training effect of the defect recognition and detection model. Therefore, it is necessary to uniformly construct a 3D model of the target object (object model) and a 3D model of the environment in which the target object resides (environment model), ensuring consistency in the spatial relationships between the object model and the environment model.

[0030] Additionally, it should be noted that the original scene surround image is panoramic image data containing the target object and its surrounding environment, directly acquired from the real environment. The original local image is high-precision local image data of the target object, directly acquired from the real environment. The object layer data is the image data used to construct the object model, determined based on the original scene surround image and the original local image; the environment layer data is the image data used to construct the environment model, determined based on the original scene surround image and the original local image. In this embodiment, the original scene surround image and the original local image are sorted according to the time frame of capture.

[0031] It is understood that in this embodiment, the preset accuracy requirement is the accuracy requirement that needs to be achieved in advance. The object layer data usually needs to achieve high accuracy, such as millimeter-level details, high resolution, and preservation of micro-texture. The environment layer data usually only needs to achieve low accuracy, such as preserving macro-geometry and not restoring all details. It can be set according to the actual situation, and there is no specific limitation on it.

[0032] In one feasible implementation, step S10 may include steps S101 to S104: Step S101: Perform distortion correction on the original scene surround image and the original local image of the target object to obtain the corrected scene image and the corrected local image; It should be noted that cameras are typically used to acquire the original scene surround image and the original local image. However, the lens design, manufacturing process, and assembly precision of the camera can easily lead to geometric nonlinear distortion, i.e., camera distortion. This distortion can cause straight lines in the image to become curved and the true coordinates of the target to deviate, affecting the accuracy of 3D reconstruction. Therefore, this embodiment performs distortion correction on the original scene surround image and the original local image. The corrected image data are the corrected scene image and the corrected local image.

[0033] It is understandable that when acquiring the original scene surround image and the original local image, a calibration board (checkerboard calibration board / encoded calibration board) will be set in the real environment. In other words, the original scene surround image and the original local image will contain the calibration board, so that the camera intrinsic parameters (focal length, principal point) and distortion coefficients can be estimated based on the calibration board in the image, and then distortion correction can be performed.

[0034] Step S102: Generate the object mask of the corrected scene image and the object mask of the corrected local image; It should be noted that an object mask is a mask that can extract the region of an image containing the target object.

[0035] In one feasible implementation, step S102 may include steps S1021 to S1024: Step S1021: Determine a seed frame image from the corrected scene image and the corrected local image, and generate prompt information based on the object foreground and background points of the seed frame image; It should be noted that seed frame images refer to representative image frames selected from the corrected scene images and corrected local images, where the outline of the target object is clear and can cover the main viewpoint changes. Multiple images can usually be selected.

[0036] Understandably, foreground points are locations within the inner region of the target object, while background points are locations outside the target object and near the boundary / easily confused areas. These can be used to define the boundary between the target object and the background. The foreground and background points on the seed frame image are used as initial cues, i.e., cue information.

[0037] Step S1022: Input the seed frame image and the prompt information of the seed frame image into the object segmentation model to obtain the initial object mask of the seed frame image; Understandably, the seed frame image and its accompanying prompts are input into the object segmentation model. Based on the prompts, the model can segment the target object region from the seed frame image, generating a corresponding mask—the initial object mask. This initial object mask requires further adjustment to obtain the final object mask.

[0038] Step S1023: Correct the initial object mask of the seed frame image to obtain the object mask of the seed frame image; It should be understood that the initial object mask of the seed frame image is obtained by performing boundary smoothing, hole filling, isolated region removal, and connected region filtering.

[0039] Step S1024: Based on the object mask of the seed frame image and the object segmentation model, determine the object mask of the corrected scene image and the object mask of the corrected local image frame by frame.

[0040] It is understandable that, based on the matching of image features between adjacent frames, the relationship between image viewpoint changes, and the consistency of object region appearance, the object mask of the seed frame image is propagated to the adjacent images before and after. During the propagation process, the object mask of the previous frame image is used as the prompt information of the next frame image. The object mask of each image is determined frame by frame using the object segmentation model, so that the object mask set of the corrected scene image and the object mask set of the corrected local image can be obtained respectively.

[0041] It should be understood that for frames with low confidence, missing boundaries, or strong object occlusion after propagation, the image of that frame is used as the image to be corrected, and the object mask of the neighboring frame is used for re-segmentation to correct the missed and mis-segmented areas.

[0042] Step S103: Based on the object mask of the corrected scene image, extract the environment region from the corrected scene image to obtain environment layer data that meets the preset accuracy requirements; It is understandable that the object mask of the scene image is used to remove the region of the target object from the scene image, so that only the background, the support surface (environment support surface), the occlusion structure and the environment contour can be retained to form a low-precision set of environment data, that is, environment layer data that meets the preset accuracy requirements.

[0043] Step S104: Based on the object mask of the corrected local image, extract the object region from the corrected local image to obtain object layer data that meets the preset accuracy requirements.

[0044] It is understandable that the object mask of the corrected local image is used to crop and extract the region of the target object in the corrected local image, remove the remaining background part, and form a high-precision data set of the object, that is, object layer data that meets the preset accuracy requirements.

[0045] Step S20: Based on the object layer data, construct an object model, and based on the environment layer data, construct an environment model; It should be noted that feature matching and relative pose estimation are performed among multiple corrected local images. The pose of the corrected local image relative to the camera is recovered using camera intrinsic parameters, i.e., the camera pose of the corrected local image. Simultaneously, feature points and texture consistency information are extracted from the object layer data.

[0046] Understandably, after obtaining the camera pose of the corrected local image, dense reconstruction is performed based on multi-view geometric consistency. Dense point clouds or object surface sampling points are generated only within the object mask constraints, preventing background points from entering the object model. Subsequently, the object dense point cloud / object surface sampling points, object mask, multi-view color observations, and camera pose are used together as optimization data. An object surface representation strategy based on planar Gaussian sputtering reconstruction is adopted, representing the object surface as a set of Gaussian elements constrained by planar geometry. The object surface is then optimized through multi-view color consistency constraints, contour consistency constraints, and geometric consistency constraints. Finally, the object mesh model is extracted based on the optimized object surface, and texture unwrapping and texture baking are performed based on the visibility relationships of each corrected local image to generate an object texture map, resulting in a high-precision object mesh model, i.e., the object model.

[0047] Additionally, it should be noted that stable features on the background, support surface, occlusion structure, and environmental contour are extracted from the environmental layer data. Multi-view matching and camera pose estimation are performed between multiple corrected scene images. The pose of the corrected scene image relative to the camera (i.e., the camera pose of the corrected scene image) and the overall sparse geometric structure of the environment are recovered through reprojection error constraints.

[0048] It should be understood that after obtaining the camera pose of the corrected scene image, a continuous scene representation of the environment is reconstructed based on multi-view observations. The object mask region is treated as an invalid region that does not participate in the environment surface reconstruction to avoid the target object geometry being incorporated into the environment layer. Then, a low-precision environment surface model is extracted from the continuous scene representation, and geometric recognition and simplification are performed on the environment support surface, occlusion structure, background contour, and environment shell. The environment support surface is fitted from a planar region in the environment point cloud that satisfies the plane consistency and area conditions and is located near the bottom projection range of the target object. The occlusion structure is determined from the environment surface region located in front of the target object's line of sight and that occludes the corresponding object mask in the corrected scene image / original scene surround image. Patch merging, hole repair, normalization, and mesh simplification are performed on the environment surface model to obtain a low-precision environment mesh model, i.e., the environment model.

[0049] In this embodiment, the object model adopts a high-precision modeling approach, pursuing object surface details, texture details, and subsequent editability, while the environment model adopts a low-precision environment modeling approach, only restoring the background, support, and occlusion relationships related to object embedding and simulation, thereby controlling the overall modeling cost.

[0050] Step S30: Map the object model and the environment model to the same reference coordinate system. Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, solve the target pose transformation matrix of the object model relative to the environment model. It is understandable that the pose relationship of the environment model relative to the unified reference coordinate system is determined by the calibration plate in the calibrated scene image / original scene surround image, and the pose relationship of the object model relative to the unified reference coordinate system is determined by the calibration plate in the calibrated object image / original local image, so that the environment model and the object model can be constrained and embedded under the same reference coordinate system.

[0051] In one feasible implementation, before step S30, the following steps may be included: obtaining the geometric center and physical dimensions of the calibration board in the original scene surround image and the original local image; determining the origin of the coordinate system based on the geometric center of the calibration board; determining the reference scale of the coordinate system based on the physical dimensions of the calibration board; and constructing a reference coordinate system based on the origin of the coordinate system and the reference scale of the coordinate system.

[0052] It is understandable that a unified reference coordinate system is established by taking the physical center (geometric center) of the calibration plate as the origin of the coordinate system, the physical dimensions of the calibration plate as the reference scale, and the coding direction of the calibration plate itself as the coordinate axis direction.

[0053] It should be understood that by mapping high-precision object models and low-precision environment models to the same reference coordinate system, both object models and environment models can be mapped to the same real-world coordinate system, which can avoid the problem of inconsistent coordinate references when objects and environments come from different acquisition processes.

[0054] It should be noted that spatial relationship constraints include scale consistency constraints, posture consistency constraints, contact consistency constraints, occlusion consistency constraints, and interpenetration penalty constraints.

[0055] After the object model and the environment model are mapped to the same reference coordinate system, a relationship can be established under the unified reference coordinate system. At this time, the object model is not directly embedded into the environment model. Instead, the optimal pose transformation between the object model and the environment model is solved based on the spatial relationship constraints. The corresponding matrix is ​​the target pose transformation matrix.

[0056] Step S40: Based on the target pose transformation matrix, embed the object model into the environment model to generate a two-layer scene asset under the defect simulation scenario.

[0057] It's important to note that the two-layer scene asset consists of an object layer, an environment layer, and a relationship layer between the object and environment layers. The object layer contains high-precision object models and textures, preserving the coordinate correspondence between object meshes and textures for subsequent defect injection and local geometry editing. The environment layer contains a low-precision environment model, as well as geometry such as support surfaces, occlusion structures, and background contours. The relationship layer contains the target pose transformation matrix of the object model relative to the environment model, the definition of a unified reference coordinate system, unit scale information, contact relationship information, and occlusion relationship information.

[0058] Understandably, by embedding the object model into the environment model according to the target pose transformation matrix obtained under spatial relationship constraints, the object and the environment are kept consistent in scale, pose, contact and occlusion, and finally a two-layer scene asset is generated, which can be directly applied to defect injection, rendering and photography and data training in defect simulation scenarios.

[0059] It should be understood that the object layer, environment layer, and relationship layer are encapsulated into a unified scene file and converted into a scene asset format suitable for loading by the simulation platform to obtain a two-layer scene asset. In this two-layer scene asset, the object layer remains independently editable, while the environment layer continuously provides background, support, and occlusion constraints, and the two are kept consistent through the relationship layer data.

[0060] This embodiment provides a hierarchical constraint modeling method for scene assets. Based on the original scene surround image and original local image of the target object, object layer data and environment layer data that meet preset accuracy requirements are constructed. Based on the object layer data, an object model is constructed, and based on the environment layer data, an environment model is constructed. The object model and environment model are mapped to the same reference coordinate system. Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, the target pose transformation matrix of the object model relative to the environment model is solved. Based on the target pose transformation matrix, the object model is embedded into the environment model to generate a two-layer scene asset under the defect simulation scene. This embodiment separates the object layer data and environment layer data from the same real-world acquisition scene, constructs a high-precision object model and a low-precision environment model, and solves the optimal pose transformation of the object model embedded into the environment model through scale consistency constraints, pose consistency constraints, contact consistency constraints, occlusion consistency constraints, and interpenetration penalty constraints. This ensures the consistency of the spatial relationship between the object model and the environment model, improves the realism of defect simulation and the quality of training data, and the output results can directly serve defect injection, simulation rendering, and training data construction.

[0061] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S30 may include steps S301 to S304: Step S301: Map the object model and the environment model to the same reference coordinate system to determine the initial pose of the object model; It should be noted that the initial pose is the preliminarily determined pose, and further optimization and adjustment are still needed.

[0062] Understandably, the corner points / encoded points of the calibration board are detected in the original scene surround image and the original local image, and the pose of the calibration board relative to the camera is solved in each frame. Based on the pose of the calibration board relative to the camera, the pose relationship between the object model and the reference coordinate system and the pose relationship between the environment model and the reference coordinate system are determined. This allows the high-precision object model and the low-precision environment model to be mapped to the same reference coordinate system, thus obtaining the initial pose of the object model and the initial pose of the environment model.

[0063] Step S302: Obtain the initial pose transformation matrix of the object model relative to the environment model; adjust the initial pose of the object model based on the initial pose transformation matrix to obtain the transformed pose of the object model. It should be noted that the initial pose transformation matrix is ​​the same as the initialized pose transformation matrix. The position and orientation of the object model in the environment model are updated according to the initial pose transformation matrix.

[0064] Step S303: Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, calculate the comprehensive constraint error corresponding to the transformed pose of the object model; In one feasible implementation, step S303 may include steps S3031 to S3032: Step S3031: Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, determine the target quantization error corresponding to the transformed pose of the object model as scale error, posture error, contact error, occlusion error and interpenetration error. It should be noted that spatial relationship constraints include scale consistency constraints, attitude consistency constraints, contact consistency constraints, occlusion consistency constraints, and interpenetration penalty constraints. Based on the spatial relationship constraints, the corresponding error is quantized, i.e., the target quantization error.

[0065] Scale Consistency Constraint: After the object model and environment model are reconstructed, the similarity transformation from the reconstructed coordinate system to the unified reference coordinate system is solved based on the calibration board corner points or coded points detected in the image. This similarity transformation includes at least a rotation factor, a translation factor, and a scale factor. The optimal scale factor is solved by minimizing the scale error of feature points in the calibration board between the reconstructed coordinate system and the unified reference coordinate system, thus uniformly transforming the object model and environment model to the same metric scale and avoiding size distortion after the object is embedded into the environment.

[0066] Pose consistency constraints: The principal orientation of the object is not directly input manually from the outside, but is calculated jointly by a high-precision object model and its image observation. First, the bottom surface region of the object is determined based on the set of candidate contact points at the bottom. Then, the principal axis orientation of the target object within the support surface is determined based on the object's mesh bounding box, principal component orientation, the major axis of the object's mask contour, and the object's surface texture orientation. When the principal component orientation is unstable due to rotational symmetry of the target object, the principal axis orientation is corrected by combining the identifiable texture region of the object's surface and the appearance orientation in the image observation. The environment support surface normal vector is obtained by plane fitting of the support surface vertex set in the low-precision environment model. The reference orientation is the horizontal reference orientation in the unified reference coordinate system, determined by the coded orientation of the calibration plate. The object's bottom surface normal vector is constrained to be consistent with the environment support surface normal vector, and the principal axis orientation of the target object within the support surface is constrained to be consistent with the reference orientation, thereby avoiding discrepancies between the object's rotational posture after embedding and the actual acquisition scene.

[0067] Contact Consistency Constraint: In real-world scenarios, target objects are typically placed on workbenches, pallets, conveyor belts, or other supporting surfaces, with a stable contact relationship between their bottom areas and the supporting surface. To ensure that the spatial state of the target object after embedding into the environment is consistent with the actual scene, the system transforms this contact relationship into a contact consistency constraint. The contact consistency constraint requires that after embedding into the environment, the bottom candidate contact area of ​​the target object should be located above the environmental supporting surface and maintain contact with it, without significant floating, sinking, or local penetration. A set of bottom candidate contact points is extracted from the object model, and the signed distances from these contact points to the environmental supporting surface are calculated. A contact tolerance is pre-defined, allowing for a small distance deviation between the bottom candidate contact points and the environmental supporting surface. When the distance is greater than the contact tolerance, the target object is considered to be floating; when the distance is less than the negative contact tolerance, the target object is considered to be sinking. During iterative optimization, the distance between the bottom candidate contact points and the supporting surface is made close to zero, and the bottom candidate contact projection falls within the effective area of ​​the environmental supporting surface, thus ensuring that the contact relationship between the target object and the environment is consistent with the real scene.

[0068] Occlusion Consistency Constraint: Utilizing the camera intrinsics and camera pose corresponding to the original scene surround image, a high-precision object model is projected onto the camera viewpoint corresponding to the original scene surround image, generating the object's projected contour and projection depth map. Simultaneously, the object mask corresponding to the original scene surround image is used as the true observed contour. The contour overlap, contour difference region, and boundary distance between the object's projected contour and the true observed contour are calculated, and the depth relationship of the environment model under the same camera viewpoint is combined to determine the front-to-back occlusion order of the object and the environment surface. If the deviation between the object's projected contour and the true observed contour is too large, or if the environment occlusion structure should be in front of the target object but the projection depth order is inconsistent, the occlusion consistency error is increased, and the target object's pose relative to the environment is adjusted in subsequent iterations.

[0069] Penetration Penalty Constraint: Object surface points (3D sampling points extracted from the object model and discretely distributed on the target object's surface) are input into the environment's signed distance field to calculate the signed distance between the object surface points and the environment surface. When an object surface point enters the environment surface, the signed distance is negative, indicating penetration (penetration), and the penetration depth is calculated accordingly. When an object surface point is outside the environment surface, the signed distance is positive, indicating no penetration. For detected penetration / penetration areas, the penetration depth, number of penetration points, and area of ​​the penetration region are statistically analyzed and constructed as a penetration penalty term. This penalty term increases with penetration depth, number of penetration points, and area of ​​the penetration region. When an object surface point is outside the environment and only in contact with the supporting surface, the penetration penalty term is zero.

[0070] In this embodiment, the target quantization error corresponding to the scale consistency constraint is the scale error, the target quantization error corresponding to the attitude consistency constraint is the attitude error, the target quantization error corresponding to the contact consistency constraint is the contact error, the target quantization error corresponding to the occlusion consistency constraint is the occlusion error, and the target quantization error corresponding to the interleaving penalty constraint is the interleaving error.

[0071] Step S3032: Based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error, determine the comprehensive constraint error.

[0072] It should be noted that the comprehensive constraint error is the error obtained by combining the scale error, attitude error, contact error, occlusion error, and interleaving error. Before determining the comprehensive constraint error, it is necessary to calculate the scale error, attitude error, contact error, occlusion error, and interleaving error.

[0073] In one feasible implementation, before step S3032, the step of calculating the scale error includes: obtaining the position coordinates of the feature points of the calibration board in the original scene surround image and the original local image in the reference coordinate system, and obtaining the position coordinates of the feature points of the calibration board in the reconstructed coordinate system; and determining the scale error based on the position coordinates of the feature points of the calibration board in the reference coordinate system and the position coordinates of the feature points of the calibration board in the reconstructed coordinate system.

[0074] It should be noted that the reconstructed coordinate system refers to the coordinate system used to construct the object model and environment model. Since the reference coordinate system is constructed with the geometric center of the calibration plate as its origin, the position coordinates of the feature points of the calibration plate in the reference coordinate system can be directly determined.

[0075] It is understandable that the error between the position coordinates of the feature points of the calibration plate in the reference coordinate system and the position coordinates of the feature points of the calibration plate in the reconstructed coordinate system is taken as the scale error.

[0076] In one feasible implementation, prior to step S3032, the step of calculating the attitude error includes: determining the object bottom surface normal vector based on the object bottom surface region in the object model; performing plane fitting on the support surface vertices in the environment model to determine the environment support surface normal vector; determining the normal alignment error based on the object bottom surface normal vector and the environment support surface normal vector; extracting the principal axis direction of the target object within the environment support surface based on the object model; determining the reference direction based on the encoding orientation of the calibration board in the original scene surround image and the original local image; determining the principal axis direction alignment error based on the principal axis direction and the reference direction; and determining the attitude error based on the normal alignment error and the principal axis direction alignment error.

[0077] Understandably, the process begins by determining the object's bottom surface region (i.e., the bottom surface of the target object) based on all candidate bottom contact points in the object model (potentially points on the bottom surface of the target object that contact the supporting surface). Then, a plane fitting is performed on the points within this bottom surface region, and the normal vector of the resulting fitted plane is the object's bottom surface normal vector. Next, a plane fitting is performed using all the vertices of the supporting surfaces in the environment model, and the normal vector of the resulting fitted plane is the environment's supporting surface normal vector. The angular deviation between the object's bottom surface normal vector and the environment's supporting surface normal vector is the normal alignment error. If the object's bottom surface normal vector and the environment's supporting surface normal vector are in the same direction, the normal alignment error is zero.

[0078] It should be noted that the principal axis direction is jointly determined by the major axis direction of the object mesh bounding box, the principal component direction, the major axis direction of the object mask contour, and the object surface texture direction. The object mesh bounding box refers to the smallest bounding box that describes the overall spatial extent of the object model, aligned with the coordinate axes of a unified reference coordinate system, with each face parallel to the reference coordinate system plane. The principal component direction refers to the feature direction corresponding to the largest eigenvalue determined after principal component analysis of the object's bottom vertices / bottom candidate contact points. The major axis direction of the object mask contour refers to the major axis direction extracted by contour fitting of the object mask in the original local image (i.e., the object mask in the corrected object image). The object surface texture direction refers to the principal texture direction extracted after performing histogram of oriented gradients (HOG) statistics based on the object texture map.

[0079] Understandably, for non-rotationally symmetric objects, extracting the consistent direction of multi-source feature results (the long axis direction of the object mesh bounding box, the principal component direction, the long axis direction of the object mask contour, and the object surface texture direction) as the final principal axis direction improves computational robustness. For rotationally symmetric objects, the principal component directions obtained through principal component analysis are unstable. In this case, the directions of the identifiable texture regions on the object surface and the appearance orientation of the original local image can be used for correction.

[0080] It should be understood that the reference direction is determined according to the coded orientation of the calibration board in the original scene surround image and the original local image. The angular error between the principal axis direction and the reference direction is the principal axis direction alignment error. The final pose error is obtained by weighted summing of the normal alignment error and the principal axis direction alignment error.

[0081] In one feasible implementation, before step S3032, the step of calculating the contact error includes: calculating the signed distance between the bottom candidate contact point and the environmental support surface based on the bottom candidate contact point in the object model and the environmental support surface in the environment model; and determining the contact error based on the signed distance between the bottom candidate contact point and the environmental support surface and the contact tolerance.

[0082] It should be noted that Signed Distance (SDF) is used to quantify the relative positional relationship between a spatial point and a geometric surface. If the bottom candidate contact point is located outside the environment surface (not occluded or penetrated), the signed distance is positive, and its value is equal to the shortest Euclidean distance from the bottom candidate contact point to the environment surface. If the bottom candidate contact point is located inside the environment surface (enclosed by the environment geometry and penetrated), the signed distance is negative, and its absolute value is equal to the shortest penetration distance from the bottom candidate contact point to the environment surface.

[0083] Understandably, if the signed distance between the bottom candidate contact point and the environmental support surface is greater than the contact tolerance, it is determined that the target object is suspended; if the signed distance between the bottom candidate contact point and the environmental support surface is less than the negative value of the contact tolerance, it is determined that the target object is sunken; if the signed distance between the bottom candidate contact point and the environmental support surface is less than or equal to the contact tolerance and greater than or equal to the negative value of the contact tolerance, it is determined that the target object and the environmental support surface have a normal contact relationship.

[0084] It should be understood that the deviation of each bottom candidate contact point is calculated based on the signed distance between the bottom candidate contact point and the environmental support surface, as well as the contact tolerance. If the signed distance between the bottom candidate contact point and the environmental support surface is greater than the contact tolerance, the difference between the signed distance and the contact tolerance is taken as the deviation of that bottom candidate contact point. If the signed distance between the bottom candidate contact point and the environmental support surface is less than or equal to the contact tolerance and greater than or equal to the negative value of the contact tolerance, the deviation of that bottom candidate contact point is 0. The final contact error is obtained by combining the deviations of all bottom candidate contact points.

[0085] In one feasible implementation, before step S3032, the step of calculating the occlusion error includes: projecting the object model onto the camera view corresponding to the original scene surround image to generate an object projection contour and a projection depth map; using the object mask corresponding to the original scene surround image as the true observation contour; determining the contour matching error based on the object projection contour and the true observation contour; determining the depth order error based on the projection depth map; and determining the occlusion error based on the contour matching error and the depth order error.

[0086] Understandably, by utilizing the camera intrinsics and camera pose corresponding to the original scene surround image, a high-precision object model is projected onto the camera viewpoint corresponding to each original scene surround image, generating the object's projected contour and projection depth map. Simultaneously, the object mask corresponding to the original scene surround image (i.e., the object mask of the corrected scene image) is used as the true observed contour. The error between the object's projected contour and the true observed contour is the contour matching error of the original scene surround image; the higher the overlap, the smaller the contour matching error.

[0087] It should be understood that when traversing all pixels of environmental occlusion structures in the original scene surround image, theoretically, the pixel depth of the environmental occlusion structure should be less than the depth of the corresponding position of the object model (the environment should be in front of the target object, so its depth is even smaller). However, in reality, the depth order is reversed, so this pixel is considered an erroneous pixel. The proportion of erroneous pixels in the original scene surround image is taken as the depth order error of the original scene surround image. The contour matching error and the depth order error are weighted and summed to determine the final occlusion error.

[0088] In one feasible implementation, before step S3032, the step of calculating the interleaving error includes: inputting the object surface point into the signed distance field of the environment, calculating the signed distance of the object surface point relative to the environment surface; determining the interleaving penalty of the object surface point based on the signed distance of the object surface point relative to the environment surface; and determining the interleaving error based on the interleaving penalty of the object surface point.

[0089] It is understandable that object surface points include object surface vertices and object surface sampling points. When the signed distance of an object surface point relative to the environment surface is negative, the object surface point is determined to be inside the environment surface, and the interpolation penalty of the object surface point is calculated based on the penetration depth, the number of penetration points, and the area of ​​the penetration region. When the signed distance of an object surface point relative to the environment surface is positive, the object surface point is determined to be outside the environment surface, and the interpolation penalty of the object surface point is set to zero. The final interpolation error (e.g., the average interpolation penalty of all object surface points) is obtained by combining the interpolation penalties of all object surface points.

[0090] Step S304: Based on the comprehensive constraint error, iteratively optimize the initial pose transformation matrix to determine the target pose transformation matrix under the spatial relationship constraint.

[0091] It is understandable that the comprehensive constraint error can be obtained by weighted summing of scale error, attitude error, contact error, occlusion error, and interleaving error. If the comprehensive constraint error is less than or equal to a set comprehensive error threshold, it is considered to have met the threshold requirement; otherwise, it is considered not to have met the threshold requirement. Alternatively, the scale error, attitude error, contact error, occlusion error, and interleaving error can be directly used as the comprehensive constraint error. In this case, if the scale error is less than or equal to the corresponding scale error threshold, the attitude error is less than or equal to the corresponding attitude error threshold, the contact error is less than or equal to the corresponding contact error threshold, the occlusion error is less than or equal to the corresponding occlusion error threshold, and the interleaving error is less than or equal to the corresponding interleaving error threshold, the comprehensive constraint error is considered to have met the threshold requirement; otherwise, it is considered not to have met the threshold requirement.

[0092] It should be understood that the initial pose transformation matrix is ​​adjusted according to the comprehensive constraint error, and the comprehensive constraint error is recalculated based on the adjusted initial pose transformation matrix. If the comprehensive constraint error does not reach the threshold requirement and other set convergence conditions are not currently met, iterative optimization continues until the comprehensive constraint error reaches the threshold requirement or other set convergence conditions are currently met.

[0093] This embodiment provides a hierarchical constraint modeling method for scene assets. It maps the object model and environment model to the same reference coordinate system to determine the initial pose of the object model. It obtains the initial pose transformation matrix of the object model relative to the environment model, adjusts the initial pose of the object model based on this matrix, and obtains the transformed pose of the object model. Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, it calculates the comprehensive constraint error corresponding to the transformed pose of the object model. Based on the comprehensive constraint error, iteratively optimizes the initial pose transformation matrix to determine the target pose transformation matrix under spatial relationship constraints. This embodiment separates the object layer data and environment layer data from the same real-world acquisition scene, constructing a high-precision object model and a low-precision environment model. It solves for the optimal pose transformation of the object model embedded in the environment model through scale consistency constraints, pose consistency constraints, contact consistency constraints, occlusion consistency constraints, and interleaving penalty constraints. This ensures the consistency of the spatial relationship between the object model and the environment model, improves the realism of defect simulation and the quality of training data, and the output results can directly serve defect injection, simulation rendering, and training data construction.

[0094] For example, to help understand the implementation process of the hierarchical constraint modeling method for scene assets obtained by combining this embodiment with the above-described embodiment two, please refer to... Figure 3 , Figure 3 A simplified flowchart illustrating a hierarchical constraint modeling method for scene assets is provided, specifically: Step 1: Input Acquisition and Two-Layer Data Construction. First, the raw input is organized uniformly, and then object-layer data and environment-layer data are generated from the same real-world scene input.

[0095] Step 2: Construction of high-precision object model and low-precision environment model. After completing the construction of the two-layer data, high-precision modeling of the object layer and low-precision modeling of the environment layer are performed separately.

[0096] Step 3: Unified Coordinate Alignment and Constraint Embedding. After establishing a relationship between the high-precision object model and the low-precision environment model under a unified reference coordinate system, the object model is not directly placed into the environment model. Instead, constraint optimization is performed on the final transformation from the object model to the environment model.

[0097] Step 4: Two-layer scene asset output. After completing unified coordinate alignment and constraint embedding, the result is output as a two-layer scene asset. The output is not a single scene mesh, but a two-layer scene asset composed of an object layer, an environment layer, and a relationship layer.

[0098] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the hierarchical constraint modeling method of scenario assets in this application. Any simple transformations based on this technical concept are within the protection scope of this application.

[0099] This application also provides a hierarchical constraint modeling system for scene assets; please refer to [reference needed]. Figure 4 The hierarchical constraint modeling system for scene assets includes: The dual-layer data construction module 10 is used to construct object layer data and environment layer data that meet preset accuracy requirements based on the original scene surround image and the original local image of the target object. The two-layer model construction module 20 is used to construct an object model based on the object layer data and an environment model based on the environment layer data. The constraint embedding module 30 is used to map the object model and the environment model to the same reference coordinate system, and solve the target pose transformation matrix of the object model relative to the environment model based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system. The scene asset output module 40 is used to embed the object model into the environment model based on the target pose transformation matrix to generate a two-layer scene asset in the defect simulation scene.

[0100] In one feasible implementation, the constraint embedding module 30 is further configured to map the object model and the environment model to the same reference coordinate system to determine the initial pose of the object model; Obtain the initial pose transformation matrix of the object model relative to the environment model, and adjust the initial pose of the object model based on the initial pose transformation matrix to obtain the transformed pose of the object model. Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, the comprehensive constraint error corresponding to the transformed pose of the object model is calculated. Based on the comprehensive constraint error, the initial pose transformation matrix is ​​iteratively optimized to determine the target pose transformation matrix under the spatial relationship constraint.

[0101] In one feasible implementation, the constraint embedding module 30 is further configured to determine the target quantization error corresponding to the transformed pose of the object model as scale error, posture error, contact error, occlusion error and interpenetration error based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system. The spatial relationship constraints include scale consistency constraints, posture consistency constraints, contact consistency constraints, occlusion consistency constraints and interpenetration penalty constraints. Based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error, the comprehensive constraint error is determined.

[0102] In one feasible implementation, the constraint embedding module 30 is further configured to obtain the position coordinates of the feature points of the calibration board in the original scene surround image and the original local image in the reference coordinate system, and to obtain the position coordinates of the feature points of the calibration board in the reconstructed coordinate system. The scaling error is determined based on the position coordinates of the feature points of the calibration plate in the reference coordinate system and the position coordinates of the feature points of the calibration plate in the reconstructed coordinate system.

[0103] In one feasible implementation, the constraint embedding module 30 is further configured to determine the object bottom surface normal vector based on the object bottom surface region in the object model; Perform plane fitting on the vertices of the supporting surfaces in the environmental model to determine the normal vectors of the environmental supporting surfaces; Based on the normal vector of the object's bottom surface and the normal vector of the environmental support surface, determine the normal alignment error; Based on the object model, the principal axis direction of the target object within the environmental support surface is extracted; A reference direction is determined based on the coded orientation of the calibration board in the original scene surround image and the original local image; Based on the spindle direction and the reference direction, determine the spindle direction alignment error; The attitude error is determined based on the normal alignment error and the principal axis alignment error.

[0104] In one feasible implementation, the constraint embedding module 30 is further configured to calculate the signed distance between the bottom candidate contact point and the environmental support surface based on the bottom candidate contact point in the object model and the environmental support surface in the environment model. The contact error is determined based on the signed distance between the bottom candidate contact point and the environmental support surface, as well as the contact tolerance.

[0105] In one feasible implementation, the constraint embedding module 30 is further configured to project the object model onto the camera view corresponding to the original scene surround image to generate an object projection outline and a projection depth map. The object mask corresponding to the original scene surrounding image is used as the real observation contour; Based on the object's projected contour and the actual observed contour, the contour matching error is determined; Based on the projected depth map, the depth order error is determined; The occlusion error is determined based on the contour matching error and the depth order error.

[0106] In one feasible implementation, the dual-layer data construction module 10 is further used to perform distortion correction on the original scene surround image and the original local image of the target object to obtain a corrected scene image and a corrected local image. Generate the object mask of the corrected scene image and the object mask of the corrected local image; Based on the object mask of the corrected scene image, the environment region is extracted from the corrected scene image to obtain environment layer data that meets the preset accuracy requirements; Based on the object mask of the corrected local image, the object region is extracted from the corrected local image to obtain object layer data that meets the preset accuracy requirements.

[0107] In one feasible implementation, the dual-layer data construction module 10 is further configured to determine a seed frame image in the corrected scene image and the corrected local image, and generate prompt information based on the object foreground point and background point of the seed frame image; The seed frame image and the prompt information of the seed frame image are input into the object segmentation model to obtain the initial object mask of the seed frame image; The initial object mask of the seed frame image is modified to obtain the object mask of the seed frame image; Based on the object mask of the seed frame image and the object segmentation model, the object mask of the corrected scene image and the object mask of the corrected local image are determined frame by frame.

[0108] The hierarchical constraint modeling system for scene assets provided in this application, employing the hierarchical constraint modeling method for scene assets in the above embodiments, can solve the technical problem of difficulty in achieving consistency in spatial relationships between object models and environment models under the premise of high object precision and low environment precision. Compared with the prior art, the beneficial effects of the hierarchical constraint modeling system for scene assets provided in this application are the same as those of the hierarchical constraint modeling method for scene assets provided in the above embodiments, and other technical features of the hierarchical constraint modeling system for scene assets are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0109] This application provides a hierarchical constraint modeling device for scene assets. The hierarchical constraint modeling device for scene assets includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the hierarchical constraint modeling method for scene assets in the above embodiment 1.

[0110] The following is for reference. Figure 5 The diagram illustrates a structural schematic of a hierarchical constraint modeling device suitable for implementing the embodiments of this application. The hierarchical constraint modeling device for scene assets in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), tablets, PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The illustrated hierarchical constraint modeling device for scene assets is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0111] like Figure 5As shown, the scene asset hierarchical constraint modeling device may include a processing system 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage system 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the scene asset hierarchical constraint modeling device. The processing system 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input systems 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output systems 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage systems 1003 including, for example, magnetic tapes, hard disks, etc.; and communication systems 1009. Communication system 1009 allows the hierarchical constraint modeling device for scene assets to communicate wirelessly or wiredly with other devices to exchange data. Although a hierarchical constraint modeling device for scene assets with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented alternatively.

[0112] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication system, or installed from storage system 1003, or installed from ROM 1002. When the computer program is executed by processing system 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0113] The hierarchical constraint modeling device for scene assets provided in this application, employing the hierarchical constraint modeling method for scene assets in the above embodiments, can solve the technical problem of difficulty in achieving consistency in spatial relationships between object models and environment models under the premise of high object precision and low environment precision. Compared with the prior art, the beneficial effects of the hierarchical constraint modeling device for scene assets provided in this application are the same as the beneficial effects of the hierarchical constraint modeling method for scene assets provided in the above embodiments, and other technical features in this hierarchical constraint modeling device for scene assets are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0114] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0115] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0116] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the hierarchical constraint modeling method for scene assets in the above embodiments.

[0117] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0118] The aforementioned computer-readable storage medium may be included in the hierarchical constraint modeling device for scene assets; or it may exist independently and not be assembled into the hierarchical constraint modeling device for scene assets.

[0119] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the scene asset hierarchical constraint modeling device, the scene asset hierarchical constraint modeling device: constructs object layer data and environment layer data that meet preset accuracy requirements based on the original scene surround image and original local image of the target object; constructs an object model based on the object layer data, and constructs an environment model based on the environment layer data; maps the object model and environment model to the same reference coordinate system, and solves the target pose transformation matrix of the object model relative to the environment model based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system; and embeds the object model into the environment model based on the target pose transformation matrix to generate a two-layer scene asset under the defect simulation scene.

[0120] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0122] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0123] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the hierarchical constraint modeling method for scene assets described above. This solves the technical problem of difficulty in achieving consistency in spatial relationships between object models and environment models when the object model has high precision but the environment model has low precision. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the hierarchical constraint modeling method for scene assets provided in the above embodiments, and will not be elaborated upon here.

[0124] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the hierarchical constraint modeling method for scene assets as described above.

[0125] The computer program product provided in this application can solve the technical problem of difficulty in achieving consistency in spatial relationships between object models and environment models when the object has high precision but the environment has low precision. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the hierarchical constraint modeling method for scene assets provided in the above embodiments, and will not be repeated here.

[0126] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A hierarchical constraint modeling method for scene assets, characterized in that, The method includes: Based on the original scene surround image and original local image of the target object, construct object layer data and environment layer data that meet the preset accuracy requirements; Based on the object layer data, an object model is constructed, and based on the environment layer data, an environment model is constructed. The object model and the environment model are mapped to the same reference coordinate system. Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, the target pose transformation matrix of the object model relative to the environment model is solved. Based on the target pose transformation matrix, the object model is embedded into the environment model to generate a two-layer scene asset in the defect simulation scenario.

2. The method as described in claim 1, characterized in that, The step of mapping the object model and the environment model to the same reference coordinate system, and solving the pose transformation matrix of the object model relative to the environment model based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, includes: The object model and the environment model are mapped to the same reference coordinate system to determine the initial pose of the object model; Obtain the initial pose transformation matrix of the object model relative to the environment model, and adjust the initial pose of the object model based on the initial pose transformation matrix to obtain the transformed pose of the object model. Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, the comprehensive constraint error corresponding to the transformed pose of the object model is calculated. Based on the comprehensive constraint error, the initial pose transformation matrix is ​​iteratively optimized to determine the target pose transformation matrix under the spatial relationship constraint.

3. The method as described in claim 2, characterized in that, The step of calculating the comprehensive constraint error corresponding to the transformed pose of the object model based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system includes: Based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system, the target quantization errors corresponding to the transformed pose of the object model are determined to be scale error, posture error, contact error, occlusion error, and interpenetration error. The spatial relationship constraints include scale consistency constraints, posture consistency constraints, contact consistency constraints, occlusion consistency constraints, and interpenetration penalty constraints. Based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error, the comprehensive constraint error is determined.

4. The method as described in claim 3, characterized in that, Before the step of determining the comprehensive constraint error corresponding to the transformed pose of the object model based on the calculated scale error, pose error, contact error, occlusion error, and interpenetration error, the method further includes: Obtain the position coordinates of the feature points of the calibration board in the original scene surround image and the original local image in the reference coordinate system, and obtain the position coordinates of the feature points of the calibration board in the reconstructed coordinate system; The scaling error is determined based on the position coordinates of the feature points of the calibration plate in the reference coordinate system and the position coordinates of the feature points of the calibration plate in the reconstructed coordinate system.

5. The method as described in claim 3, characterized in that, Before the step of determining the comprehensive constraint error based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error, the method further includes: Based on the bottom surface region of the object in the object model, determine the object's bottom surface normal vector; Perform plane fitting on the vertices of the supporting surfaces in the environmental model to determine the normal vectors of the environmental supporting surfaces; Based on the normal vector of the object's bottom surface and the normal vector of the environmental support surface, determine the normal alignment error; Based on the object model, the principal axis direction of the target object within the environmental support surface is extracted; A reference direction is determined based on the coded orientation of the calibration board in the original scene surround image and the original local image; Based on the spindle direction and the reference direction, determine the spindle direction alignment error; The attitude error is determined based on the normal alignment error and the principal axis alignment error.

6. The method as described in claim 3, characterized in that, Before the step of determining the comprehensive constraint error based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error, the method further includes: Based on the bottom candidate contact points in the object model and the environmental support surface in the environment model, calculate the signed distance between the bottom candidate contact points and the environmental support surface; The contact error is determined based on the signed distance between the bottom candidate contact point and the environmental support surface, as well as the contact tolerance.

7. The method as described in claim 3, characterized in that, Before the step of determining the comprehensive constraint error based on the calculated scale error, attitude error, contact error, occlusion error, and interpenetration error, the method further includes: The object model is projected onto the camera view corresponding to the original scene surround image to generate the object projection outline and projection depth map; The object mask corresponding to the original scene surrounding image is used as the real observation contour; Based on the object's projected contour and the actual observed contour, the contour matching error is determined; Based on the projected depth map, the depth order error is determined; The occlusion error is determined based on the contour matching error and the depth order error.

8. The method as described in claim 1, characterized in that, The steps of constructing object layer data and environment layer data that meet preset accuracy requirements based on the original scene surround image and original local image of the target object include: Distortion correction is performed on the original scene surround image and the original local image of the target object to obtain the corrected scene image and the corrected local image; Generate the object mask of the corrected scene image and the object mask of the corrected local image; Based on the object mask of the corrected scene image, the environment region is extracted from the corrected scene image to obtain environment layer data that meets the preset accuracy requirements; Based on the object mask of the corrected local image, the object region is extracted from the corrected local image to obtain object layer data that meets the preset accuracy requirements.

9. The method as described in claim 8, characterized in that, The steps of generating the object mask of the corrected scene image and the object mask of the corrected local image include: A seed frame image is determined from the corrected scene image and the corrected local image, and a prompt message is generated based on the object foreground and background points of the seed frame image; The seed frame image and the prompt information of the seed frame image are input into the object segmentation model to obtain the initial object mask of the seed frame image; The initial object mask of the seed frame image is modified to obtain the object mask of the seed frame image; Based on the object mask of the seed frame image and the object segmentation model, the object mask of the corrected scene image and the object mask of the corrected local image are determined frame by frame.

10. A hierarchical constraint modeling system for scene assets, characterized in that, The system includes: The dual-layer data construction module is used to construct object-layer data and environment-layer data that meet preset accuracy requirements based on the original scene surround image and the original local image of the target object. A two-layer model building module is used to build an object model based on the object layer data and an environment model based on the environment layer data. The constraint embedding module is used to map the object model and the environment model to the same reference coordinate system, and solve the target pose transformation matrix of the object model relative to the environment model based on the spatial relationship constraints between the object model and the environment model in the reference coordinate system. The scene asset output module is used to embed the object model into the environment model based on the target pose transformation matrix to generate a two-layer scene asset in the defect simulation scene.