Three-dimensional scene layout method and system, storage medium, equipment and program product

By constructing a pose optimization function, combining constraints and collision cost functions, the problem of handling multiple constraints in three-dimensional scene layout is solved, achieving better layout effect and user experience.

CN120147549APending Publication Date: 2025-06-13AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510320131.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing three-dimensional scene layout technology is difficult to effectively deal with multiple constraints, resulting in insufficient layout rationality and poor user experience.

Method used

By determining the constraint cost function based on the position constraint relationship of the target object set, building a pose optimization function with the collision cost function, solving the object pose pose and generating three-dimensional scene information.

Benefits of technology

It realizes the effective handling of multiple constraints in complex scenarios, improves the rationality and aesthetics of the layout, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147549A_ABST
    Figure CN120147549A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a three-dimensional scene layout method and system, a storage medium, equipment and a program product, and the method comprises the steps: determining a constraint cost function corresponding to a target object set based on a position constraint relation corresponding to the target object set in a target region; according to the collision cost function and the constraint cost function corresponding to the target object set, constructing a pose optimization function corresponding to the target object set; solving the poses of the objects in the target object set by using the pose optimization function corresponding to the target object set; and generating three-dimensional scene information of the target area based on the pose of the object in the target area. According to the embodiment of the invention, various constraint conditions of three-dimensional scene layout can be effectively processed, the position relationship of various objects can be met, a better layout effect can be realized when a complex scene is processed, the requirements of a user on layout rationality and attractiveness can be met, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical fields of computer graphics and three-dimensional scene layout, and particularly to a three-dimensional scene layout method and system, a storage medium, a device, and a program product. Background Art

[0002] The three-dimensional scene layout technology has a wide range of applications in various fields. A reasonable three-dimensional scene layout can not only improve the authenticity and rationality of visual presentation, but also optimize the interactivity of the scene, enabling users to understand scene information more naturally and intuitively. Some related methods rely on a large amount of data for training, but the data acquisition cost is high and the generalization ability is limited, making it difficult to adapt to changing application scenarios. Other related methods can understand the scene layout, but often lack fine spatial constraints when performing three-dimensional spatial layout, resulting in insufficient layout rationality.

[0003] Based on this, the embodiments of the present application provide a three-dimensional scene layout method and system, a storage medium, a device, and a program product to improve related technologies. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a three-dimensional scene layout method and system, a storage medium, a device, and a program product, which can effectively handle various constraint conditions of three-dimensional scene layout and achieve better layout effects when dealing with complex scenes.

[0005] The purpose of the embodiments of the present application is achieved by adopting the following technical solutions:

[0006] In a first aspect, the embodiments of the present application provide a three-dimensional scene layout method, which includes: determining a constraint cost function corresponding to a set of target objects in a target area based on the corresponding position constraint relationship; constructing a pose optimization function corresponding to the set of target objects according to the corresponding collision cost function and constraint cost function of the set of target objects; using the pose optimization function corresponding to the set of target objects to solve the poses of the objects in the set of target objects; and generating three-dimensional scene information of the target area based on the poses of the objects in the target area.

[0007] In some embodiments, the set of target objects is used to form one or more object pairs, each object pair includes a relative active object and a passive object, and the position constraint relationship of the object pair includes at least one of the following: the active object is located on the passive object; the active object is located inside the passive object; the active object is adjacent to the passive object.

[0008] In some embodiments, when the main object is located on the passive object, the constraint cost function of the position constraint relationship includes a penalty cost function and a distance cost function, and the penalty cost function is determined according to the bottom feature points of the main object and the top feature points of the passive object in the corresponding position constraint relationship.

[0009] In some embodiments, when the main object is located inside the passive object, the constraint cost function of the position constraint relationship includes a penalty cost function and a distance cost function, and the penalty cost function is determined according to the top feature points of the main object and the top and bottom feature points of the passive object in the corresponding position constraint relationship.

[0010] In some embodiments, when the main object and the passive object are adjacent in position, the constraint cost function of the position constraint relationship is determined according to the distance between the main object and the passive object, the longest radius of the main object on the first plane, and the longest radius of the passive object on the first plane.

[0011] In some embodiments, the set of target objects has a position constraint relationship with the objects already placed in the target area.

[0012] In some embodiments, solving for the poses of the objects in the set of target objects using the pose optimization function corresponding to the set of target objects includes: solving for the poses of the objects in the set of target objects according to the pose optimization function corresponding to the set of target objects and the poses of the already placed objects.

[0013] In some embodiments, the already placed objects are ground objects.

[0014] In some embodiments, the process of determining the poses of the ground objects includes: when the number of ground objects is greater than 1, determining the poses of the ground objects in sequence according to the priorities of the ground objects; the priorities of the ground objects are determined according to at least one of the following: whether there is a position constraint relationship with other ground objects; the size on the layout plane; priority setting information; the scene type of the target area.

[0015] In some embodiments, the process of determining the poses of the ground objects further includes: determining the minimum bounding rectangle of at least two ground objects with a position constraint relationship on the layout plane; using the minimum bounding rectangle as a new ground object to optimize the poses of some or all of the ground objects.

[0016] In some embodiments, the process of obtaining the position constraint relationship of the objects in the target area includes: in response to user interaction information, determining the objects in the target area and the relative position relationship between the objects; the user interaction information includes one or more of the natural language information, interactive drawing information, and picture information corresponding to the target area; and determining the position constraint relationship of each object according to the objects in the target area and the relative position relationship between the objects.

[0017] In some embodiments, when the user interaction information includes the natural language information corresponding to the target area, the step of, in response to the user interaction information, determining the objects in the target area and the relative position relationship between the objects includes: using a large language model to parse the natural language information to generate scene description information; and performing semantic analysis on the scene description information to obtain the objects in the target area and the relative position relationship between the objects.

[0018] In some embodiments, when the user interaction information includes the interactive drawing information corresponding to the target area, the step of, in response to the user interaction information, determining the objects in the target area and the relative position relationship between the objects includes: in response to the interactive drawing information, performing an object drawing operation on a two-dimensional plane; and in response to an editing operation on the objects in the two-dimensional plane, editing the objects in the two-dimensional plane to obtain the objects in the target area and the relative position relationship between the objects.

[0019] In some embodiments, when the user interaction information includes the picture information corresponding to the target area, the step of, in response to the user interaction information, determining the objects in the target area and the relative position relationship between the objects includes: using an image processing model to process the picture information to determine the objects in the picture information; using a detection model to detect the positions of the objects in the picture information; generating corresponding representations of the objects on a two-dimensional plane based on the positions of the objects in the picture information; and in response to an editing operation on the objects in the two-dimensional plane, editing the objects in the two-dimensional plane to obtain the objects in the target area and the relative position relationship between the objects.

[0020] In some embodiments, the step of using the pose optimization function corresponding to the target object set to solve the poses of the objects in the target object set includes: solving the poses of the objects in the target object set according to the pose optimization function corresponding to the target object set and the object information; wherein the object information is retrieved from an object asset library, or the object information includes one or more of a three-dimensional model and dimension information.

[0021] In a second aspect, an embodiment of the present application provides a three-dimensional scene layout system, which includes a control module for executing any one of the above methods.

[0022] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any one of the above methods.

[0023] In a fourth aspect, an embodiment of the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements any one of the above methods.

[0024] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program that, when executed by a processor, implements any one of the above methods.

[0025] The embodiments of the present application provide a three-dimensional scene layout method, system, storage medium, device, and program product. The method includes: determining a constraint cost function corresponding to a set of target objects based on the corresponding position constraint relationships in a target area; constructing a pose optimization function corresponding to the set of target objects according to the corresponding collision cost function and constraint cost function of the set of target objects; using the pose optimization function corresponding to the set of target objects to solve the poses of the objects in the set of target objects; and generating three-dimensional scene information of the target area based on the poses of the objects in the target area. In the embodiments of the present application, the object poses are solved first, and then the object poses are used for three-dimensional scene layout, which is not only efficient but also reduces the dependence on scene layout data and is applicable to diverse scene layouts. When constructing the pose optimization function, the collision cost and constraint cost are considered, and the constraint cost function is constructed based on the position constraint relationships, which can effectively handle various constraint conditions of three-dimensional scene layout, meet the position relationships of various objects, achieve better layout effects when dealing with complex scenes, meet the user's requirements for layout rationality and aesthetics, and improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The embodiments of the present application will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0027] Figure 1 is a flowchart of a three-dimensional scene layout method provided by an embodiment of the present application.

[0028] Figure 2 is a flowchart of a three-dimensional scene layout provided by an embodiment of the present application.

[0029] Figure 3aIt is a top view of a carrier, a box, and a cup provided by an embodiment of the present application.

[0030] Figure 3b It is a side view of a carrier, a box, and a cup provided by an embodiment of the present application.

[0031] Figure 3c It is a schematic diagram of the bottom feature points of a cup provided by an embodiment of the present application.

[0032] Figure 3d It is a schematic diagram of the sampling points of a cup provided by an embodiment of the present application.

[0033] Figure 3e It is a schematic diagram of the top feature points of a box provided by an embodiment of the present application.

[0034] Figure 3f It is a schematic diagram of the sampling points of a box provided by an embodiment of the present application.

[0035] Figure 4a It is a schematic diagram of the positions of a bearing surface, a bowl, and an apple provided by an embodiment of the present application.

[0036] Figure 4b It is a schematic diagram of the positions of a microwave oven, a baking tray, and a pizza provided by an embodiment of the present application.

[0037] Figure 5 It is another schematic diagram of the three-dimensional scene layout process provided by an embodiment of the present application.

[0038] Figure 6 It is a structural block diagram of a three-dimensional scene layout system provided by an embodiment of the present application.

[0039] Figure 7 It is a structural block diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the embodiments of the present application.

[0041] In the description of the embodiments of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0042] The related 3D (Three Dimensions) scene layout methods include the following two types of technical paths. The first type is the model-based training and inference method, which collects a large amount of scene layout data and realizes the layout of the scene through model training. However, it is difficult to obtain and process 3D scene layout data; at the same time, the trained model is limited by the data distribution and is difficult to generalize to scene types not covered by the training set (such as migrating from home layout to industrial equipment assembly).

[0043] The second type of technology is the method of combining key scene information generation by a large model with optimization solving. Through the large model's understanding of the scene layout, it outputs the objects suitable for placement in the scene, selects the appropriate objects through a retrieval method, and solves the positions of the objects in the scene. Although this method performs well in 2D plane layout (such as the top view arrangement of indoor furniture), it is insufficient in dealing with 3D space layout, difficult to handle the complex positional relationships of multiple objects, such as the scene of placing food inside a microwave oven, and lacks an intelligent optimization mechanism, unable to effectively handle multiple constraint conditions, resulting in an unsatisfactory layout result. In addition, this method is inefficient in performing complex scene layout and is difficult to meet the user's requirements for layout rationality and aesthetics. These lead to limitations in the user experience when related technologies handle complex 3D scene layout, and there is an urgent need for improvement.

[0044] See Figure 1 and Figure 2 , Figure 1 are the schematic flowcharts of a 3D scene layout method provided by the embodiments of the present application, Figure 2 is the schematic flowchart of a 3D scene layout process provided by the embodiments of the present application.

[0045] To improve the related technologies, the embodiments of the present application provide a 3D scene layout method, and the method includes steps S101 to S104.

[0046] Step S101: Determine the constraint cost function corresponding to the target object set based on the corresponding position constraint relationship in the target area.

[0047] Step S102: Construct a pose optimization function corresponding to the set of target objects according to the collision cost function and the constraint cost function corresponding to the set of target objects.

[0048] Step S103: Use the pose optimization function corresponding to the set of target objects to solve the poses of the objects in the set of target objects.

[0049] Step S104: Generate three-dimensional scene information of the target area based on the poses of the objects in the target area.

[0050] In some embodiments, the above method can be executed on a control module of a three-dimensional scene layout system.

[0051] The above embodiments do not limit the target area, which may be, for example, an indoor area or an outdoor area. Among the objects in the target area, some or all of the objects may have position constraint relationships. The above embodiments do not limit the number of objects in the target area, which may be, for example, at least two.

[0052] In some embodiments, the objects can be used to form one or more object groups, each object group including at least two objects, and the position constraint relationship of the object group is used to indicate the position constraint conditions between the objects in the object group. The number of objects in different object groups may be the same or different. For example, 1 conference table and 10 chairs can be used as an object group, and the chairs are evenly placed around the conference table and the distance is not greater than a specified distance threshold.

[0053] In some embodiments, the set of target objects is used to form one or more object pairs, each object pair including a relative active object and a passive object, and the position constraint relationship of the object pair includes at least one of the following: the active object is located on the passive object; the active object is located inside the passive object; the active object is adjacent to the passive object in position.

[0054] In the above embodiments, the corresponding position constraint relationships of the target object set may include the position constraint relationships of all object pairs formed by the target object set. Each object pair includes two objects, and there is a position constraint relationship between these two objects to constrain the relative positions of the two objects. An object pair has a relative active object and passive object, and the corresponding position constraint relationship may include at least one of the following three types: the active object a is located on the passive object b (a on b); the active object a is located inside the passive object b (a in b); the active object a is adjacent to the passive object b (a near b). That is to say, within an object pair, the two objects with a position constraint relationship are distinguished, and the objects before and after the relational preposition are respectively used as the active object and the passive object. Moreover, the active object and the passive object are relative. For example, in the position constraint relationship of "the plate is located on the table", the active object is the plate and the passive object is the table; but in the position constraint relationship of "the apple is located inside the plate", the active object is the apple and the passive object is the plate. It can be seen that the same object can be included in one or more object pairs. For example, the object "plate" is in the object pair of "plate and table" and the object pair of "apple and plate". And the same object can be used as the active object and the passive object respectively in different object pairs (such as the "plate" in the above text), or can also be used as the passive object (for example, "the plate is located on the table", "the bookshelf is adjacent to the table", and at this time, "table" is used as the passive object in both object pairs), or can also be used as the active object (for example, "the bookshelf is adjacent to the table", "the bookshelf is adjacent to the sofa", and at this time, "bookshelf" is used as the active object in both object pairs), and the above embodiments do not limit this.

[0055] In some embodiments, using the pose optimization function corresponding to the target object set to solve the poses of the objects in the target object set may include: solving the poses of the objects in the target object set according to the pose optimization function corresponding to the target object set and the object information.

[0056] The above embodiments do not limit the method for obtaining object information. In some embodiments, the object information may be retrieved from an object asset library. The above embodiments do not limit the object information. In some embodiments, the object information may include one or more of a three-dimensional model and dimensional information. Among them, the three-dimensional model may be, for example, a triangle mesh model (Triangle Mesh), a quadrilateral mesh model, a three-dimensional CAD model, etc., and the above embodiments do not limit this. The full name of CAD is Computer Aided Design, computer-aided design. The mesh model is suitable for real-time detection, and the three-dimensional CAD model can retain the surface accuracy. The dimensional information may include, for example, parameters such as the length, width, height, diameter, radius of the object, or parametric dimensional constraint information, and the above embodiments do not limit this.

[0057] By retrieving the corresponding layout object from the object asset library, basic information such as the size (i.e., dimensions) of the object can be obtained as object information. For example, the object asset library may be a structured database containing three-dimensional models and their metadata, covering classification objects such as furniture and industrial parts (such as coffee tables and mechanical components in a CAD model library), and storing attribute data such as dimensions, materials, and physical properties. As an example, the object asset library may support formats such as triangle mesh models and three-dimensional CAD models to meet the requirements of real-time collision detection and engineering-level layout verification. During retrieval, retrieval matching can be performed through dimensional ranges, keywords (such as "office furniture", "bookshelf", etc.), and the above embodiments do not limit this. In some embodiments, the object information may include a three-dimensional model. By parsing the relevant parameters of the three-dimensional model (for example, the vertex coordinates of the triangle mesh model or reading the CAD preset parameters), the dimensional information of the object can be automatically generated, providing a basic input for pose optimization solving.

[0058] During the object pose solving process, a pose optimization function can be used for optimization solving. Whether the object will collide needs to be considered during the object pose solving process. Therefore, during the solving process, it is necessary to ensure that there is no collision between objects and at the same time meet the relevant constraint conditions required for the layout. The above embodiments consider a collision cost function when constructing the pose optimization function to quantify the collision cost. In addition, the above embodiments consider a constraint cost function when constructing the pose optimization function to quantify the constraint cost. It should be noted that in the expression of the pose optimization function, corresponding weights can also be set for the collision cost function and the constraint cost function respectively, and the specific values of the weights can be selected, set, or adjusted according to the needs in actual applications, and the above embodiments do not limit this.

[0059] In the above embodiments, a constraint cost function is designed according to the corresponding position constraint relationship of the target object set. The constraint cost function corresponds to the position constraint relationship. Therefore, it can effectively handle various position constraint relationships that the objects need to satisfy and obtain a better three-dimensional layout result.

[0060] First, the calculation process of the collision cost function will be described below. The collision cost function corresponding to the target object set may include the collision cost function of each object pair formed by the target object set.

[0061] In some embodiments, for each object pair, the object can be represented by an object 3D box to calculate the function value of the collision cost function to ensure that the objects do not collide. Among them, the object 3D box is, for example, an axis-aligned bounding box (Axis-Aligned Bounding Box, AABB) defined based on a three-dimensional coordinate system, which is used to represent the spatial position, size, and orientation of the object. This bounding box can be quickly constructed through the calculation of extreme points and is used for object collision detection and pose estimation, but it is difficult to accurately describe complex shapes (such as nested objects). For example, the collision cost calculation method based on the object 3D box has difficulty meeting the requirements for the position constraint relationship where the active object a is located inside the passive object b, and at the same time, the solution distance between the objects will be relatively large.

[0062] In some other embodiments, surface sampling of all objects in the target area can be performed first. As an example, N points are sampled for each object respectively, and N is, for example, 2000. Next, the passive object in the object pair can be represented by using SDF (Signed Distance Field) data in combination with the sampling points. As an example, using the triangular mesh model of the object, a signed distance field is generated by constructing a ray casting scene and calculating the distance from each grid point to the grid surface. An interpolation function sdf_fun is created by using RegularGridInterpolator in the interpolate sub-library of SciPy (as an example of an open-source library), which converts the discrete SDF data into a continuous function for more accurate query and operation in three-dimensional space. Next, the collision cost function collision_cost is used to calculate the collision cost between the transformed point cloud of the active object and the passive object in the current pose. For example, first, the point cloud data of the active object is transformed into the coordinate system of the passive object, and then the transformed point cloud data of the active object is flattened into a two-dimensional array, where each row represents the three-dimensional coordinates of a corresponding point of the active object. The interpolation function sdf_fun created above is used to query the signed distance signed_distance of each point. Then the collision cost function is used to calculate the collision cost. The collision cost function can be expressed as, for example, collision_cost = -np.sum(np.maximum(0, threshold - signed_distance)). The resulting effect is that if the signed distance signed_distance of the point is less than the preset distance threshold threshold, a penalty will be incurred. Specifically, it can be seen from the collision cost function that for each point, threshold - signed_distance is calculated respectively. If the calculation result is positive, it means the point is within the collision range and needs to be penalized, and the penalty value of the point is the calculation result; otherwise, no penalty is required, that is, the penalty value of the point is set to zero. Finally, the penalty values of all points are summed and then the negative value is taken as the collision cost. In the example of the above collision cost function, functions of the numpy library are used, including np.sum and np.maximum, and "np" is the abbreviation of the numpy library. np.maximum is used to compare two numerical values or corresponding elements in arrays and return the maximum value, and np.sum is used to calculate the sum of array elements.

[0063] For an object subject to a position constraint, in addition to considering the collision cost between objects, it is also necessary to consider the constraint cost corresponding to the position constraint relationship, and the corresponding cost function is the constraint cost function. During the pose solution process, different constraint cost functions can be designed for different position constraint relationships. The following gives examples of the corresponding constraint cost functions for different types of position constraint relationships. It should be noted that when the position constraint relationship corresponds to an object pair, when solving the poses of the active object and the passive object in the object pair based on this position constraint relationship, the collision cost function and the constraint cost function used can be the same. When solving the poses for the target object set, the constraint cost function corresponding to the target object set includes the constraint cost functions for each position constraint relationship.

[0064] In some embodiments, when the active object is located on the passive object, the constraint cost function of the position constraint relationship may include a penalty cost function and a distance cost function, and the penalty cost function is determined according to the bottom feature points of the active object and the top feature points of the passive object in the corresponding position constraint relationship.

[0065] See Figures 3a to 3f , Figure 3a is a top view of a carrier, a box, and a cup provided by an embodiment of the present application, Figure 3b is a side view of a carrier, a box, and a cup provided by an embodiment of the present application, Figure 3c is a schematic diagram of the bottom feature points of a cup provided by an embodiment of the present application, Figure 3d is a schematic diagram of the sampling points of a cup provided by an embodiment of the present application, Figure 3e is a schematic diagram of the top feature points of a box provided by an embodiment of the present application, Figure 3f is a schematic diagram of the sampling points of a box provided by an embodiment of the present application. As can be seen from Figures 3a to 3f , the box is located on the carrier, and the cup is located on the box. Although Figure 3b the surface of the carrier in Figure 3b is not horizontal, however, the above embodiments do not limit this, and the surface of the passive object for carrying the active object can be in a horizontal state or a non-horizontal state. In addition, although Figures 3c to 3f the contact surface between the cup and the box in

[0066] is not horizontal, however, the above embodiments do not limit this, and the contact surface between the active object and the passive object can be in a horizontal state or a non-horizontal state. In Figures 3c to 3f , Coordinate is the coordinate, X Coordinate is the X coordinate, Y Coordinate is the Y coordinate, Z Coordinate is the Z coordinate, where X, Y, and Z respectively correspond to the X-axis, Y-axis, and Z-axis, and the X-axis, Y-axis, and Z-axis can be the coordinate axes in the Cartesian coordinate system.

[0066] For example, asFigure 3a and Figure 3b As shown, assume that the main object a is a cup and the passive object b is a box, and the cup is located on the box. First, the feature points of both can be obtained respectively. For example, the bottom feature points of the cup (here the main object, assumed to be M1 in number) and the top feature points of the box (here the passive object, assumed to be M2 in number) can be obtained. Next, sample the feature points. For example, sample M points from the bottom feature points of the cup (e.g., uniform sampling or non-uniform sampling), and sample M points from the top feature points of the box (e.g., uniform sampling or non-uniform sampling). M is, for example, 10, and M is less than M1 and M is less than M2. After obtaining the sampled points, the mean feature points can be calculated. For example, randomly select M3 (M3 is less than M) points from the M sampled points of the main object and the passive object respectively for mean processing to obtain the mean feature point ac_bot of the main object (i.e., the cup) and the mean feature point pa_top of the passive object (i.e., the box). After determining the mean feature points of the main object and the passive object, the penalty cost function and the distance cost function can be calculated respectively. As an example, when calculating the penalty cost function cost_add, check the heights of these two mean feature points. If the mean feature point of the main object is lower than the feature point of the passive object, it is determined that the function value of the penalty cost function is 10. When calculating the distance cost function distance_cost, the Euclidean distance, Manhattan distance, etc. can be used, and the above embodiments do not limit this. As an example, the Euclidean distance between ac_bot and pa_top can be calculated to obtain the function value of the distance cost function. In the case where the main object is located on the passive object, when calculating the constraint cost function, the function values of the penalty cost function and the distance cost function can be directly summed, or weighted summation can be performed on the two, and the above embodiments do not limit this. As an example, assume that the calculation of the constraint cost function adopts the direct summation method, then the expression of the constraint cost function constraint_cost is, for example, constraint_cost = cost_add + distance_cost.

[0067] In some other embodiments, in the case where the main object is located inside the passive object, the constraint cost function of the position constraint relationship may include a penalty cost function and a distance cost function, and the penalty cost function is determined according to the top feature points of the main object, the top feature points and the bottom feature points of the passive object in the corresponding position constraint relationship.

[0068] See Figure 4a and Figure 4b , Figure 4a is a schematic diagram of the positions of a bearing surface, a bowl, and an apple provided by an embodiment of the present application. Figure 4bIt is a schematic diagram of the positions of a microwave oven, a baking tray, and a pizza provided by an embodiment of the present application. As can be seen from Figure 4a , the bowl is located on the bearing surface, and the apple is located inside the bowl. As can be seen from Figure 4b , the baking tray is located inside the microwave oven, and the pizza is located on the baking tray.

[0069] For example, as shown in Figure 4a and Figure 4b , assuming that the main object a is an apple and the passive object b is a bowl, and the apple is located inside the bowl. When the main object a (i.e., the apple) is located inside the passive object b (i.e., the bowl), first, obtain the bottom feature points of the main object, as well as the top and bottom feature points of the passive object. Then sample the bottom feature points of the main object (for example, random sampling), and after obtaining N1 (N1 is 3 for example) bottom feature points, perform mean processing to obtain the mean feature point of the main object; sample the top feature points of the passive object, and after obtaining N1 (N1 is 3 for example) top feature points, perform mean processing to obtain the top mean feature point of the passive object; sample the bottom feature points of the passive object, and after obtaining N2 (N2 is 10 for example) bottom feature points, perform mean processing to obtain the bottom mean feature point of the passive object. After determining the mean feature point of the main object, as well as the top and bottom mean feature points of the passive object, the penalty cost function and the distance cost function can be calculated respectively. As an example, when calculating the penalty cost function cost_add, it can be checked whether the mean feature point of the main object is higher than the top mean feature point of the passive object. If so, set the function value of the penalty cost function to 10. When calculating the distance cost function distance_cost, the Euclidean distance between the mean feature point of the active object and the bottom mean feature point of the passive object can be calculated to obtain the function value of the distance cost function. When the main object is located inside the passive object, when calculating the constraint cost function, the function values of the penalty cost function and the distance cost function can be directly summed, or weighted summation can be performed on the two. The above embodiments do not limit this. As an example, assuming that the calculation of the constraint cost function adopts the direct summation method, the expression of the constraint cost function constraint_cost is, for example, constraint_cost = cost_add + distance_cost.

[0070] In some other embodiments, when the main object and the passive object are adjacent in position, the constraint cost function of the position constraint relationship is determined according to the distance between the main object and the passive object in the corresponding position constraint relationship, the longest radius of the main object on the first plane, and the longest radius of the passive object on the first plane.

[0071] In the above embodiments, the distance between the active object and the passive object is not limited. For example, it can be the spatial distance between their centers (which can be calculated using three-dimensional coordinate data), or it can be the central distance between them on the same plane (which can be calculated using two-dimensional coordinate data). For example, when the central distance between the active object a and the passive object b on the first plane is greater than the distance threshold limit_distance, the cost increases by 10. The above embodiments do not limit the method for determining the distance threshold, which can be selected or set according to the needs in actual applications. As an example, limit_distance = a_max_size + b_max_size, where a_max_size and b_max_size respectively represent the longest radii of the active object a and the passive object b on the first plane. Here, the longest radius refers to, for example, projecting the object onto the first plane (such as the ground or horizontal plane) and calculating the maximum extension distance of its two-dimensional bounding box in each direction of the plane coordinate system. For example, for a table with a rectangular projection, the longest radius is half of the long side of the bounding box; for a cup with a circular projection, the longest radius is the radius of the circle. In addition, if the object has an asymmetric shape (such as an L-shaped bookshelf), algorithms such as principal component analysis can be used to extract the principal direction axis of the projection, and the maximum extension length along this axis can be calculated as the longest radius.

[0072] In the process of solving the pose of the target object set in the above embodiments, for the corresponding position constraint relationships of the target object set, a pose optimization function including a constraint cost function can be constructed, and the object pose can be solved using the target pose optimization algorithm. The constraint cost functions corresponding to different position constraint relationships are different. Therefore, the pose optimization functions corresponding to different position constraint relationships are also different. The above embodiments do not limit the target pose optimization algorithm, which can include, for example, the Dual Annealing algorithm, the Particle Swarm Optimization (PSO) algorithm, the Simulated Annealing algorithm, etc., and can be selected according to actual needs to improve the adaptability and flexibility of pose calculation.

[0073] In some embodiments, when there are many objects in the target area, the objects in the target area can be divided into a target object set, and the poses of all objects can be solved at once.

[0074] In other embodiments, such as Figure 2As shown, a hierarchical method can be adopted to optimize and solve the pose of an object. At each level, the pose of the corresponding partial object (rather than all objects in the target area) at that level is solved. Then, the objects whose poses have been solved are used as the placed objects, and the pose of the unplaced objects (i.e., the objects whose poses have not been solved) that have a position constraint relationship with the placed objects is solved by using the poses of the placed objects. It should be noted that after the pose of the placed object is solved, its pose will no longer change.

[0075] The above embodiments do not limit the hierarchical division method. For example, it can be divided according to whether the object is in contact with the ground (first solve the pose of the object in contact with the ground, and then solve the pose of the object not in contact with the ground), or it can also be divided according to the distance of the object from the edge of the area (the entire target area is divided into multiple annular sub-areas according to the distance from the edge of the area, such as divided into an outer ring, a middle ring, and an inner ring, and solved layer by layer in the order from the outside to the inside). The above embodiments do not limit the number of levels, which can be, for example, two levels, three levels, etc. Thus, the number of objects whose poses need to be solved at each level can be reduced, the computational complexity can be reduced, the computational amount can be saved, and the computational speed and computational efficiency can be improved.

[0076] In some embodiments, the set of target objects may have a position constraint relationship with the placed objects in the target area.

[0077] In some embodiments, solving the poses of the objects in the set of target objects by using the pose optimization function corresponding to the set of target objects may include: solving the poses of the objects in the set of target objects according to the pose optimization function corresponding to the set of target objects and the poses of the placed objects.

[0078] In some embodiments, the placed object may be a ground object.

[0079] In the process of solving the pose of the ground object, if the number of ground objects is 1, the pose of the ground object is directly solved; if the number of ground objects is greater than 1, the poses of these ground objects can be solved in a random order, or the poses of these ground objects can be solved in the order of decreasing priority.

[0080] In some embodiments, the process of determining the pose of the ground object may include: when the number of ground objects is greater than 1, determining the poses of the ground objects in sequence according to the priorities of the ground objects.

[0081] In some embodiments, the priority of the ground object is determined according to at least one of the following: whether it has a position constraint relationship with other ground objects; the size on the layout plane; the priority setting information; the scene type of the target area.

[0082] In the above embodiments, for the convenience of description, an object in contact with the ground (e.g., a table, a refrigerator, a sofa, etc. located on the ground) can be referred to as a ground object, the object whose pose is solved is set as the placed object, and the object whose pose is not solved is set as the unplaced object. It can be understood that the number of unplaced objects and placed objects is dynamically changing. As the unplaced objects are placed, these unplaced objects will become placed objects. Therefore, the number of unplaced objects will gradually decrease, and the number of placed objects will gradually increase. As an example, a placement status flag can be used to distinguish the placement status of objects. For example, before starting to solve the object pose, the placement status flags of all objects are initialized to "0". As the solving process progresses, the poses of some objects are solved, and the placement status flags of these objects whose poses are solved (i.e., the placed objects) are set to "1". After solving the poses of all objects, the placement status flags of all objects are "1", indicating that the pose solving is completed. In addition to using the placement status flag to distinguish placed objects and unplaced objects, the placed objects and unplaced objects can also be stored in the form of a placed object set and an unplaced object set respectively. Before starting to solve the object pose, all objects are stored in the unplaced object set. As the solving process progresses, objects are continuously taken out from the unplaced object set and their poses are solved, and then these objects whose poses are solved are put into the placed object set. Finally, the unplaced object set is empty, and all objects are stored in the placed object set, indicating that the pose solving is completed. After all objects complete a pose optimization solving, the pose optimization solving process can be executed again, and initialization settings (i.e., setting all objects as unplaced objects) are made before each start of solving. The above embodiments do not limit this.

[0083] The following will first describe the process of solving the pose of the ground object, and then describe the process of solving the pose of the corresponding target object set based on the pose of the placed object.

[0084] The above embodiments do not limit the method for determining the priority of ground objects. In some embodiments, the priority of a ground object having a position constraint relationship with other ground objects may be higher than that of a ground object having no position constraint relationship with other ground objects. That is to say, among ground objects, if some ground objects have a position constraint relationship with other ground objects, while some ground objects have no position constraint relationship with other ground objects, then the pose of the ground objects having a position constraint relationship with other ground objects can be solved first, and then the pose of the ground objects having no position constraint relationship with other ground objects can be solved. Since these ground objects are all in contact with the ground, if there is a position constraint relationship between these ground objects, this position constraint relationship can be "near". For example, the ground objects include a bookshelf and a bed, and the bookshelf is adjacent to the bed.

[0085] Put all the ground objects having a position constraint relationship into a ground object set. When solving the pose of the ground object set having a position constraint relationship, the corresponding constraint cost function of the ground object set can be determined based on the corresponding position constraint relationship of the ground object set, and then the corresponding pose optimization function of the ground object set can be constructed according to the corresponding collision cost function and constraint cost function of the ground object set, and the pose of the ground object set can be solved using the corresponding pose optimization function of the ground object set.

[0086] For a ground object having no position constraint relationship with other ground objects, the positions in the placeable area can be traversed. For example, it can be placed at any position in the placeable area. As an example, for a ground object having no position constraint relationship with other ground objects, at least one candidate pose of the ground object in the placeable area can be determined according to the size of the ground object in the layout plane and a preset angular interval; the pose of the ground object can be determined from the at least one candidate pose; wherein, the placeable area is located in the layout plane and does not overlap with the placed area. The above embodiments do not limit the preset angular interval, which can be, for example, 30 degrees. Each candidate pose corresponds to a combination of a candidate position and a candidate attitude. However, each candidate position can correspond to one or more candidate attitudes. For example, for each candidate position, 0 to 360 degrees can be divided at intervals of 30 degrees to obtain 12 candidate attitudes.

[0087] In other embodiments, the priority of each ground object can be determined according to the order of size from large to small on the layout plane. The layout plane can be flexibly selected or set according to actual needs, such as the ground or horizontal plane, which is not limited in the above embodiments. Thus, the ground objects can be sorted according to their size on the layout plane, and larger ground objects can be placed first. This priority setting method is simple and effective. By giving priority to placing larger ground objects (such as sofas and bookcases), their occupied areas can be quickly eliminated, so that the candidate positions of the remaining ground objects are greatly reduced, and the amount of subsequent optimization calculations is significantly reduced. Specifically, large objects have a strong spatial anchoring effect, and priority positioning can reduce layout conflicts caused by small objects occupying key areas (such as in a conference room scene, indoor landscape plants occupy the center of the conference room and force the conference table to be placed against the wall). This layout method meets the user's intuitive expectations of space partitioning (such as the position of the dining table determines the restaurant area), and improves the rationality of the layout. In addition, large objects usually require fixed positions (such as bookcases against the wall to prevent tipping), and priority solving can ensure higher stability.

[0088] In some other embodiments, the priority of each ground object can be determined according to priority setting information. The priority setting information is, for example, a priority number set by a user, where a smaller number indicates a higher priority, and vice versa, a larger number indicates a lower priority.

[0089] In some other embodiments, the priority of each ground object can be determined according to the scene type of the target area and the scene semantics. For example, assuming the scene type is a conference room, the priority of the conference table is set higher than that of the chairs.

[0090] In some other embodiments, the priority of each ground object can be determined based on at least one of the following: whether there is a position constraint relationship with other ground objects; the size on the layout plane; the priority setting information; the scene type of the target area. For the above multiple influencing factors, corresponding weight coefficients can be set to characterize the importance of these influencing factors to the priority sorting, so as to more accurately meet the user needs in various scenarios.

[0091] To accelerate the solution speed, the planar position (e.g., the central position of an object in the layout plane) and the yaw angle (e.g., yaw) can be used as solution variables to solve, and the pose of the ground object can be determined based on the solution variables. For the case where some ground objects have position constraint relationships with multiple other ground objects (e.g., multiple chairs are placed around a conference table), the pose solution result may not have a good layout effect. To further optimize the poses of these ground objects, the optimization solution process can be executed again after the first optimization solution. In some embodiments, the pose determination process of the ground object may further include: determining the minimum bounding rectangle of at least two ground objects with position constraint relationships on the layout plane according to the poses of the at least two ground objects; using the minimum bounding rectangle as a new ground object to optimize the poses of some or all of the ground objects.

[0092] Illustratively, for the convenience of calculation, during the pose solution process, the layout plane can be meshed, for example, divided at intervals of 40 cm to generate a plurality of grid points. During the pose solution process of the ground object, the grid points can be divided into unoccupied grid points and occupied grid points, and the occupied grid points can be excluded from the placeable area, and only unoccupied grid points are allowed to place objects. To distinguish between unoccupied grid points and occupied grid points, an occupancy flag can be set for the grid points. For example, when the grid point is unoccupied, the occupancy flag of the grid point can be set to "0"; when the grid point is occupied, the occupancy flag of the grid point can be set to "1". For two or more ground objects with position constraint relationships (which need to correspond to the same position constraint relationship), combining the poses of these ground objects that have been solved, their minimum bounding rectangle can be obtained as a new ground object for the next optimization solution process. It should be noted that since the size of the new ground object changes, the corresponding priority ranking may also change. During each optimization solution process, the pose solution can be performed according to the current priority ranking of each ground object. Compared with the random order solution, the pose solution method based on priority can accurately meet the personalized needs of users and has a wide range of application prospects.

[0093] After solving the poses of the objects on the ground, the poses of the objects not in contact with the ground can be solved. It should be noted that after solving the poses of the objects on the ground, these objects on the ground (e.g., tables, desks, cupboards) are set as placed objects. For the objects not in contact with the ground, if there are unplaced objects having position constraint relationships with the objects on the ground (e.g., the green plants on the table, the plates on the desk, the microwave ovens on the cupboard), then the poses of the corresponding unplaced objects can be solved based on the poses of the objects on the ground and the corresponding position constraint relationships. In practical applications, the unplaced objects in the target area can be divided into one or more target object sets having position constraint relationships with the objects on the ground, and the poses of all the objects in the corresponding target object sets can be solved based on the poses of the placed objects. For example, assume that the objects on the ground include a dining table and a desk, there are a plate and a bowl placed adjacent to each other on the dining table, and there are books, a pen holder and a vase placed on the desk (and the books and the pen holder are adjacent in position), and there is a flower arrangement in the vase. The plate and the bowl can be taken as a target object set corresponding to the dining table, and the books, the pen holder, the vase and the flower arrangement can be taken as another target object set corresponding to the desk.

[0094] It can be seen that the position constraint relationships between the unplaced objects and the objects on the ground can be direct or indirect, and through the direct or indirect position constraint relationships, the placed objects related to each unplaced object can be determined. In other words, for each unplaced object, the corresponding placed object can be determined through one or more position constraint relationships, and the unplaced object can be put into the corresponding target object set of the placed object, and it is determined that the target object set has a position constraint relationship with the placed object. In practical applications, multiple unplaced objects can also be processed at one time. For example, assume that the flower arrangement has a position constraint relationship with the vase (the flower arrangement is inside the vase), the vase has a position constraint relationship with the desk (the vase is on the desk), and the desk is an object on the ground, then it can be determined that the corresponding object on the ground for the flower arrangement and the vase is the desk, and the flower arrangement and the vase are put into the corresponding target object set of the desk. Similarly, assume that the cupboard is an object on the ground, there is a microwave oven placed on the cupboard, there is a baking tray placed inside the microwave oven, and there is a slice of bread placed on the baking tray. Since the slice of bread has a position constraint relationship with the baking tray (the slice of bread is on the baking tray), the baking tray has a position constraint relationship with the microwave oven (the baking tray is inside the microwave oven), and the microwave oven has a position constraint relationship with the cupboard (the microwave oven is on the cupboard), and the cupboard is an object on the ground, thus, it can be determined that the corresponding object on the ground for the slice of bread, the baking tray and the microwave oven is the cupboard, and the slice of bread, the baking tray and the microwave oven are put into the corresponding target object set of the cupboard. It should be noted that each unplaced object corresponds to a unique object on the ground and cannot be put into the corresponding target object sets of two objects on the ground at the same time.

[0095] In some embodiments, the position constraint relationships of the objects in the target object set may include one or more of "on", "in", and "near". For example, assuming the table is a ground object, a book, a plate, and a green plant are placed on the table, glasses are placed on the book, and apples (quantity: 3) are placed in the plate. Then the corresponding target object set of the table includes the book, the glasses, the green plants (quantity: 2), the plate, and the apples (quantity: 3). There are 3 corresponding position constraint relationships for this target object set. The position constraint relationship between the glasses and the book is "on" (the glasses are on the book), the position constraint relationship between the two green plants is "near" (green plant 1 and green plant 2 are adjacent), and the position constraint relationship between the apples and the plate is "in" (the apples are in the plate). As an example, a constraint cost function and a collision cost function can be constructed for each position constraint relationship, and then a pose optimization function including all the constraint cost functions and collision cost functions can be constructed, and the target pose optimization algorithm can be used to solve the poses of the objects in the target object set.

[0096] As can be seen from the above embodiments, the process of solving the object pose requires obtaining the position constraint relationships of the objects. The above embodiments do not limit the process of obtaining the position constraint relationships of the objects. In some embodiments, the process of obtaining the position constraint relationships of the objects in the target area may include: receiving the position constraint relationships of each object specified by the user. That is, the user can directly specify the position constraint relationships of the objects to meet the user's quick specification requirements. Moreover, the user can specify manually or in the way of importing information, and the above embodiments do not limit this.

[0097] In some other embodiments, the process of obtaining the position constraint relationships of the objects in the target area may include: in response to user interaction information, determining the objects in the target area and the relative position relationships between the objects; the user interaction information includes one or more of the natural language information, interactive drawing information, and picture information corresponding to the target area; and determining the position constraint relationships of each object according to the objects in the target area and the relative position relationships between the objects. In this way, it is convenient for the user to flexibly select a suitable interaction method according to the requirements in the actual application and their own interaction preferences to determine the objects to be laid out in the target area and the relative position relationships between the objects, and then determine the position constraint relationships of each object.

[0098] Users can provide user interaction information in the form of language input. In some embodiments, when the user interaction information includes natural language information corresponding to the target area, the determining of the objects in the target area and the relative position relationship between the objects in response to the user interaction information may include: parsing the natural language information using a large language model to generate scene description information; performing semantic analysis on the scene description information to obtain the objects in the target area and the relative position relationship between the objects.

[0099] For example, users can input scene design requirements through human natural language. As an example, the natural language information input by the user can be: "Help me design a specific living room that can be used for relaxation and reading on a daily basis." The above embodiments can use a large language model (LLM) to parse the natural language information input by the user and extract key requirements, such as the function, style, and layout preferences of the target area. The parsed information can be converted into specific spatial relationship instructions that define the relative position relationship between objects, such as "near", "on", "in", etc. According to the spatial relationship instructions, scene description information including object information and its relative position description can be generated, and semantic analysis is performed on the scene description information to obtain the objects required for the scene layout (such as sofas, bookshelves, etc.) and the relative position relationship between the objects. The semantic analysis result can also include the number of objects (such as 2 sofas, 1 bookshelf, etc.).

[0100] Users can provide user interaction information in the form of interactive drawing. In addition, there is a problem of user interaction limitations in related technologies, and most interaction methods do not support users to dynamically adjust the layout, such as adding, deleting, or modifying objects. In some embodiments, when the user interaction information includes interactive drawing information corresponding to the target area, the determining of the objects in the target area and the relative position relationship between the objects in response to the user interaction information may include: performing an object drawing operation on a two-dimensional plane in response to the interactive drawing information; performing editing on the objects in the two-dimensional plane in response to an editing operation on the objects in the two-dimensional plane to obtain the objects in the target area and the relative position relationship between the objects.

[0101] For example, users can draw different-sized graphics (such as squares, triangles, circles, trapezoids, etc.) on a two-dimensional plane through the interactive drawing function to represent objects, and the graphic positions can represent the positions of the objects. The above embodiments support editing operations such as adding, deleting, and modifying these objects on the two-dimensional plane, generating the position information of the objects on the two-dimensional plane, so as to determine the relative position relationship between the objects.

[0102] The user can provide user interaction information in a picture-style input manner. In some embodiments, when the user interaction information includes picture information corresponding to the target area, the determining of the objects in the target area and the relative position relationship between the objects in response to the user interaction information may include: processing the picture information using a picture processing model to determine the objects in the picture information; using a detection model to detect the positions of the objects in the picture information; generating corresponding representations of the objects on a two-dimensional plane based on the positions of the objects in the picture information; and editing the objects in the two-dimensional plane in response to an editing operation on the objects in the two-dimensional plane to obtain the objects in the target area and the relative position relationship between the objects.

[0103] The picture processing model in the above embodiments is not limited, and it may be, for example, a Visual Language Model or a multimodal large model (such as Deepseek, GPT4o, etc.). In the above embodiments, the user interaction information may include picture information, and the user can select a reference picture as the picture information for interaction and perform scene layout according to their own preferences. The above embodiments can process the objects in the picture information through a picture processing model (such as scene understanding), and then use a detection model (such as Grounding dino, a general detection model) to obtain the positions of the objects in the picture information (expressed in the form of two-dimensional coordinates). Then, corresponding representations of the objects are generated on a two-dimensional plane (such as graphics like rectangles, triangles, circles, trapezoids, etc.), and then the objects on the two-dimensional plane are edited to determine the objects in the scene and the relative position relationship between the objects. As an example, using the picture information input method, a position layout consistent with the layout result of the interactive drawing information can be obtained.

[0104] See Figure 5 , Figure 5It is another schematic diagram of the three-dimensional scene layout process provided by the embodiments of the present application. Users can select input methods such as language input (input natural language information), interactive drawing (input interactive drawing information), picture style input (input picture information), etc. to input user interaction information. The input user interaction information is processed correspondingly to obtain the objects required for the scene layout and the relative position relationship between the objects. When solving the object pose, a constraint cost function can be constructed according to the position constraint relationship, and a pose optimization function can be constructed based on the collision cost function and the constraint cost function. Retrieve the objects for scene layout from the object asset library, and according to the retrieved object information, use the pose optimization function to solve the pose of the objects to obtain the pose of the objects for three-dimensional scene layout (that is, the final object pose information). Next, based on the poses of the objects in the scene, three-dimensional scene information can be generated, for example, it can be represented in the form of an effect diagram, or it can also be represented in the form of point cloud data information, etc.

[0105] In some specific application scenarios, multiple input methods can be used in an organic combination. For example, after the user inputs picture information, a picture processing model is used to determine the objects in the picture information, and then the interactive drawing function is used to draw a sketch, and the relative position relationship between the objects is obtained through sketch recognition. This method eliminates the use of the detection model and provides the relative position relationship by means of the user drawing a sketch. Through the above embodiments, the object layout constraints (that is, the position constraint relationships that each object needs to satisfy) with unified expression can be obtained through different interaction methods, the object information is provided by calling the object asset library, and then the pose of the object in the target area is solved (that is, the spatial pose solution), and then the final three-dimensional scene layout result is obtained. Thus, by integrating multiple input methods and the spatial pose solution function, an efficient and accurate scene layout design is realized, meeting the diverse needs of users.

[0106] The above embodiments propose a 3D scene layout method based on hierarchical optimization, which supports layout optimization under complex constraint conditions and significantly improves the efficiency and rationality of three-dimensional scene layout. This scalable system architecture is convenient for designing the corresponding target object set according to the actual requirements and computing power level, supporting the generation of multiple scenes and diverse user interaction requirements. In addition, a user independent adjustment function is provided during the interaction process to enhance the user experience and satisfaction.

[0107] See Figure 6 , Figure 6 It is a structural block diagram of a three-dimensional scene layout system provided by the embodiments of the present application.

[0108] The present application also provides a three-dimensional scene layout system, and the system includes a control module, and the control module is used to execute any one of the above methods.

[0109] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned method according to any one of the embodiments is implemented.

[0110] An embodiment of the present application also provides a computer program product including a computer program, and when the computer program is executed by a processor, the above-mentioned method according to any one of the embodiments is implemented.

[0111] The computer program product may be a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the computer program product of the embodiments of the present application is not limited thereto, and the computer program product may adopt any combination of one or more computer-readable media.

[0112] See Figure 7 , Figure 7 is a structural block diagram of a computer device provided by an embodiment of the present application.

[0113] An embodiment of the present application also provides a computer device including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the above-mentioned method according to any one of the embodiments is implemented.

[0114] The computer device in the embodiments of the present application is not limited, and it may be, for example, a local computer device, a cloud computer device, a distributed computer device, etc.

[0115] The computer device may include: a memory 110, a processor 120, and a communication interface 130. Among them, the memory 110, the processor 120, and the communication interface 130 are connected through an internal connection path.

[0116] The memory 110 is used to store a computer program. In some implementation manners, the computer program may include code for implementing the method of the embodiments of the present application.

[0117] The processor 120 is used to execute the computer program stored in the memory 110 to control the communication interface 130 to receive input data and information and output operation result data, etc. In some implementation manners, when implementing the solution of the embodiments of the present application through software or firmware, the computer program for implementing the solution of the embodiments of the present application may be stored in the processor 120 and executed by the processor 120.

[0118] The memory 110 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM). It should be noted that the memory 110 described herein is intended to include, but is not limited to, any of these and other suitable types of memories. As an example, the memory 110 includes a random access memory (RAM), a cache memory, and a read-only memory (ROM). Among them, the memory 110 stores a computer program, and the computer program can be executed by the processor 120, so that the processor 120 implements the steps of any of the above methods.

[0119] The processor 120 may be a central processing unit (CPU), and the processor 120 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor 120 may also be any conventional processor, etc.

[0120] In the implementation process, each step of the above method may be completed by the integrated logic circuit of the hardware in the processor 120 or the instructions in the form of software. The method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor 120. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 110, and the processor 120 reads the information in the memory 110 and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0121] In some implementations, in addition to the hardware units described above, a computer device may further include software modules. Among them, the software modules may be, for example, an operating system, a Basic Input Output System (BIOS), application software, etc.

[0122] The operating system is used to manage the hardware and / or software resources of a computer device and is the core and cornerstone of the computer device. The operating system needs to handle basic tasks such as managing and configuring memory, determining the priority order of system resource supply and demand, controlling input and output devices, operating the network, and managing the file system. To facilitate user operation, most operating systems provide an operation interface for users to interact with the system.

[0123] The BIOS is used to run hardware initialization during the power-on boot phase and provide runtime services for the operating system and application programs. In some implementations, the BIOS can also monitor the display processor temperature and perform functions such as adjusting the temperature protection strategy.

[0124] Application software, also known as an application program, can be understood as software written for a specific application purpose of users and is one of the main classifications of computer software. For example, the application software can be a program for achieving purposes such as power control and temperature management.

[0125] It can be understood that the specific examples in the embodiments of the present application are only to help those skilled in the art better understand the implementation manners of the embodiments of the present application, rather than limiting the protection scope of the embodiments of the present application.

[0126] It can be understood that in various implementation manners of the embodiments of the present application, the magnitudes of the sequence numbers of the various processes do not mean the order of execution, and the execution order of the various processes should be determined according to their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0127] It can be understood that the various implementation manners described in the embodiments of the present application can be implemented alone or in combination, and the embodiments of the present application do not limit this.

[0128] Unless otherwise specified, all technical and scientific terms used in the embodiments of this application have the same meanings as those commonly understood by those skilled in the technical field of the embodiments of this application. The terms used in the embodiments of this application are only for the purpose of describing specific embodiments, and are not intended to limit the scope of the embodiments of this application. The term "and / or" used in the embodiments of this application includes any and all combinations of one or more of the related listed items. The singular forms "a", "above-mentioned", and "the" used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0129] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of this application.

[0130] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described embodiments can refer to the corresponding processes in other embodiments, and will not be elaborated herein.

[0131] In several embodiments provided by the embodiments of this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0132] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the technical solution of the embodiments of this application.

[0133] In addition, in each embodiment of the embodiments of this application, the functional units can be integrated into one processing unit, or each unit exists physically alone, or two or more units are integrated into one unit.

[0134] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0135] The above is only the specific implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the embodiments of the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application shall be subject to the protection scope of the claims.

Claims

1. A three-dimensional scene layout method, characterized in that: The method comprises: Determining a constraint cost function corresponding to the target object set based on a position constraint relationship corresponding to the target object set in the target area; Constructing a posture optimization function corresponding to the target object set according to a collision cost function and a constraint cost function corresponding to the target object set; Using the pose optimization function corresponding to the target object set, solving the pose of the objects in the target object set; Based on the positions and postures of the objects in the target area, three-dimensional scene information of the target area is generated.

2. The three-dimensional scene layout method according to claim 1, characterized in that: The target object set is used to form one or more object pairs, each object pair includes an active object and a passive object relative to each other, and the position constraint relationship of the object pairs includes at least one of the following: the active object is located on the passive object; The active object is located inside the body of the animal; The active object is located adjacent to the passive object.

3. The three-dimensional scene layout method according to claim 2, characterized in that: In the case where the active object is located on the passive object, the constraint cost function of the position constraint relationship includes a penalty cost function and a distance cost function, and the penalty cost function is determined according to the bottom feature point of the active object and the top feature point of the passive object in the corresponding position constraint relationship; or, In the case where the active object is located in the body of the animal, the constraint cost function of the position constraint relationship includes a penalty cost function and a distance cost function, and the penalty cost function is determined according to the top feature point of the active object and the top feature point and the bottom feature point of the passive object in the corresponding position constraint relationship; or, When the active object is adjacent to the passive object, the constraint cost function of the position constraint relationship is determined according to the distance between the active object and the passive object in the corresponding position constraint relationship, the longest radius of the active object on the first plane, and the longest radius of the passive object on the first plane.

4. The three-dimensional scene layout method according to claim 1, characterized in that: The target object set has a position constraint relationship with the objects placed in the target area; The using the posture optimization function corresponding to the target object set to solve the posture of the objects in the target object set includes: The poses of the objects in the target object set are solved according to the pose optimization functions corresponding to the target object set and the poses of the placed objects.

5. The three-dimensional scene layout method according to claim 4, characterized in that: The placed object is a ground object, and the process of determining the position and posture of the ground object includes: When the number of ground objects is greater than 1, the position and posture of each ground object is determined in turn according to the priority of each ground object; The priority of the ground object is determined according to at least one of the following: whether it has a position constraint relationship with other ground objects; its size on the layout plane; priority setting information; and the scene type of the target area.

6. The three-dimensional scene layout method according to claim 5, characterized in that: The process of determining the position and posture of the ground object further includes: Determining, according to the positions and postures of at least two ground objects having a position constraint relationship, a minimum circumscribed rectangle of the at least two ground objects on the layout plane; The minimum circumscribed rectangle is used as a new ground object to optimize the position and posture of part or all of the ground objects.

7. The three-dimensional scene layout method according to claim 1, characterized in that: The process of acquiring the position constraint relationship of the objects in the target area includes: In response to user interaction information, determine the objects in the target area and the relative position relationship between the objects; the user interaction information includes one or more of natural language information, interactive drawing information and picture information corresponding to the target area; The position constraint relationship of each object is determined according to the objects in the target area and the relative position relationship between the objects.

8. The three-dimensional scene layout method according to claim 7, characterized in that: In a case where the user interaction information includes natural language information corresponding to the target area, determining objects in the target area and relative positional relationships between objects in response to the user interaction information includes: Parsing the natural language information using a large language model to generate scene description information; The scene description information is semantically analyzed to obtain objects in the target area and relative positional relationships between the objects.

9. The three-dimensional scene layout method according to claim 7, characterized in that: In a case where the user interaction information includes the interactive drawing information corresponding to the target area, determining the objects in the target area and the relative positional relationship between the objects in response to the user interaction information includes: In response to the interactive drawing information, performing an object drawing operation on a two-dimensional plane; In response to an editing operation on the object in the two-dimensional plane, the object in the two-dimensional plane is edited to obtain the objects in the target area and the relative positional relationship between the objects.

10. The three-dimensional scene layout method according to claim 7, characterized in that: In a case where the user interaction information includes picture information corresponding to the target area, determining objects in the target area and relative positional relationships between objects in response to the user interaction information includes: Process the image information using an image processing model to determine an object in the image information; Using a detection model to detect the position of the object in the image information; generating a corresponding representation of the object on a two-dimensional plane based on a position of the object in the image information; In response to an editing operation on the object in the two-dimensional plane, the object in the two-dimensional plane is edited to obtain the objects in the target area and the relative positional relationship between the objects.

11. The three-dimensional scene layout method according to claim 1, characterized in that: The using the posture optimization function corresponding to the target object set to solve the posture of the objects in the target object set includes: Solving the poses of objects in the target object set according to the pose optimization function and object information corresponding to the target object set; The object information is retrieved from an object asset library, or the object information includes one or more of a three-dimensional model and size information.

12. A three-dimensional scene layout system, characterized in that: The system comprises a control module, and the control module is used to execute the method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

14. A computer device, characterized in that: The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 11 when executing the computer program.

15. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.