Interaction intention driven three-dimensional asset sub-database retrieval method for embodied scenario generation

CN122817439APending Publication Date: 2026-09-25CHONGQING RES INST OF HARBIN UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611026275.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]针对现有三维资产检索方法的检索结果在物理交互能力属性上与自然语言描述中的动作需求之间存在失配的问题,本申请提供一种面向具身场景生成的交互意图驱动三维资产分库检索方法

Benefits of technology

[0039]本申请的有益效果,本申请首先对自然语言中的动作意图进行结构化解析,通过预定义映射函数将动作词转化为可供性标签集合与关节类型集合;其次根据是否具有交互需求执行分库路由,使具有交互需求的对象优先从可交互资产库检索,避免大规模静态资产的系统性压制;随后对可交互资产候选执行交互能力硬过滤,排除关节类型与目标动作不匹配的伪交互资产;进而针对两类资产的特性差异采用差异化评分函数进行排序;最后通过尺度适配过滤排除几何尺度显著不合理的候选,输出满足语义、交互能力与尺度约束的Top-K候选集合。其中:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817439A_ABST
    Figure CN122817439A_ABST
Patent Text Reader

Abstract

The application discloses an interaction intention driving three-dimensional asset sub-database retrieval method for embodied scene generation, solves the problem that the existing retrieval result is mismatched between the physical interaction capability attribute and the action demand in the natural language description, and belongs to the cross technical field of information retrieval and three-dimensional scene generation. The application comprises the following steps: acquiring a natural language scene description and analyzing the natural language scene description to determine object demands and corresponding interaction instruction variables; mapping the interaction action demands of the object demands into interaction constraint conditions, including a target availability label set and a joint type set; performing routing according to the interaction instruction variables to obtain an initial candidate set, the availability label set of the candidate objects in the initial candidate set completely covers the target availability label set, and the candidate objects whose joint types belong to the joint type set are determined to obtain a legal candidate subset; and scoring and sorting the candidate objects in the legal candidate subset according to different scoring functions according to the asset sources of the candidate objects to obtain a retrieval result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to an interactive intent-driven 3D asset retrieval method for embodied scene generation, belonging to the interdisciplinary field of information retrieval and 3D scene generation. Background Technology

[0002] Language-driven 3D scene generation is an important research direction in the fields of artificial intelligence and 3D content generation. Its core task is to automatically retrieve and organize 3D assets based on natural language descriptions to construct indoor scenes. Existing mainstream scene generation systems suffer from the following three key technical limitations in asset retrieval methods:

[0003] (1) Existing methods generally adopt a two-stage paradigm of object category recognition plus semantic similarity ranking. Object category words are extracted from natural language, and candidates are then selected based on text or image-text similarity. This process assumes that the key attributes of an object are determined by its category and appearance, and does not parse the action semantics in natural language. When the user description includes "a cabinet that can be opened", the system still only searches based on the category word "cabinet", and the returned candidates are static mesh models that have a suitable appearance but do not have joint structures.

[0004] (2) In practical applications, it is necessary to utilize both large-scale static asset libraries (such as Objaverse, with an order of millions) and interactive asset libraries (such as PartNet-Mobility, with an order of thousands). Existing methods directly put the two types of assets into a unified candidate pool for ranking. Static assets have an overwhelming advantage in quantity. Even if interactive assets are highly matched to the target action requirements, they still score lower in the ranking competition due to their quantity disadvantage. Objects with interactive requirements are ultimately retrieved as static models that cannot be operated.

[0005] (3) There are cases where the joint types, joint constraints, and availability tags of objects in the interactive asset library do not match the specific action requirements. Existing methods take the entire interactive asset library as a candidate source and sort it only by category and similarity. Objects with joint types that do not match the target action, abnormal joint constraints, or missing availability tags are ranked high in the sorting because of their similar appearance or text description. Even after being selected, they still cannot truly support the target interactive task. Summary of the Invention

[0006] To address the mismatch between the physical interaction capabilities and the action requirements described in natural language in existing 3D asset retrieval methods, this application provides a 3D asset database retrieval method driven by interactive intents for embodied scene generation.

[0007] This application presents a method for interactive intent-driven 3D asset retrieval based on embodied scene generation, comprising:

[0008] Natural language scene descriptions are obtained and parsed to determine the needs of each object and the corresponding interaction indicator variables. The interaction indicator variables are used to indicate whether the object's needs include interactive action requirements.

[0009] For object requirements where the interaction indicator variable is true, the interaction action requirement is mapped to an interaction constraint, which includes a set of target availability labels and a set of joint types that match the interaction action requirement.

[0010] Routing is performed based on the interaction indicator variable: if the interaction indicator variable is true, candidate objects are retrieved from the interactive asset library; if the interaction indicator variable is false, candidate objects are retrieved from the static asset library to obtain an initial candidate set.

[0011] Perform an interaction capability legality determination on the candidate objects in the initial candidate set: if the actual availability tags of the candidate object completely cover the target availability tag set, and the joint type of the candidate object belongs to the joint type set, and the joint movement limit range is legal, the determination is passed; otherwise, it is eliminated, and a legal candidate subset is obtained.

[0012] For the candidate objects in the legal candidate subset, different scoring functions are used to score and sort them according to their asset source; the candidate objects after scoring and sorting are output in descending order of score to obtain the search results.

[0013] Preferably, mapping the interaction action requirements to interaction constraints includes:

[0014] When the interaction action requirement is an opening and closing action, the set of mapped target availability tags includes opening and closing availability tags, and the set of mapped joint types includes rotary joints or sliding joints.

[0015] When the interaction action requirement is a push-pull action, the mapped target availability tag set includes push-pull type availability tags, and the mapped joint type set includes sliding joints;

[0016] When the interaction action requirement is a rotation action, the set of mapped target availability tags includes rotation-type availability tags, and the set of mapped joint types includes rotation joints;

[0017] When the interaction action requirement is a switch toggle action, the mapped target availability tag set includes switch toggle type availability tags, and the mapped joint type set includes rotary joints.

[0018] As a preferred option, the object requirements are:

[0019] ;

[0020] in For object category labels, For the number of objects, As a scenario prior, The target scale vector. A set of keywords for interactive action requirements;

[0021] Based on the set of keywords for interactive action requirements Calculate interaction indicator variables :

[0022] ;

[0023] Generate object-level query structure , The set of availability tags for the target;

[0024] Execute routing based on the interaction indicator variables:

[0025] If the interaction indicator variable The candidate pool is routed to the interactive asset library. If the interactive indicator variable The candidate pool is routed to the static asset repository. ;

[0026] In the candidate pool after routing, based on the object category label Perform basic category filtering based on the aforementioned scenario priors. Execution scenario filtering based on the category tags The corresponding placement type is inferred, and the support surface type label is matched and filtered based on the inference result to obtain an initial candidate set.

[0027] Preferably, when the number of candidate objects that meet the screening criteria in the interactive asset library is lower than a preset threshold, candidate objects are added from the static asset library, and the added candidate objects are marked as fallback assets.

[0028] Preferably, the interactive asset library is derived from a 3D asset library including joint structures, while the static asset library is derived from a general 3D asset library.

[0029] As a preferred approach, ranking assets based on their source using different scoring functions includes:

[0030] For candidate objects derived from the static asset library, the scoring function includes a textual semantic similarity term, a visual similarity term, and a geometric scale penalty term;

[0031] For candidate objects derived from the interactive asset library, the scoring function includes a textual semantic similarity term and a geometric scale penalty term;

[0032] The geometric scale penalty term is based on the target scale vector. Penalty for deviations from the actual size of the candidate object.

[0033] Preferably, the method of this application further includes scale-adaptive filtering after the scoring and sorting:

[0034] Calculate the actual size of the candidate object and the target scale vector. The relative error between them is used to eliminate candidate objects whose relative error exceeds a preset threshold;

[0035] The scale-adapted filtered candidates are output in descending order of their scores.

[0036] Preferably, the output search results include K candidate objects corresponding to each object requirement, where K is greater than 1.

[0037] Preferably, the output information for each candidate object in the search results includes:

[0038] The unique identifier and source mark of the asset, category information, comprehensive score and modal similarity decomposition value, availability tag and joint type and motion limit, 3D bounding box and 2D footprint shape, and whether it is a catch-all asset.

[0039] The beneficial effects of this application are as follows: First, it performs structured parsing of action intentions in natural language, transforming action words into a set of availability tags and a set of joint types using a predefined mapping function. Second, it performs database routing based on whether there is an interaction requirement, prioritizing the retrieval of objects with interaction requirements from the interactive asset library, avoiding the systematic suppression of large-scale static assets. Then, it performs hard filtering of interactive capability on the interactive asset candidates, eliminating pseudo-interactive assets whose joint types do not match the target action. Furthermore, it uses a differentiated scoring function to rank the two types of assets based on their characteristic differences. Finally, it uses scale adaptation filtering to eliminate candidates with significantly unreasonable geometric scales, outputting a Top-K candidate set that satisfies semantic, interactive capability, and scale constraints. Wherein:

[0040] First, by explicitly transforming the action intent in natural language into programmable interactive constraints (target availability tag set and joint type set), the problem of existing retrieval methods being unable to perceive interactive needs is fundamentally solved.

[0041] Second, the database routing mechanism uses the presence of interactive requirements as the primary criterion for selecting candidates, fundamentally avoiding the systematic suppression of interactive assets by large-scale static assets and ensuring that objects with interactive requirements are preferentially retrieved from the interactive asset database.

[0042] Third, the differentiated scoring function for the two types of assets is designed with different scoring weight strategies to avoid the scoring distortion caused by the heterogeneity of feature distribution in a unified scoring space, thereby improving the interpretability and stability of the scoring results.

[0043] Fourth, the hard filtering mechanism for interactive capabilities accurately eliminates pseudo-interactive assets in the interactive library, ensuring the authenticity of the final candidate functions, rather than just meeting the matching requirements at the appearance or category level.

[0044] Fifth, the fallback mechanism ensures the system's fault tolerance when there is insufficient coverage of interactive assets, maintaining the integrity and continuity of scene generation.

[0045] This application enables reliable asset retrieval based on embodied needs on a large-scale heterogeneous asset database, and is applicable to fields such as the construction of embodied intelligent simulation environments and the configuration of robot task scenarios. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating the method described in this application;

[0047] Figure 2 This is a schematic diagram illustrating the principle of the database sharding routing and differentiated dual-track scoring strategy in this application;

[0048] Figure 3 This is a logical diagram illustrating the hard filtering of interactive capabilities and the scale adaptation filtering in the embodiments of this application. Detailed Implementation

[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0050] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0051] The present application will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the application.

[0052] In existing 3D asset retrieval methods, there is a mismatch between the physical interactivity attributes of the search results and the action requirements in the natural language description. Specifically: in retrieval processes ranked by category and appearance similarity, action semantics are not involved in the retrieval decision; static assets with matching categories and appearances but lacking joint structures are confused with objects requiring actions; when large-scale static assets compete with interactive assets in a unified ranking space, assets with interactive attributes rank lower due to their numerical disadvantage; in search results from interactive asset libraries, objects with joint types inconsistent with action requirements, abnormal joint constraints, or missing availability tags are still ranked highly due to appearance or textual similarity. These phenomena result in a failure to establish a correspondence between the physical interactivity attributes of objects in the search results and the action requirements in the natural language description.

[0053] This embodiment provides a method for retrieving 3D assets based on interactive intents generated for embodied scenarios. This method addresses the technical problem of mismatch between the physical interaction capability attributes of the retrieval results and the action requirements in the natural language description in existing 3D asset retrieval methods.

[0054] The overall execution flow of this method is as follows: First, the user-input natural language scene description is obtained. This description is then parsed into a structured set of object requirements using a large language model. Interaction indicator variables are determined for each object requirement to distinguish whether the object has interactive action requirements. For objects with interactive action requirements, action keywords are mapped to a set of availability tags and a set of joint types using a predefined mapping function, forming programmable interaction constraints. Based on the interaction indicator variables, a database routing process is performed, directing the search candidate pool to either an interactive asset library or a static asset library. Category filtering and scenario prior filtering are then performed in the routed asset library to obtain an initial candidate set. For each candidate object in the initial candidate set, an interaction capability legality determination is performed. Through joint verification of availability tag coverage conditions, joint type matching conditions, and joint movement limit legality conditions, candidate objects that do not meet the interaction constraints are eliminated, resulting in a legal candidate subset. Candidate objects are ranked and scored using a differentiated scoring function based on their asset source. The scoring function for candidate objects from the interactive asset library does not include a visual similarity term. Finally, the K candidate assets corresponding to each object requirement are output in descending order of score as the search results for use by the downstream layout optimization module.

[0055] This method is applicable to application scenarios that require accurate retrieval of interactive objects from heterogeneous 3D asset libraries, such as the construction of embodied intelligent simulation environments, language-driven 3D scene generation, and robot operation task scene configuration.

[0056] The interactive intent-driven 3D asset retrieval method for embodied scene generation in this embodiment is characterized by including:

[0057] Step 1: Obtain and parse the natural language scene description to determine the needs of each object and the corresponding interaction indicator variables. The interaction indicator variables are used to indicate whether the object's needs include interactive action requirements.

[0058] The system receives natural language scene descriptions input by users. The scene descriptions are textual expressions of indoor scene requirements, such as "a kitchen with cabinet doors that can be opened" or "a living room with a pull-out drawer and a floor lamp".

[0059] The system inputs a natural language scene description into a large language model. This large language model is a pre-trained generative language model with natural language understanding and structured output capabilities. The system configures the large language model with structured output instructions, requiring it to output parsing results according to a preset structured data format. The large language model performs semantic parsing on the scene description, identifies the needs of each object included in the scene, and generates corresponding quintuple data for each object need, thus completing the object need set. The parsing process. Each object requires a 5-tuple. The definition of is:

[0060]

[0061] in, Label the object category (e.g., cabinet, drawer, faucet). For the number of objects, For the scenario prior (e.g., kitchen, bathroom, living room). The target scale range or size of the object is a priori. This represents the interaction requirements of an object; when an object has no explicit interaction, it is an empty set.

[0062] The language model is output in structured JSON format, with all five fields for each object generated at once to ensure semantic consistency between fields. The system receives and parses the structured output to obtain the set of object requirements. .

[0063] For each object requirement in the object requirement set The system extracts its action requirement keyword set. And based on the keyword set for action requirements Calculate the corresponding interaction indicator variable Interactive indicator variables The determination rule is as follows:

[0064]

[0065] When action requirement keyword set When it is a non-empty set, This indicates that the object's requirements include interactive action requirements; further, the target action set is determined. : Among them, open and close correspond to the opening and closing actions of objects such as cabinet doors and doors; pull corresponds to the pulling action of drawer objects; rotate corresponds to locally rotating objects such as faucet and knob; and switch corresponds to small-angle rotation or state switching of switch objects.

[0066] When action requirement keyword set empty set hour, This indicates that the requirements for this object do not include interactive action requirements.

[0067] Step 2: For object requirements where the interaction indicator variable is true, map the interaction action requirements to interaction constraints. The interaction constraints include a set of target availability labels and a set of joint types that match the interaction action requirements, and generate a structured query.

[0068] for Each object requirement The system extracts its action requirement keyword set. and input it into a predefined mapping function. Predefined mapping functions This includes a predefined semantic mapping rule table, which establishes the correspondence between natural language action keywords and the physical interaction attributes of 3D assets. Mapping function. Set of action requirement keywords The mapping is to interactive constraints, and the output format is:

[0069]

[0070] in For the target availability tag set, A set of joint types that match the needs of the movement.

[0071] When action requirement keyword set When the code includes the action words "open" or "close", the mapping function... Convert it to: This includes the availability tags "openable" or "closable". This includes joint types such as "revolute" or "prismatic".

[0072] When action requirement keyword set When the text includes the action words "pull" or "push", the mapping function... Convert it to: This includes the availability tags "pullable" or "pushable". This includes the joint type "prismatic".

[0073] When action requirement keyword set When the action word "rotate" is included, the mapping function... Convert it to: Including the availability tag "rotatable", This includes the joint type "revolute".

[0074] When action requirement keyword set When the action word "switch" is included, the mapping function... Convert it to: Including the availability tag "switchable", This includes the joint type "revolute".

[0075] In the above mapping rules, the applicable object categories for each action are as follows: open / close applies to object categories with opening and closing structures such as cabinet, door, and frame; pull / push applies to object categories with drawer structures such as drawer and drawer_cabinet; rotate applies to object categories with local rotation structures such as faucet, knot, and handle; and switch applies to object categories with switch structures such as light_switch and socket. The applicable object categories have already been defined as category labels in step 1. This is retrieved but not used as the output of the mapping function.

[0076] After completing category identification ( ), interaction judgment ( ) and action mapping ( After that, the object requirements are standardized into a retrieval-oriented object-level query structure:

[0077]

[0078] Not written directly Instead, it is used as an independent judgment condition for subsequent interactive capability filtering, in order to maintain the separation of responsibilities between query structure and filtering logic.

[0079] Step 3: Execute routing based on interaction indicator variables: If the interaction indicator variable is true, retrieve candidate objects from the interactive asset library; if the interaction indicator variable is false, retrieve candidate objects from the static asset library to obtain the initial candidate set.

[0080] For each object requirement The system uses its corresponding interaction indicator variables. Execute routing, define the routing function:

[0081]

[0082] like The system will route the candidate pool to the interactive asset library. Interactive asset library This is a 3D asset library containing joint structures, where each asset is labeled with an availability tag, joint type, and joint movement limit information; if The system will route the search candidate pool to the static asset repository. The static asset library is a large-scale, general-purpose 3D asset library used to construct static scene elements.

[0083] Within the asset library determined by the route, the system uses the category tags specified in the object requirements. Perform category filtering:

[0084] Label the category of each asset in the asset pool with Compare and retain only the category label and... Consistent candidate objects. Based on this, the system considers the scenario priors in the object requirements. Prior screening of execution scenarios:

[0085] Extract the room_priors field for each candidate object and compare it with the scene priors. Perform an intersection check, retaining only the `room_priors` field and... Candidate objects that intersect. Based on this, the system performs placement type filtering according to the support surface type label of the candidate objects:

[0086] According to category labels The system infers the corresponding placement type requirements and then matches and filters the support surface type labels of candidate objects, eliminating those with mismatched support surface types. Specifically, the candidate list is further narrowed down using the `support_type` field. The `support_type` field is a predefined object placement support surface type label in the asset library, used to characterize the appropriate placement location category of the object in the scene, including but not limited to floor, wall, table, or shelf.

[0087] After the above routing decisions and filtering, the system obtains an initial candidate set.

[0088] When the number of candidate objects that meet the filtering criteria and are routed to the interactive asset library falls below a preset minimum threshold (default 3), the system supplements the candidate objects from the static asset library to ensure the integrity of the scene generation. The supplemented candidate objects are marked as fallback assets in the output information. The fallback mechanism does not change the basic principle that objects with interactive requirements are retrieved first from the interactive asset library.

[0089] Step 4: Perform interaction capability legality judgment on candidate objects in the initial candidate set: if the actual availability tags of the candidate object completely cover the target availability tag set, and the candidate object's joint type belongs to the joint type set, and the joint movement limit range is legal, the judgment is passed; otherwise, it is eliminated, and a legal candidate subset is obtained.

[0090] The scoring phase only performs soft preference ranking, which cannot guarantee that high-scoring candidates meet functional constraints. Explicit filtering is required to execute the final legality determination. (See below) Figure 3 For each candidate object in the initial candidate set The system performs a validity check on the interactive capabilities. The function for checking the validity of interactive capabilities is as follows:

[0091]

[0092] First, availability label coverage conditions: the system extracts candidate objects. The actual set of available tags Combine it with the target availability tag set generated in step 2. A comparison will be performed. Candidates must meet the following conditions to pass this evaluation. This means that the actual set of available labels possessed by the candidate object completely covers all labels in the target set of available labels.

[0093] Second, joint type matching conditions: the system extracts candidate objects. The actual joint types are compared with the set of joint types generated in step 2. A comparison is performed. For a candidate object to pass this condition, its joint type must belong to the set of joint types.

[0094] Third, the legality conditions for joint movement limitation: the system extracts candidate objects. The actual range of joint movement is determined, and its validity is assessed. If the candidate joint is a rotational joint, its range of joint movement must be within a preset valid angle range. The preset valid angle range is... If the joint type of the candidate object is a sliding joint, its range of motion must be greater than 0. Candidate objects without a joint structure are directly judged as not meeting the legality conditions for joint movement limit.

[0095] For each candidate object in the initial candidate set, the system performs a joint judgment based on the above three conditions: the candidate object passes the interaction capability validity judgment if and only if all three conditions are met; if any one of the three conditions is not met, the system removes the candidate object from the candidate set. After this hard filtering, the candidate objects retained by the system constitute a valid candidate subset. This step often eliminates pseudo-interactive assets that come from interactive libraries but whose joint types do not match the target action—for example, a cabinet from the PM library that only has fixed joints and cannot perform opening and closing actions.

[0096] Step 5: For the candidate objects in the legitimate candidate subset, score and rank them according to their asset sources using different scoring functions;

[0097] For legitimate candidate subsets For each candidate object in the database, the system first identifies that the candidate object originates from the interactive asset library. Still a static asset library Then, based on its source, the corresponding scoring function is used to calculate the comprehensive score.

[0098] For candidate objects derived from the static asset repository, the system uses the following scoring function:

[0099]

[0100] in This refers to the text semantic similarity calculated based on a text embedding model. Visual similarity is calculated based on an image embedding model. This is a geometric scale penalty term, which is based on the target scale vector. Penalty for deviations from the actual size of the candidate object. , , These are text semantic weight (default 0.4), visual weight (default 0.5), and size penalty weight (default 0.1).

[0101] For candidate objects from the interactive asset library, the system uses the following scoring function:

[0102]

[0103] in, For interactive asset text semantic weights, This is a penalty weight for the size of interactive assets. The interactive asset scoring function does not include a visual similarity term.

[0104] The system selects a subset of legitimate candidates. After all candidates have completed their scoring calculations, they are sorted in descending order of their overall scores.

[0105] After the scoring and ranking but before the output, the system performs scale-adaptive filtering: for the valid candidate subset... Each candidate object in Calculate the three-axis relative error between it and the object requirements:

[0106]

[0107] in, The target scale vector for the object requirements. This is the actual size vector of the candidate object. Traverse the three spatial axes x, y, and z. ( Candidates are filtered based on a scale threshold (default 0.2), removing those with significantly unreasonable scales (such as oversized faucets or undersized wardrobes).

[0108] This step is a coarse-grained preliminary screening, the goal of which is to eliminate significantly unreasonable objects; the subsequent layout optimization stage further refines the layout under the constraints of specific room geometry and object relationships. The two constitute a relationship of preliminary contraction and subsequent optimization, rather than repetitive design.

[0109] Step 6: Output the candidate objects after rating sorting in descending order of rating to obtain the search results.

[0110] The system extracts the top K candidate objects from the qualified candidate subset that has been scored and sorted and filtered by scale adaptation, in descending order of score, to form the Top-K candidate set corresponding to each object requirement:

[0111]

[0112] The system aggregates the Top-K candidate sets corresponding to each object requirement, forming an object-asset candidate binding set, which is then output as the search result. The results are selected as the optimal candidates. The output information for each candidate includes: a unique asset identifier and source tag, category information, a comprehensive score and modal similarity decomposition values, availability tags, joint types and motion limits, a 3D bounding box and a 2D footprint shape, and an adaptation tag indicating whether it is a fallback asset. The search results can be directly used as input data for subsequent scene layout generation.

[0113] K is a preset integer greater than 1. The reason for retaining Top-K (K>1) instead of just Top-1 is that in the subsequent layout stage, due to spatial conflicts or local constraints, it may be necessary to replace the best candidate with the second-best candidate. Retaining multiple high-quality alternatives can significantly reduce the backtracking cost of scene generation.

[0114] The candidate set of all objects is aggregated to obtain the object-asset candidate binding set U, which serves as the input for the subsequent layout optimization module. Based on this, the subsequent layout optimization module further considers the spatial constraints and action clearance requirements of the objects to complete the final scene space arrangement.

[0115] While this application has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of this application. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of this application as defined by the appended claims. It should be understood that different dependent claims and features herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other embodiments.

Claims

1. A method for retrieving 3D assets based on interactive intents generated for embodied scenarios, characterized in that: include: Natural language scene descriptions are obtained and parsed to determine the needs of each object and the corresponding interaction indicator variables. The interaction indicator variables are used to indicate whether the object's needs include interactive action requirements. For object requirements where the interaction indicator variable is true, the interaction action requirement is mapped to an interaction constraint, which includes a set of target availability labels and a set of joint types that match the interaction action requirement. Routing is performed based on the interaction indicator variable: if the interaction indicator variable is true, candidate objects are retrieved from the interactive asset library; if the interaction indicator variable is false, candidate objects are retrieved from the static asset library to obtain an initial candidate set. Perform an interaction capability legality determination on the candidate objects in the initial candidate set: if the actual availability tags of the candidate object completely cover the target availability tag set, and the joint type of the candidate object belongs to the joint type set, and the joint movement limit range is legal, the determination is passed; otherwise, it is eliminated, and a legal candidate subset is obtained. For the candidate objects in the legal candidate subset, different scoring functions are used to score and sort them according to their asset source; the candidate objects after scoring and sorting are output in descending order of score to obtain the search results.

2. The interactive intent-driven 3D asset retrieval method for embodied scene generation according to claim 1, characterized in that, Mapping the interaction action requirements to interaction constraints includes: When the interaction action requirement is an opening and closing action, the set of mapped target availability tags includes opening and closing availability tags, and the set of mapped joint types includes rotary joints or sliding joints. When the interaction action requirement is a push-pull action, the mapped target availability tag set includes push-pull type availability tags, and the mapped joint type set includes sliding joints; When the interaction action requirement is a rotation action, the set of mapped target availability tags includes rotation-type availability tags, and the set of mapped joint types includes rotation joints; When the interaction action requirement is a switch toggle action, the mapped target availability tag set includes switch toggle type availability tags, and the mapped joint type set includes rotary joints.

3. The interactive intent-driven 3D asset retrieval method for embodied scene generation according to claim 1, characterized in that, The requirements for the object are as follows: ; in For object category labels, For the number of objects, As a scenario prior, The target scale vector. A set of keywords for interactive action requirements; Based on the set of keywords for interactive action requirements Calculate interaction indicator variables : ; Generate object-level query structure , The set of availability tags for the target; Execute routing based on the aforementioned interaction indicator variables: If the interaction indicator variable The candidate pool is routed to the interactive asset library. If the interactive indicator variable The candidate pool is routed to the static asset repository. ; In the candidate pool after routing, based on the object category label Perform basic category filtering based on the aforementioned scenario priors. Execution scenario filtering based on the category tags The corresponding placement type is inferred, and the support surface type label is matched and filtered based on the inference result to obtain an initial candidate set.

4. The interactive intent-driven 3D asset retrieval method for embodied scene generation according to claim 3, characterized in that, When the number of candidate objects that meet the screening criteria in the interactive asset library is lower than a preset threshold, candidate objects are added from the static asset library, and the added candidate objects are marked as fallback assets.

5. The interactive intent-driven 3D asset retrieval method for embodied scene generation according to claim 3, characterized in that, The interactive asset library is derived from a 3D asset library including joint structures, while the static asset library is derived from a general 3D asset library.

6. The interactive intent-driven 3D asset retrieval method for embodied scene generation according to claim 3, characterized in that, The method of ranking and scoring based on different scoring functions according to the source of assets includes: For candidate objects derived from the static asset library, the scoring function includes a textual semantic similarity term, a visual similarity term, and a geometric scale penalty term; For candidate objects derived from the interactive asset library, the scoring function includes a textual semantic similarity term and a geometric scale penalty term; The geometric scale penalty term is based on the target scale vector. Penalty for deviations from the actual size of the candidate object.

7. The interactive intent-driven 3D asset retrieval method for embodied scene generation according to claim 6, characterized in that, The method further includes scale-adaptive filtering after the score ranking: Calculate the actual size of the candidate object and the target scale vector. The relative error between them is used to eliminate candidate objects whose relative error exceeds a preset threshold; The scale-adapted filtered candidates are output in descending order of their scores.

8. The interactive intent-driven 3D asset retrieval method for embodied scene generation according to claim 1, characterized in that, The output search results include K candidate objects corresponding to each object requirement, where K is greater than 1.

9. The interactive intent-driven 3D asset retrieval method for embodied scene generation according to claim 1, characterized in that, The output information for each candidate object in the search results includes: The unique identifier and source mark of the asset, category information, comprehensive score and modal similarity decomposition value, availability tag and joint type and motion limit, 3D bounding box and 2D footprint shape, and whether it is a catch-all asset.

10. A three-dimensional asset retrieval system driven by interactive intents for embodied scene generation, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the interactive intent-driven 3D asset database retrieval method for embodied scene generation as described in any one of claims 1 to 9.