Desktop arrangement method and system for reasoning human preference based on large model

By using Gaussian sputtering method and large-model inference to generate target scene diagrams and action sequences in the desktop sorting system, the problems of insufficient response and hallucination in the existing technology are solved, and more efficient and accurate desktop sorting is achieved.

CN120067362APending Publication Date: 2025-05-30HUNAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510537947.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing large-scale models reason that human preferences are insufficient in response to handling user temporary instructions and new item requests, and are prone to hallucinations, resulting in inaccurate collation results.

Method used

The Gaussian sputtering method is used to model the candidate area scene, combine user instructions and preferred scene maps, and generate the CSS-format two-dimensional layout coordinates of the target scene map and the item through large-scale inference, convert it into actual placement coordinates, and generate an action sequence to complete the desktop organization task.

Benefits of technology

It improves the robot's response to temporary changes, enhances the accuracy of desktop sorting, reduces the inference failure problem caused by overlapping item boundaries, and significantly improves the accuracy of position inference results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067362A_ABST
    Figure CN120067362A_ABST
Patent Text Reader

Abstract

The invention discloses a desktop arrangement method and system for reasoning human preference based on a large model, and the method comprises the steps: firstly, carrying out the modeling of a candidate region scene through employing a Gaussian sputtering method, and obtaining a Gaussian sputtering model; then, obtaining a preference scene graph of a user for the desktop arrangement task, and generating a target state of the desktop arrangement task by using large model reasoning in combination with a Gaussian sputtering model, a user instruction and the preference scene graph, the target state including a target article and a corresponding desktop coordinate; secondly, task planning is carried out based on the target state, and an action sequence from empty desktop arrangement to the target state is generated; then, the system sequentially executes the action sequence according to the planning sequence, and a personalized desktop arrangement task is completed; and after finishing arrangement, storing the target scene graph meeting the user requirement and the corresponding user instruction in pairs in a database. Compared with an existing method, the method has the advantages that the capability of responding to temporary instruction changes is improved, and the accuracy of position reasoning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model inference, and particularly to a desktop arrangement method and system based on large model inference of human preferences. Background Art

[0002] Desktop service robots are mainly aimed at task execution in desktop scenarios and are widely used in environments such as homes and offices. They can not only complete companion tasks such as education and entertainment, but also achieve intelligent arrangement of desktop items. With the rapid development of artificial intelligence technology, desktop arrangement robots combined with human preferences have become a hot research and application direction. Existing human preference inference methods include traditional task-specific deep learning methods and methods based on pre-trained large models in recent years. Traditional preference inference methods mainly rely on deep learning models for personalized inference. However, this method requires building a large personalized database before the task, with high data annotation costs and limited generalization ability, making it difficult to adapt to changing user needs. With the emergence of large models, more and more research has begun to focus on using large models for preference inference. Large models are trained based on large-scale data, have powerful inference capabilities and higher generalization abilities, and can infer more accurate personalized results based on users' historical data, preferences, and behavior patterns.

[0003] However, although existing systems such as 'DegustaBot' or 'TidyBot' introduce large models to handle desktop arrangement tasks, their inference mechanisms are mostly static and difficult to handle the immediate preference changes or new item requests expressed in user instructions. Specifically, the following technical problems exist: (1) Lack of dynamic instruction response. Existing human preference inference methods mainly rely on historical preference data, and their input data often does not cover users' real-time instructions or cannot adjust the desktop layout in a timely manner through user instructions. In specific scenarios, users may temporarily change the desktop layout requirements, and existing methods cannot dynamically adjust the inference results, resulting in the final arrangement plan not meeting the current needs of users.

[0004] (2) The problem of large model inference hallucinations leads to inaccurate arrangement results. Although large models have advantages in inference speed, they are prone to hallucination problems in desktop arrangement tasks. For example, due to co-occurrence biases in the training data, the model may wrongly infer items that "should exist". For example, when arranging a dining table, due to the presence of a fork, it may wrongly add an unnecessary knife; at the same time, in the absence of clear spatial relationship information, the large model may arrange items too densely, resulting in physically infeasible placement plans (such as two items completely overlapping). Summary of the Invention

[0005] To solve the problems that the existing desktop organization using large models to infer human preferences cannot respond to temporary changes and the hallucination problem caused by item co-occurrence during the inference process, the present invention provides a desktop organization method and system based on large model inference of human preferences, which improves the robot's response ability to temporary changes and the accuracy of desktop organization.

[0006] To achieve the above technical objectives, the present invention adopts the following technical solutions: A desktop organization method based on large model inference of human preferences, applied to a desktop organization robot, includes: S1. Collect multi-view images of the candidate area scene, and use the Gaussian sputtering method to model the candidate area scene to obtain the Gaussian sputtering model of the candidate area scene ; S2. Obtain the preference scene map of the user for the desktop organization task ; S3. Take the user instruction , the preference scene map and the Gaussian sputtering model as inputs, call the large model that receives visual information, infer and generate the target scene map, and further infer and generate the two-dimensional layout coordinates in CSS format of each item therein, convert the two-dimensional layout coordinates into the actual placement coordinates in the target desktop; then combine the actual placement coordinates with the item ID to form the target desktop state of the desktop organization task. S4. According to the Gaussian sputtering model , perform task planning based on the target desktop state generated in S3, and generate an action sequence for organizing the desktop from an empty desktop to the target desktop state. S5. Use the robotic arm to execute the action sequence generated in S4, complete the picking and placing of each item in the target desktop state, and complete the desktop organization task.

[0007] Further, after completing the desktop organization task in S5, it also includes: S6. Pair the target scene map obtained in S3 and its corresponding user instruction, and store them in the database of the current user in pairs to update the desktop organization task record in the database.

[0008] Further, the specific process of using the Gaussian sputtering method to model the candidate area scene is as follows: S11: Collect multi-view images of the current desktop scene; S12: Use a large model that receives visual information and, based on multi-view images, infer and generate the categories of various items in the current desktop scene; then, for each item, use the large model to generate a differential name based on the category features; then encode the differential names of each item to obtain the text embedding features of the item; the difference is reflected in: under the inferred item category, describe the feature differences between the sub-category to which the item belongs and other sub-categories. S13: Use the SAM segmentation model to perform semantic segmentation on the multi-view images, and use the CLIP image encoder to extract the image embedding features of each item in the multi-view images. S14: Adopt a multi-modal Transformer model, and use the cross-attention mechanism to fuse the text embedding features and image embedding features of each item in the candidate region scene to obtain multi-modal features. S15: Map the multi-modal features to a low-dimensional space through an autoencoder, and by calculating the mean vector and covariance matrix corresponding to each item, use the mean vector and covariance matrix to describe the language-visual distribution features of each item in the low-dimensional space, and obtain a Gaussian sputtering model with personalized language embedding for the current scene. 。

[0009] Further, in step S2, if it is the first time to perform desktop organization, perform initialization settings based on the organization rules set by the user to obtain a preferred scene graph. ; if it is not the first time to perform desktop organization, then combine the user's instructions and database inference to obtain a preferred scene graph. 。

[0010] Further, combining the user's instructions and database inference to obtain a preferred scene graph The specific process is as follows: Screen several instructions in the database with the highest semantic similarity to the user's instructions to obtain a preset number of scene graphs corresponding to the instructions as example scene graphs. Input the obtained example scene graphs and preset induction prompt words into a large language model, and generate a structured preferred scene graph by inducing item categories, attributes, and spatial relationships. ; in the preferred scene graph, each node represents a category of item and its corresponding attribute label, and the edge represents the relative spatial relationship between the nodes.

[0011] Further, use the BERT model to calculate the semantic similarity, specifically including: Input the user's instructions and the instructions in the database into the BERT model respectively to obtain their semantic vector representations. By calculating the cosine similarity between semantic vectors, several instructions with the highest semantic similarity to the user's instruction are selected; the expression for calculating the cosine similarity between two instructions is as follows: ; Among them, represents the cosine similarity, represents the instruction in the database, represents the user instruction input by the current user.

[0012] Furthermore, the specific process of step S3 is as follows: (1) Use the large model grounding algorithm to obtain the target scene graph , according to the Gaussian sputtering model , the user instruction and the preference scene graph ; (2) Give the large model example data of the position relationship; (3) Input the obtained target scene graph into the large model, and the large model infers the two-dimensional layout representing the absolute position in the CSS format of the item according to the example data of the position relationship ; left and top respectively represent the distances of the rectangle frame from the left and above of the canvas, and width and height are used to define the width and height of the rectangle frame, both in pixels; (4) Map the two-dimensional layout in CSS format to the position coordinates of the item placed on the desktop , and the specific formula is as follows: ; ; Among them, and are the width and height of the CSS canvas, and are the width and height of the target desktop, both in pixels; (5) Combine the item ID and the desktop position coordinates in the obtained target scene graph to form the target desktop state of the desktop sorting task.

[0013] Furthermore, in the target scene graph, it includes the item ID, the position of the item in the candidate area scene, the boundary information of the item, and the relative position relationship between the item and the neighboring items in the target desktop; the relative position relationship between the item and the neighboring items includes: front, back, left, right; the neighboring item refers to another item with the closest position.

[0014] Further, in step S4, domain planning language is used for task planning, and the optimal action sequence for sorting the desktop from an empty desktop to the target desktop state is inferred and generated according to the target desktop state.

[0015] A desktop sorting system based on large model inference of human preferences, which is applied to a desktop sorting robot, includes: an image acquisition module, a Gaussian sputtering modeling module, a preference scene graph acquisition module, a large model module, an action sequence generation module, and an execution module; The image acquisition module is used to: acquire multi-view RGB-D images of the candidate area scene and output image data in a unified format for subsequent processing; The Gaussian sputtering modeling module is used to: perform visual semantic fusion modeling on the items in the acquired images, and generate a Gaussian distribution representation of each item in the low-dimensional space, where the representation includes a mean vector and a covariance matrix; The preference scene graph acquisition module is used to: acquire the preference scene graph of the user for the desktop sorting task ; The large model module is used to: infer and generate a target scene graph according to the Gaussian sputtering model, user instructions, and preference scene graph; The large model module is also used to: take user instructions , preference scene graph and Gaussian sputtering model as inputs, infer and generate a target scene graph, and further infer and generate the two-dimensional layout coordinates in CSS format of each item therein, convert the two-dimensional layout coordinates into the actual placement coordinates in the target desktop; then combine the actual placement coordinates with the item ID to form the target desktop state of the desktop sorting task; The action sequence generation module is used to: perform task planning according to the Gaussian sputtering model and the target desktop state to obtain the action sequence for sorting the desktop from an empty desktop to the target desktop state;

[0016] Compared with the prior art, the beneficial effects of the present invention include: 1. The present invention allows the temporary changes represented by instructions to modify the desktop layout, fills the gap that temporary changes were not considered in the previous inference process, endows the robot with the ability to cope with temporary changes, increases the richness of desktop task input, and improves the desktop sorting efficiency.

[0017] 2. The present invention is based on the scene graph as a method for representing user preferences. This representation method makes part of the robot's inference process visual, increasing the controllability during the inference process. At the same time, using the formatted scene graph can increase the accuracy of inference.

[0018] 3. The present invention uses a novel method for representing positional relationships for desktop organization tasks and uses layout as a method for representing the absolute positions of items. In the layout, multiple CSS-style rectangular boxes are used to represent the absolute positions of different items. Compared with the traditional method of directly using the center points of items as the positional representation, the present invention is beneficial to improving the reasoning accuracy, specifically manifested in: effectively reducing the problem of reasoning failure caused by the overlap of item boundaries.

[0019] 4. The present invention adopts a progressive position reasoning method from relative positional relationships to absolute positional relationships. Among them, the relative positional relationships are represented by a scene graph, and the absolute positional relationships are represented by a layout. The present invention completes the progressive reasoning process of positions through these two representation methods. This progressive reasoning is beneficial to reducing the situation of redundant items caused by the hallucination problem of large models and effectively improving the accuracy of position reasoning results. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a flowchart of the desktop organization method according to an embodiment of the present invention.

[0021] Figure 2 It is a framework diagram of the desktop organization method according to an embodiment of the present invention.

[0022] Figure 3 It is an example diagram of a preferred scene graph according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The following makes a detailed description of the embodiments of the present invention. This embodiment is carried out based on the technical solution of the present invention, and gives a detailed implementation manner and specific operation process, and further explains the technical solution of the present invention.

[0024] Embodiment 1

[0025] This embodiment provides a desktop organization method for inferring human preferences based on a large model, which is applied to a desktop organization robot and can be used for automatically organizing various different types of desktops, such as desks, dining tables, kitchen countertops, workbenches, etc. The method of the present invention can be used for desktop organization of different desktops.

[0026] Refer to Figure 1 、 2 、3, the desktop organization method of this embodiment includes the following steps: S1. Collect multi-view images of the candidate area scene, and use the Gaussian sputtering method to model the candidate area scene to obtain the Gaussian sputtering model of the candidate scene .

[0027] As used in this invention, the "Gaussian sputtering method" refers to a method of constructing a Gaussian distribution representation for each item in a low-dimensional embedding space by fusing the language features and image features of the item. This representation consists of a mean vector and a covariance matrix, which can characterize the distribution characteristics of the item in the semantic-visual joint space, thereby enhancing the system's ability to differentially perceive items in the candidate region. Specifically, it includes: (1) Collect multi-view images of the candidate region scene using a multi-view image acquisition method for Gaussian sputtering processing to ensure the accuracy and adaptability of subsequent calculations. By fixing the camera, collect RGB-D images of the candidate region scene from multiple angles. The viewing angle range includes at least the front view, top view, and multiple side views, and the image resolution is not less than 1920×1080 to ensure the complete visibility of the items in the candidate region and avoid incomplete item modeling due to occlusion.

[0028] (2) Use a large model (GPT-4o) that receives visual information and, based on the multi-view images, infer the categories of each item in the current candidate region scene; then, for each item, use the large model to generate a differential name; and then use BERT to encode the differential names of each item to obtain the text embedding of the item. Among them, input the multi-view images into the large model. The large model first classifies the items in the images. For example, the items are classified into 、 and other multiple categories. Then, use the large model to obtain the differential titles of each item. For example: The large model filters out the "cup" category , and this category includes and and other various sub-categories of cups. Regarding the attribute differences between the sub-categories and , after giving the large model appropriate prompts, the large model can extract meaningful differences based on the images of the same category of items and generate differential text descriptions. For example, for the category, obtain and .

[0029] (3) Use the SAM segmentation model (SegmentAnythingModel) to perform semantic segmentation on the multi-view images, accurately extract the contours of each item, and obtain accurate images containing only each item, avoiding recognition errors caused by background interference. Use the CLIP image encoder to extract visual embeddings from the segmented item images.

[0030] (4) For the text embeddings of each item in the candidate region scene, use a multi-modal Transformer model, and use the cross-attention mechanism to fuse the text embedding features and image embedding features to obtain multi-modal features. The specific formula is as follows: ; where is the text embedding (BERT output), is the visual embedding (CLIP output), , , are trainable weight matrices, is the finally fused multi-modal feature, and d represents the dimension of the vector.

[0031] (5) Map the multi-modal feature to a low-dimensional space through an autoencoder, and describe the language-visual distribution characteristics of each item in the low-dimensional space by calculating the mean vector and covariance matrix corresponding to each item, and obtain the Gaussian sputtering model with personalized language embedding for the current scene , so as to improve the discrimination ability of similar items in the semantic space.

[0032] S2. Obtain the preferred scene graph of the desktop organization task for the user .

[0033] 1. The first desktop organization task. If the user performs desktop organization for the first time, the system will guide the user to input the items to be placed and their feature information (including category, color, shape, use, etc.). For example: If the user wants a pure white coffee cup with a handle to hold coffee in the current desk scene, the user needs to input: {cup category; white, with a handle; for holding coffee}. In addition, the user needs to specify the position information of the items, that is, the relative position relationship between the items or between the items and the desktop. For example: I like to place the coffee cup on the right side of the table, the laptop is in the center of the table, and I need a pen to be placed on the right side of the laptop. According to the initial preferences input by the user, provide prompts for preference scene graph reasoning to the large model to obtain the preference scene graph .

[0034] 2. Non-first desktop organization task. If it is not the first time to perform desktop organization, then according to the user's instruction search for example scene graphs in the database and process them to obtain the preference scene graph . Specifically include: (1) Screen several instructions with the highest similarity to the user's instruction in the database, and obtain a preset number of scene graphs as example scene graphs.

[0035] In this embodiment, based on the similarity between the current instruction and the database instructions as the screening criterion, BERT is used. The user instruction and the instructions in the database are respectively input into the BERT model to obtain their semantic vector representations. By calculating the cosine similarity between the semantic vectors, the 5 scene graphs corresponding to the instructions with the highest similarity in the database are selected as example scene graphs. If there are less than 5 scene graphs in the current database, all the existing scene graphs in the current database are used as example scene graphs. Among them, the formula for calculating the cosine similarity between two instructions is as follows: ; Among them, represents the cosine similarity, represents the user instruction received for query in the database, represents the query instruction input by the current user.

[0036] (2) Input the obtained example scene graph and the preset induction prompt words into the large language model. By inducing the item categories, attributes, and spatial relationships, a structured preferred scene graph is generated . In the preferred scene graph, each node represents a class of items and their corresponding attribute labels, and the edges represent the relative spatial relationships between the nodes.

[0037] In this embodiment, by inducing the example scene graph, a unified preferred scene graph representing the user is generated . For this reason, for the induction of the preferred example scene graph, prompt words are designed and input into the large model for inductive reasoning of the example scene graph. The item selection and position placement rules of the user in the desktop sorting task are summarized from multiple example scene graphs in the database, and a preferred scene graph is formed. The nodes of the preferred scene graph no longer represent specific items, but rather summarize the overall preferences of the user for a certain type of item. For example: in the dining table scene, if the example scene graphs all contain different types of plates, the large model will induce the common preference and obtain {plate class; black; ceramic material; used for serving main courses}, as Figure 3 shown. The preferences are stored in the nodes of the preferred scene graph. This preferred scene graph reflects the sorting preferences of the user in the current scene and can be used to optimize the execution of subsequent desktop sorting tasks.

[0038] S3. Take the user instruction , the preferred scene graph and the Gaussian sputtering model as inputs, call the large model that receives visual information, infer and generate the target scene graph, and further infer and generate the two-dimensional layout coordinates in CSS format for each item in it. Convert the two-dimensional layout coordinates into the actual placement coordinates in the target desktop; then combine the actual placement coordinates with the item IDs to form the target desktop state of the desktop sorting task.

[0039] In the target scene graph Among them, the nodes are the ID corresponding to the target object in the candidate area, the position coordinates of the target object in the candidate area, and its bounding box information. The edges are the positional relationships between the objects on the target table and the neighboring objects, including: front, back, left, and right; the neighboring object refers to the other object with the closest position. The object ID, position, and bounding box information are all obtained by querying in the Gaussian sputtering modeling of the candidate area using the large model grounding algorithm. This step uses the large model grounding algorithm to infer the object ID, position coordinates, and bounding box information, which is prior art. For specific details, please refer to the introduction of "LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent".

[0040] The user instructions in the present invention can be instructions that do not reflect temporary changes, such as "Help me place the dining table", or instructions that allow modifying the desktop layout to point to temporary changes, such as "I want to add a fork today".

[0041] When using the large model to infer the absolute coordinates of objects, the layout of each object is represented by a rectangular box expressed in CSS (Cascading Style Sheets) format to represent the position of the object.

[0042] The specific steps are as follows: S301: Input the Gaussian sputtering model and the preference scenario graph into the large model, and retrieve and query node by node in the candidate area using the large model grounding algorithm to obtain the object ID that best matches the preference scenario graph and its position coordinates in the candidate area. At the same time, the grounding algorithm also outputs the bounding box information corresponding to the object. These information are stored in the nodes of the target scenario graph Node by node means node by node of the preference scenario graph, that is, classify and retrieve and query according to the category of the object to obtain the target object that best meets the user's preference. Specifically, through the node information of the preference scenario graph, the large model obtains a description of this type of object as a query. For example, according to the plate class node in the preference scenario graph {plate class; black; ceramic material; used for serving main courses}, the query "black ceramic plate" is obtained. By giving the large model hints, construct queries and eliminate irrelevant preference descriptions according to the preference scenario graph. Then input this query as the input of the large model grounding algorithm, and the large model grounding algorithm can obtain the object ID, three-dimensional bounding box, and absolute coordinates in the candidate area that best correspond to this query.

[0043] After obtaining the target object information, the large model will according to the user instructions Check the nodes of the target scene graph. If the instruction does not point to modifying the target item, then assign the relative position relationship of such items in the preferred scene graph to the target scene graph correspondingly. For example, if the instruction is: "Help me place this desk", then the target item and its position relationship will not be changed, so the relative position relationship in the target scene graph is the same as that in the preferred scene Figure 1 graph. If the instruction points to modifying the target item, then the large model needs to modify the target scene graph according to the instruction and the preferred scene graph. Such modification includes adding or deleting target items and modifying the relative position relationship. For example, the user instruction "I don't want a fork today" means deleting fork-like items. If such items exist in the target scene graph, then the corresponding nodes in the target scene graph need to be deleted and then the remaining relative position relationships in the preferred scene graph are assigned to the target scene graph correspondingly.

[0044] The specific process can be expressed as: ; In the formula represents the inference process of the large model, represents the target scene graph, including the ID of the item, its position in the candidate area, the boundary information of the item, and the relative position relationship between the item and its neighboring items on the target desktop. Among them, the relative position relationship is in the form of a triple to represent, where o i and o j represent the item and its neighboring item respectively. In this embodiment, four position relationships, right, left, above, and behind, are used as the optional relative position relationships relationship. When describing the relative position of the target item, it is stipulated that only the relative position relationship with the other target item closest to its position is used.

[0045] S302: Give the large model example data of the position relationship to increase the accuracy of the large model's position inference. The specific method is to input a certain amount of example data as a prompt into the large model so that it can learn the relative position relationship and thus more accurately infer the layout of the target item. The content of the example data includes the coordinates of the target item and the relative position relationship between the items, and all item coordinates are given in CSS format.

[0046] S303: Input the target scene graph obtained in S301 into the large model. The large model infers and obtains the two-dimensional layout output of the target item in CSS format to represent the absolute position. Among them corresponds to the left margin in CSS, indicating the distance from the left edge of the canvas to the left boundary of the item; corresponds to the top margin in CSS, indicating the distance from the top of the canvas to the upper boundary of the item; Represents the width of the item, and its value can be obtained from the bounding box information in the target scene graph; Represents the height of the item, which is also obtained based on the bounding box information of the target scene graph. Using the CSS format as the output result of the large model increases the large model's perception ability of the item's boundaries and improves the accuracy of reasoning.

[0047] The canvas is a standardized coordinate space defined artificially in this embodiment, used to describe the layout information of the target item during the large model reasoning process. In this solution, the main role of the canvas is to provide a unified two-dimensional coordinate reference framework, facilitating the large model to understand and process the item positions. By setting the upper left corner as (0,0), with the positive X-axis direction to the right and the positive Y-axis direction downwards, the canvas can standardize the item positions in the scene, enabling target desktops of different sizes and image data of different resolutions to be calculated and transformed in the same coordinate system. By defining the canvas, the consistency of the item layout data can be maintained in different scenarios, improving the generalization ability of the model and providing a basis for subsequent coordinate transformation (such as mapping from CSS coordinates to desktop coordinates).

[0048] S304: Conversion from CSS format layout coordinates to target desktop coordinates. Use coordinate transformation to map the CSS format layout to the corresponding on the target desktop, where it refers to the position of the center point of the item on the desktop when placed. To accurately map the CSS format layout coordinates to the target desktop coordinates, coordinate normalization, origin transformation, and size scaling are required. First, normalize the position parameters of the item in the CSS coordinate system based on the width and height of the CSS canvas, so that its coordinate values are transformed to between [0,1]. Since the origin of the CSS coordinate system is at the upper left corner, while the origin of the target desktop coordinate system is at the center of the desktop, a central symmetry transformation needs to be performed on the normalized coordinates to ensure the consistency of the coordinate system. Finally, by multiplying the normalized coordinate values by the physical width and height of the target desktop respectively, the actual placement coordinates of the item on the desktop can be obtained, ensuring the accuracy and consistency of the layout mapping.

[0049] The specific conversion formula is as follows: ; ; where, and are the width and height of the CSS canvas. In the present invention, the canvas size is specified as 512x512, and are the width and height of the target desktop, both in pixels.

[0050] S305: Store the target object ID obtained in step S301 and the position coordinates on the desktop obtained in S304 as the target desktop state of the desktop organization task.

[0051] S4. Based on the target state generated in S3, use PDDL for task planning to obtain the action sequence for organizing the desktop from an empty desktop to the target scene graph.

[0052] In this embodiment, the domain planning language PDDL (Planning Domain Definition Language) is used, combined with the target state information, and the FastDownward solver is used for task planning to infer the action sequence that the robot needs to execute.

[0053] In the PDDL planning framework, the domain file defines the operation rules, actions, and conditions of the robot, while the problem file specifically describes the initial state, target state, and their constraint conditions of the task. For the desktop organization task, the PDDL domain file of the present invention defines executable actions (such as picking up objects, placing objects, etc.) and the states of objects (such as the position and type of objects). In the problem file, the initial state of the objects is specified, and the object ID and the absolute position of the objects are used as the target state. By combining the domain file and the problem file, the PDDL planner can infer a series of action steps according to the current desktop layout and the target scene graph to effectively guide the robot to complete the desktop organization task.

[0054] S5. Execute the action sequence generated in S4 through the robotic arm to complete the picking and placing of each object in the target desktop state and complete the desktop organization task.

[0055] In this step, the system executes sequentially according to the action sequence generated by the PDDL planning to ensure the feasibility and rationality of the task. The specific steps are as follows: S501: Parse the target action sequence and convert it into executable instructions, including the starting position, target position, execution method, etc. of the target object, and adapt it in combination with the environmental information to ensure that the instructions are executable in the current desktop state.

[0056] S502: Gradually execute the organization task in the planned order, sequentially complete operations such as picking up and placing objects, and update the state after each step to ensure the logic and coherence of the actions and avoid task failure or incorrect object placement caused by incorrect operations.

[0057] S6. After S5 completes the desktop organization task, the target scene graph obtained in step S3 Store it in the database. This step aims to record the organized desktop state for subsequent querying of example scenario diagrams.

[0058] Specifically, it includes the following steps: S601: After the system completes all organization actions, it requests the user to confirm whether the current desktop meets the user's instructions and the user's preferences are consistent to ensure that the target desktop that meets both the temporary changes and past preferences is obtained.

[0059] S602: After the user confirms compliance, pair the target scenario diagram generated in S3 with the corresponding user instructions and store them in the database in a paired manner for subsequent data analysis. If the user is not satisfied with the current desktop organization result, the current target scenario diagram will not be stored.

[0060] The method of the embodiment of the present invention integrates the user's historical preferences and temporary instructions, dynamically adjusts the inference result, thereby enhancing the response ability to temporary changes in the desktop layout and achieving efficient and accurate execution of the desktop organization task. At the same time, this method effectively alleviates the hallucination problem in large model inference and significantly improves the accuracy of position inference by introducing structured position representation (such as layout description in CSS format) and progressive position inference (gradually transitioning from relative position relationships to absolute position relationships).

[0061] Embodiment 2

[0062] This embodiment provides a desktop organization system based on large model inference of human preferences, which is applied to a desktop organization robot and includes: an image acquisition module, a Gaussian sputtering modeling module, a preference scenario diagram acquisition module, a large model module, an action sequence generation module, and an execution module; The image acquisition module is used to: acquire multi-view RGB-D images of the candidate area scene and output image data in a unified format for subsequent processing; The Gaussian sputtering modeling module is used to: perform visual semantic fusion modeling on the items in the acquired image to generate a Gaussian distribution representation of each item in the low-dimensional space, and the representation includes a mean vector and a covariance matrix; The preference scenario diagram acquisition module is used to: acquire the preference scenario diagram of the user for the desktop organization task ; The large model module is used to: infer and generate a target scenario diagram according to the Gaussian sputtering model, user instructions, and preference scenario diagram; The large model module is also used to: pair the user instructions with the preference scenario diagram and the Gaussian sputtering model As input, it infers and generates a target scene graph, further infers and generates the two-dimensional layout coordinates in CSS format for each item therein, converts the two-dimensional layout coordinates into the actual placement coordinates in the target desktop; then combines the actual placement coordinates with the item IDs to form the target desktop state of the desktop sorting task; The action sequence generation module is used to: according to the Gaussian sputtering model and the target desktop state to perform task planning, and obtain the action sequence for sorting the desktop from an empty desktop to the target desktop state; The execution module, using a robotic arm, is used to execute the generated action sequence to complete the picking and placing of each item in the target desktop state and complete the desktop sorting task.

[0063] The specific implementation manners of the constituent modules of the desktop sorting system in this embodiment are the same as those in Embodiment 1.

[0064] The above embodiments are the preferred embodiments of the present application. Those of ordinary skill in the art can also make various transformations or improvements on this basis. Without departing from the general concept of the present application, these transformations or improvements should all fall within the scope of protection required by the present application.

Claims

1. A desktop organization method based on large model reasoning of human preferences, characterized in that: Applied to desktop organizing robots, including: S1, collect multi-view images of the candidate area scene, use the Gaussian sputtering method to model the candidate area scene, and obtain the Gaussian sputtering model of the candidate area scene ; S2, obtain the user's preference scenario graph for the desktop organization task ; S3, the user instruction , preference scene graph And Gaussian sputtering model As input, the large model that receives visual information is called to infer and generate the target scene graph, and further infer and generate the CSS format two-dimensional layout coordinates of each item in it, and convert the two-dimensional layout coordinates into the actual placement coordinates in the target desktop; then the actual placement coordinates are combined with the item ID to form the target desktop state of the desktop organization task; S4, according to the Gaussian sputtering model , based on the target desktop state generated by S3, task planning is performed to generate an action sequence for arranging the desktop from an empty desktop to a target desktop state; S5, the robot arm executes the action sequence generated by S4 to complete the picking and placement of each object in the target desktop state, thus completing the desktop tidying task.

2. The desktop tidying method according to claim 1, characterized in that: After S5 completes the desktop arrangement task, the following steps are also included: S6, converting the target scene graph obtained in S3 and its corresponding user instructions The pairs are stored in the database of the current user, and the desktop organization task records in the database are updated.

3. The desktop tidying method according to claim 1, characterized in that: The specific process of modeling the candidate area scene using the Gaussian sputtering method is as follows: S11: Collect the current desktop scene to obtain a multi-view image; S12: using the large model receiving visual information and based on the multi-view images, inferring and generating the categories of each item in the current desktop scene; then using the large model to generate a difference name for each item based on the category feature; then encoding the difference name of each item to obtain the text embedding feature of the item; the difference is reflected in: under the inferred item category, describing the feature difference between the subcategory to which the item belongs and other subcategories; S13: Use the SAM segmentation model to perform semantic segmentation on multi-view images, and use the CLIP image encoder to extract image embedding features of each object in the multi-view images; S14: Adopting the multimodal Transformer model and the cross-attention mechanism, the text embedding features of each object in the candidate region scene are fused with the image embedding features to obtain multimodal features; S15: The multimodal features are mapped to a low-dimensional space through an autoencoder, and the mean vector and covariance matrix corresponding to each item are calculated. The mean vector and covariance matrix are used to describe the language-visual distribution characteristics of each item in the low-dimensional space, and a Gaussian sputtering model with personalized language embedding is obtained for the current scene. .

4. The desktop tidying method according to claim 1, characterized in that: In step S2, if it is the first time to organize the desktop, the initialization setting is performed based on the organization rules set by the user to obtain the preferred scene map. ; If this is not the first time you are organizing your desktop, combine it with the user's instructions Reasoning with the database to obtain a preference scene graph .

5. The desktop tidying method according to claim 4, characterized in that: Combined with user instructions Reasoning with the database to obtain a preference scene graph The specific process is as follows: Filtering and user instructions in the database A number of instructions with the highest semantic similarity are used to obtain a preset number of scene graphs corresponding to the instructions as example scene graphs; The obtained example scene graph and the preset summary prompt words are input into the large language model to generate a structured preference scene graph by summarizing the item categories, attributes and spatial relationships. ; In the preference scene graph, each node represents a category of items and its corresponding attribute labels, and the edges represent the relative spatial relationship between nodes.

6. The desktop tidying method according to claim 5, characterized in that: The BERT model is used to calculate semantic similarity, including: Input the user instructions and instructions in the database into the BERT model respectively to obtain their semantic vector representations; By calculating the cosine similarity between semantic vectors, several instructions with the highest semantic similarity to the user instructions are screened out; the expression for calculating the cosine similarity of two instructions is as follows: ; in, represents the cosine similarity, Represents the instructions in the database, Indicates the user command currently entered by the user.

7. The desktop tidying method according to claim 1, characterized in that: The specific process of step S3 is as follows: (1) Using the large model grounding algorithm, according to the Gaussian sputtering model , User instructions and preference scene graph Get the target scene graph ; (2) Provide sample data of the positional relationship of the large model; (3) Obtain the target scene graph Input the big model, and the big model infers the two-dimensional layout representing the absolute position of the item in CSS format based on the position relationship example data. ; left and top represent the distance of the rectangle box relative to the left and top of the canvas respectively, width and height are used to define the width and height of the rectangle box, both in pixels; (4) Convert the two-dimensional layout in CSS format Mapped to the coordinates of the location where the item is placed on the desktop , the specific formula is as follows: ; ; in, and is the width and height of the CSS canvas, and is the width and height of the target desktop, both in pixels; (5) Obtain the target scene graph Item ID and desktop location coordinates in , together they constitute the target desktop state of the desktop organization task.

8. The desktop tidying method according to claim 1, characterized in that: The target scene graph includes the object's ID, its position in the object candidate area scene, the object's boundary information, and the relative position relationship between the object and its neighboring objects on the target desktop; the relative position relationship between the object and its neighboring objects includes: front, back, left, and right; the neighboring object refers to another object that is closest in position.

9. The desktop tidying method according to claim 1, characterized in that: Step S4 uses domain planning language to perform task planning, and generates an optimal action sequence for the desktop organization task from an empty desktop to a target desktop state based on reasoning of the target desktop state.

10. A desktop organization system based on large model reasoning of human preferences, characterized in that: Applied to desktop tidying robots, including: image acquisition module, Gaussian sputtering modeling module, preference scene graph acquisition module, large model module, action sequence generation module and execution module; The image acquisition module is used to: acquire multi-view RGB-D images of the candidate area scene and output image data in a unified format for subsequent processing; The Gaussian sputtering modeling module is used to: perform visual semantic fusion modeling on the objects in the collected images, and generate a Gaussian distribution representation of each object in a low-dimensional space, wherein the representation includes a mean vector and a covariance matrix; The preference scene graph acquisition module is used to: acquire the user's preference scene graph for the desktop organization task ; The large model module is used to: generate a target scene graph by reasoning according to a Gaussian sputtering model, user instructions and a preferred scene graph; The large model module is also used to: , preference scene graph And Gaussian sputtering model As input, the target scene graph is generated by inference, and the CSS format two-dimensional layout coordinates of each item in it are further generated by inference, and the two-dimensional layout coordinates are converted into actual placement coordinates in the target desktop; the actual placement coordinates are then combined with the item ID to form the target desktop state of the desktop organization task; The action sequence generation module is used to: Perform task planning with the target desktop state to obtain an action sequence for organizing the desktop from an empty desktop to the target desktop state; The execution module uses a robotic arm to execute the generated action sequence to complete the picking and placement of each object in the target desktop state and complete the desktop tidying task.

Citation Information

Cited By

  • GPU-accelerated 3D Gaussian splash three-dimensional coordinate high-precision real-time pickup method and system

    CN121095466A

  • Gpu accelerated 3d gaussian splatting three-dimensional coordinate high-precision real-time picking method and system

    CN121095466B