Weight-based hierarchical collaborative method, apparatus, equipment, and medium for clothing stacking.

CN122572801APending Publication Date: 2026-08-14FIBOCOM WIRELESS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]本申请提供了基于权重的分级协作衣物叠放方法、装置、设备及介质,以解决现有通过单个机器人进行衣物叠放存在部署成本高,以及通过多工位流水线式折叠存在自适应调度性及协同性差,易造成衣物起皱、滑移、翻边,折叠效率差的问题

Benefits of technology

本申请提供的基于权重的分级协作衣物叠放方法,通过搭建一个基于分级异构节点、多模态底座共享及权重分离子网络的架构,在共用统一多模态底座的同时,将取衣、叠衣、储衣按节点切分动作空间与权重分离,单节点动作空间更小、误匹配概率更低,更有利于系统兼顾泛化与低成本部署能力,并降低端到端大模型在单机上学习全流程的难度;在叠衣VLA节点,通过在VLA动作解码中将布料的局部切向位移与法向期望力同时纳入端到端输出,使得多元组高斯混合动作头输出轮廓、力与动作的操作策略,将压平、熨烫与折线生成在同一策略下统一优化,能够显著减少起皱、滑移、折边不齐的问题,更有利于增强柔性体折叠的稳定性与可控性,提高折叠效率;在多工位协作过程中,实现无需手工规则的语言路由与子任务分解,调度性更强,且将最大的节点权重分配至叠衣VLA节点以匹配叠衣VLA节点的叠衣难度,使得更关注薄弱节点,多工位协同性更高。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122572801A_ABST
    Figure CN122572801A_ABST
Patent Text Reader

Abstract

This application relates to a weighted hierarchical collaborative clothing stacking method, apparatus, device, and medium. The method includes: generating sub-tasks for each VLA node based on a natural language task using a shared multimodal base; configuring a weighted separation sub-network and node weights for each VLA node; retrieving and transferring flexible clothing to be stacked based on the first weighted separation sub-network and first interval routing response of the clothing retrieval VLA node for the clothing retrieval sub-task; generating a stacking action vector superimposed with local tangential displacement and normal expected force using a multivariate Gaussian mixture motion head based on the second weighted separation sub-network and second interval routing response of the clothing stacking VLA node for the clothing stacking sub-task; and stacking the clothing based on the stacking action vector; and moving the folded clothing to the target storage location based on the third weighted separation sub-network and third interval routing response of the clothing storage VLA node for the clothing storage sub-task. This application enables low-cost deployment and reduces learning difficulty, while enhancing the stability and controllability of flexible body folding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence control technology, and in particular to a weight-based hierarchical collaborative method, apparatus, device and medium for stacking clothes. Background Technology

[0002] As life becomes increasingly intelligent and users' needs gradually improve, robots are gaining acceptance as a replacement for people in handling daily household chores, including cleaning, cooking, retrieving items, washing and folding clothes.

[0003] In terms of clothing folding, existing clothing sorting / folding systems typically employ a single robot (single-arm or dual-arm) to handle the entire process of "retrieving clothes—folding—putting into the cabinet / drawer." Control strategies are often rule-based (key point detection, preset folding sequences, and gripper trajectory templates) decomposition or a single end-to-end strategy network directly outputting actions. This requires massive amounts of training data and high-performance computing and hardware, resulting in high deployment costs. To address this, a multi-station pipeline folding system is proposed, where multiple stations perform grasping, flattening, folding, and storage. However, the stations rely heavily on fixed manual rules / mechanical fixtures for interaction, lacking adaptive routing based on user natural language tasks and cross-station learning-based collaborative scheduling. Ultimately, executable cross-station sub-tasks are generated directly from the user's natural language task, resulting in weak generalization capabilities. Especially for folding flexible clothing, the fabric is approximated as a "gripable but kinematically negligible" object. The fold line is estimated based on the visual contour / key point, and the folding is mainly performed by position control. However, there is usually a lack of explicit modeling of the gripper contact force / normal flattening force, which leads to wrinkling, slippage, and edge-folding failure.

[0004] It is evident that existing methods of folding clothes using a single robot have high deployment costs, and multi-station assembly line folding suffers from poor adaptive scheduling and coordination, which can easily cause clothes to wrinkle, slip, or fold over, resulting in poor folding efficiency. Summary of the Invention

[0005] This application provides a weight-based hierarchical collaborative clothing stacking method, apparatus, equipment, and medium to solve the problems of high deployment costs in existing clothing stacking methods using a single robot, and poor adaptive scheduling and coordination in multi-station assembly line folding, which easily leads to clothing wrinkling, slippage, and folding, resulting in poor folding efficiency.

[0006] According to one aspect of the embodiments of this application, this application provides a weight-based hierarchical collaborative clothing stacking method, the method comprising: generating subtasks corresponding to each VLA node based on an acquired natural language task through a multimodal base shared by multiple VLA nodes, wherein each VLA node is configured with a weight separation subnetwork and node weights; retrieving flexible clothing to be stacked and transferring it to a folding station based on a first weight separation subnetwork and a first interval routing response to the clothing retrieval subtask of the clothing retrieval VLA node; generating a stacking action vector superimposed with local tangential displacement and normal expected force through a multivariate Gaussian mixture motion head based on a second weight separation subnetwork and a second interval routing response to the clothing stacking subtask of the clothing stacking VLA node, and stacking the clothing at the folding station based on the stacking action vector; and transferring the folded clothing from the folding station to a target storage location based on a third weight separation subnetwork and a third interval routing response to the clothing storage subtask of the clothing storage VLA node, wherein the node weight of the clothing stacking VLA node is the largest.

[0007] Optionally, after the third weight separation sub-network and the third interval routing response of the clothing storage VLA node transfer the folded clothing from the folding station to the target storage station, the method further includes: obtaining the task processing time and total task time for executing the clothing retrieval sub-task, the clothing folding sub-task, and the clothing storage sub-task, respectively; if the total task time exceeds a preset total time constraint, then based on the processing time of each task, the node weights of the clothing retrieval VLA node, the clothing folding VLA node, and the clothing storage VLA node are reassigned through the weight gating configured during model training, wherein the reassigned node weights satisfy that the sum of the node weights of the clothing folding VLA node, the clothing storage VLA node, and the clothing retrieval VLA node is 1 and the node weights decrease sequentially.

[0008] Optionally, the multimodal base shared by multiple VLA nodes generates subtasks corresponding to each VLA node based on the acquired natural language task, including: acquiring the natural language task input by the user and inputting the natural language task into the multimodal base; acquiring an image of the drying area of ​​the flexible clothing to be folded based on the visual encoder in the multimodal base according to the natural language task, and encoding language instructions based on the language encoder in the multimodal base; and performing feature fusion through the cross-modal fusion network in the multimodal base based on the image of the drying area of ​​the flexible clothing to be folded and the language instructions to generate the clothing retrieval subtask, clothing folding subtask, and clothing storage subtask corresponding to the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node.

[0009] Optionally, the first weighted separation subnetwork and first interval routing based on the garment retrieval VLA node respond to the garment retrieval sub-task, retrieve the flexible garments to be folded and transfer them to the folding station, including: obtaining the garment retrieval sub-task generated by the multimodal base based on the first interval routing corresponding to the garment retrieval VLA node; responding to the garment retrieval sub-task, performing hierarchical weighted projection transformation based on a first bottleneck dimension through the first weighted separation subnetwork of the garment retrieval VLA node, and performing garment edge detection on the drying area of ​​the flexible garments to be folded based on the drying area image to output a garment retrieval action vector; grabbing the flexible garments to be folded according to the garment retrieval action vector, and transferring the grabbed flexible garments to be folded to the folding station through a mobile device deployed by the garment retrieval VLA node.

[0010] Optionally, the second weight separation subnetwork and the second interval routing based on the folding VLA node respond to the folding subtask, generating a folding action vector superimposed with local tangential displacement and normal expected force through a multivariate Gaussian mixture action head, and folding clothes at the folding station based on the folding action vector, including: obtaining the folding subtask generated by the multimodal base based on the second interval routing corresponding to the folding VLA node; responding to the folding subtask, performing a hierarchical weight projection transformation based on the second bottleneck dimension through the second weight separation subnetwork of the folding VLA node, estimating the self-occlusion area of ​​the flexible clothing to be folded, generating the folding action vector through the multivariate Gaussian mixture action head, wherein the folding action vector includes the local tangential displacement and normal expected force of the flexible clothing to be folded at the folding station; and folding clothes at the folding station through the folding VLA node based on the folding action vector.

[0011] Optionally, in response to the folding sub-task, the second weight separation sub-network of the folding VLA node performs a hierarchical weight projection transformation based on the second bottleneck dimension, estimates the self-occlusion area of ​​the flexible clothing to be folded, and generates the folding action vector through the multivariate Gaussian mixture action head. This includes: using the folding sub-task as input features of the second weight separation sub-network of the folding VLA node; performing hierarchical calculations of down-projection weights, nonlinear activation, and up-projection weights on the input features based on the second bottleneck dimension, and outputting key features corresponding to the folding sub-task; identifying the clothing outline and fold boundaries of the flexible clothing to be folded, and estimating the self-occlusion area of ​​the flexible clothing to be folded; inputting the key features and the self-occlusion area into the multivariate Gaussian mixture action head, generating a corresponding folding action strategy based on the selected Gaussian components, including generating the local tangential displacement for flattening or folding the clothing and the normal expected force for flattening the clothing; and outputting the folding action vector based on the generated folding action strategy.

[0012] Optionally, the third weight separation sub-network and third interval routing based on the clothing storage VLA node respond to the clothing storage sub-task, transferring the folded clothing from the folding station to the target storage location, including: obtaining the clothing storage sub-task generated by the multimodal base based on the third interval routing corresponding to the clothing storage VLA node and identifying the target storage location; responding to the clothing storage sub-task, performing hierarchical weight projection transformation based on the third bottleneck dimension through the third weight separation sub-network of the clothing retrieval VLA node to generate a clothing storage action vector; and transferring the folded clothing from the folding station to the target storage location based on the clothing storage action vector.

[0013] According to another aspect of the embodiments of this application, this application provides a weight-based hierarchical collaborative clothing stacking system, the system comprising: a task splitting module, used to generate sub-tasks corresponding to each VLA node based on the acquired natural language task through a multimodal base shared by multiple VLA nodes, wherein each VLA node is configured with a weight separation sub-network and node weights; a clothing retrieval module, used to retrieve flexible clothing to be stacked and transfer it to a folding station based on the first weight separation sub-network and the first interval routing response of the clothing retrieval VLA node for the clothing retrieval sub-task; a clothing stacking module, used to generate a clothing stacking action vector superimposed with local tangential displacement and normal expected force through a multivariate Gaussian mixture action head based on the second weight separation sub-network and the second interval routing response of the clothing stacking VLA node for the clothing stacking sub-task, and to stack clothing at the folding station based on the clothing stacking action vector; and a clothing storage module, used to transfer the folded clothing from the folding station to a target storage location based on the third weight separation sub-network and the third interval routing response of the clothing storage VLA node for the clothing storage sub-task.

[0014] According to another aspect of the embodiments of this application, this application provides a clothing folding device, including: a processor, a memory, and a network interface. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory through the network interface, and the processor executes the machine-readable instructions to perform the steps of the weighted hierarchical collaborative clothing folding method.

[0015] According to another aspect of the embodiments of this application, this application provides a storage medium having processor-executable non-volatile program code, the program code causing the processor to perform the steps of the weighted hierarchical collaborative clothing stacking method.

[0016] Compared with related technologies, the technical solutions provided in this application have the following advantages: The weight-based hierarchical collaborative clothing stacking method provided in this application constructs an architecture based on hierarchical heterogeneous nodes, shared multimodal base, and weight-separated subnetworks. While sharing a unified multimodal base, it separates the action space and weights of clothing retrieval, folding, and storage by node. This results in a smaller single-node action space, a lower probability of mismatch, and is more conducive to the system's ability to generalize and deploy at low cost. It also reduces the difficulty of learning the entire process of an end-to-end large model on a single machine. In the clothing stacking VLA node, the local tangential displacement and normal expected force of the fabric are simultaneously incorporated during VLA action decoding. The end-to-end output enables the multi-group Gaussian hybrid motion head to output contour, force, and motion operation strategies. Flattening, ironing, and fold line generation are uniformly optimized under the same strategy, which can significantly reduce wrinkling, slippage, and uneven fold edges. This is more conducive to enhancing the stability and controllability of flexible body folding and improving folding efficiency. In the process of multi-station collaboration, language routing and subtask decomposition without manual rules are realized, which is more scheduling-oriented. The largest node weight is allocated to the folding VLA node to match the folding difficulty of the folding VLA node, so that more attention is paid to weak nodes and the multi-station collaboration is higher. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0018] Figure 1 This is a schematic diagram of an optional weight-based hierarchical collaborative clothing stacking method provided according to an embodiment of this application; Figure 2 This is a schematic diagram illustrating an optional VLA node and data stream collaboration process according to an embodiment of this application; Figure 3 This is a flowchart illustrating an optional weight-based hierarchical collaborative clothing stacking method provided according to an embodiment of this application. Figure 4 This is a schematic diagram illustrating an optional subtask partitioning method based on a multimodal base according to an embodiment of this application; Figure 5 This is a schematic diagram of an optional weight-based hierarchical collaborative clothing stacking system provided according to an embodiment of this application; Figure 6 This is a schematic diagram of an optional electronic device structure provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] To address the problems mentioned in the background art, according to one aspect of the embodiments of this application, an embodiment of a weight-based hierarchical collaborative clothing stacking method is provided.

[0021] It should be noted that the weight-based hierarchical collaborative clothing stacking method provided in this application embodiment is generally executed by a server and / or a terminal device. Correspondingly, the weight-based hierarchical collaborative clothing stacking system is generally set up in the server / terminal device.

[0022] like Figure 1 As shown, Figure 1 A flowchart of a weighted hierarchical collaborative clothing stacking method provided in an embodiment of the present invention. Taking the weighted hierarchical collaborative clothing stacking method executed by a server as an example, the weighted hierarchical collaborative clothing stacking method includes the following steps: Step S102: Through the multimodal base shared by multiple VLA nodes, subtasks corresponding to each VLA node are generated according to the acquired natural language task. Each VLA node is configured with a weight separation subnetwork and node weights.

[0023] Combination Figure 2 As shown, Figure 2This is a flowchart of another optional weight-based hierarchical collaborative clothing stacking method provided in this embodiment. The multimodal base includes a visual encoder (ViT, VisionTransformer) for vision encoding, a language encoder (BERT, Bidirectional EncoderRepresentations from Transformers) for language encoding, and a fusion layer (Cross-Attn) for cross-modal fusion. Each VLA (Vision-Language-Action) node shares the same multimodal base, which serves as the backbone network, and its parameters can be fixed. Each VLA node is configured with a lightweight adapter (lightweight parameter adaptation module) for task-level or node-level fine-tuning on the shared multimodal base, forming weight-separated sub-networks. These weight-separated sub-networks collaborate hierarchically without interfering with each other. When adjusting parameters and training the model, only the weight-separated sub-networks need to be adjusted, thereby reducing the amount of training data and lowering model deployment costs.

[0024] The weighted separation subnetwork is shown in equation (1) below: (1) in, For the output of the multimodal base, As a bottleneck adapter, residual connection ensures that the base features are not lost. VLA node The multimodal input includes at least one of the following: image, depth, state, language token, etc. The parameter quantity is controlled to be within 3% of the total parameters of the multimodal base to meet the real-time inference requirements of the edge side.

[0025] Furthermore, the user layer provides input to the terminal and outputs only a description of the natural language task. For example, fold the shirts on the balcony and put them in drawer B. (Natural Language Task) After multimodal base language encoding, a token sequence for subtasks is generated, including subtask token sequences for the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node. Based on the weighted separation subnetwork of each VLA node, when executing a subtask, each VLA node only performs attention on its corresponding token sequence interval, realizing language routing, that is, automatically assigning tasks to the corresponding nodes according to the task description, without the need for manual rules.

[0026] Each VLA node is assigned a node weight, and the node weight of each VLA node is taken as follows. Node weights of VLA nodes And the node weight of the VLA node satisfy And the size relationship satisfies The node weights are matched based on the training difficulty and task execution difficulty of different VLA nodes, allowing the model to focus more on the more critical parts. The learnable parameters for the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node are as follows: , , ; Step S104: Based on the first weighted separation sub-network and the first interval routing response of the clothing retrieval VLA node, retrieve the flexible clothing to be stacked and transfer it to the folding station.

[0027] In this context, language routing can be represented in a multi-station or multi-module robot system as performing semantic understanding, intent parsing, task decomposition, and dynamic allocation of user natural language commands. This automatically distributes high-level commands to appropriate workstations, sub-modules, or policy networks, achieving intelligent scheduling of language commands, executable subtasks, and corresponding execution units, rather than relying on hard-coded fixed rules.

[0028] Specifically, for the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node, the weight separation subnetworks configured in sequence are the first weight separation subnetwork, the second weight separation subnetwork, and the third weight separation subnetwork. The corresponding interval routes are the first interval route, the second interval route, and the third interval route, so as to allocate the corresponding subtasks to the corresponding VLA nodes respectively.

[0029] Combination Figure 2 and Figure 3 As shown, the complete user task instruction is semantically segmented and decomposed based on the first interval routing (first token interval routing), accurately identifying the semantic token sequence interval corresponding to the clothing retrieval sub-task. The semantic token sequence interval and the corresponding features required for clothing retrieval output by the multimodal base are distributed to the first weight separation sub-network of the clothing retrieval VLA node for feature transformation through the first token interval routing. Based on the transformed features, the action head of the clothing retrieval VLA node outputs the executable instructions of the clothing retrieval VLA node's execution terminal, allowing the clothing retrieval VLA node to accurately locate the optimal grab point of the flexible clothing to be stacked.

[0030] In some examples, the garment retrieval VLA node can be fixedly mounted on the clothes rack rail. A 2D camera and a single robotic arm can be deployed within the garment retrieval VLA node, enabling 2D image acquisition and robotic arm status monitoring. The acquired 2D images are then used by the robotic arm to grasp the garments to be folded and transfer them to the folding station. This provides well-organized pre-operational conditions for the subsequent garment folding VLA node to perform garment flattening and folding operations, thus completing the garment retrieval sub-task process based on hierarchical collaboration and weighted separation adaptation.

[0031] In some examples, the robotic arm can have obstacle avoidance capabilities. During the transfer to the folding station, it can adjust and plan the preset route according to the road conditions to avoid collisions during the movement.

[0032] Step S106: Based on the second weight separation sub-network and the second interval routing response of the folding VLA node, the folding action vector with superimposed local tangential displacement and normal expected force is generated by the multivariate Gaussian mixture action head, and the folding action vector is used to fold clothes at the folding station.

[0033] Among them, local tangential displacement refers to the displacement component in the local coordinate system plane of the fabric, and normal expected force refers to the expected contact force perpendicular to the fabric surface, used for flattening, ironing, or inhibiting slippage. The multivariate Gaussian mixture motion head refers to a motion head that receives task-specific features through a Gaussian Mixture Model (GMM) and outputs multiple sets of multidimensional Gaussian distribution parameters. The multivariate Gaussian mixture motion head includes a "contour-force-motion" triplet Gaussian mixture motion head, which can output folding motion vectors including local tangential displacement and normal expected force in an end-to-end manner, improving folding accuracy.

[0034] Combination Figure 2 and Figure 3 As shown, similarly, for the clothing folding VLA subtask, the user's natural language task is semantically segmented and decomposed through the second interval routing (second token interval routing), accurately identifying the semantic token sequence interval corresponding to the clothing folding subtask. Then, through the second token interval routing, the semantic token sequence interval and the corresponding features required for clothing folding output by the multimodal base are distributed to the second weight separation subnetwork of the clothing folding VLA node for feature transformation, and then output to the multivariate Gaussian mixture action head to generate the clothing folding action vector according to "contour-force-action".

[0035] In some examples, the garment folding VLA node can be placed at the folding station, including a depth camera, dual robotic arms, and an ironing or pressing plate, capable of acquiring RGB-D image data, end-effector force / tactile feedback, and the status of the two arms. The garment folding VLA node can perform flattening, folding, and pressing actions on the flexible garment to be folded at the folding station based on the garment folding motion vector.

[0036] Step S108: Based on the third weight separation sub-network and the third interval routing response of the clothing storage VLA node, the folded clothing is transferred from the folding station to the target storage location, wherein the node weight of the folding VLA node is the largest.

[0037] For the clothing storage VLA subtask, the user's natural language task is semantically segmented and decomposed using third-interval routing (third-token interval routing). This accurately identifies the semantic token sequence interval corresponding to the clothing storage subtask. The semantic token sequence interval, along with the features required for clothing storage output by the multimodal base, is then distributed to the third weight separation subnetwork of the clothing storage VLA node for feature transformation. The transformed features are then output to the action head of the clothing storage VLA node to obtain the final action executed by the robotic arm. Based on this action, the folded clothing is transferred from the folding station to the target storage location. The clothing storage VLA node can be located on a guide rail trolley above the drawer cabinet or at the end effector of the robotic arm. It can also be equipped with a camera to capture images of the drawer area and the pose status during movement.

[0038] In some embodiments, the overall system corresponding to the weight-based hierarchical collaborative clothing stacking method described above is a hierarchical collaborative system with three levels of heterogeneous VLA nodes. It includes: a user layer (L0) that receives terminal input and outputs only a natural language task description. The clothing retrieval VLA node (L1) is fixedly installed on the clothes rack rail, including a 2D camera and a low-cost single arm, to perform the retrieval of clothes from the clothes rack and their transfer to the indoor folding station. The learnable parameters are recorded as follows. The node weight is The garment folding VLA node (L2) is placed on an indoor folding table, including a depth camera, dual arms, and an ironing / pressing plate, to perform the unfolding, folding, and pressing of flexible garments. The learnable parameters are recorded as follows: The node weight is The VLA node (L3) for clothing storage is located on the drawer carriage or at the end of the robotic arm above the drawer unit. It performs tasks such as placing items into designated drawers / slots and arranging them for positioning. The learnable parameters are... The node weight is .

[0039] In some examples, the node weights of the VLA nodes are... Maximum, the node weight of the VLA node for clothing retrieval The node weight can be greater than or less than that of the VLA node. The node weights of the three VLA nodes satisfy the following conditions: Preferably, the solidified configuration after convergence satisfies .

[0040] In this embodiment of the invention, an architecture based on hierarchical heterogeneous nodes, shared multimodal base, and weighted sub-networks is constructed. While sharing a unified multimodal base, the action space and weights of clothing retrieval, folding, and storage are separated by nodes. This results in a smaller single-node action space, a lower probability of mismatch, and is more conducive to the system's ability to balance generalization and low-cost deployment. It also reduces the difficulty of learning the entire process of an end-to-end large model on a single machine. In the folding VLA node, the local tangential displacement and normal expected force of the fabric are simultaneously incorporated into the end-to-end output during VLA action decoding. This allows the multi-group Gaussian hybrid motion head to output contour, force, and motion operation strategies, unifying and optimizing flattening, ironing, and fold line generation under the same strategy. This significantly reduces wrinkling, slippage, and uneven folding edges, enhancing the stability and controllability of flexible body folding and improving folding efficiency. In multi-station collaboration, it enables language routing and subtask decomposition without manual rules, resulting in stronger scheduling capabilities. Furthermore, it allocates the highest node weight to the folding VLA node to match the folding difficulty of the folding VLA node, allowing for greater focus on weaker nodes and higher multi-station collaboration.

[0041] In some alternative embodiments, after step S108 described above, the method further includes: S110, respectively obtain the task processing time and total task time for executing the clothing retrieval task, clothing folding task and clothing storage task; S112, if the total duration of the task exceeds the preset total duration constraint, then based on the processing time of each task, the node weights of the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node are redistributed through the weight gating configured during model training. The redistributed node weights satisfy that the sum of the node weights of the clothing folding VLA node, clothing storage VLA node, and clothing retrieval VLA node is 1 and the node weights decrease sequentially.

[0042] In some examples, learnable or adjustable weight gating can be configured during model training. The losses of the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node are injected into the overall objective function according to the node weights, and a timeout redistribution rule is introduced to redistribute the weights, so as to automatically increase the weights of weak nodes during the real trajectory collaborative fine-tuning stage.

[0043] The weighted gating loss and constraints are implemented as follows, and the total loss is defined as: (2) In equation (2), , , These represent the branch losses for retrieving clothes, folding clothes, and storing clothes, respectively. This represents the penalty coefficient for the L1 regularization term (a regularization hyperparameter, a fixed constant preset by the user). This represents the complete set of trainable parameters to be optimized, including the node weights of the weight separation subnetwork for each VLA node. The learnable parameters are as follows: , , All distribution parameters of the multivariate Gaussian mixture action head; Represents the parameter set The L1 norm is used to achieve sparse decoupling of weights and suppress overfitting.

[0044] In the above formula (2): ; ; ; .

[0045] in, , , The first weighted separation subnetwork of the clothing retrieval VLA node, the third weighted separation subnetwork of the clothing storage VLA node, and the second weighted separation subnetwork of the clothing folding VLA node, respectively, represent the task-specific features output by these subnetworks. Indicated by For conditions, the action vector of picking up clothes The probability distribution of clothing selection based on the observed values; Indicated by For conditions, storage action vector The probability distribution of the observed clothing values; Indicated by For conditions, folding clothes action vector The probability distribution of folded clothes for the observed values; This is a regularization term for the rate of change of the self-covering area of ​​the fabric, used to suppress violent pulling and excessive rolling: For a moment Self-occlusion area estimation, The projected area of ​​the fabric. It is a numerically stable term.

[0046] Node weight constraints satisfy: and .

[0047] Among them, the node weight constraint can reflect the intuition that folding clothes is the most difficult and taking clothes is the easiest.

[0048] Specific implementation methods include: using the softmax function to convert unconstrained parameters The initial weights are mapped as shown in equation (3) below: (3) Further achieve this through penalty constraints. For example, increasing penalties for samples that violate inequalities. .

[0049] In this embodiment, in order to enable the system to converge to an importance coefficient that matches the task difficulty without manual parameter tuning, the task processing time of each node (the time spent by the node in executing the subtask) is obtained. And obtain the total time taken for the system to perform natural language tasks. ,like (Given a preset total duration constraint), node weights are redistributed using the aforementioned weight gating. The optimal weight redistribution step size is specified. , can Select from the range. Among them, This is the average total task duration over a historical preset time period, which can be... Select from a range, for example, the historical preset time period is the previous 10 times.

[0050] In some examples, it can be The slowest node (the node with the longest execution time) has its node weight increased by 0.05, while the weights of the other two nodes are each decreased by 0.025, to maintain a total weight of 1. Of course, the above is only one possible approach, and the specific values ​​can be adjusted. After adjustment, the weight gating constraint must be met, and the node weight adjustment amount must be guaranteed to be 1.

[0051] In this embodiment, by automatically redistributing node weights based on task processing time, the system automatically focuses on weak nodes and converges to an importance coefficient that matches the task difficulty without manual parameter tuning. The converged weight ratio can be solidified into the firmware to form an importance coefficient configuration that can be reviewed, compared, and protected.

[0052] In some optional embodiments, step S102 specifically includes: S1021, Obtain the natural language task input by the user, and input the natural language task into the multimodal base; S1022, according to the natural language task, the drying area image of the flexible clothing to be stacked is acquired based on the visual encoder in the multimodal base, and language instructions are encoded based on the language encoder in the multimodal base. S1023, based on the drying area image of the flexible clothing to be stacked and the language command, feature fusion is performed through the cross-modal fusion network in the multimodal base to generate the clothing retrieval VLA node, clothing folding VLA node and clothing storage VLA node corresponding to the clothing retrieval VLA node, clothing folding VLA node and clothing storage VLA node.

[0053] In this embodiment, when a user issues a natural language task, the system first acquires the user's real-time input of the natural language task during the task execution. This task represents the user's customized, complete operational requirements for the flexible clothing to be folded, such as removing cotton T-shirts from the drying area, flattening and folding them before storing them in the upper drawer, or grabbing a drying shirt, neatly folding it, and placing it in a wardrobe storage compartment. Unlike traditional fixed-program single-operation instructions, issuing tasks via natural language allows for personalized operational needs tailored to different clothing types, folding specifications, and storage locations.

[0054] Furthermore, after acquiring the natural language task input from the user, the input is fed into a multimodal base shared by multiple VLA nodes in the system. The multimodal base serves as the basic module for common feature extraction in the entire clothing folding system. It keeps the backbone parameters frozen throughout the process, adapts only to the lightweight backend sub-network, and can uniformly accept both visual acquisition and language parsing inputs, providing general feature support for subsequent cross-modal fusion and task decomposition.

[0055] Furthermore, the multimodal base's built-in language encoder performs word-by-word and sentence-by-sentence semantic encoding and intent parsing on the user's input natural language task. This completes instruction segmentation, keyword extraction, and task semantic decomposition, accurately identifying the core elements of the task. These elements include the target items for folding flexible clothing (T-shirts, shirts, pants, etc.), the actions to be performed (picking up clothes, flattening, folding, and storing), the work specifications (half-folding, square folding, and neat storage), and the target storage location (drawers, wardrobes, and storage compartments). The system outputs structured semantic features of the language instructions, realizing the conversion of unstructured natural language instructions into machine-recognizable semantic features. Simultaneously, a visual encoder captures images of the drying area for folding flexible clothing and extracts visual features from the images.

[0056] Furthermore, semantic features can be obtained by encoding language instructions based on the language encoder. After independently extracting visual and semantic features, bidirectional feature alignment and deep fusion processing are performed based on semantic and visual features through a cross-modal fusion network in the multimodal base. The cross-modal fusion network establishes a one-to-one mapping between the visual features of clothing and the semantic features of natural language tasks, linking the real-time drying status of clothing with the user's customized folding requirements. Further, based on the fused multimodal global features and combined with the pre-defined VLA node task partitioning logic, the global task is hierarchically decomposed, accurately generating refined sub-tasks corresponding to the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node. For example, a single end-to-end model needs to operate within a unified action space. The entire learning process in China, and the above is achieved by breaking down tasks into... , , The effective search space is then determined by Downgraded to This addresses the combined learning problem, reducing the difficulty of end-to-end learning. Specifically, the pre-defined VLA node task division logic includes: the "Retrieve Clothes" sub-task, which addresses the needs for clothing grasping and positioning, posture adaptation, and precise picking; the "Fold Clothes" sub-task, which addresses the needs for clothing flattening, wrinkle elimination, posture correction, and standardized folding; and the "Store Clothes" sub-task, which addresses the needs for spatial positioning, orderly placement, and stacking of folded clothing.

[0057] In some examples, the aforementioned visual encoder is a ViT-B visual encoder, which divides the RGB image (the captured image of the drying area of ​​the flexible clothing to be stacked) into fixed-size patches, for example... The token sequence, obtained through linear projection, is input into a 12-layer Transformer Encoder. Each layer contains multi-head self-attention and two feedforward layers; for example, each layer contains 12 self-attention layers, and the hidden dimension of the feedforward network is 3072. GELU is used as the activation function, and the residual connections and LayerNorm adopt the standard Transformer structure, resulting in a final output dimension of 768. The ViT-B visual encoder employs a Transformer-based visual feature extraction architecture, possessing global context modeling and long-range dependency capture capabilities. Compared to traditional convolutional networks, it is more suitable for feature extraction scenarios involving non-rigid, easily deformable objects such as flexible clothing. The ViT-B visual encoder, through image block embedding, multi-head self-attention mechanism, and multi-layer coding transformation, fully captures the size and shape of the flexible clothing to be stacked, the density and distribution of folds, the deformation posture during drying and hanging, the position of edge contours, local self-occlusion areas, and the corresponding texture visual features of different fabrics. It can effectively distinguish the clothing itself, fold shadows, overlapping occlusion areas, and background interference, avoid background interference in complex drying environments, and output high-dimensional, fine-grained, and globally consistent clothing visual feature maps. This can provide high-precision visual feature support for subsequent estimation of clothing self-occlusion area, determination of deformation state, and positioning of folding baseline.

[0058] In some examples, the language encoder described above is a BERT-based language encoder. It employs a 12-layer Transformer Encoder with 768 hidden dimensions, 12 attention heads, and GELU activation. The input consists of a natural language task and a sequence of constructed subtask tokens. Semantic decomposition using the BERT-based language encoder provides powerful bidirectional semantic modeling capabilities. Compared to traditional unidirectional encoding models, it can comprehensively capture the semantic relationships and fine-grained constraints within natural language instructions, making it highly suitable for diverse and customized natural language task instructions in scenarios such as clothing folding.

[0059] In some examples, the cross-modal fusion network described above can employ cross-attention layers, using language tokens as queries and visual tokens as key / value pairs or bidirectional interactions to obtain task-related visual representations. The number of cross-attention layers can range from 2 to 4, or other numbers are also possible.

[0060] In this embodiment, a visual encoder can accurately capture images of the drying area to extract fine-grained visual features of the flexible clothing to be folded. A language encoder bidirectionally parses the semantic features in the natural language task. Then, a cross-modal fusion network is used to achieve deep alignment and binding of image visual features and text instruction semantic features, generating independent sub-tasks corresponding to each VLA node. This overcomes the shortcomings of traditional single visual recognition, which can only identify objects and cannot understand the user's personalized folding needs. Furthermore, the complete task is decomposed into hierarchical sub-tasks that can be adapted to the partitioned execution of multiple robotic arms, enabling the system to take into account both generalization and low-cost deployment. The single node has a smaller action space and a lower probability of mismatch, which reduces the end-to-end learning difficulty and improves the success rate of clothing folding.

[0061] In some optional embodiments, step S104 above includes: S1041, Obtain the clothing retrieval sub-task generated by the multimodal base based on the first interval route corresponding to the clothing retrieval VLA node; S1042, in response to the clothing retrieval sub-task, the first weight separation sub-network of the clothing retrieval VLA node performs hierarchical weight projection transformation based on the first bottleneck dimension, and performs clothing edge detection on the drying area of ​​the flexible clothing to be stacked based on the drying area image, so as to output the clothing retrieval action vector. S1043, the flexible garment to be folded is grasped according to the garment retrieval action vector, and the grasped flexible garment to be folded is transferred to the folding station through the mobile device deployed by the garment retrieval VLA node.

[0062] Among them, the action vector for retrieving clothing can be defined as a low-dimensional continuous action, a pick (pose + gripper opening and closing), where x, y, and z are the Cartesian coordinates of the end effector, which are the three-dimensional positions of the end effector of the robotic arm in the base coordinate system, used to specify the spatial position of the gripping point. This indicates the end-effector attitude angle, used to specify the gripping direction, such as the direction of clothing folds. This represents the gripper's opening and closing degree, typically a continuous value from 0 to 1: 0 indicates fully closed, and 1 indicates fully open, used to control the gripping force and gripping state. Of course, the above clothing retrieval motion vector can also be defined as a discrete gripping point and a continuous offset.

[0063] Furthermore, combined Figure 3 As shown, after the clothing retrieval subtask generated by the multimodal base is sent to the clothing retrieval VLA node based on the first interval routing, it is processed through the first weight separation sub-network of the clothing retrieval VLA node.

[0064] In some examples, the Adapter structure of the first weight separation subnetwork, the second weight separation subnetwork, and the third weight separation subnetwork above adopts a bottleneck-type two-layer MLP for hierarchical weight projection transformation, as shown in the following equation (4): h=W down z, , (4) Among them, W down Let z be the downprojection (dimensionality reduction matrix), and z be the output feature of the multimodal base, which is also the input feature vector of the Adapter structure. The bottleneck feature after activation is the result of nonlinear activation of the dimensionality-reduced feature. W is the activation function. up Let z be the upward projection weight matrix (upgraded dimension matrix), and z′ be the output feature vector of the Adapter.

[0065] In some examples, the aforementioned first bottleneck dimension d b1 =64 or d b1 =96, activation function Use ReLU or GELU. Through control... Make the number of Adapter parameters less than 3% of the total parameters.

[0066] Furthermore, 2D features are extracted from the images of the drying area acquired by the visual encoder. Based on the extracted 2D features, an edge detection network is used to detect the edges of the clothes in the drying area with stacked flexible clothing, and finally, a low-dimensional clothing retrieval action vector a is output. pick Based on the action vector a of retrieving clothes pick After picking up the clothing, the clothing is transferred to a designated area at the L2 folding station via a mobile device at the clothing retrieval VLA node. This mobile device can be a single robotic arm, a mobile cart, etc. The clothing retrieval node does not involve a high-dimensional GMM motion head.

[0067] In this embodiment, the targeted distribution of the clothing retrieval sub-task is achieved through the first interval routing. Based on the first weighted separation sub-network and the first bottleneck dimension, a lightweight hierarchical weight projection transformation is completed. Clothing retrieval-related features are extracted without changing the frozen multimodal backbone parameters. Simultaneously, the clothing edge is accurately detected by combining the drying area image to clarify the retrieval reference position, and the clothing retrieval action vector is output. This avoids the retrieval positioning deviation caused by the mixing of multi-task features and reduces the computational power consumption of repeated inference of large models. Furthermore, the first bottleneck dimension controls the parameter amount of the Adapter of the first weighted separation sub-network, reducing training and inference overhead. This enables real-time inference at the edge, reduces hardware costs, and facilitates multi-node expansion.

[0068] In some optional embodiments, step S106 above includes: S1061, Based on the second interval route corresponding to the folding VLA node, obtain the folding sub-task generated by the multimodal base; S1062, in response to the folding sub-task, the second weight separation sub-network of the folding VLA node performs a hierarchical weight projection transformation based on the second bottleneck dimension, estimates the self-occlusion area of ​​the flexible clothing to be folded, and generates the folding motion vector through the multivariate Gaussian mixture motion head, wherein the folding motion vector includes the local tangential displacement and the normal expected force of the flexible clothing to be folded at the folding station; S1063, based on the folding action vector, folding is performed at the folding station through the folding VLA node.

[0069] In this embodiment, for the folding process, the folding VLA node obtains the folding sub-task through the second interval routing, and the second bottleneck dimension d of the second weight separation sub-network Adapter2 is... b2 =128, the number of parameters is about 4.8% of the multimodal base, and the largest among the three VLA nodes. The hierarchical weight projection transformation of the second weight separation subnetwork Adapter2 is implemented based on the above equation (4).

[0070] The self-occlusion area of ​​the flexible clothing to be stacked can be estimated based on the point cloud or height map of the depth camera, and the stacking action vector can be output through the triple Gaussian mixture head. The stacking action vector is generated based on the multivariate Gaussian mixture action head, and the clothing is stacked by the dual robotic arms and the pressure plate. The stacking VLA node introduces the "contour-force-action" triple Gaussian mixture head, and the strategy is as follows (5): (5) in, The mixing weights of the k-th component; This refers to the local tangential displacement and rotation of the end effector in the local coordinate system of the fabric. The expected normal force at the end is used for flattening, ironing, and inhibiting slippage; This refers to the change in the opening and closing amount of the grippers or the gripper spacing; K can be... Within the range of optimization, K=5 is preferred to cover five types of folding elements: flattening, folding in half, folding in thirds, turning up sleeves, and flattening. , These are the mean vector and covariance matrix of the k-th component, respectively; the whole is obtained by superimposing multiple Gaussian weighted sums.

[0071] Considering the difficulty of the clothing stacking VLA node, the clothing retrieval VLA node and clothing storage VLA node do not use the high-dimensional action head of the clothing stacking VLA node. Therefore, the action space dimension is reduced by about 40% compared to the clothing stacking VLA node, in order to reduce cross-task mismatches.

[0072] In this embodiment, a multivariate Gaussian mixture motion head that integrates "contour-force-motion" is used for the VLA node, which has a higher execution difficulty in folding clothes. In VLA motion decoding, the local tangential displacement and normal expected force of the fabric are simultaneously included in the end-to-end output. This allows flattening, ironing, and fold line generation to be uniformly optimized under the same strategy, which can significantly reduce wrinkling, slippage, and uneven folds, and enhance the stability and controllability of flexible body folding. Furthermore, by controlling the parameter quantity of the Adapter of the second weight separation sub-network through the second bottleneck dimension, the training and inference overhead can be reduced, the hardware cost can be lowered, and multi-node expansion can be facilitated.

[0073] In some optional embodiments, step S1062 specifically includes: The folding clothes sub-task is used as the input feature of the second weight separation sub-network of the folding clothes VLA node. Based on the second bottleneck dimension, the input feature is subjected to hierarchical calculation of down-projection weight, nonlinear activation and up-projection weight, and the key features corresponding to the folding clothes sub-task are output. Identify the outline and fold boundaries of the flexible garment to be stacked, and estimate the self-covering area of ​​the flexible garment to be stacked; The key features and the self-occlusion area are input into the multivariate Gaussian mixture motion head, and a corresponding folding action strategy is generated according to the selected Gaussian component, including generating the local tangential displacement for flattening or folding clothes and the normal expected force for flattening clothes. The generated clothing folding action strategy is used to output the clothing folding action vector.

[0074] In this embodiment, after the multimodal task is decomposed and the folding clothes sub-task is generated, the multimodal fusion features representing the folding clothes requirement are used as the input features of the second weighted separation sub-network corresponding to the folding clothes VLA node. The second weighted separation sub-network is configured with a dedicated second bottleneck dimension, which differs from the bottleneck dimension parameters of the clothes retrieval and storage nodes, and can adapt to the refined extraction requirements of deformation features, wrinkle features, and occlusion features in clothing folding scenarios.

[0075] Furthermore, the second weighted separation sub-network sequentially performs hierarchical weighted projection transformation operations on the input features. Combining the above equation (4), the high-dimensional fusion features are first compressed using the lower projection weight matrix to filter redundant background features and irrelevant task noise, retaining the core basic features related to the folding shape of clothing. Then, a nonlinear activation function is used to perform a nonlinear mapping transformation on the compressed bottleneck features, effectively mining the complex, non-rigid latent feature associations such as the deformation of flexible clothing folds and overlapping layers, compensating for the deficiency that linear transformation cannot characterize the nonlinear deformation law of fabric. Finally, the features are restored using the upper projection weight matrix, reconstructing key features adapted to the folding sub-task and possessing both fine-grained details and global semantic information, achieving accurate purification and adaptive transformation of common multimodal features into folding-specific task features.

[0076] Furthermore, RGB images of the garments to be folded at the folding station are acquired using a depth camera. Based on these RGB images, garment contours are extracted and fold boundaries are detected, precisely pinpointing the outer contour edges, fold boundaries, and fabric layer overlap areas of the flexible garments to be folded, thus identifying the current irregular deformation state and spatial overlap distribution of the garments. Based on this, the self-occlusion area of ​​the flexible garments to be folded can be quantitatively estimated by combining the fold density, fabric layer coverage, and the proportion of contour overlap areas. This method quantifies the uncertainties in visual occlusion, fabric overlap, and deformation that occur during clothing folding, providing a quantifiable state basis for adaptive fine-tuning of subsequent folding actions. It effectively solves the problem of inaccurate state perception caused by the transparency, easy deformation, and easy overlap of flexible clothing.

[0077] Furthermore, the calculated self-occlusion area and extracted key features are input into a multivariate Gaussian Mixture Model (GMM) motion head for fusion analysis. Leveraging the multi-component probabilistic modeling characteristics of the Gaussian Mixture Model, the limitations of traditional single-action output are overcome. The multivariate GMM motion head adaptively selects and matches the optimal Gaussian components based on the input occlusion and deformation features. It models the motion distribution of folding operations using the mixing coefficients, mean, and covariance parameters of each Gaussian component, generating a folding action strategy adapted to the current garment state. This folding action strategy covers at least five folding primitives: flattening, folding in half, folding in thirds, turning up sleeves, and flattening. Different primitives correspond to different Gaussian components k. Therefore, after analysis based on key features and self-occlusion area, the customer's folding requirements can be confirmed, allowing for the selection of the corresponding Gaussian component. The folding action strategy includes at least two core control dimensions: local tangential displacement for flattening, aligning, and / or folding operations, and normal expected force for flattening, smoothing wrinkles, and / or suppressing fabric slippage. Among them, the local tangential displacement can adaptively adjust the fabric stretching, translation and folding trajectory according to the distribution of wrinkles and the position of obstruction, so as to ensure that the folding path fits the real-time deformation state of the garment; the normal expected force can dynamically adapt the pressing force according to the size of the self-obstruction area. When the obstruction overlap area increases, the pressing force is appropriately increased to eliminate loose wrinkles in multiple layers of fabric, and the contact force is reduced in thin, unobstructed areas to avoid squeezing and damaging the fabric.

[0078] Furthermore, the multi-group GMM motion head completes motion sampling and parameter integration based on the selected Gaussian components, integrates adaptive tangential displacement trajectory and normal force control parameters, and outputs a folding motion vector containing displacement information and force control information. This enables adaptive, refined, and force-controlled folding operation output that addresses the deformation and self-occlusion interference of flexible clothing, ensuring that flexible clothing of different materials, different fold states, and different degrees of overlap can be folded neatly, stably, and without damage.

[0079] In this embodiment, by extracting key features based on the folding sub-task and performing self-occlusion area analysis on the flexible clothing to be folded, the multivariate GMM motion head can generate the optimal folding motion strategy by combining the task with the current state of the clothing. Furthermore, by injecting local tangential displacement and normal expected force into the folding motion strategy of the multivariate GMM motion head, flattening, ironing, and fold line generation are uniformly optimized under the same strategy, significantly reducing problems such as wrinkling, slippage, and uneven folds, thus achieving efficient and accurate folding.

[0080] In some optional embodiments, step S108 above includes: S1081, Based on the third interval route corresponding to the clothing storage VLA node, obtain the clothing storage sub-task generated by the multimodal base and identify the target storage location; S1082, in response to the clothing storage sub-task, the clothing retrieval VLA node performs hierarchical weight projection transformation based on the third bottleneck dimension through the third weight separation sub-network to generate the clothing storage action vector. S1083, Based on the clothing storage action vector, the folded clothing is transferred from the folding station to the target storage station.

[0081] The action vector for storing clothes can be defined as follows: ,in, This is the token for the third-interval route. The VLA node only notices this through the third-interval route. The third bottleneck dimension d of the third weighted sub-network Adapter3, which is the interval (clothing storage interval). b3 =64.

[0082] Furthermore, the clothing storage sub-task based on visual and language analysis includes the user's clothing storage needs, such as placing folded clothes into the first drawer of the first row. Therefore, the target storage location can be extracted from the clothing storage sub-task. Based on the above equation (4), in response to the clothing storage sub-task, the feature-fused clothing storage sub-task is input into the third weight separation sub-network and a hierarchical weight projection transformation is performed based on the third bottleneck dimension. Finally, a clothing storage action vector is generated. Based on the clothing storage action vector, the clothes on the folding station can be transferred to the target storage location by a guide rail trolley or a robotic arm. Among them, after extracting the target storage location, the clothing storage VLA node can scan the clothing storage area with a camera, perform path planning based on the scanned target storage location, and transfer and store the clothes according to the planned path.

[0083] In some examples, the clothing storage VLA node can capture images using a camera and combine them with classification algorithms to classify and identify clothing. If the user's natural language task does not include a target storage location for clothing, the clothing storage VLA node can search historical storage data based on the identified clothing type to categorize similar or identical clothing to estimate the target storage location before storing the clothing. For example, if the customer's natural language task is: "Put away the socks in the drying area," the clothing storage VLA node, upon reaching the folding station, can capture images of the socks using a camera, classify them, search for the storage locations of the three most recently received socks, and then place the socks accordingly. It can even differentiate between items for adults and children.

[0084] In this embodiment, segmented semantics of language routing tokens are used to distinguish independent task areas for clothing storage operations, achieving precise task diversion and accurate workstation positioning for multiple processes such as clothing retrieval, folding, and storage. This effectively avoids interference from mixed features of multiple processes affecting the accuracy of storage location identification. A separate third-weight separation sub-network performs hierarchical weight projection transformation based on the third bottleneck dimension. This allows for analysis based on a dedicated trainable network with completely weighted separation from the clothing retrieval and folding nodes, building upon the frozen backbone network of the multimodal base. This enables the generation of targeted storage solutions adapted to clothing transfer, alignment, and stacking. The clothing motion vector retains the general perception and generalization capabilities of the multimodal large model, and through the hierarchical adaptation mechanism of node interval routing and weight separation sub-network, it realizes lightweight, customized, and high-precision feature adaptation and motion generation for clothing storage tasks. This significantly improves the positional accuracy, posture regularity, and scene adaptability of clothing transfer and storage in multi-station assembly line operations, while reducing the overall model training cost and hardware deployment cost. It effectively solves the problems of poor adaptability of fixed actions, large storage location recognition error, skewed storage of stacked clothing, disordered stacking, and crosstalk of features from multiple processes in traditional clothing storage operations.

[0085] In some alternative embodiments, combined with Figure 3 As shown, the entire process of collecting, folding, and storing clothing can be monitored, and any abnormal events that occur can be reported to the customer, including through voice and text.

[0086] In some optional embodiments, prior to step S102 above, the method further includes: training the model, specifically including: Step 1: Offline pre-trained multimodal base It uses approximately 2 million video-text pairs of "rigid object" grabbing and placement. Preprocessing includes: image normalization to... Standardized according to ImageNet statistics; resolution unified to Text segmentation (WordPiece / BPE) is performed and truncated to 128 tokens. The Transformer layer is initialized with Xavier, and the Adam optimizer parameters are... , Learning rate (With learning rate warm-up and cosine annealing), weight decay is... The batch size is 32, and the training iterations are performed for 100 rounds.

[0087] Step 2: Flexible Data Distillation: Train only the folding action and the folding action head. Use NVIDIA Flex physics simulation to generate 100,000 pseudo-labeled trajectories of "fabric-language-action". Freeze the weights of the clothing retrieval and storage nodes. The learning rate range is [range missing]. The best model is selected based on the success rate of folding the validation set.

[0088] Step 3: Perform three-node online training with real device collaborative fine-tuning, recording the complete trajectory after each real user command is completed. The weight redistribution rule is as follows: record the time consumed by each node. If the total time ( If the average value is historical, then the weight of the slowest node will be increased. The remaining two nodes each decreased Keeping the sum constant at 1. After repeating this process approximately 5000 times, the weight ratio converges to approximately... .

[0089] During training, data anonymization measures include detecting and occluding or blurring faces / sensitive regions, and adding differential privacy noise to pixels in sensitive regions when necessary. Privacy budgets include... Regarding the handling of anomalous samples, samples with severe occlusion, strong reflection, camera defocus, or robotic arm exceeding limits were labeled as anomalous and statistically analyzed separately for robustness assessment and regression testing.

[0090] To better illustrate the weight-based hierarchical collaborative clothing stacking method provided in this application, the following description is based on some optional specific embodiments.

[0091] Combination Figure 4 As shown, in some examples, the user input natural language task T is: fold the shirt on the balcony and put it in drawer B.

[0092] The first step is for user layer L0 to... Input-shared multimodal base The multimodal base uses ViT-B (12 layers, 768 dimensions) to process the balcony area image, and generates a subtask token sequence after cross-modal fusion by encoding language instructions through BERT-base (12 layers, 768 dimensions): .

[0093] The second step involves the VLA node L1, which only notices the token range routing. In this range, its lightweight Adapter1 adopts a bottleneck dimension. The number of parameters is approximately 2.5% of that of the base. Based on the 2D features output by the visual encoder ViT, clothing edge detection is performed in the clothes rack area, and low-dimensional grasping actions are output. The clothing is then transferred to the designated area of ​​the folding platform at the clothing folding VLA node. The clothing retrieval VLA node does not involve a high-dimensional GMM motion head.

[0094] Third, the VLA node L2, through token range routing, only noticed... The bottleneck dimension of its Adapter2 within the range The number of parameters is approximately 4.8% of the base, the largest among the three nodes. The self-occlusion area is estimated based on the depth camera point cloud or height map. ; via GMM motion head ( Component) Output action sequence: Step 1: Select the "Pave" primitive (corresponding to Gaussian component) Weight The end effector outputs tangential displacement. Lay the shirt flat; Step 2: Select the "Fold in Half" primitive ( ), fold along the center line; step 3 select the "flip sleeve" element ( ) Process the cuffs; Step 4: Select the "Flatten" module ( Apply normal expectation force Complete the flattening process. The folding action strategy incorporates the above formula (2). The regularization term suppresses the pulling, and the spatial dimension of the entire motion is about 40% higher than that of the clothing retrieval node.

[0095] Fourth, the VLA node L3 in the clothing storage area only notices the token range routing. The interval, its Adapter3 bottleneck dimension After parsing the target drawer ID, the guide rail trolley is driven to the top of the drawer, and the clothing storage motion vector is output. Complete the second sorting and close the drawer.

[0096] Step 5, if the total time of this task is... If this occurs, a weight redistribution will be triggered: the weight of the timed-out node will be increased. The remaining two nodes each decreased ,Keep For example, if folding clothes takes a long time on the first execution, the weight is adjusted accordingly. After convergence of 5000 trajectories, the typical weight ratio is: Satisfying constraints This reflects that retrieving clothes is the easiest task, while folding them is the most difficult. This sample record is used for subsequent fine-tuning (learning rate) of the Adam (Adaptive Moment Estimation) optimizer. ).

[0097] In other examples, the user input natural language task T is: fold the bath towel into thirds and place it in the second compartment from the left. The first step is for the user layer L0 to... Input shared dock (ViT-B / BERT-base, 768-dimensional), after cross-modal fusion, a subtask token sequence is generated: .in, This indicates that the fold type is three-fold, and the base encodes this semantic and routes it to the folding VLA node L2.

[0098] The second step involves the VLA node L1, which only notices the token range routing. The interval, its Adapter1 bottleneck dimension Approximately 2.5% of the base parameters. Corner detection is performed on the bath towel based on the 2D features output by ViT, outputting a low-dimensional grasping action. The process involves small swinging and releasing motions to shake and flatten the towel before transferring it to the folding table. The motion space dimension of the VLA node for retrieving clothes is approximately 5 dimensions, and it does not involve the GMM high-dimensional motion head.

[0099] Third, the VLA node L2, through token range routing, only noticed... The interval, its Adapter2 bottleneck dimension The self-occlusion area is estimated based on RGB-D data collected by a depth camera, with approximately 3.6% of the base parameters. GMM Action Head ( In the composition, the "triple-fold" primitive corresponds to the Gaussian component. (through mixed weights) (Automatic selection). During execution, the end effector outputs a five-dimensional motion representation of "profile-force-motion". The first step is to fold the fold along its length by 1 / 3 (primarily tangential displacement); the second step is to fold it again by 1 / 3 to form a three-fold; a normal expected force is applied at each fold line. (Thicker bath towels require more pressure) to flatten. Throughout the process, the self-occlusion changes are monitored using the regularization term in the above formula (2) to prevent excessive rolling. The action space dimension of the folding clothes VLA node is approximately 9 dimensions, which is about 40% higher than that of the retrieval clothes VLA node.

[0100] Fourth, the VLA node L3 in the clothing storage area only notices the token range routing. The interval, its Adapter3 bottleneck dimension The target grid ID is resolved as follows: The rear-drive guide rail trolley moves to the designated position, outputting the placement action. Placement complete. Visually confirm that the stack height does not exceed the threshold. If necessary, perform secondary flattening.

[0101] Fifth step: Record the time taken at each node after the task is completed. Assume this time... Total time Exceed (assuming historical average) If the VLA node is the slowest, then a weight redistribution will be triggered: the VLA node will be the slowest, and its node weight will be... Increase The remaining two nodes each decreased The updated weight ratio tends to ,satisfy Constraints: This sample is included in the Adam fine-tuning batch (learning rate). (with cosine annealing).

[0102] According to another aspect of the embodiments of this application, such as Figure 5 As shown, corresponding to the weight-based hierarchical collaborative clothing stacking method in the above embodiments, this embodiment provides a weight-based hierarchical collaborative clothing stacking system, the system comprising: The task splitting module 501 is used to generate subtasks corresponding to each VLA node based on the acquired natural language task through a multimodal base shared by multiple VLA nodes. Each VLA node is configured with a weight separation subnetwork and node weights. The garment retrieval module 503 is used to retrieve garment sub-tasks based on the first weighted separation sub-network and the first interval routing response of the garment retrieval VLA node, retrieve the flexible garments to be folded and transfer them to the folding station. The folding module 505 is used to respond to the folding sub-task based on the second weight separation sub-network and the second interval routing based on the folding VLA node. It generates a folding action vector that superimposes local tangential displacement and normal expected force through a multivariate Gaussian mixture action head, and performs folding at the folding station based on the folding action vector. The clothing storage module 507 is used to respond to the clothing storage sub-task based on the third weight separation sub-network and the third interval routing of the clothing storage VLA node, and to transfer the folded clothing from the folding station to the target storage location.

[0103] It should be noted that in this embodiment, the task splitting module 501 can be used to execute step S102 in this application embodiment, the clothes retrieval module 503 in this embodiment can be used to execute step S104 in this application embodiment, the clothes folding module 505 in this embodiment can be used to execute step S106 in this application embodiment, and the clothes storage module 507 in this embodiment can be used to execute step S108 in this application embodiment.

[0104] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should also be noted that the above modules, as part of a device, can operate in environments such as... Figure 1 The hardware environment shown can be implemented either through software or through hardware.

[0105] It should be noted that the suffixes such as "module" used to indicate elements in the above-described apparatus are only for the purpose of illustrative purposes and have no specific meaning in themselves. Therefore, they can be used in combination.

[0106] According to another aspect of the embodiments of this application, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the weighted hierarchical collaborative clothing stacking method in any of the above embodiments.

[0107] According to another aspect of the embodiments of this application, this application also provides a clothing folding device, such as... Figure 6 As shown, the device includes a memory 601, a processor 603, and a network interface 605. The memory 601 stores a computer program that can run on the processor 603. The memory 601 and the processor 603 communicate through the network interface 605 and a communication bus 607. When the electronic device is running, the processor 603 communicates with the memory 601 through the network interface 605. When the processor 603 executes the computer program, it implements the steps of the weight-based hierarchical collaborative clothing folding method described above. The clothing folding device described above can execute the weight-based hierarchical collaborative clothing folding method of this application embodiment through the weight-based hierarchical collaborative clothing folding system, thereby achieving a combination of... Figure 1 This describes a weight-based hierarchical collaborative method for stacking clothing.

[0108] The clothing folding device provided in this application embodiment can specifically be a module capable of communication or a terminal device containing such a module; the module can specifically be a wireless communication module, such as any one of 2G communication module, 3G communication module, 4G communication module, 5G communication module, NB-IoT communication module, etc.

[0109] The memory and processor in the aforementioned electronic device communicate with each other via a communication bus and a communication interface. The communication bus can be a peripheral component interconnect standard (PCI) bus or an extended industry standard structure (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. The memory can include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor. The aforementioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0110] It is understood that the embodiments described herein can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits, digital signal processors, digital signal processing devices, microprocessors, and other electronic units or combinations thereof for performing the functions described herein. For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. Software code can be stored in memory and executed by a processor.

[0111] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0112] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or unit, and may be electrical, mechanical or other forms.

[0113] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0114] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or parts of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0115] It should be noted that, in this document, relational terms such as first, second, etc., are used only to distinguish one entity or operation from another entity or operation. The terms include, encompass, or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0116] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A weight-based hierarchical collaborative method for stacking clothing, characterized in that, The method includes: Through a multimodal base shared by multiple VLA nodes, subtasks corresponding to each VLA node are generated based on the acquired natural language task. Each VLA node is configured with a weight separation subnetwork and node weights. Based on the first weighted separation subnetwork and the first interval routing response of the clothing retrieval VLA node, the clothing retrieval sub-task is carried out to retrieve the flexible clothing to be folded and transfer it to the folding station; Based on the second weight separation subnetwork and the second interval routing response of the folding subtask of the folding VLA node, a folding action vector with superimposed local tangential displacement and normal expected force is generated by a multivariate Gaussian mixture action head, and folding is performed at the folding station based on the folding action vector. Based on the third weight separation subnetwork and the third interval routing response of the clothing storage VLA node, the folded clothes are transferred from the folding station to the target storage location. Among them, the node weight of the folding VLA node is the largest.

2. The weight-based hierarchical collaborative clothing stacking method according to claim 1, characterized in that, After the third weighted separation subnetwork and the third interval routing response clothing storage subtask based on the clothing storage VLA node transfer the folded clothing from the folding station to the target storage location, the method further includes: Obtain the task processing time and total task time for executing the clothing retrieval task, clothing folding task, and clothing storage task, respectively; If the total duration of the task exceeds the preset total duration constraint, then based on the processing time of each task, the node weights of the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node are redistributed through the weight gating configured during model training. The redistributed node weights satisfy that the sum of the node weights of the clothing folding VLA node, clothing storage VLA node, and clothing retrieval VLA node is 1, and the node weights decrease sequentially.

3. The weight-based hierarchical collaborative clothing stacking method according to claim 1, characterized in that, The multimodal base shared by multiple VLA nodes generates subtasks corresponding to each VLA node based on the acquired natural language task, including: Obtain the natural language task input by the user and input the natural language task into the multimodal base; According to the natural language task, the drying area image of the flexible clothing to be stacked is acquired based on the visual encoder in the multimodal base, and language instructions are encoded based on the language encoder in the multimodal base. Based on the image of the drying area of ​​the flexible clothing to be folded and the language instructions, feature fusion is performed through the cross-modal fusion network in the multimodal base to generate the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node corresponding to the clothing retrieval VLA node, clothing folding VLA node, and clothing storage VLA node.

4. The weight-based hierarchical collaborative clothing stacking method according to claim 3, characterized in that, The first weighted separation subnetwork and the first interval routing response garment retrieval sub-task based on the garment retrieval VLA node retrieve the flexible garments to be folded and transfer them to the folding station, including: The clothing retrieval sub-task generated by the multimodal base is obtained based on the first interval route corresponding to the clothing retrieval VLA node; In response to the clothing retrieval sub-task, the first weight separation sub-network of the clothing retrieval VLA node performs hierarchical weight projection transformation based on the first bottleneck dimension, and performs clothing edge detection on the drying area of ​​the flexible clothing to be stacked based on the drying area image, so as to output the clothing retrieval action vector. The garment is grasped according to the garment-grabbing motion vector and then transferred to the folding station via a mobile device deployed at the garment-grabbing VLA node.

5. The weight-based hierarchical collaborative clothing stacking method according to claim 1, characterized in that, The second weighted separation subnetwork and the second interval routing response folding subtask based on the folding VLA node generate a folding action vector superimposed with local tangential displacement and normal expected force through a multivariate Gaussian mixture action head, and perform folding at the folding station based on the folding action vector, including: The folding sub-task generated by the multimodal base is obtained based on the second interval route corresponding to the folding VLA node; In response to the folding sub-task, the second weight separation sub-network of the folding VLA node performs a hierarchical weight projection transformation based on the second bottleneck dimension, estimates the self-occlusion area of ​​the flexible clothing to be folded, and generates the folding motion vector through the multivariate Gaussian mixture motion head. The folding motion vector includes the local tangential displacement and the normal expected force of the flexible clothing to be folded at the folding station. Based on the folding action vector, folding is performed at the folding station through the folding VLA node.

6. The weight-based hierarchical collaborative clothing stacking method according to claim 5, characterized in that, In response to the folding subtask, the second weight separation subnetwork of the folding VLA node performs a hierarchical weight projection transformation based on the second bottleneck dimension, estimates the self-occlusion area of ​​the flexible clothing to be folded, and generates the folding action vector through the multivariate Gaussian mixture action head, including: The folding clothes sub-task is used as the input feature of the second weight separation sub-network of the folding clothes VLA node. Based on the second bottleneck dimension, the input feature is subjected to hierarchical calculation of down-projection weight, nonlinear activation and up-projection weight, and the key features corresponding to the folding clothes sub-task are output. Identify the outline and fold boundaries of the flexible garment to be stacked, and estimate the self-covering area of ​​the flexible garment to be stacked; The key features and the self-occlusion area are input into the multivariate Gaussian mixture motion head, and a corresponding folding action strategy is generated according to the selected Gaussian component, including generating the local tangential displacement for flattening or folding clothes and the normal expected force for flattening clothes. The generated clothing folding action strategy is used to output the clothing folding action vector.

7. The weight-based hierarchical collaborative clothing stacking method according to any one of claims 1 to 6, characterized in that, The third weighted separation subnetwork and third interval routing response to the clothing storage subtask based on the clothing storage VLA node transfer the folded clothing from the folding station to the target storage location, including: Based on the third interval route corresponding to the clothing storage VLA node, the clothing storage sub-task generated by the multimodal base is obtained and the target storage location is identified; In response to the clothing storage sub-task, the third weight separation sub-network of the clothing retrieval VLA node performs hierarchical weight projection transformation based on the third bottleneck dimension to generate a clothing storage action vector. Based on the clothing storage motion vector, the folded clothing is transferred from the folding station to the target storage location.

8. A weight-based hierarchical collaborative clothing stacking system, characterized in that, The system includes: The task splitting module is used to generate subtasks corresponding to each VLA node based on the acquired natural language task through a multimodal base shared by multiple VLA nodes. Each VLA node is configured with a weight separation subnetwork and node weights. The garment retrieval module is used to retrieve garments from the stacked flexible garments and transfer them to the folding station based on the first weighted separation subnetwork and the first interval routing response of the garment retrieval VLA node. The folding module is used to respond to the folding sub-task based on the second weight separation sub-network and the second interval routing based on the folding VLA node. It generates a folding action vector that superimposes local tangential displacement and normal expected force through a multivariate Gaussian mixture action head, and performs folding at the folding station based on the folding action vector. The clothing storage module is used to respond to the clothing storage sub-task based on the third weight separation sub-network and the third interval routing of the clothing storage VLA node, and to transfer the folded clothing from the folding station to the target storage location.

9. A clothing folding device, comprising: The device includes a processor, a memory, and a network interface, wherein the memory stores machine-readable instructions executable by the processor, characterized in that: when the garment folding device is in operation, the processor communicates with the memory via the network interface, and the processor executes the machine-readable instructions to perform the steps of the weighted hierarchical collaborative garment folding method as described in any one of claims 1 to 7.

10. A storage medium having processor-executable non-volatile program code, characterized in that, The program code causes the processor to execute the steps of the weighted hierarchical collaborative clothing stacking method according to any one of claims 1 to 7.