A humanoid robot garment wrinkle detection and posture fine-tuning method and system

By simultaneously acquiring multi-view images and tension data, and combining multi-head cross-modal attention fusion, the system achieves accurate detection and automatic adjustment of clothing folds on humanoid robots, solving the problem of reliance on manual intervention in existing technologies and improving display quality and efficiency.

CN122241442APending Publication Date: 2026-06-19SHENZHEN PINKUO INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN PINKUO INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-05-21
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

In existing technologies, after humanoid robots switch poses, their clothing develops flaws such as wrinkles, pile-ups, and skewing, leading to a decrease in shooting quality. Furthermore, single-modal perception solutions cannot accurately determine the type of wrinkles and adjust their direction, requiring extensive manual intervention.

Method used

By simultaneously acquiring multi-view images and tension data, and combining multi-head cross-modal attention fusion, visual wrinkle features and tension distribution features are extracted, classified into different types of wrinkles, and the joint posture fine-tuning is calculated to achieve automatic adjustment.

Benefits of technology

It enables precise judgment and automatic adjustment of clothing folds, reducing manual intervention and improving the efficiency and consistency of display preparation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122241442A_ABST
    Figure CN122241442A_ABST
Patent Text Reader

Abstract

This invention relates to the field of visual inspection technology, and discloses a method and system for detecting and fine-tuning the posture of clothing on a humanoid robot. The method involves: acquiring multi-view images of the humanoid robot wearing clothing and tension data for each body region; extracting visual wrinkle features and tension distribution features for each body region, aligning and fusing the visual wrinkle features and tension distribution features to obtain a wrinkle state vector for each body region; classifying each body region into normal drooping wrinkles, slight pile wrinkles, obvious stretching wrinkles, and severely skewed wrinkles, and determining the areas to be adjusted; calculating the joint posture fine-tuning amount for the areas to be adjusted and performing the adjustment until all body regions are either normal drooping wrinkles or slight pile wrinkles, and then outputting a shooting ready signal. This invention solves the technical problem that single-modal perception schemes cannot simultaneously acquire wrinkle morphology information and clothing force information, eliminating the reliance on repeated manual intervention during clothing display preparation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual inspection technology, and in particular to a method and system for detecting wrinkles in clothing and fine-tuning the posture of a humanoid robot. Background Technology

[0002] In the apparel e-commerce industry, the quality of product display images directly impacts consumers' purchasing decisions. Humanoid robots, as an emerging solution to replace live models, have been widely used in apparel display photography. However, when humanoid robots change their display postures, the distribution of clothing on their body surface changes, inevitably resulting in display defects such as wrinkles, bunching, and skewing, severely affecting the quality of the photos.

[0003] One existing solution uses machine vision to detect the state of garment folds, but this purely visual method cannot distinguish between normal folds caused by natural drape and flawed folds that affect the display effect, resulting in a high false positive rate. Another solution uses tension sensors to detect the stress state of the garment, but this purely tension sensing method cannot obtain the spatial shape, direction, and position information of the folds, and cannot accurately guide the posture adjustment direction. Both of these single-modal sensing solutions struggle to accurately determine the garment display state. Furthermore, it cannot automatically verify whether the robot's posture adjustment effectively eliminates fold defects, and it cannot automatically revert to alternative directions when the adjustment is ineffective, ultimately requiring significant manual intervention. This leads to low efficiency and poor quality consistency in garment display preparation. Summary of the Invention

[0004] The main objective of this invention is to provide a method and system for detecting and fine-tuning the posture of humanoid robot clothing wrinkles. This invention solves the technical problem that single-modal perception schemes cannot simultaneously acquire wrinkle morphology information and clothing force information, and eliminates the reliance on repeated manual intervention in the clothing display preparation process.

[0005] To achieve the above objectives, the present invention provides a method for detecting clothing wrinkles and fine-tuning the posture of a humanoid robot, comprising the following steps: S1: Simultaneously collect multi-view images of the humanoid robot wearing clothing and tension data of various body areas; S2: Extract visual wrinkle features of each body region from the multi-view display image, calculate the tension distribution features of each body region based on the tension data, and align and fuse the visual wrinkle features with the tension distribution features to obtain the wrinkle state vector of each body region. S3: Based on the wrinkle state vector, classify each body region into normal overhang wrinkles, slight pile wrinkles, obvious stretching wrinkles, and severe skewed wrinkles, and determine the areas to be adjusted; S4: Calculate the joint posture fine-tuning amount of the area to be adjusted and send the joint posture fine-tuning amount to the humanoid robot through the robot control interface to perform the adjustment. After the clothing fabric is relaxed and stable, repeat steps S1-S3 until all body areas are normal hanging wrinkles or slight pile wrinkles, and output the shooting ready signal.

[0006] Optionally, in a first implementation of the first aspect of the present invention, step S1 includes: S11: Multiple cameras positioned on the front, sides, and back of the humanoid robot are simultaneously triggered by hardware trigger signals to capture multi-view images of the robot wearing clothing. S12: The contact tension values ​​of each sensor node are collected by the sensor array arranged in each body area of ​​the humanoid robot, and the contact tension values ​​of each sensor node are time-stamped and aligned according to the acquisition timestamp of the multi-view display image to obtain the tension data of each body area.

[0007] Optionally, in a second implementation of the first aspect of the present invention, step S2 includes: S21: Perform human body region segmentation on the multi-view display image, and extract the fold direction angle, fold density and fold depth of each body region as visual fold features. S22: Perform regional aggregation on the tension data, and extract the regional average tension, tension standard deviation and tension gradient of adjacent regions of each body region as tension distribution features; S23: Using the visual wrinkle features of each body region as the query vector and the tension distribution features of each body region as the key vector, a multi-head cross-modal attention mechanism is used to fuse the wrinkle state vectors of each body region.

[0008] Optionally, in a third implementation of the first aspect of the present invention, step S23 includes: Multi-head cross-modal attention is calculated and concatenated using the visual wrinkle features of each body region as query vectors and the tension distribution features of each body region as key vectors to obtain cross-modal attention features; The visual wrinkle features and the tension distribution features are concatenated to calculate the gating weights. Based on the gating weights, the cross-modal attention features are weighted and fused to obtain the wrinkle state vectors of each body region.

[0009] Optionally, in a fourth implementation of the first aspect of the present invention, step S3 includes: S31: Calculate the rule classification probability of each body region based on the region average tension, fold direction angle and fold density in the fold state vector of each body region; S32: Input the wrinkle state vectors of each body region into the fully connected classification network for forward inference to obtain the model classification probability of each body region; S33: The classification probability of the rule and the classification probability of the model are weighted and fused based on the weighting coefficient to obtain the fused classification probability. The classification result of each body region is determined according to the fused classification probability. The classification result is normal drooping folds, slight stacking folds, obvious stretching folds or severe skewed folds. S34: The body areas that are classified as having obvious stretching wrinkles or severely skewed wrinkles are identified as areas to be adjusted.

[0010] Optionally, in a fifth implementation of the first aspect of the present invention, step S31 includes: When the average tension of each body region is within the preset normal tension range, the fold direction angle is less than the first direction angle threshold, and the fold density is less than the first density threshold, the rule determination result of the corresponding body region is determined to be a normal overhang fold. When the average tension of the region is lower than a preset multiple of the lower limit of the normal tension range, the fold direction angle is less than the first direction angle threshold, and the fold density is between the first density threshold and the second density threshold, the rule determination result of the corresponding body region is determined to be slight accumulation folds. When the average tension of each body region is higher than a preset multiple of the upper limit of the normal tension range, the fold direction angle is greater than the second direction angle threshold, and the fold density is greater than the first density threshold, the rule judgment result of the corresponding body region is determined to be obvious stretching folds. When the tension gradient of adjacent regions of each body region exceeds the tension gradient threshold, the fold direction angle dispersion exceeds the direction disorder threshold, and the fold density is greater than the second density threshold, the rule judgment result of the corresponding body region is determined to be a severely skewed fold. The probability of the category corresponding to the rule judgment result is recorded as 1, and the probability of the other categories is recorded as 0, so as to obtain the rule classification probability of each body region.

[0011] Optionally, in a sixth implementation of the first aspect of the present invention, step S4 includes: S41: Determine the set of joints corresponding to the region to be adjusted, and calculate the joint posture fine-tuning amount of the set of joints corresponding to the region to be adjusted; S42: The joint posture fine-tuning amount is sent to the joint set of the humanoid robot through the robot control interface and executed. During the execution, if the contact tension value of any body area exceeds the safe tension threshold, the execution is stopped immediately. After the clothing fabric is relaxed and stabilized, steps S1-S3 are executed again to calculate the first wrinkle comprehensive score before adjustment and the second wrinkle comprehensive score after adjustment. S43: If the second fold comprehensive score is higher than the first fold comprehensive score and there are no newly added severely skewed fold areas, then the adjustment is confirmed to be effective. When the classification results of all body areas are normal hanging folds or slight piled folds, output the shooting ready signal.

[0012] Optionally, in a seventh implementation of the first aspect of the present invention, S41 includes: S411: Map the region to be adjusted to the corresponding joint set; S412: For the adjustment area with obvious stretching and wrinkling, calculate the angle increment of the corresponding joint based on the deviation between the average tension of the area to be adjusted and the target tension and the deviation of the wrinkle direction angle, and obtain the joint posture fine adjustment amount. S413: For the adjustment area with severe skewed folds, construct the desired garment displacement correction vector based on the fold direction angle deviation of the adjustment area and the tension gradient of the adjacent area. Solve the inverse kinematics of the desired garment displacement correction vector based on the Jacobian matrix of the joint set to obtain the joint posture fine adjustment amount.

[0013] Optionally, in an eighth implementation of the first aspect of the present invention, step S413 includes: The garment displacement correction direction is determined based on the fold direction angle deviation of the area to be adjusted, and the garment displacement correction amount is determined based on the excessive tension gradient of the adjacent areas of the area to be adjusted, and the desired garment displacement correction vector is constructed. Based on the kinematic model of each joint in the joint set, the Jacobian matrix corresponding to the joint set is calculated analytically. The desired clothing displacement correction vector is input into a sequential quadratic programming optimizer constrained by the Jacobian matrix to perform inverse kinematics solution, thereby obtaining the joint posture fine-tuning amount.

[0014] This invention also provides a system for detecting clothing wrinkles and fine-tuning posture of a humanoid robot, comprising: The acquisition module is used to simultaneously acquire multi-view images of the humanoid robot wearing clothing and tension data of various body areas; The fusion module is used to extract visual wrinkle features of each body region from the multi-view display image, calculate the tension distribution features of each body region based on the tension data, and align and fuse the visual wrinkle features with the tension distribution features to obtain the wrinkle state vector of each body region. The classification module is used to classify each body region into normal overhanging folds, slight stacking folds, obvious stretching folds, and severe skew folds based on the fold state vector and to determine the areas to be adjusted. The adjustment module is used to calculate the joint posture fine-tuning amount of the area to be adjusted and perform the adjustment. After the clothing fabric is relaxed and stabilized, the operation steps from the acquisition module to the classification module are repeated until all body areas are normal hanging wrinkles or slight pile wrinkles, and then the shooting ready signal is output.

[0015] In summary, this invention solves the technical problem of single-modal perception schemes being unable to simultaneously acquire fold morphology information and clothing stress information by simultaneously acquiring multi-view images and aligning them with timestamps. It uses visual fold features as query vectors and tension distribution features as key vectors for multi-head cross-modal attention fusion to obtain fold state vectors for each body region. Based on these fold state vectors, a weighted fusion of rule-based classification probabilities and model-based classification probabilities is used to finely classify folds in each body region into four categories: normal draped folds, slight piled folds, obvious stretching folds, and severely skewed folds. Posture adjustment is triggered only for obvious stretching folds and severely skewed folds, avoiding excessive intervention in normal draped folds and slight piled folds in existing technologies, thus achieving differentiated and refined processing. For different types of areas to be adjusted, single-joint angle increment calculation based on tension deviation and fold direction angle deviation and multi-joint collaborative fine-tuning based on Jacobian matrix inverse kinematics solution are respectively adopted. Combined with low-speed angular velocity execution and real-time monitoring of safety tension threshold, convergence verification is performed by comparing the fold comprehensive score after the fabric is relaxed and stabilized. If invalid, the joint angle is automatically retracted and re-executed from the alternative fine-tuning direction, eliminating the reliance on repeated manual intervention in the garment display preparation process. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the steps of a method for detecting clothing wrinkles and fine-tuning posture of a humanoid robot in one embodiment of the present invention; Figure 2 This is a schematic diagram of image and tension acquisition in an embodiment of the present invention; Figure 3 This is a schematic diagram of feature alignment and fusion in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the classification of body regions in an embodiment of the present invention; Figure 5 This is a schematic diagram of joint posture fine-tuning in an embodiment of the present invention; Figure 6 This is a block diagram of the humanoid robot clothing wrinkle detection and posture fine-tuning system in an embodiment of the present invention.

[0017] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0019] Reference Figure 1 This embodiment provides a method for detecting clothing wrinkles and fine-tuning the posture of a humanoid robot, including the following steps: S1: Simultaneously collect multi-view images of the humanoid robot wearing clothing and tension data of various body areas; S2: Extract visual wrinkle features of each body region from the multi-view image, calculate the tension distribution features of each body region based on the tension data, and align and fuse the visual wrinkle features with the tension distribution features to obtain the wrinkle state vector of each body region. S3: Based on the fold state vector, classify each body region into normal overhang folds, slight stacking folds, obvious stretching folds, and severe skew folds, and determine the areas to be adjusted; S4: Calculate the joint posture fine-tuning amount of the area to be adjusted and send the joint posture fine-tuning amount to the humanoid robot through the robot control interface to perform the adjustment. After the clothing fabric is relaxed and stable, repeat steps S1-S3 until all body areas are normal hanging wrinkles or slight pile wrinkles, and output the shooting ready signal.

[0020] In one example, such as Figure 2 Step S1 includes: S11: Multiple cameras positioned on the front, sides, and back of the humanoid robot are simultaneously triggered by hardware trigger signals to capture multi-view images of the robot wearing clothing. S12: The contact tension values ​​of each sensor node are collected by a sensor array arranged in each body area of ​​the humanoid robot, and the contact tension values ​​of each sensor node are time-stamped and aligned according to the acquisition timestamp of the multi-view image to obtain the tension data of each body area.

[0021] In this embodiment, a humanoid robot is fixed on a rotating display platform or gait demonstration station, and multiple industrial cameras are arranged on the front, left and right sides, and back. The optical axis of each camera is oriented towards the humanoid robot's clothing display area. The front camera covers the overall outline of the clothing and the front folds, the side cameras cover the shoulder, waist, sleeves, and trouser legs in a side-fitting state, and the back camera covers the back, buttocks, and the draping state of the hem area. A unified hardware trigger pulse is output by a synchronization control unit. After isolation and driving, the hardware trigger pulse is connected to the external trigger input terminal of each camera, so that each camera enters the exposure state on the same trigger edge. The synchronization control unit can be an FPGA, a synchronization trigger control board, or an embedded controller with deterministic input / output timing. For example, the hardware trigger pulse width can be set to about 10μs, and the trigger period can be set to about 33.3ms according to the display frame rate, corresponding to a multi-view image acquisition rhythm of about 30 frames per second. To minimize inconsistencies in viewing angles caused by exposure differences, each camera employs an external trigger mode, uniform exposure time, uniform gain, and uniform white balance parameters. The exposure time can be set within the range of 2ms to 5ms based on the supplementary lighting intensity, minimizing noticeable motion blur during slight swaying or turning of the humanoid robot while preserving fabric texture, edge contours, and wrinkle shadow information. The synchronization control unit synchronously latches the trigger timestamp when issuing a hardware trigger pulse and writes the trigger sequence number, camera number, exposure parameters, and acquisition completion flag into the image acquisition buffer, ensuring that multi-view images from the front, side, and back are captured in the same batch. After each camera completes exposure, the image frames are transmitted to the image buffer via Gigabit Ethernet, USB 3.0, or MIPI interface. The image receiving module performs integrity verification, distortion correction, and viewing angle identification binding on the image frames to prevent misalignment of images from different viewing angles in the buffer queue.

[0022] Simultaneously, flexible tension or contact pressure sensor arrays are deployed in body areas such as the shoulders, chest, waist, hips, elbows, knees, and ankles of the humanoid robot. Sensor nodes can be embedded between the body-hugging base layer, the inner lining of the clothing, or the flexible covering layer of the robot's shell to reflect the local fit, tension constraints, and compression states between the clothing and the humanoid robot's surface. Each body area corresponds to an independent set of acquisition channels. The analog voltage or resistance changes output by the sensor nodes are first processed through constant current excitation, bridge conditioning, low-noise amplification, and analog-to-digital conversion to obtain the contact tension value of each sensor node. If the sensor array uses a digital bus output, the acquisition controller polls and reads the data according to the node address, and adds a local sampling timestamp to each node. For example, the sensor sampling frequency can be set to around 200Hz, which can cover the low-frequency deformation changes of the clothing during the humanoid robot's demonstration movements, while being higher than the multi-view image acquisition frequency, facilitating the selection of adjacent tension sampling values ​​near the image acquisition time point for time alignment. To reduce the impact of transient noise and local contact jitter on tension data, the acquisition controller performs zero-point subtraction, sensitivity coefficient conversion, and sliding filtering on the raw contact tension values ​​of each sensor node. The processed contact tension values ​​are then written to the tension data cache according to body region, node number, and sampling timestamp. After acquiring multi-view display images, the time alignment module reads the acquisition timestamp corresponding to each set of multi-view display images and retrieves two sets of sensor sampling records within the same body region before and after that timestamp from the tension data cache. Linear interpolation of the contact tension values ​​is performed based on the time distance, ensuring that each sensor node obtains a tension estimate corresponding to the instant the image was acquired. For sensor nodes experiencing packet loss or abnormal jumps near the acquisition timestamp, a combination of adjacent node spatial compensation and the previous valid sample-and-hold method is used for correction, preventing distortion of the entire body region's tension distribution due to a single node anomaly. After timestamp interpolation alignment, the multi-view images of the front, side and back under the same trigger number are bound to the tension data of body areas such as shoulders, chest, waist, hips, elbows, knees and ankles as a set of synchronous samples, so that the appearance of clothing, local wrinkle status, edge droop status and body contact tension distribution are under the same time reference.

[0023] In one example, such as Figure 3 Step S2 includes: S21: Perform human body region segmentation on multi-view display images, and extract the fold direction angle, fold density and fold depth of each body region as visual fold features; S22: Perform regional aggregation on the tension data, and extract the regional average tension, tension standard deviation and tension gradient of adjacent regions of each body region as tension distribution features; S23: Using the visual wrinkle features of each body region as the query vector and the tension distribution features of each body region as the key vector, a multi-head cross-modal attention mechanism is used to fuse the wrinkle state vectors of each body region.

[0024] In this embodiment, multi-view images from the front, side, and back are input into the image preprocessing module. Distortion correction, brightness equalization, background suppression, and viewpoint label binding are performed on the images from each viewpoint to ensure that the clothing outline, limb boundaries, and fabric texture captured by different cameras are under relatively consistent imaging conditions. A human body region segmentation model is used to perform pixel-level segmentation on the image of the humanoid robot wearing clothing. The segmentation results can be used to create region masks according to body regions such as head and neck, shoulders and chest, waist and abdomen, back, buttocks, upper arms, forearms, thighs, and calves. The segmentation boundaries are constrained by the humanoid robot's skeletal joints or preset structural dimensions to prevent region mismatches caused by clothing occlusion, loose hems, and changes in limb posture. Within each body region mask, grayscale texture, edge response, and local shadow variations are jointly analyzed. The principal direction of fabric folds is extracted using a directional filter bank or structural tensor to obtain the fold direction angle. Then, fold density is calculated based on the correspondence between the number of fold lines and the effective area of ​​the region; fold density represents the number of fold lines per unit effective area. Simultaneously, after calibrating the intrinsic and extrinsic parameters of each camera, the fold depth characterization value is estimated by combining fold shadow intensity, local brightness gradient, and parallax changes or edge undulation consistency of the same body region under multiple viewpoints. This forms visual fold features that reflect the stretching, stacking, draping, and local compression states of the clothing surface. For the same body region visible from multiple viewpoints, the visual fold features are fused according to camera viewpoint weights and the visible area of ​​the region. For areas with significant occlusion from the side or back, viewpoint features with larger visible areas and higher texture clarity are used, and the visibility confidence is recorded along with the visual fold features to reduce misjudgments caused by occlusion.

[0025] Simultaneously, the tension data processing module aggregates the contact tension values ​​after timestamp interpolation and alignment according to the body region to which the sensor nodes belong, and removes node readings that significantly exceed the sensor's range or exhibit abnormal short-term abrupt changes. Then, it performs regional aggregation on the contact tension values ​​of valid nodes within the same body region, extracting the regional average tension to characterize the overall fit or pressure level of the clothing on the body region; it also extracts the tension standard deviation to characterize the uniformity of tension distribution within the same body region; and it calculates the tension gradient between adjacent regions based on their spatial adjacency to characterize whether there are abrupt changes in tension transmission from the shoulder to the chest, from the waist to the hips, or from the thigh to the calf. For regions with uneven sensor node distribution, aggregation weights are set according to node coverage area or node reliability to prevent areas with a large number of nodes from being unreasonably amplified in the overall evaluation, while ensuring that key stress locations are adequately represented in the tension distribution characteristics. After extracting visual wrinkle features and tension distribution features, the wrinkle orientation angle, wrinkle density, and wrinkle depth of each body region are combined to form a visual feature sequence, and the region's average tension, tension standard deviation, and tension gradient of adjacent regions are combined to form a tension feature sequence. Scale normalization and dimensionality mapping are then performed on both types of features to ensure that image texture information and contact force information are integrated into a unified feature space. In the multi-head cross-modal attention mechanism, the visual wrinkle features of each body region serve as query vectors to express the appearance state requirements indicated by the wrinkles on the clothing surface; the tension distribution features of each body region serve as key vectors to provide mechanical response information regarding the contact between the clothing and the humanoid robot surface. The multi-head attention branches focus on different types of relationships, such as the relationship between wrinkle density and local tension concentration, the relationship between wrinkle orientation angle and tension gradient of adjacent regions, and the relationship between wrinkle depth and region's average tension. The fusion results of multiple attention branches are then processed through residual connections and nonlinear mapping to obtain the wrinkle state vectors of each body region. The generated wrinkle state vector contains both the wrinkle shape at the visual level and the contact tension distribution at the sensor level. It can be used to determine whether the clothing is unbalanced at the shoulders, whether the waist and abdomen are excessively tight, whether there are stacked wrinkles on the back, whether there is local stretching in the elbow and knee bending areas, and the fit of the entire clothing in the humanoid robot's display posture.

[0026] In one example, step S23 includes: Multi-head cross-modal attention is calculated and concatenated using the visual wrinkle features of each body region as query vectors and the tension distribution features of each body region as key vectors to obtain cross-modal attention features; The visual wrinkle features and tension distribution features are concatenated to calculate the gating weights. Based on the gating weights, the cross-modal attention features are weighted and fused to obtain the wrinkle state vectors of each body region.

[0027] In this embodiment, each body region is organized into a region visual feature sequence according to the topological relationship of the humanoid robot skeleton, and the tension distribution features corresponding to each body region are organized into a region tension feature sequence. For cases where multiple local image patches or multiple sensor nodes exist within the same body region, the local image patch features and sensor node features are further organized into a local feature sequence, enabling the multi-head cross-modal attention mechanism to establish correlations between regions, nodes, and between the visual and tension modalities. Since the dimensions and numerical ranges of visual wrinkle features and tension distribution features are different, they are normalized before entering the fusion module, and both types of features are mapped to the same latent space dimension through a linear mapping layer, allowing the wrinkle morphology information on the visual side and the contact force information on the tension side to participate in attention calculation at the same feature scale. The visual wrinkle features of each body region are used as query vectors, expressing the wrinkle morphology requirement of the current body region in the image; the tension distribution features of each body region are used as key vectors and value vectors, with the key vectors used to calculate the matching degree between the visual wrinkle state and the tension distribution state, and the value vectors used to provide tension response content that can be selected and absorbed by visual features. In multi-head cross-modal attention computation, the fusion module configures independent linear projection parameters for each attention head. Different attention heads learn the correspondences between wrinkle orientation angle and tension gradient, wrinkle density and regional tension dispersion, and wrinkle depth and regional average tension, respectively, and generate attention weights based on the similarity between the query vector and the key vector. The higher the attention weight, the stronger the ability of the tension distribution features of the corresponding body region to interpret the current visual wrinkle state. The fusion module then assigns a higher response weight to the value vector, thereby obtaining the cross-modal response results for each attention head. The outputs of each attention head are concatenated along the feature dimension and compressed to a unified dimension through an output mapping layer to form cross-modal attention features. These cross-modal attention features simultaneously contain both local and inter-regional correlations between visual wrinkle representation and tension response.

[0028] The original visual wrinkle features and tension distribution features are concatenated along the feature dimension and input into a gating computation layer to obtain gating weights. These gating weights can be understood as adaptive coefficients determining whether the current body region should rely more on visual texture judgment or tension sensing judgment. When the clothing surface wrinkles are clear, the lighting is stable, and the visible area of ​​the body region is large, the gating weights can increase the visual contribution. When there are occlusions, reflections, or indistinct wrinkle textures in the image, but the tension sensor readings are stable, the gating weights can increase the tension contribution. Based on the gating weights, cross-modal attention features are weighted and fused, and residual connections are used to retain basic appearance information from the visual wrinkle features. Then, a normalization layer and a nonlinear mapping layer are used to output the wrinkle state vectors for each body region. The wrinkle state vectors can characterize the wrinkle direction, wrinkle concentration, wrinkle depth variation, local tension concentration, and abrupt changes in stress in adjacent areas for each body region.

[0029] In one example, such as Figure 4 Step S3 includes: S31: Calculate the rule classification probability of each body region based on the region average tension, fold direction angle and fold density in the fold state vector of each body region; S32: Input the wrinkle state vectors of each body region into the fully connected classification network for forward inference to obtain the model classification probability of each body region; S33: The rule classification probability and the model classification probability are weighted and fused based on the weighted coefficient to obtain the fused classification probability. The classification result of each body region is determined according to the fused classification probability. The classification result is normal overhang folds, slight stacking folds, obvious stretching folds or severe skew folds. S34: Body areas with obvious stretching wrinkles or severely skewed wrinkles are identified as areas to be adjusted.

[0030] In this embodiment, the wrinkle state vector of each body region is used as a unified input for classification and determination. The wrinkle state vector retains the average tension of the region, the wrinkle direction angle, the wrinkle density, and the implicit state components formed after cross-modal fusion, so that the rule determination branch and the model inference branch can establish a consistent data foundation around the same body region. For calculating the probability of rule classification, a rule scoring logic is established based on indicators with clear physical meaning in the clothing display state. Among them, the regional average tension is used to reflect the degree of overall adhesion or stretching of the clothing to the surface of the humanoid robot, the fold direction angle is used to reflect the deflection state of the fold texture relative to the main axis direction of the body area, and the fold density is used to reflect the degree of aggregation of fold lines in a unit area. When the regional average tension is within the preset normal tension range and maintains small fluctuations, the fold direction angle is close to the natural drooping direction of the body area, and the fold density is kept within a relatively uniform range, the rule branch increases the probability of normal hanging folds. When the fold density increases but the regional average tension does not increase significantly, the rule branch tends to increase the probability of slightly piled folds. When the regional average tension increases and the fold direction angle shows a concentrated deflection along the tension transmission direction, the rule branch increases the probability of obviously stretched folds. When the fold direction angle deviates significantly from the natural drooping direction of the body area, accompanied by abnormally concentrated fold density or uneven distribution of regional average tension, the rule branch increases the probability of severely skewed folds.

[0031] Meanwhile, the fully connected classification network receives the wrinkle state vectors of each body region and completes forward inference through the input mapping layer, the hidden feature extraction layer, and the output probability layer. The input mapping layer performs dimensional transformation on the vector after fusing visual wrinkle features and tension distribution features. The hidden feature extraction layer performs nonlinear expression on the local wrinkle combination relationship, the force transmission relationship between adjacent regions, and the multi-view texture difference, which are difficult to describe directly by rules. The output probability layer generates four types of model classification probabilities: normal overhanging wrinkles, slight stacking wrinkles, obvious stretching wrinkles, and severely skewed wrinkles.

[0032] The system weights and fuses the rule-based classification probabilities and the model-based classification probabilities using preset weighting coefficients. The rule branch provides interpretable engineering constraints, while the model branch provides adaptability to complex wrinkle shapes. When there are few training samples or significant fluctuations in lighting, the weight of the rule branch can be appropriately increased; when there is sufficient sample accumulation and stable multi-view image quality, the weight of the model branch can be increased. After generating the fused classification probabilities, the classification decision module selects the category with the highest probability as the classification result for each body region and writes the classification result of each body region, along with the body region number, probability confidence level, and corresponding trigger timestamp, into the display status record. For body regions with obvious stretching wrinkles or severely skewed wrinkles, the corresponding body region is marked as an area to be adjusted and passed to the garment finishing, posture correction, or pattern adaptation suggestion module. For normal draping wrinkles and slight pile wrinkles, the system only retains the display status record to avoid over-adjusting natural draping shapes within acceptable limits.

[0033] In one example, step S31 includes: When the average tension of each body region is within the preset normal tension range, the fold direction angle is less than the first direction angle threshold, and the fold density is less than the first density threshold, the rule judgment result of the corresponding body region is determined to be a normal overhang fold. When the average tension of the region is lower than a preset multiple of the lower limit of the normal tension range, the fold direction angle is less than the first direction angle threshold, and the fold density is between the first density threshold and the second density threshold, the rule judgment result of the corresponding body region is determined to be slight accumulation folds. When the average tension of each body region is higher than a preset multiple of the upper limit of the normal tension range, the fold direction angle is greater than the second direction angle threshold, and the fold density is greater than the first density threshold, the rule judgment result of the corresponding body region is determined to be obvious stretching folds. When the tension gradient of adjacent regions of each body region exceeds the tension gradient threshold, the fold direction angle dispersion exceeds the direction disorder threshold, and the fold density is greater than the second density threshold, the rule judgment result of the corresponding body region is determined to be a severely skewed fold. The probability of the category corresponding to the rule judgment result is recorded as 1, and the probability of the other categories is recorded as 0, so as to obtain the rule classification probability of each body region.

[0034] In this embodiment, a unified rule judgment table is established for each body region in the rule classification module, and the region's average tension, wrinkle direction angle, wrinkle density, tension gradient between adjacent regions, and wrinkle direction angle dispersion are used as input fields. The region's average tension can be measured in Newtons. For example, the preset normal tension range can be set to 0.8N to 2.5N, which is suitable for describing the contact state where the clothing and the humanoid robot's surface are in close contact but without significant pressure. The preset multiple below the lower limit of the normal tension range can be set to 0.75 times, that is, when the region's average tension is below 0.6N, it is considered that the clothing has a tendency to loosen and accumulate in the corresponding body region. The preset multiple above the upper limit of the normal tension range can be set to 1.20 times, that is, when the region's average tension is above 3.0N, it is considered that the clothing has a tendency to stretch and constrain in the corresponding body region. The multiple setting retains a certain tolerance, which can reduce misjudgments caused by instantaneous fluctuations in the sensor. The fold direction angle is used to characterize the degree of deviation of the main fold direction from the natural draping direction of the body area. For example, the first direction angle threshold can be set to 15°, which corresponds to the slight deflection allowed by the natural draping fold; the second direction angle threshold can be set to 35°, which corresponds to the state where the fold deflects significantly along the direction of force after being stretched. Fold density can be calculated based on the number of fold skeleton lines or the length of the fold skeleton within a unit effective area. For example, the first density threshold can be set to 0.20 lines / cm. 2 The second density threshold can be set to 0.55 lines / cm. 2 A lower threshold is used to distinguish between smooth drapes and slight stacking, while a higher threshold is used to identify areas with obvious fold aggregation or disordered orientation. The tension gradient threshold between adjacent regions can be set to 1.0 N / cm, and the orientation disorder threshold can be set to 25° to identify severe skewness states where tension changes abruptly between adjacent body regions are accompanied by fold orientation dispersion.

[0035] When determining the rules, the feature records of the corresponding body region at the same timestamp are read, and the severe skewed fold condition is judged first, because severe skewed folds are simultaneously manifested as tension abrupt change, direction disorder and fold concentration. If the severe skewed fold condition is not met, the obvious stretching fold condition is judged. When the average tension of the region exceeds the upper limit multiple, the fold direction angle exceeds the second direction angle threshold and the fold density is higher than the first density threshold, the rule judgment result is set to obvious stretching fold. Then the slight accumulation fold condition is judged. When the average tension of the region is lower than the lower limit multiple, the fold direction angle is still less than the first direction angle threshold and the fold density is between the first density threshold and the second density threshold, the rule judgment result is set to slight accumulation fold. When the average tension of the region is within the normal tension range, the fold direction angle is less than the first direction angle threshold and the fold density is less than the first density threshold, the rule judgment result is set to normal overhanging fold. After the rule determination result is determined, the rule classification probability is output in the form of one-hot encoding. That is, in the four categories of normal drooping wrinkles, slight stacking wrinkles, obvious stretching wrinkles and severe skewed wrinkles, the probability of the rule determination result corresponding to the category is written as 1, and the probability of the other categories is written as 0. The body area number, trigger timestamp, trigger rule name and each input feature value are simultaneously written into the rule determination record.

[0036] In one example, step S4 includes: S41: Determine the set of joints corresponding to the region to be adjusted, and calculate the joint posture fine-tuning amount of the set of joints corresponding to the region to be adjusted; S42: Send the joint posture fine-tuning amount to the joint set of the humanoid robot through the robot control interface and execute it. During the execution, if the contact tension value of any body area exceeds the safe tension threshold, the execution will stop immediately. After the clothing fabric is relaxed and stabilized, steps S1-S3 will be executed again to calculate the first wrinkle comprehensive score before adjustment and the second wrinkle comprehensive score after adjustment. S43: If the overall score of the second fold is higher than that of the first fold and there are no new severely skewed fold areas, then the adjustment is confirmed to be effective. When the classification results of all body areas are normal hanging folds or slight piled folds, output a shooting ready signal.

[0037] In this embodiment, after obtaining the area to be adjusted, the corresponding joint set is determined based on the mapping relationship between the body area and the humanoid robot's skeletal structure. For example, the area to be adjusted in the shoulder and chest is associated with the shoulder joint, upper arm rotation joint, and trunk pitch joint; the area to be adjusted in the waist and abdomen is associated with the waist rotation joint, hip joint, and trunk tilt joint; and the area to be adjusted near the elbow or knee is associated with the adjacent flexion-extension joint and the corresponding limb rotation joint. When determining the joint set, the control module combines the classification results of the area to be adjusted and the wrinkle state vector to determine the adjustment direction. Obvious pulled wrinkles indicate that the clothing fabric is subjected to excessive traction, and the joint posture fine-tuning amount should be generated in the direction of reducing local stretching and releasing contact tension; severely skewed wrinkles indicate that there is unbalanced traction between adjacent areas or clothing boundary offset, and the joint posture fine-tuning amount should be generated in the direction of restoring the relatively symmetrical posture of the body area and reducing the tension gradient of adjacent areas. The amount of joint posture fine-tuning should not be too large at once. For example, the single fine-tuning angle can be set in the range of 0.5° to 2°, and the angular velocity can be limited to within 5° / s. The angle range is used to avoid sudden pulling of the clothing, and the angular velocity limit is used to allow the fabric folds to gradually release with the change of posture.

[0038] After writing the joint posture fine-tuning values ​​into the joint control queue, the control module sends them to the servo driver sequentially or synchronously according to the priority of each joint in the joint set. During execution, it continuously reads the contact tension values ​​of sensor nodes in each body area. For example, the safe tension threshold can be set to 4.5N. 4.5N is higher than the boundary for obvious stretching and wrinkle recognition, which can reserve necessary adjustment space while avoiding excessive local load on the clothing fabric or the robot's flexible covering layer. When the contact tension value of any effective sensor node in any body area exceeds the safe tension threshold, the control module immediately freezes the joint control queue, sends a stop or torque hold command to the servo driver, and records the body area that triggered the stop, the sensor node number, the joint posture, and the timestamp. After stopping execution, the robot maintains its current posture or returns to the previous safe posture and waits for the clothing fabric to relax and stabilize naturally. For example, the relaxation and stabilization waiting time can be set to about 1.5s, taking into account the rebound and drape recovery process of the fabric after slight stretching. After stabilization, the multi-view image acquisition, tension data acquisition, wrinkle state vector generation, and classification process are repeated. Based on the two synchronized samples before and after adjustment, a first wrinkle comprehensive score and a second wrinkle comprehensive score are calculated respectively. The first wrinkle comprehensive score characterizes the overall display quality of each body region before adjustment, while the second wrinkle comprehensive score characterizes the overall display quality after joint posture fine-tuning. The scoring comparison considers not only the proportion of normal overhanging wrinkles and slightly accumulated wrinkles, but also the penalty weights for obvious stretching wrinkles and severely skewed wrinkles, as well as whether the contact tension is close to the safe boundary.

[0039] If the overall score of the second fold is higher than that of the first fold, and no new severely skewed fold areas are found in the reassessment results, the system confirms the effectiveness of this adjustment and writes the joint posture fine-tuning amount, the scores before and after the adjustment, the changed area, and the tension change trend into the adjustment record. If the overall score of the second fold does not improve, or if a new severely skewed fold area appears, the system determines the adjustment is invalid, cancels the joint posture fine-tuning amount, or reduces the fine-tuning amplitude before regenerating the adjustment strategy. After one or more effective adjustments, the system checks the classification results of all body areas again. When all body areas are determined to be normal drooping folds or slightly piled folds, it indicates that the clothing does not have obvious stretching or severe skewing in the current posture of the humanoid robot. The control module outputs a shooting ready signal to the shooting equipment, lighting control module, and display process management module, enabling the multi-view shooting or display image archiving process to enter the formal execution state.

[0040] In one example, S41 includes: S411: Map the region to be adjusted to the corresponding set of joints; S412: For the adjustment area with obvious stretching and wrinkling, calculate the angle increment of the corresponding joint based on the deviation between the average tension of the area to be adjusted and the target tension, as well as the deviation of the wrinkle direction angle, to obtain the joint posture fine adjustment amount. S413: For the adjustment area with severe skewed folds, construct the desired garment displacement correction vector based on the fold direction angle deviation of the adjustment area and the tension gradient of the adjacent area. Solve the inverse kinematics of the desired garment displacement correction vector based on the Jacobian matrix of the joint set to obtain the joint posture fine adjustment amount.

[0041] In this embodiment, a mapping table between body regions and robot joint sets is pre-established in the control module. The mapping table is configured according to the force transmission path of the clothing and the kinematic chain of the robot skeleton, so that the adjustment areas such as the shoulders, chest and back, waist and abdomen, hips, elbows, and knees can be mapped to the joint sets with posture adjustment functions. After the adjustment area is input, the control module reads the adjustment area number, wrinkle classification result, average tension of the area, wrinkle direction angle, tension gradient of adjacent areas, and the natural drooping direction of the corresponding body region, and retrieves the proximal joints and auxiliary joints connected to the adjustment area from the mapping table. The proximal joints undertake the main adjustment function, and the auxiliary joints are used to reduce the cascading effect of local movements on adjacent areas. For obvious stretching wrinkles, the target tension is set to the median of the normal tension range. For example, when the normal tension range is 0.8N to 2.5N, the target tension can be 1.65N. Taking the median value can avoid the adjustment being close to the boundary of being too loose or too tight. At the same time, the target wrinkle direction angle is set to a small deflection angle near the natural drooping direction of the corresponding body region, for example, within a range of no more than 15°. Then, based on the deviation between the regional average tension and the target tension, and the deviation between the fold direction angle and the target fold direction angle, the joint angle increment is calculated: in, Indicates the first The angle increment of each joint, in degrees, is used to form the joint posture fine-tuning amount; This indicates the upper limit of a single angle increment, in degrees, and can be set to 2°. The numerical setting is used to limit the amplitude of a single movement and reduce the risk of sudden stretching of clothing fabric. Indicates the first The angle adjustment coefficient for tension deviation of each joint, in degrees per Newton; Indicates the first The average tension of the region to be adjusted is expressed in Newtons. The target tension is expressed in Newtons. Indicates the first The angle adjustment coefficient for each joint relative to the direction angle deviation is expressed in degrees / degrees and is used as a proportional coefficient in the calculation to the increase in joint angle from the direction angle deviation. Indicates the first The fold direction angle of the area to be adjusted is in degrees; Indicates the direction angle of the target fold, in degrees; This represents the saturation limiting function, used to restrict the calculated angle increment within the allowable fine-tuning range. The direction sign in the formula is determined by the joint-body region coupling relationship. When the average tension of the region is higher than the target tension and the fold direction angle deviates from the target direction, the control module combines the joint rotation axis direction, the local coordinate system of the body region, and the direction of force release of the clothing to determine the corresponding positive or negative angle increment of the joint, so that the joint adjustment direction is generated in the direction of releasing tension and reducing directional deviation. For severely skewed folds, a desired clothing displacement correction vector is constructed based on the fold direction angle deviation and the tension gradient of adjacent regions, so that the clothing surface of the area to be adjusted produces a small displacement in the direction of decreasing tension gradient and restoring the natural drape of the fold direction. in, Indicates the first The expected garment displacement correction vector for the area to be adjusted, in millimeters, describes the direction and magnitude of the garment surface that is to be released or returned to center. The conversion factor from angular deviation to displacement is expressed in millimeters per degree. Indicates the first A unit direction vector used to correct the fold direction deviation in the area to be adjusted; This represents the conversion factor from tension gradient to displacement, and the unit can be set in millimeters squared per Newton. Indicates the first The tension gradient between adjacent regions of the region to be adjusted is expressed in Newtons per centimeter. Indicates the first The unit direction vector used to reduce the tension gradient in the region to be adjusted. After obtaining the desired garment displacement correction vector, the control module calls the Jacobian matrix corresponding to the joint set to convert the small displacement of the garment surface into a joint angle fine-tuning amount: in, Indicates the first The joint pose fine-tuning vector of the joint set corresponding to each region to be adjusted, in degrees; Indicates the first The Jacobian matrix of the joint set corresponding to each region to be adjusted is used to describe the effect of small changes in joint angle on the surface displacement of the garment in the region to be adjusted. This represents the damping coefficient, which can be set to 0.01 to reduce the solution instability when the Jacobian matrix approaches a singular configuration; The unit matrix is ​​represented. The joint posture fine-tuning amount obtained from the solution needs to be checked by joint limit, velocity limit and collision constraint. When the single fine-tuning amount exceeds the allowable range, it is scaled to the safe range according to the ratio. The near joint adjustment amount and the auxiliary joint adjustment amount are combined into the final control command, so that the body area corresponding to obvious stretching wrinkles will release local tension first, and the body area corresponding to severe skew wrinkles will reduce orientation confusion and tension change in adjacent areas simultaneously.

[0042] In one example, step S413 includes: The direction of garment displacement correction is determined based on the fold direction angle deviation of the area to be adjusted, and the amount of garment displacement correction is determined based on the excessive tension gradient of the adjacent areas of the area to be adjusted. The desired garment displacement correction vector is then constructed. Based on the kinematic model of each joint in the joint set, the Jacobian matrix corresponding to the joint set is calculated analytically. The expected clothing displacement correction vector is input into a sequential quadratic programming optimizer constrained by the Jacobian matrix to solve the inverse kinematics and obtain the joint posture fine adjustment amount.

[0043] In this embodiment, the fold direction angle of the area to be adjusted is compared with the natural drooping reference direction of the corresponding body area to obtain the fold direction angle deviation. The direction of garment displacement correction is determined based on the sign and magnitude of the deviation. For different body areas such as the shoulder, waist, abdomen, hip, elbow, or knee, the system pre-saves a local coordinate system for each body area. This local coordinate system includes a first direction along the longitudinal direction of the body and a second direction along the transverse direction of the body. When the fold direction angle deviation is clockwise, the garment displacement correction direction is generated along the opposite local tangential direction; when the fold direction angle deviation is counterclockwise, the garment displacement correction direction is generated along the corresponding reverse tangential direction, allowing the micro-displacement of the garment surface to weaken the directional deflection of the skewed folds. The tension gradient between the area to be adjusted and adjacent body areas is read and compared with a tension gradient threshold to obtain the tension gradient exceedance. The tension gradient threshold is 1.0 N / cm, which is used to distinguish between normal stress transition and local abrupt stress. The larger the tension gradient exceedance, the greater the expected garment displacement correction. The expected garment displacement correction vector is: in, Indicates the first The expected garment displacement correction vector for each region to be adjusted, in millimeters, is used to describe the small release displacement that the garment surface needs to produce. Indicates the sequence number of the area to be adjusted; This represents the conversion factor from the excess tension gradient to the displacement correction, expressed in square centimeters per Newton, and can be set to 0.8 cm. 2 / N, so that the excess amplitude of 1N / cm corresponds to a release scale of about 0.8cm, which is about 8mm after conversion. This can avoid new wrinkle shift caused by excessive single displacement. Indicates the first The tension gradient between adjacent regions of the region to be adjusted is expressed in Newtons per centimeter. This represents the tension gradient threshold, expressed in Newtons per centimeter. This represents the dimensionless unit correction direction vector determined by the fold direction angle deviation. After obtaining the desired garment displacement correction vector, the control module establishes a kinematic model based on the link length, rotation axis direction, current joint angle, and joint coordinate transformation relationship of each joint in the joint set. It then calculates the partial derivative of the garment surface reference point in the adjustment area with respect to each joint angle for each joint, forming the Jacobian matrix corresponding to the joint set. Each column of the Jacobian matrix corresponds to the influence of a small angle change of a joint on the displacement of the garment surface reference point, and each row corresponds to a displacement component in the local coordinate system. The desired garment displacement correction vector is input into a sequential quadratic programming optimizer. The optimizer uses the Jacobian matrix as a local linear constraint and simultaneously incorporates joint angle limits, single-step fine-tuning limits, joint velocity limits, and safety tension constraints to construct a quadratic approximate subproblem near the current posture: in, Indicates the first The joint pose fine-tuning vector of the joint set corresponding to each region to be adjusted, in degrees; Indicates the first The Jacobian matrix of the joint set corresponding to each region to be adjusted is used to describe the linear approximate relationship between small changes in joint posture and displacement of the clothing surface. This represents the attitude fine-tuning regularization coefficient, which, for example, can be set to 0.05, to suppress unnecessary large joint changes. The sequential quadratic programming optimizer solves for candidate values ​​for joint attitude fine-tuning in each iteration and recalculates the displacement of the reference point on the garment surface and the Jacobian matrix using the kinematic model until the displacement error is lower than the preset convergence error or the maximum number of iterations is reached. For example, the convergence error can be set to 1mm, and the maximum number of iterations can be set to 10. The final joint attitude fine-tuning is then scaled and safety-checked to ensure that the single adjustment angle of a single joint remains within the range of 0.5° to 2°, and before being sent to the servo driver, it is confirmed that the adjustment direction will not increase the abrupt change in contact tension between the area to be adjusted and adjacent areas.

[0044] Reference Figure 6 This embodiment provides a humanoid robot clothing wrinkle detection and posture fine-tuning system 600, including: The acquisition module 601 is used to simultaneously acquire multi-view images of the humanoid robot wearing clothing and tension data of various body areas; The fusion module 602 is used to extract visual wrinkle features of each body region from the multi-view display image, calculate the tension distribution features of each body region based on the tension data, and align and fuse the visual wrinkle features with the tension distribution features to obtain the wrinkle state vector of each body region. Classification module 603 is used to classify each body region into normal overhanging folds, slight stacking folds, obvious stretching folds and severe skew folds based on the fold state vector and determine the region to be adjusted. The adjustment module 604 is used to calculate the joint posture fine-tuning amount of the area to be adjusted and perform the adjustment. After the clothing fabric is relaxed and stabilized, the operation steps from the acquisition module to the classification module are re-executed until all body areas are normal hanging wrinkles or slight pile wrinkles, and then the shooting ready signal is output.

[0045] In this embodiment, the specific implementation of each unit in the above system embodiment is described in the above method embodiment, and will not be repeated here.

[0046] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, system, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, system, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, system, article, or method that includes that element.

[0047] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for detecting wrinkles in clothing and fine-tuning the posture of a humanoid robot, characterized in that, include: S1: Simultaneously collect multi-view images of the humanoid robot wearing clothing and tension data of various body areas; S2: Extract visual wrinkle features of each body region from the multi-view display image, calculate the tension distribution features of each body region based on the tension data, and align and fuse the visual wrinkle features with the tension distribution features to obtain the wrinkle state vector of each body region. S3: Based on the wrinkle state vector, classify each body region into normal overhang wrinkles, slight pile wrinkles, obvious stretching wrinkles, and severe skewed wrinkles, and determine the areas to be adjusted; S4: Calculate the joint posture fine-tuning amount of the area to be adjusted and send the joint posture fine-tuning amount to the humanoid robot through the robot control interface to perform the adjustment. After the clothing fabric is relaxed and stable, repeat steps S1-S3 until all body areas are normal hanging wrinkles or slight pile wrinkles, and output the shooting ready signal.

2. The method for detecting wrinkles and fine-tuning the posture of humanoid robot clothing according to claim 1, characterized in that, Step S1 includes: S11: Multiple cameras positioned on the front, sides, and back of the humanoid robot are simultaneously triggered by hardware trigger signals to capture multi-view images of the robot wearing clothing. S12: The contact tension values ​​of each sensor node are collected by the sensor array arranged in each body area of ​​the humanoid robot, and the contact tension values ​​of each sensor node are time-stamped and aligned according to the acquisition timestamp of the multi-view display image to obtain the tension data of each body area.

3. The method for detecting wrinkles and fine-tuning the posture of humanoid robot clothing according to claim 1, characterized in that, Step S2 includes: S21: Perform human body region segmentation on the multi-view display image, and extract the fold direction angle, fold density and fold depth of each body region as visual fold features. S22: Perform regional aggregation on the tension data, and extract the regional average tension, tension standard deviation and tension gradient of adjacent regions of each body region as tension distribution features; S23: Using the visual wrinkle features of each body region as the query vector and the tension distribution features of each body region as the key vector, a multi-head cross-modal attention mechanism is used to fuse the wrinkle state vectors of each body region.

4. The method for detecting wrinkles and fine-tuning the posture of humanoid robot clothing according to claim 3, characterized in that, Step S23 includes: Multi-head cross-modal attention is calculated and concatenated using the visual wrinkle features of each body region as query vectors and the tension distribution features of each body region as key vectors to obtain cross-modal attention features; The visual wrinkle features and the tension distribution features are concatenated to calculate the gating weights. Based on the gating weights, the cross-modal attention features are weighted and fused to obtain the wrinkle state vectors of each body region.

5. The method for detecting wrinkles and fine-tuning the posture of humanoid robot clothing according to claim 1, characterized in that, Step S3 includes: S31: Calculate the rule classification probability of each body region based on the region average tension, fold direction angle and fold density in the fold state vector of each body region; S32: Input the wrinkle state vectors of each body region into the fully connected classification network for forward inference to obtain the model classification probability of each body region; S33: The classification probability of the rule and the classification probability of the model are weighted and fused based on the weighting coefficient to obtain the fused classification probability. The classification result of each body region is determined according to the fused classification probability. The classification result is normal drooping folds, slight stacking folds, obvious stretching folds or severe skewed folds. S34: The body areas that are classified as having obvious stretching wrinkles or severely skewed wrinkles are identified as areas to be adjusted.

6. The method for detecting wrinkles and fine-tuning the posture of humanoid robot clothing according to claim 5, characterized in that, Step S31 includes: When the average tension of each body region is within the preset normal tension range, the fold direction angle is less than the first direction angle threshold, and the fold density is less than the first density threshold, the rule determination result of the corresponding body region is determined to be a normal overhang fold. When the average tension of the region is lower than a preset multiple of the lower limit of the normal tension range, the fold direction angle is less than the first direction angle threshold, and the fold density is between the first density threshold and the second density threshold, the rule determination result of the corresponding body region is determined to be slight accumulation folds. When the average tension of each body region is higher than a preset multiple of the upper limit of the normal tension range, the fold direction angle is greater than the second direction angle threshold, and the fold density is greater than the first density threshold, the rule judgment result of the corresponding body region is determined to be obvious stretching folds. When the tension gradient of adjacent regions of each body region exceeds the tension gradient threshold, the fold direction angle dispersion exceeds the direction disorder threshold, and the fold density is greater than the second density threshold, the rule judgment result of the corresponding body region is determined to be a severely skewed fold. The probability of the category corresponding to the rule judgment result is recorded as 1, and the probability of the other categories is recorded as 0, so as to obtain the rule classification probability of each body region.

7. The method for detecting wrinkles and fine-tuning the posture of humanoid robot clothing according to claim 6, characterized in that, Step S4 includes: S41: Determine the set of joints corresponding to the region to be adjusted, and calculate the joint posture fine-tuning amount of the set of joints corresponding to the region to be adjusted; S42: The joint posture fine-tuning amount is sent to the joint set of the humanoid robot through the robot control interface and executed. During the execution, if the contact tension value of any body area exceeds the safe tension threshold, the execution is stopped immediately. After the clothing fabric is relaxed and stabilized, steps S1-S3 are executed again to calculate the first wrinkle comprehensive score before adjustment and the second wrinkle comprehensive score after adjustment. S43: If the second fold comprehensive score is higher than the first fold comprehensive score and there are no newly added severely skewed fold areas, then the adjustment is confirmed to be effective. When the classification results of all body areas are normal hanging folds or slight piled folds, output the shooting ready signal.

8. The method for detecting wrinkles and fine-tuning the posture of humanoid robot clothing according to claim 7, characterized in that, S41 includes: S411: Map the region to be adjusted to the corresponding joint set; S412: For the adjustment area with obvious stretching and wrinkling, calculate the angle increment of the corresponding joint based on the deviation between the average tension of the area to be adjusted and the target tension and the deviation of the wrinkle direction angle, and obtain the joint posture fine adjustment amount. S413: For the adjustment area with severe skewed folds, construct the desired garment displacement correction vector based on the fold direction angle deviation of the adjustment area and the tension gradient of the adjacent area. Solve the inverse kinematics of the desired garment displacement correction vector based on the Jacobian matrix of the joint set to obtain the joint posture fine adjustment amount.

9. The method for detecting wrinkles and fine-tuning the posture of humanoid robot clothing according to claim 7, characterized in that, Step S413 includes: The garment displacement correction direction is determined based on the fold direction angle deviation of the area to be adjusted, and the garment displacement correction amount is determined based on the excessive tension gradient of the adjacent areas of the area to be adjusted, and the desired garment displacement correction vector is constructed. Based on the kinematic model of each joint in the joint set, the Jacobian matrix corresponding to the joint set is calculated analytically. The desired clothing displacement correction vector is input into a sequential quadratic programming optimizer constrained by the Jacobian matrix to perform inverse kinematics solution, thereby obtaining the joint posture fine-tuning amount.

10. A system for detecting wrinkles in clothing and fine-tuning the posture of a humanoid robot, characterized in that, The steps for implementing the humanoid robot clothing wrinkle detection and posture fine-tuning method according to any one of claims 1 to 9 include: The acquisition module is used to simultaneously acquire multi-view images of the humanoid robot wearing clothing and tension data of various body areas; The fusion module is used to extract visual wrinkle features of each body region from the multi-view display image, calculate the tension distribution features of each body region based on the tension data, and align and fuse the visual wrinkle features with the tension distribution features to obtain the wrinkle state vector of each body region. The classification module is used to classify each body region into normal overhanging folds, slight stacking folds, obvious stretching folds, and severe skew folds based on the fold state vector and to determine the areas to be adjusted. The adjustment module is used to calculate the joint posture fine-tuning amount of the area to be adjusted and perform the adjustment. After the clothing fabric is relaxed and stabilized, the operation steps from the acquisition module to the classification module are repeated until all body areas are normal hanging wrinkles or slight pile wrinkles, and then the shooting ready signal is output.