A method and device for processing user fitting data
Patent Information
- Application Number
- CN202610877290.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-22
AI Technical Summary
[0006]然而,现有虚拟试衣技术难以在单目低成本硬件条件下,仍无法满足智能试衣镜、互动导购屏、移动端试衣等场景对高稳定性、高适配性、高流畅度交互的需求,因此亟需一种单目图像条件下的多模态的试衣方法,以确保单目低成本硬件条件下试衣场景中的高稳定性、高适配性、高流畅度交互
[0018]采用本发明,通过对单目摄像头获取用户图像数据,生成用户服装数据的服装展示控制指令;并获取用户输入的自然语言文本,生成第一候选服装集合;最终依据所述服装展示控制指令和第一候选服装集合,对第一候选服装集合中的服装数据进行用户试衣数据的生成和展示,从而实现试衣场景中的用户服装试穿到用户时,能确保稳定性、适配性、以及流畅的交互;本发明中交互稳定度判定带来的展示流畅性提升;连续语义梯度调整带来的渐进式推荐体验;零样本语义映射有免标注、免训练的优势。
Smart Images

Figure CN122799050A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method and apparatus for processing user fitting data. Background Technology
[0002] With the rapid popularization of new retail, smart shopping guides, virtual try-on, and multimodal interaction technologies, interactive clothing recommendation and try-on display systems based on visual perception and natural language understanding have become core functions of offline stores, online e-commerce, and mobile applications. Existing smart try-on, interactive shopping guide, AR try-on, and multimodal recommendation systems are typically based on monocular cameras, voice input, text understanding, and recommendation algorithms to complete human body state recognition, user intent parsing, and content display control. Their mainstream technical approaches can be mainly divided into three categories.
[0003] The first category is human posture recognition solutions based on single-frame images. This solution uses a monocular camera to capture a single frame image of the user, and utilizes a human keypoint detection model to identify key points such as the head, shoulders, hands, waist, and feet. Based on the size of the human detection box, the coordinates of the center point, or the relative positions of the key points, it determines the user's state, such as moving closer, further away, turning around, raising an arm, or remaining still, and triggers clothing switching, virtual try-on display refresh, or interactive prompts accordingly. This type of solution is simple to implement and has a fast inference speed, and is widely used in lightweight interactive terminals.
[0004] The second category is semantic recognition solutions based on voice or text input. This approach obtains user descriptions through voice recognition or text input, such as requests like "more mature," "brighter but not too exaggerated," "more slimming," or "not too formal." It then uses keyword matching, tag retrieval, intent classification models, or large language models to analyze the user's intent. Finally, it retrieves matching items from clothing, styling, or content libraries to output recommendation results and adjust fitting parameters. This type of solution directly responds to users' subjective needs, enhancing proactive interaction.
[0005] The third category is recommendation schemes based on multimodal fusion. This approach simultaneously extracts visual features from images, textual features from speech, user profiles, and historical behavioral data. It integrates multimodal information through feature concatenation, weighted fusion, or similarity calculation, ultimately outputting clothing recommendations, outfit suggestions, virtual try-on images, or display control commands. This type of scheme is more comprehensive in terms of information dimensions, improving the accuracy and rationality of recommendations to some extent.
[0006] However, existing virtual try-on technology still cannot meet the requirements of high stability, high adaptability, and high smoothness of interaction in scenarios such as smart try-on mirrors, interactive shopping guide screens, and mobile try-on under the condition of low-cost monocular hardware. Therefore, there is an urgent need for a multimodal try-on method under monocular image conditions to ensure high stability, high adaptability, and high smoothness of interaction in try-on scenarios under the condition of low-cost monocular hardware. Summary of the Invention
[0007] This invention provides a method and apparatus for processing user fitting data, which can achieve highly stable, highly adaptable, and highly smooth interaction in fitting scenarios.
[0008] In a first aspect, embodiments of the present invention provide a method for processing user fitting data, comprising: acquiring first user image data through a monocular camera; acquiring first spatial geometric features of the user based on the first user image data; determining the user's first spatial state by combining the first spatial geometric features of consecutive frames with a visual perspective projection model; generating a spatial perception trigger signal based on the first spatial state, and generating a clothing display control instruction for clothing data; acquiring first natural language text input by the user, and converting the first natural language text into a first semantic feature vector within the latent space of the clothing data using a zero-sample semantic mapping operator; performing a similarity search in a clothing inventory vector set based on the first semantic feature vector to generate a first candidate clothing set; and generating and displaying user fitting data for clothing data in the first candidate clothing set based on the clothing display control instruction and the first candidate clothing set.
[0009] Optionally, in the user fitting data processing method of this embodiment of the invention, the step of obtaining the user's first spatial geometric features based on the first user image data includes: performing human body key point detection on the first user image data acquired by a monocular camera to obtain a set of key points including the head, shoulders, hips, left foot, right foot, left ankle, and right ankle; determining foot reference points based on the left foot key point, right foot key point, or ankle key point, and performing time smoothing processing on the foot reference points; calculating the pixel distance from the foot reference point to the bottom edge of the image, the human body imaging height, the lateral offset of the foot, the longitudinal offset of the foot, and the change in the bottom contour of the human body.
[0010] Optionally, in the user fitting data processing method of this embodiment of the invention, the step of determining the user's first spatial state by combining the visual perspective projection model with the first spatial geometric features of consecutive frames includes: acquiring N consecutive frames of images obtained by a monocular camera to form a detection time window, where N is 8 to 15 frames; calculating the rate of change of human body imaging height and the rate of change of distance between the feet and the bottom edge pixels between adjacent frames; and constructing a spatial displacement determination formula based on the visual perspective projection model. in, The height of the human body image in the current frame. The height of the human body image in the previous frame; The distance from the foot to the bottom edge in the current frame is 1 pixel. The distance in pixels from the foot to the bottom edge in the previous frame; For highly normalized rate of change in human body imaging, The normalized rate of change of the pixel distance at the bottom edge of the foot; a fixed threshold for spatial state determination: height fluctuation threshold. Feature static threshold Limb movement threshold If satisfied and It determines that the user is in a close proximity state; if the condition is met... and The system determines that the user is in a distant state; if the lateral displacement of the foot continuously increases or decreases, and The system determines that the user is in a lateral movement state; if the rate of change of the human body imaging height, the lateral offset of the foot, and the longitudinal offset of the foot are all less than [a certain value], the system will determine that the user is in a lateral movement state. The system determines that the user is in a stationary state; it calculates the average offset of the shoulder and hand keypoint coordinates, and determines the user's status when the average offset of the limb keypoints is greater than a certain value. And the rate of change of foot features is less than At that time, it is determined that the user is in a state of partial limb movement.
[0011] Optionally, in the user fitting data processing method of this embodiment of the invention, the step of generating a spatial perception trigger signal based on a first spatial state and generating a clothing display control instruction for clothing data includes: when the user is in a close-up state, generating a local detail display instruction to magnify the clothing fabric, texture, and neckline; when the user is in a far-away state, generating an overall matching display instruction to display the full-body fitting effect or the effect of multiple clothing combinations; when the user is in a lateral movement state, generating a multi-angle switching instruction to display the side view effect or multi-view comparison; when the user is in a stationary state and the duration exceeds a preset threshold, generating an active recommendation or interactive guidance instruction; when the user is in a partial limb movement state, blocking the spatial trigger signal and not generating a display switching instruction.
[0012] Optionally, in the user fitting data processing method of this embodiment of the invention, the step of obtaining the first natural language text input by the user and converting the first natural language text into a first semantic feature vector in the latent space of clothing data through a zero-shot semantic mapping operator includes: pre-constructing a clothing-specific five-dimensional feature semantic space, wherein the five feature dimensions are color, fit, material, style, and pattern; performing hierarchical semantic decomposition on the first natural language text to identify the intent anchors, degree modifiers, and negation constraint words in the first natural language text; and configuring independent initial semantic weights for each feature dimension. , j is the feature dimension index; for feature dimensions with negation constraints, the negation weight decay formula is used: Weight correction is performed, where the attenuation coefficient is... Intensity adjustment coefficients are configured based on degree modifiers. This is used to amplify or weaken the intensity of corresponding feature expressions; through the feature fusion formula: After fusing all corrected features and performing L2 normalization, the first semantic feature vector is generated. .
[0013] Optionally, in the user fitting data processing method of this embodiment of the invention, the step of performing similarity retrieval in the clothing inventory vector set based on the first semantic feature vector to generate a first candidate clothing set includes: encoding clothing images, text descriptions, style tags, material tags, and scene tags into clothing inventory vectors; calculating the cosine similarity between the first semantic feature vector and each clothing inventory vector, and sorting them from high to low similarity to obtain a preliminary clothing set; and performing a secondary sorting of the preliminary clothing set in combination with the user space state, current display state, and historical interaction information to obtain the first candidate clothing set.
[0014] Optionally, in the user fitting data processing method of this embodiment of the invention, the step of generating and displaying user fitting data for clothing data in the first candidate clothing set based on the clothing display control instruction and the first candidate clothing set includes: determining the display granularity and display perspective according to the clothing display control instruction, and establishing a display granularity weight mapping rule: the close state corresponds to the local detail weight α1∈[0.7, 0.9], the far state corresponds to the overall matching weight α2∈[0.6, 0.8], and the lateral movement state corresponds to the multi-angle switching weight α3∈[0.5, 0.7]; based on the first candidate clothing set and the user's human body posture features, a multi-constraint fusion fitting generation model is used to output the fitting image I_out, and the generation formula is: I_out =ω1·I_cloth + ω2·I_body + ω3·I_bg where I_cloth is the clothing texture feature map, I_body is the human body structure feature map, and I_bg is the background environment feature map; ω1, ω2, and ω3 are adaptive fusion weights, satisfying the normalization constraint: ω1 + ω2 + ω3 = 1, and ω1∈[0.4, 0.6], ω2∈[0.3, 0.45], ω3∈[0.1, 0.2]; Introduce the continuity constraint factor β to calculate the smooth update loss of the fitting results: L_smooth = |I_current − I_prev|≤ β, where β=0.05; Combine the current display state and interaction context, output the fitting data according to the principle of minimizing L_smooth to realize the fitting display.
[0015] Optionally, the user fitting data processing method in this embodiment of the invention further includes: performing multimodal fusion stability calculation on the current interactive fitting display state, and constructing an interaction stability evaluation formula: ,in, The interaction stability coefficient, The average rate of change of human body height in consecutive frames. The average rate of change of foot distance in consecutive frames. The cosine similarity between the current semantic vector and the semantic vector of the previous frame; These are modal weighting coefficients, and they satisfy... The values range from 0.35 to 0.5, 0.25 to 0.4, and 0.15 to 0.3, respectively; a stability threshold is set. ;when When the user interaction is deemed stable, the current first candidate clothing set is locked and kept statically displayed; when... When the user is in a dynamic interaction process, the candidate clothing sorting is updated and the granularity of the try-on display is adjusted.
[0016] Optionally, the method for processing user fitting data in this embodiment of the invention further includes: acquiring second natural language text input by the user, and generating a second semantic feature vector through a zero-shot semantic mapping operator. The clothing inventory vector corresponding to the first candidate clothing set is denoted as... Construct the formula for adjusting the offset of the continuous semantic gradient: ,in, This is the first semantic feature vector. It is a semantic adjustment coefficient and its value range is , This is the historical interaction semantic offset vector. The context weight coefficient has a value range of 1. For vectors Apply cosine similarity normalization constraints to make , The continuity threshold is set to a value of 1. Based on the first set of candidate garments, according to Similarity is used for secondary retrieval and ranking to generate a second set of candidate clothing, thus realizing continuous clothing recommendation with semantic gradient constraints.
[0017] Secondly, embodiments of the present invention provide a user fitting data processing device, comprising: an image acquisition module configured to acquire first user image data via a monocular camera; a geometric feature extraction module configured to acquire first spatial geometric features of the user based on the first user image data; a spatial state determination module configured to determine the user's first spatial state by using a visual perspective projection model combined with the first spatial geometric features of consecutive frames; a display control module configured to generate a spatial perception trigger signal based on the first spatial state, and generate a clothing display control instruction for the clothing data; a semantic mapping module configured to acquire first natural language text input by the user, and convert the first natural language text into a first semantic feature vector within the latent space of the clothing data using a zero-shot semantic mapping operator; a candidate retrieval module configured to perform a similarity retrieval in a clothing inventory vector set based on the first semantic feature vector, and generate a first candidate clothing set; and a fitting display module configured to generate and display user fitting data for the clothing data in the first candidate clothing set based on the clothing display control instruction and the first candidate clothing set.
[0018] This invention acquires user image data using a monocular camera to generate clothing display control instructions for user clothing data; it also acquires natural language text input by the user to generate a first candidate clothing set; finally, based on the clothing display control instructions and the first candidate clothing set, it generates and displays user fitting data for the clothing data in the first candidate clothing set, thereby ensuring stability, adaptability, and smooth interaction when the user tries on clothing in a fitting scenario. The invention also features improved display smoothness due to interaction stability determination; a progressive recommendation experience through continuous semantic gradient adjustment; and the advantages of zero-sample semantic mapping, which requires no annotation or training. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating a method for processing user fitting data according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the user fitting data processing device provided in Embodiment 2 of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.
[0023] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”
[0024] Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0025] In this embodiment of the invention, user image data is acquired by a monocular camera to generate clothing display control instructions for user clothing data; natural language text input by the user is acquired to generate a first candidate clothing set; finally, based on the clothing display control instructions and the first candidate clothing set, user fitting data is generated and displayed for the clothing data in the first candidate clothing set; thus, by combining the clothing display control instructions and the candidate clothing set, stability, adaptability, and smooth interaction can be ensured when the user tries on clothing in the fitting scenario.
[0026] Example 1: Figure 1 This is a flowchart illustrating a method for processing user fitting data according to Embodiment 1 of the present invention. The method for processing user fitting data in Embodiment 1 includes the following steps: Step 100: Acquire first user image data using a monocular camera. In this first embodiment, the data processing method for fitting rooms is applied to a smart fitting mirror device in an offline physical clothing store. This fitting mirror is a vertical, integrated terminal with a 55-inch 4K high-definition display screen. A monocular RGB high-definition camera is embedded in the center of the top of the screen. The camera uses a fixed-focus wide-angle lens with a focal length of 3.6mm, a horizontal field of view of 72°, a vertical field of view of 55°, and an image acquisition resolution of 1920×1080. The acquisition frame rate is fixed at 25fps to ensure smooth continuous frame images and accurate geometric feature extraction. When a user enters the effective interaction area of 0.5 to 3 meters in front of the fitting mirror, the monocular camera is automatically awakened by the ambient light sensor and begins to continuously acquire the user's real-time image data, forming a continuous video stream. The first user image data is the RGB image output frame by frame in this video stream. After acquisition, the image data is directly transmitted to the edge computing unit built into the fitting mirror for real-time processing, without cloud forwarding, ensuring user privacy and security and an interaction response latency of less than 150ms, meeting the real-time fitting display needs of offline stores. Preferably, in this embodiment, to improve image processing efficiency and feature detection stability, the system performs a preprocessing pipeline on the raw images captured by the monocular camera: First, lens distortion correction is performed to eliminate barrel distortion at the edges caused by wide-angle lenses; second, adaptive brightness and contrast adjustment is performed to adapt to the mixed lighting environment of fluorescent lights, spotlights, and natural light in the store, avoiding key point detection failure due to excessive darkness or overexposure; simultaneously, Gaussian filtering is performed to reduce noise and suppress interference from image noise on human contours and key point localization. After preprocessing, the system performs effective region cropping on the image, retaining only the central 80% area as the human detection area, and removing invalid background on the left and right sides to reduce the computational load of subsequent algorithms and improve overall running speed.
[0027] Step 101: Obtain the user's first spatial geometric features based on the first user image data. In this embodiment, the first spatial geometric feature is the basic data for realizing user spatial behavior perception, and it is all calculated locally in real time on the fitting mirror. First, the system uses a lightweight human pose estimation algorithm to detect human key points on the preprocessed user image, and outputs a fixed set of key points, including: head vertex, left shoulder key point, right shoulder key point, hip center point, left foot key point, right foot key point, left ankle key point, and right ankle key point. All key points are represented in pixel coordinates, with the upper left corner of the image as the origin, the horizontal axis as the X-axis, and the vertical axis as the Y-axis. Second, a foot reference point is determined based on the relevant foot key points: the average of the X coordinates of the left and right foot key points is taken as the X coordinate of the foot reference point, and the average of the Y coordinates of the left and right ankle key points is taken as the Y coordinate of the foot reference point, forming a unique stable foot reference point. To avoid reference point jitter caused by user limb micro-movements and single-frame detection errors, the system performs time-sliding window smoothing on the foot reference point. The window length is set to 7 frames. By filtering jitter through weighted averaging of multiple consecutive frames, the coordinate changes of the foot reference point are made continuous and stable. After obtaining a stable foot reference point, the system calculates five primary spatial geometric features: 1) Pixel distance D from the foot reference point to the bottom edge of the image: Subtracting the Y coordinate of the foot reference point from the image height value yields the vertical pixel distance from the foot to the bottom edge of the image, reflecting the physical distance between the user and the dressing mirror; 2) Human body imaging height H: Subtracting the Y coordinate of the foot reference point from the Y coordinate of the head vertex yields the overall pixel height of the user in the image, following the perspective projection rule of "nearer objects appear larger, farther objects appear smaller"; 3) Lateral foot offset: The absolute difference in the X coordinate of the foot reference point between consecutive frames, used to determine the user's left-right movement range; 4) Vertical foot offset: The absolute difference in the Y coordinate of the foot reference point between consecutive frames, used to determine the user's forward-backward movement range; 5) Change in the human body bottom contour: Sampling the lower body contour line by line and calculating the contour overlap between consecutive frames, used to distinguish between whole-body displacement and local upper limb movements. All of the above geometric features are output in numerical form, providing accurate input for subsequent visual perspective projection models.
[0028] Step 102: Determine the user's first spatial state by combining the visual perspective projection model with the first spatial geometric features of consecutive frames. In this embodiment, since the monocular camera of the smart fitting mirror is fixed in position, the user's actions such as forward and backward movement, left and right movement, stillness, and waving in the physical space will form a stable and quantifiable geometric change pattern on the imaging plane. Therefore, the system uses a visual perspective projection model as the underlying basis for spatial state determination to accurately distinguish five user spatial states. First, the system acquires N consecutive frames of images to form a detection time window. In this embodiment, N is 10 frames, which balances real-time performance and anti-interference capability, avoiding misjudgments caused by momentary jitter. In this invention, N can be 8 to 15 frames. Second, a spatial displacement determination formula is constructed based on the visual perspective projection model to calculate the normalized rate of change between adjacent frames:
[0029] in, The height of the human body image in the current frame. The height of the human body image in the previous frame ; represents the pixel distance from the foot to the bottom edge in the current frame. ΔH is the pixel distance from the foot to the bottom edge in the previous frame; ΔH is the normalized rate of change of human height, and ΔD is the normalized rate of change of foot distance. Both are dimensionless relative values, which can eliminate system biases caused by user height, posture, and camera installation height. This embodiment presets a fixed judgment threshold: height fluctuation threshold. Feature static threshold Limb movement threshold The aforementioned thresholds were calibrated through extensive real-world user interaction tests in offline stores. The system determines the user's primary spatial state according to the following rules: 1) Approaching state: ΔH > 0 and ΔD < 0, i.e., the body size increases and the feet are close to the bottom edge, indicating the user is approaching the fitting room mirror; 2) Moving away state: ΔH < 0 and ΔD > 0, i.e., the body size decreases and the feet are away from the bottom edge, indicating the user is moving away; 3) Lateral movement state: the lateral displacement of the feet changes continuously, and That is, high stability and left-right movement; 4) stationary state: the rate of change of body height and the lateral / longitudinal offset of the feet are both less than 5) Local limb movement status: The average deviation of the key points of the shoulder and hand is greater than 100%. Furthermore, the foot features remain stable and are identified as localized actions such as waving, raising a hand, or adjusting hair. Based on these rules, the system can achieve high-precision spatial state determination using only monocular images without the need for a depth camera.
[0030] Step 103: Generate a spatial perception trigger signal based on the first spatial state, and generate clothing display control instructions based on clothing data. In this embodiment, five spatial states are mapped to corresponding display control commands, realizing a natural interaction where "the fitting mirror displays whatever the user goes," perfectly matching offline shopping habits. 1) When the user is close, the system generates a command to display details, automatically magnifying the mirror to show the fabric texture, seams, collar, cuffs, buttons, and other details of the garment, allowing the user to observe the quality up close. 2) When the user is far away, the system generates a command to display the overall outfit, switching to a full-body look and showing the overall effect of the top, pants, and shoes, highlighting body proportions. 3) When the user is moving laterally, the system generates a command to switch between multiple angles, automatically displaying the fit from the front, left side, and right side, satisfying the user's need for a comprehensive view. 4) When the user remains still for more than 2.5 seconds, the system generates a proactive recommendation or interactive guidance command, automatically popping up similar styles and popular outfits, or prompting the user to input their needs via voice. 5) When the user is making partial body movements, the system blocks spatial trigger signals and does not switch the display screen, avoiding frequent screen jumps caused by waving or raising hands, thus improving display stability. This mechanism allows the mirror to control the display content without the user clicking the screen, simply through natural body movements, significantly improving the smoothness of offline shopping.
[0031] Step 104: Obtain the first natural language text input by the user, and convert the first natural language text into a first semantic feature vector in the latent space of the clothing data using a zero-shot semantic mapping operator. In this embodiment, the user inputs first natural language text through the built-in microphone of the fitting mirror, such as "I want a light blue, loose-fitting, plain cotton casual top." The system converts the text into a clothing-specific semantic feature vector, achieving zero-sample accurate understanding. First, the system pre-constructs a five-dimensional semantic space for clothing features, with the five dimensions fixed as: color, fit, material, style, and pattern. Second, the natural language text undergoes hierarchical semantic decomposition to identify: Intended anchor points: light blue, loose fit, cotton, casual, no pattern; Degree modifier: None (in this example); Negation constraint term: No pattern (negating the pattern dimension). Then, initial semantic weights are assigned to the five dimensions. , ∈[0.15,0.45], where j is the feature dimension index: Color w_1 = 0.30; Pattern size w_2 = 0.40; Material w_3=0.35; Style w_4 = 0.25; Pattern w_5 = 0.20. Due to the presence of the negation constraint "no pattern", the negation weight decay formula is used for correction: Preferably, in this embodiment, σ ∈ [0.2, 0.6] is taken as σ = 0.5, after pattern dimension correction. =0.10. The degree adjustment coefficient τ is set to 1.0. In this invention, the degree adjustment coefficient... The initial vector is obtained through the feature fusion formula: j=1~5 Let j be the normalized basic feature vector corresponding to the j-th clothing feature dimension, where j=1,2,3,4,5, corresponding to the five inherent feature vectors of clothing: color, pattern, material, style, and design. Each basic feature vector is pre-mapped to the same clothing latent semantic space, with consistent vector dimensions, and is used to complete the weighted fusion calculation of multiple feature dimensions. Finally, Perform L2 normalization to obtain the final first semantic feature vector. This vector will be used for subsequent clothing retrieval.
[0032] Step 105: Based on the first semantic feature vector, perform similarity retrieval in the clothing inventory vector set to generate the first candidate clothing set. In this embodiment, the system is suitable for the rapid retrieval of tens of thousands of garments in offline stores. First, the store's backend system uniformly encodes the images, titles, styles, materials, and scene tags of all garments into a garment inventory vector. Vector dimension and Completely identical. Secondly, calculation. With all The cosine similarity scores are used to rank the items, and the top 30 items are selected to form an initial selection set. Finally, a second ranking is performed, incorporating three pieces of information: 1) the user's current spatial state (near / far / lateral movement); 2) the current display state of the fitting room mirror (details / overall / multi-angle); and 3) the user's historical click, pause, and skip records. The top 10 items after this second ranking are selected as the first candidate set, ensuring that the recommendations both meet semantic requirements and are suitable for the current interaction state.
[0033] Step 106: Based on the clothing display control instructions and the first candidate clothing set, generate and display user fitting data for the clothing data in the first candidate clothing set. In this embodiment, the fitting room generation and display process realizes spatially perceptive display and multi-constraint fitting room synthesis.
[0034] (1) Display granularity and weight allocation Assign display weights based on the display instructions corresponding to the spatial status: Proximity state: Local detail weight α1 = 0.80; Distant state: Overall combination weight α2 = 0.70; Lateral movement state: multi-angle switching weight α3=0.60.
[0035] In this invention, the local detail weight α1∈[0.7, 0.9] corresponding to the approach state, the overall combination weight α2∈[0.6, 0.8] corresponding to the distance state, and the multi-angle switching weight α3∈[0.5, 0.7] corresponding to the lateral movement state are all permissible and are within the protection scope of this invention.
[0036] (2) Multi-constraint fusion fitting room generation model Based on the first candidate clothing set and the user's human posture features, a multi-constraint fusion virtual fitting generation model is used to output the virtual fitting image I_out: I_out = ω1・I_cloth + ω2・I_body + ω3・I_bg where: I_cloth: clothing texture feature map; I_body: human body structure feature map; I_bg: virtual fitting mirror background environment feature map; the weights satisfy the normalization constraint: ω1+ω2+ω3=1. In this embodiment, we take: ω1=0.50, ω2=0.35, ω3=0.15.
[0037] In this invention, ω1∈[0.4, 0.6], ω2∈[0.3, 0.45], and ω3∈[0.1, 0.2] are all valid values and fall within the scope of this invention. A continuity constraint factor β is introduced to calculate the smooth update loss of the fitting results: L_smooth = |I_current − I_prev| ≤ β, where β = 0.05. Combining the current display state and the interaction context, the fitting data is output according to the principle of minimizing L_smooth, thus realizing the fitting display.
[0038] Preferably, during the virtual try-on data display interaction in this embodiment one, a multimodal fusion stability calculation is further performed on the current interactive virtual try-on display state, wherein an interaction stability evaluation formula is constructed as follows: in, The interaction stability coefficient, The average rate of change of human body height in consecutive frames. The average rate of change of foot distance in consecutive frames. The cosine similarity between the current semantic vector and the semantic vector of the previous frame; These are modal weighting coefficients, and they satisfy... The values range from 0.35 to 0.5, 0.25 to 0.4, and 0.15 to 0.3, respectively. Preferably, θ1 = 0.40, θ2 = 0.30, and θ3 = 0.30, satisfying θ1 + θ2 + θ3 = 1, which is a set of preferred modal weight coefficients. Set stability threshold ;when When the user interaction is deemed stable, the current first candidate clothing set is locked and kept statically displayed; when... When the user is in a dynamic interaction process, the candidate clothing sorting is updated and the granularity of the try-on display is adjusted.
[0039] Preferably, in this first embodiment, the entire method realizes monocular visual spatial perception, zero-sample semantic understanding, multi-constraint fitting generation, and dynamic stability determination on an offline smart fitting mirror. The entire process does not require manual operation by the user. Immersive virtual fitting can be completed simply by walking, standing, or speaking, achieving the requirements of high stability, high adaptability, and high smoothness of interaction.
[0040] Example 2: To more clearly illustrate the structure of the user fitting data processing device involved in this invention, Figure 2 This is a schematic diagram of the user fitting data processing device provided in Embodiment 2 of the present invention.
[0041] Embodiment 2 of the present invention discloses a user fitting data processing device 200, which includes: an image acquisition module 201, a geometric feature extraction module 202, a spatial state determination module 203, a display control module 204, a semantic mapping module 205, a candidate retrieval module 206, and a fitting display module 207. The image acquisition module 201 is configured to acquire first user image data through a monocular camera; the geometric feature extraction module 202 is configured to acquire the user's first spatial geometric features based on the first user image data; the spatial state determination module 203 is configured to determine the user's first spatial state by using a visual perspective projection model combined with the first spatial geometric features of continuous frames; the display control module 204 is configured to generate a spatial perception trigger signal based on the first spatial state and generate clothing display control instructions for clothing data; the semantic mapping module 205 is configured to acquire the first natural language text input by the user and convert the first natural language text into a first semantic feature vector in the latent space of clothing data through a zero-shot semantic mapping operator; the candidate retrieval module 206 is configured to perform similarity retrieval in the clothing inventory vector set based on the first semantic feature vector to generate a first candidate clothing set; and the fitting display module 207 is configured to generate and display user fitting data for clothing data in the first candidate clothing set based on the clothing display control instructions and the first candidate clothing set.
[0042] In this embodiment, the user fitting data processing device 200, and its components including image acquisition module 201, geometric feature extraction module 202, spatial state determination module 203, display control module 204, semantic mapping module 205, candidate retrieval module 206, and fitting display module 207, will be described in detail with a practical and complete example.
[0043] This embodiment differs from Embodiment 1 in that the user fitting data processing device 200 described in this embodiment uses a smartphone APP as its hardware carrier and is applied to scenarios such as self-service shopping in offline clothing stores, remote fitting at home, and mobile outfit recommendations. Users only need to use the front / rear monocular RGB camera, touch screen, and microphone of their mobile phone to complete the entire interaction process, without the need for a dedicated fitting mirror device, thus possessing higher versatility, portability, and widespread applicability. The user fitting data processing device 200 runs on Android or iOS systems, and all algorithm modules support local inference on the terminal, enabling stable operation in environments without or with weak networks. It achieves a complete functional closed loop of monocular visual spatial perception, zero-sample semantic understanding, multi-constraint fitting generation, interaction stability determination, and continuous semantic gradient adjustment.
[0044] In this embodiment, the image acquisition module 201 serves as the input terminal of the user fitting data processing device 200. Its core function is to utilize the built-in monocular RGB camera of the mobile phone to acquire, preprocess, and manage the user's image data. Specifically, the image acquisition module 201 obtains access to the mobile phone's camera via the system API, and by default enables the front-facing camera for user portrait capture, meeting the user's selfie-style fitting needs. Users can also manually switch to the rear-facing camera, suitable for full-body shots and long-distance fitting scenarios. The camera parameters are configured as follows: resolution 1280×720 or 1920×1080, frame rate 20-30fps, fixed focal length, automatic exposure, and automatic white balance, ensuring clear and stable image frames can be output under various lighting conditions, including indoor store lighting, home lighting, and natural light. The image acquisition module 201 performs real-time preprocessing on the acquired raw image: first, it performs lens distortion correction to eliminate edge distortion caused by the wide-angle camera of the mobile phone; second, it performs Gaussian filtering for noise reduction and adaptive brightness equalization to suppress the interference of image noise, overexposure, and underexposure on subsequent feature extraction; finally, it performs effective region cropping, retaining 75%-85% of the human body area in the center of the image and removing invalid background to reduce the computational load of subsequent modules. Preferably, in this embodiment, the image acquisition module 201 uses a frame buffer queue to manage the continuous image sequence, with a buffer length set to 15 frames, ensuring that the spatial state determination module 203 can acquire continuous and temporally complete image data, while avoiding excessive memory usage. The image acquisition module 201 outputs the preprocessed image frame by frame in RGB format to the geometric feature extraction module 202, completing the first step of the data link transmission. In this invention, the frame buffer queue buffer length of the image acquisition module 201 can be set to 8 to 15 frames, all of which are within the protection scope of this invention.
[0045] In this embodiment, the geometric feature extraction module 202 is the core foundational module for realizing user spatial behavior perception, responsible for extracting steady-state and quantized first spatial geometric features from user images. The geometric feature extraction module 202 receives a single-frame image from the image acquisition module 201, first runs a lightweight human keypoint detection model, and outputs a fixed set of keypoints, including: head vertex, left shoulder, right shoulder, hip center point, left foot keypoint, right foot keypoint, left ankle keypoint, and right ankle keypoint. All keypoints are stored in pixel coordinates, with the coordinate system having the upper left corner of the image as the origin, the horizontal axis as the X-axis, and the vertical axis as the Y-axis. After obtaining the keypoints, the geometric feature extraction module 202 calculates foot reference points based on the left foot, right foot, or ankle keypoints: the average of the X-coordinates of the left and right foot keypoints is taken as the X-coordinate of the foot reference point, and the average of the Y-coordinates of the left and right ankle keypoints is taken as the Y-coordinate of the foot reference point, forming a unique steady-state anchor point. To eliminate single-frame detection errors and coordinate jumps caused by minor limb tremors, the geometric feature extraction module 202 performs time-sliding window smoothing on the foot reference point. The window length is set to 7 frames, and the reference point changes are made continuous and stable through multi-frame weighted averaging. In this invention, the window length can be 8 to 15 frames, all of which are within the scope of protection of this invention. Based on the stable foot reference point, the geometric feature extraction module 202 calculates five first spatial geometric features and outputs them all to the spatial state determination module 203: The pixel distance D from the foot reference point to the bottom edge of the image: image height minus the Y coordinate of the foot reference point, directly reflecting the physical distance between the user and the phone camera; Human body imaging height H: The Y coordinate of the head vertex minus the Y coordinate of the foot reference point, representing the vertical size of the human body in the picture, following the perspective projection rule of "nearer is larger and farther is smaller"; Lateral foot offset: The absolute difference in the X coordinate of the foot reference point between consecutive frames, used to determine the magnitude of left and right movement; Foot longitudinal offset: The absolute difference in the Y coordinate of the foot reference point between consecutive frames, used to determine the magnitude of forward and backward movement; Human body bottom contour variation: The lower body contour is sampled line by line, and the pixel overlap of the contour in consecutive frames is calculated to distinguish between global displacement and local limb movements. The geometric feature extraction module 202 uses foot anchoring extraction, abandoning the traditional full-body bounding box detection, which greatly improves feature stability and provides accurate input for subsequent visual perspective projection models.
[0046] In this embodiment, the spatial state determination module 203, based on a visual perspective projection model and continuous frame geometric features, accurately determines five spatial states of the user, serving as a core hub connecting visual perception and display control. The spatial state determination module 203 first obtains continuous temporal features from the geometric feature extraction module 202, constructing a detection temporal window with N=10 frames to balance real-time performance and anti-interference capability. Based on the visual perspective projection model, the spatial state determination module 203 calculates the normalized rate of change using the following formula: in, The height of the human body image in the current frame. The height of the human body image in the previous frame; The distance from the foot to the bottom edge in the current frame is 1 pixel. ΔH represents the pixel distance from the foot to the bottom edge in the previous frame; ΔH and ΔD are dimensionless relative change values, which can eliminate system deviations caused by user height, standing posture, and phone grip height. The spatial state determination module 203 has a built-in fixed determination threshold: height fluctuation threshold. Feature static threshold Limb movement threshold All parameters were calibrated through real-world scenario testing. The spatial state determination module 203 outputs a unique spatial state to the display control module 204. Proximity state: ΔH>0 and ΔD<0, the human body becomes larger, the feet are close to the bottom edge, and the user walks closer to the phone; Remote state: ΔH < 0 and ΔD > 0, the human body shrinks, the feet move away from the bottom edge, and the user moves backward; Lateral movement state: The lateral displacement of the foot changes continuously, and |ΔH|≤T_H, the user moves left and right; Dwelling status: Both height and lateral / vertical offset change rates are < The user stands still; Local limb movement status: Shoulder / hand key point deviation > And the foot is stable (or the rate of change of foot characteristics is less than 1%). (At the time), the user's upper limb movements such as waving or raising their hand. The spatial state determination module 203 can determine the spatial state using only monocular images, without the need for additional hardware such as depth cameras or radar, and has extremely strong terminal adaptability.
[0047] In this embodiment, the display control module 204 is responsible for converting the spatial state into specific clothing display control instructions, realizing a natural linkage between "behavior and display". The display control module 204 receives the state result from the spatial state determination module 203, generates instructions according to a fixed mapping rule, and sends them to the fitting display module 207. Proximity state: Generates a command to display local details, controlling the phone screen to zoom in and show details such as fabric, texture, collar, cuffs, and buttons; Away from the main outfit: Generates a command to display the overall outfit, switching to a full-body try-on effect, showcasing the complete outfit including the top, bottoms, shoes, and bag; Horizontal movement state: Generates multi-angle switching commands, automatically displays front, left side, and right side views, and supports multi-view side-by-side comparison; Dwell time: If the dwell time is >2.5 seconds, generate proactive recommendations / interactive guidance commands, pop up similar styles, popular combinations, or prompt for voice input; Localized limb movement state: Spatial trigger signals are blocked, and the screen is not switched to avoid display jumps caused by non-displacement movements. Preferably, in this embodiment, the display control module 204 supports a command anti-shake mechanism: the same state must be maintained for 3-5 frames before the command is triggered to avoid frequent switching caused by momentary state jitter and improve the smoothness of interaction. The commands output by the display control module 204 include parameters such as display granularity, display angle, switching speed, and whether to lock, to fully control the subsequent try-on display process.
[0048] In this embodiment, the semantic mapping module 205 performs zero-sample conversion from natural language text to clothing semantic feature vectors, and is the core module for understanding the user's dressing intentions. The semantic mapping module 205 acquires the user's first natural language text via the phone's microphone, for example: "I want a light gray, loose-fitting, plain knit top," and performs hierarchical semantic parsing. Construct a five-dimensional semantic space for clothing features: color, cut, material, style, and pattern; Semantic decomposition: Identify intent anchors (light gray, loose, knitted, simple, no pattern), degree modifiers, and negative constraint words (no pattern); Assign initial semantic weights to the five dimensions , j represents the feature dimension number: color 0.30, layout 0.40, material 0.35, style 0.25, pattern 0.20; Perform weight decay on the negative constraint dimension: σ=0.5, pattern weight correction is 0.10; attenuation coefficient in this invention All of these are within the scope of protection of this invention.
[0049] The intensity adjustment coefficient τ ∈ [0.8, 1.2] is configured, and in this embodiment τ = 1.0; Execution feature fusion: j = 1~5; Let j be the normalized basic feature vector corresponding to the j-th clothing feature dimension, where j=1,2,3,4,5, corresponding to the five inherent feature vectors of clothing: color, pattern, material, style, and pattern. Each basic feature vector is pre-mapped to the same clothing latent semantic space, with consistent vector dimensions, and is used to complete the weighted fusion calculation of multiple feature dimensions.
[0050] right Perform L2 normalization to obtain the final first semantic feature vector. The output of semantic mapping module 205 With fixed dimensions and smooth distribution, it can be directly used for clothing vector retrieval. No pre-training or labeled data is required, achieving true zero-sample generalization and fully adapting to the lightweight operation requirements of mobile terminals.
[0051] In this embodiment, the candidate retrieval module 206 is responsible for quickly retrieving and reordering the clothing library based on semantic vectors, outputting a precise candidate set. The candidate retrieval module 206 has a built-in clothing inventory vector library, uniformly encoding the images, titles, styles, materials, and scene tags of all clothing into a single vector. Same-dimensional inventory vector It supports real-time retrieval of tens of thousands of vector data points locally. Specific process: calculate With all Cosine similarity; Sort the items in descending order of similarity and select the top 30 to form an initial selection of clothing items; Perform a second reordering, incorporating three constraints: User's current spatial state (near / far / horizontal); Current display status of the mobile phone (details / overall / multiple angles); User's historical click, dwell, and skip behaviors; After reordering, the top 10 items are selected as the first candidate clothing set and output to the fitting display module 207. This module uses semantic retrieval and spatial state reordering to ensure that the recommendation results are consistent with the textual intent and adapted to the current interaction behavior, solving the problems of traditional recommendations being rigid and not relevant to the scenario.
[0052] In this embodiment, the fitting display module 207 is the output end of the fitting data processing device 200, responsible for completing the fitting image generation, display control, stability determination, and continuous semantic iteration. It is the final presentation module of the entire technical solution.
[0053] (I) Basic Fitting Generation and Display The fitting room display module 207 receives the display control command and the first set of candidate garments, and first determines the display granularity weight based on the spatial status: Proximity: Local detail weight α1 = 0.80; Distance: Overall combination weight α2 = 0.70; Lateral movement: Multi-angle weight α3=0.60.
[0054] In this invention, the proximity state corresponds to a local detail weight α1 ∈ [0.7, 0.9], the distance state corresponds to an overall matching weight α2 ∈ [0.6, 0.8], and the lateral movement state corresponds to a multi-angle switching weight α3 ∈ [0.5, 0.7]. Values within these ranges are all within the scope of protection of this invention. In the fitting display module 207, based on the first candidate clothing set and the user's human body posture features, a multi-constraint fusion fitting generation model is used to output the fitting image I_out: I_out = ω1・I_cloth + ω2・I_body + ω3・I_bg where: I_cloth: Clothing texture feature map; I_body: Human body structural feature map; I_bg: Background environment feature map; weights satisfy normalization: ω1+ω2+ω3=1, in this embodiment ω1=0.50, ω2=0.35, ω3=0.15. In this invention, ω1∈[0.4, 0.6], ω2∈[0.3, 0.45], ω3∈[0.1, 0.2], and values within these ranges are within the protection scope of this invention. In this embodiment, a continuity constraint factor β is introduced to calculate the smooth update loss of the fitting result: L_smooth = |I_current − I_prev| ≤ β, β=0.05. The fitting display module 207 outputs the screen according to the principle of minimizing L_smooth to avoid switching flickering and discontinuity, and ensure smooth display.
[0055] (II) Multimodal interaction stability determination The fitting room display module 207 calculates the interaction stability coefficient C in real time: C = θ1・|ΔH| + θ2・|ΔD| + θ3・cos (V1,V_last) where θ1=0.40, θ2=0.30, θ3=0.30, and the sum is 1; the stability threshold C_th=0.18. In this invention... The values range from 0.35 to 0.5, 0.25 to 0.4, and 0.15 to 0.3, respectively, and all of these values are within the protection scope of this invention.
[0056] C ≤ C_th: The user's state is stable, the candidate set is locked, and static display is maintained; C > C_th: Dynamic user interaction, automatically updating sorting and display granularity. This mechanism avoids frequent screen refreshes when the user shakes or moves, improving the mobile experience.
[0057] (III) Continuous semantic gradient adjustment Preferably, in this second embodiment, when the user initially inputs "light gray loose-fitting knitted simple top without pattern", the first semantic vector is obtained. After the fitting demonstration, the user inputs again via voice: "A darker gray, a slightly more fitted style," which is the second natural language text. The fitting demonstration module 207 fully implements the following technical solution: The semantic mapping module 205 processes the second text to obtain the second semantic feature vector. ; Take the inventory vector corresponding to the first candidate garment ; Execution vector offset formula: Where: λ=0.5; δ=0.2; ΔV_hist is the historical interaction semantic offset vector; in, This is the first semantic feature vector. This is the semantic offset vector for historical interactions; It is a semantic adjustment coefficient and its value range is In this embodiment, a value of 0.5 is preferred. The context weight coefficient has a value range of 1. In this embodiment, a value of 0.2 is preferred; in this invention, The range of values is , The range of values is The values taken within these ranges are all within the protection scope of this invention.
[0058] For vectors Apply cosine similarity normalization constraints to make , The continuity threshold is set to a value of 1. This ensures that the recommendations are consistent and seamless. according to and The similarity is reordered to generate a second set of candidate clothing; Based on the new set, the virtual try-on process continues to generate and display recommendations, achieving a progressive, smooth, and continuous recommendation iteration. Through this process, users can gradually refine their needs through multiple rounds of natural language processing without having to re-search, fully conforming to the lightweight and continuous interactive usage habits of mobile devices.
[0059] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing user fitting data, characterized in that, include: Acquire the first user's image data using a monocular camera; The user's first spatial geometric features are obtained based on the first user's image data; The user's first spatial state is determined by combining the visual perspective projection model with the first spatial geometric features of consecutive frames. Based on the first spatial state, a spatial perception trigger signal is generated, and clothing display control instructions based on clothing data are generated. The first natural language text input by the user is obtained, and the first natural language text is converted into a first semantic feature vector in the latent space of clothing data through a zero-shot semantic mapping operator. Based on the first semantic feature vector, a similarity search is performed in the clothing inventory vector set to generate a first candidate clothing set; Based on the clothing display control instructions and the first candidate clothing set, user fitting data is generated and displayed for the clothing data in the first candidate clothing set.
2. The method according to claim 1, characterized in that, The step of obtaining the user's first spatial geometric features based on the first user image data includes: Human key point detection is performed on the first user image data acquired by a monocular camera to obtain a set of key points including the head, shoulders, hips, left foot, right foot, left ankle, and right ankle; foot reference points are determined based on the left foot key points, right foot key points, or ankle key points, and time smoothing is performed on the foot reference points; the pixel distance from the foot reference points to the bottom edge of the image, the human body imaging height, the lateral offset of the foot, the longitudinal offset of the foot, and the change in the bottom contour of the human body are calculated.
3. The method according to claim 1, characterized in that, The method of determining the user's first spatial state by combining a visual perspective projection model with the first spatial geometric features of consecutive frames includes: The detection time window is formed by acquiring N consecutive frames of images from a monocular camera, where N ranges from 8 to 15 frames. Calculate the rate of change of human body image height and the rate of change of distance between the feet and the bottom edge pixels between adjacent frames, and construct a spatial displacement determination formula based on the visual perspective projection model: in, The height of the human body image in the current frame. The height of the human body image in the previous frame; The distance from the foot to the bottom edge in the current frame is 1 pixel. The distance in pixels from the foot to the bottom edge in the previous frame; For highly normalized rate of change in human body imaging, The normalized rate of change of pixel distance at the bottom edge of the foot; Set a fixed threshold for spatial state determination: height fluctuation threshold Feature static threshold Limb movement threshold ; If satisfied and It determines that the user is in a close proximity state; If satisfied and The system determines that the user is in a remote state. If the lateral displacement of the foot continuously increases or decreases, and The system determines that the user is in a lateral movement state. If the rates of change of human body imaging height, foot lateral offset, and foot longitudinal offset are all less than The system determines that the user is in a dormant state. Calculate the mean offset of the shoulder and hand key points coordinates. When the mean offset of the limb key points is greater than... And the rate of change of foot features is less than At that time, it is determined that the user is in a state of partial limb movement.
4. The method according to claim 1, characterized in that, The clothing display control instructions for generating clothing data based on the first spatial state to generate a spatial perception trigger signal include: When the user is close, a command to display local details is generated to zoom in on the clothing fabric, texture, and neckline. When the user is far away, a command to display the overall outfit is generated to show the full-body try-on effect or the effect of multiple clothing combinations. When the user is moving horizontally, a command to switch between multiple angles is generated to show the side view effect or multi-view comparison. When the user is stationary for a period of time exceeding a preset threshold, a command to actively recommend or guide interaction is generated. When the user is in a partial body movement state, spatial trigger signals are blocked, and no display switching command is generated.
5. The method according to claim 4, characterized in that, The process of obtaining the first natural language text input by the user and converting it into a first semantic feature vector within the latent space of the clothing data using a zero-shot semantic mapping operator includes: A five-dimensional semantic space for clothing features is pre-constructed, with the five feature dimensions being color, cut, material, style, and pattern. Hierarchical semantic decomposition of the first natural language text was performed to identify intent anchors, degree modifiers, and negation constraint words in the first natural language text. Configure independent initial semantic weights for each feature dimension. , j is the feature dimension index; for feature dimensions with negation constraints, the negation weight decay formula is used: Weight correction is performed, where the attenuation coefficient is... Intensity adjustment coefficients are configured based on degree modifiers. This is used to amplify or weaken the intensity of the corresponding feature expression; Through feature fusion formula: After fusing all corrected features and performing L2 normalization, the first semantic feature vector is generated. .
6. The method according to claim 1, characterized in that, The step of generating a first candidate clothing set by performing similarity retrieval in the clothing inventory vector set based on the first semantic feature vector includes: The clothing images, text descriptions, style tags, material tags, and scene tags are encoded into clothing inventory vectors. The cosine similarity between the first semantic feature vector and each clothing inventory vector is calculated, and the initial clothing set is obtained by sorting the similarity from high to low. The initial clothing set is then sorted a second time by combining the user space state, current display state, and historical interaction information to obtain the first candidate clothing set.
7. The method according to claim 1, characterized in that, The step of generating and displaying user fitting data based on the clothing display control instructions and the first candidate clothing set includes: Based on the clothing display control instructions, the display granularity and viewing angle are determined, and a display granularity weight mapping rule is established: the close state corresponds to the local detail weight α1∈[0.7, 0.9], the far state corresponds to the overall matching weight α2∈[0.6, 0.8], and the lateral movement state corresponds to the multi-angle switching weight α3∈[0.5, 0.7]. Based on the first candidate clothing set and the user's human body posture features, a multi-constraint fusion fitting image generation model is used to output the fitting image I_out. The generation formula is: I_out = ω1·I_cloth + ω2·I_body + ω3·I_bg, where I_cloth is the clothing texture feature map, I_body is the human body structure feature map, and I_bg is the background environment feature map; ω1, ω2, and ω3 are adaptive fusion weights, satisfying the normalization constraint: ω1 + ω2 + ω3 = 1, and ω1∈[0.4, 0.6], ω2∈[0.3, 0.45], ω3∈[0.1, 0.2]; Introduce the continuity constraint factor β to calculate the smooth update loss of the fitting results: L_smooth = |I_current − I_prev| ≤ β, where β=0.05; Combine the current display state and interaction context, output the fitting data according to the principle of minimizing L_smooth to realize the fitting display.
8. The method according to claim 7, characterized in that, Further includes: Multimodal fusion stability calculation is performed on the current interactive virtual try-on display state, and an interaction stability evaluation formula is constructed: in, The interaction stability coefficient, The average rate of change of human body height in consecutive frames. The average rate of change of foot distance in consecutive frames. The cosine similarity between the current semantic vector and the semantic vector of the previous frame; These are modal weighting coefficients, and they satisfy... The values range from 0.35 to 0.5, 0.25 to 0.4, and 0.15 to 0.3, respectively. Set stability threshold ;when When the user interaction is deemed stable, the current first candidate clothing set is locked and kept statically displayed; when... When the user is in a dynamic interaction process, the candidate clothing sorting is updated and the granularity of the try-on display is adjusted.
9. The method according to claim 1, characterized in that, Further includes: Obtain the second natural language text input by the user, and generate a second semantic feature vector through a zero-shot semantic mapping operator. The clothing inventory vector corresponding to the first candidate clothing set is denoted as... Construct the formula for adjusting the offset of the continuous semantic gradient: in, This is the first semantic feature vector. It is a semantic adjustment coefficient and its value range is , This is the historical interaction semantic offset vector. The context weight coefficient has a value range of 1. For vectors Apply cosine similarity normalization constraints to make , The continuity threshold is set to a value of 1. Based on the first set of candidate garments, according to Similarity is used for secondary retrieval and ranking to generate a second set of candidate clothing, thus realizing continuous clothing recommendation with semantic gradient constraints.
10. A user fitting data processing device, comprising: The image acquisition module is configured to acquire the first user's image data through a monocular camera; The geometric feature extraction module is configured to obtain the user's first spatial geometric features based on the first user image data; The spatial state determination module is configured to determine the user's first spatial state by combining the visual perspective projection model with the first spatial geometric features of continuous frames. The display control module is configured to generate a spatial perception trigger signal based on the first spatial state, and generate clothing display control instructions based on clothing data. The semantic mapping module is configured to acquire the first natural language text input by the user and convert the first natural language text into a first semantic feature vector in the latent space of the clothing data through a zero-shot semantic mapping operator. The candidate retrieval module is configured to perform similarity retrieval in the clothing inventory vector set based on the first semantic feature vector to generate a first candidate clothing set. The fitting room display module is configured to generate and display user fitting room data based on the clothing display control instructions and the first candidate clothing set.