Multi-mode AI character and clothing commodity image replacement vending machine and method

Through multi-modal AI vending machines and virtual reality technology, personalized clothing images and text generation is achieved, which solves the problem of traffic shrinking and single experience of offline retail, improves consumer participation and conversion rates, and achieves an efficient closed loop of consumer decision-making.

CN120299128APending Publication Date: 2025-07-11ZHONGHE TECHNOLOGY (HEILONGJIANG) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510712356.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional offline retail faces problems such as shrinking traffic, single experience, and delayed consumption decisions, which are difficult to stimulate consumer vitality and conversion rate.

Method used

Vending machines that combine multimodal AI technology with virtual reality can realize voice or gesture interaction through terminals such as electronic screens, VR glasses, smart cars, etc., generate personalized clothing images and text content, support instant preview, payment and fulfillment, and build a full-link closed loop.

Benefits of technology

It has improved the conversion rate, consumer participation and repurchase rate of offline merchants, shortened the decision-making link, achieved efficient virtual to physical consumption conversion, and reshaped the core competitiveness of offline consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299128A_ABST
    Figure CN120299128A_ABST
Patent Text Reader

Abstract

The invention belongs to the cross technical field of artificial intelligence, computer vision and intelligent retail, and discloses a multi-modal AI character and clothing commodity image replacement vending machine, which is a multi-modal AI character and clothing commodity image replacement vending machine, is an innovative retail device fusing a leading-edge multi-modal AI technology, and is a multi-modal AI character and clothing commodity image replacement vending machine. The corresponding clothing commodity images or characters are quickly and accurately matched and generated by using the AI and are displayed on the screen, so that the appearance of the commodity is intuitively presented. Under e-commerce impact and consumption habit change, traditional offline merchants are faced with traffics such as sharp passenger flow reduction, the multi-mode AI vending machine injects vitality in a full-scene consumption entrance mode, a shopping scene is extended to multiple scenes through an electronic screen, VR glasses, an intelligent mobile phone, an intelligent automobile and an intelligent robot multi-element terminal, geographical limitation is broken, and the shopping experience is improved. A mobile consumption network is constructed, a fragmented traffic conversion channel is opened up, and static display is upgraded into dynamic virtual fitting and intelligent copywriting generation by AI remodeling commodity display logic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field at the intersection of artificial intelligence, computer vision and intelligent retail, and specifically relates to a vending machine and method for replacing multi-modal AI text with clothing product images. Background Art

[0002] In the industry dilemma where physical retail faces online impact and homogeneous consumption scenarios, the multi-modal AI graphic and text replacement vending machine reconstructs the offline business ecosystem driven by technology, opening up a new path for traditional merchants to break the deadlock. First of all, it solves the problem of shrinking offline traffic. By integrating all-scenario entrances such as electronic screens, VR glasses, smart cars, smart robots and mobile phones, the shopping scenario is extended from fixed stores to life touchpoints such as shopping malls, scenic spots, commercial blocks, transportation hubs, communities, etc., injecting a "mobile consumption base station" for offline merchants, activating the consumption potential in fragmented scenarios, effectively alleviating the dilemma of shrinking physical business formats, and making stores, stations, in-vehicle terminals, etc. all become traffic conversion nodes.

[0003] Secondly, the subversive experience reshapes the consumption link. Different from the static display of traditional retail, this device is built with artificial intelligence and virtual reality technologies: customers can foresee the dressing effects through AI virtual fitting without having to try on clothes, and can let intelligent copywriting customize exclusive content for the scenario without having to conceive. Multi-modal interactions such as voice commands and gesture controls break down age and technology barriers, turning consumption from "passive selection" into "active creation". This strong sense of participation experience mode not only increases the customer's stay time and repurchase intention, but also stimulates users to spread spontaneously through the generation of social-media-style graphics and texts, forming a positive cycle of "experience - sharing - drainage".

[0004] Finally, it reconstructs the consumption decision-making logic. Relying on the "what you see is what you get" instant fulfillment ability of the device - instant printing of high-definition photos and extremely fast distribution of intelligent warehousing of clothing products, the virtual experience is directly converted into physical consumption, shortening the decision-making chain from interest to payment. Especially for the trust pain point of "buyer show and seller show" in the clothing industry, the wearing effects generated by AI's accurate matching of body size data, combined with real-time logistics tracking, balance impulsive consumption and rational decision-making. This business form that integrates emotional value and efficiency value is redefining the core competitiveness of offline consumption. Summary of the Invention

[0005] To achieve the above object, the present invention provides the following technical solution: A vending machine for replacing multi-modal AI text with clothing product images, which is an innovative retail device integrating cutting-edge multi-modal AI technology. By using AI to quickly and accurately match and generate corresponding clothing product images and display them on the screen, the appearance of the product is intuitively presented to help consumers efficiently lock in their favorite clothing.

[0006] Preferably, the specific steps of the vending machine and method for replacing multimodal AI text with clothing product images are as follows:

[0007] S1: Entrance selection: Trigger the device camera and code scanning function through five types of terminal entrances: traditional indoor and outdoor electronic display screens, virtual reality VR glasses and mobile phones, smart car screens, and smart robot display screens;

[0008] S2 interactive start: click the button to replace the product on the electronic screen and smart robot screen to directly display the operation on the screen, or scan the QR code with your mobile phone to enter the mobile page for operation; there is also scanning and perception of physical objects, through VR glasses, cars, mobile phones and smart robots as carrier paths to directly enter the interactive page, select clothing replacement or text generation category.

[0009] S3: Command input: Input requirements by voice or typing, such as "black slim dress size L" or "retro travel copywriting", and the device will simultaneously collect the user's body shape data, height, shoulder width, and the scene required by the copywriting you like to generate images and text;

[0010] S4: AI generation: Multimodal AI combined with virtual reality technology generates images of wearing virtual clothing in real time, embeds customized copy in the design template, and supports dynamic preview of zoom and rotation;

[0011] S5: Confirm payment: After being satisfied with the generated results, users can choose high-definition pictures and photos, support local printing or generate electronic files, meet the needs of multiple scenarios, and payment covers all scenarios such as scanning codes, in-vehicle contactless payment, VR eye tracking, etc., to achieve "one-click ordering". After the transaction, the clothing products are delivered by express or smart cabinets, and the pictures and photos are pushed to the terminal in real time. The whole process is centered on "what you see is what you get", breaking through the barriers between virtual and physical consumption and building an efficient and intelligent new retail closed loop;

[0012] S6: Fulfillment and delivery: Offline devices print photos instantly, and clothing products are shipped through smart warehousing and picked up by designated merchants, realizing a full-link closed loop of "what you see is what you get".

[0013] Preferably, the specific steps of the entry selection in S1 are as follows:

[0014] A: Traditional indoor and outdoor electronic display screens: electronic screen operation or mobile phone scanning code activation and category selection, click the electronic screen opening page to replace the product button, scan the code to enter the mobile page and on the electronic screen, select "clothing image replacement", "text content synthesis" category; input clothing size, color, number parameters through voice or typing, design the copy theme and style requirements of the layout; AI generates virtual clothing wearing effects in real time, customized copy, and place an order after satisfactory preview, and can choose to purchase the generated pictures and photos or the actual clothing and complete the payment;

[0015] B: VR glasses and mobile phones for virtual reality: Use VR glasses or mobile phone cameras to shoot real objects around you, such as your own image or scene images, automatically trigger the virtual reality page, and select the "clothing image replacement" and "text content synthesis" categories; enter clothing size, color, and number parameters by voice or typing, and specify the copy theme, style, and layout design requirements; AI combines real-time images to generate virtual clothing wearing effects, and intelligent copy is embedded in the layout. After previewing and adjusting to your satisfaction, place an order to purchase the generated pictures, photos, and clothing and complete the payment;

[0016] C: Smart car screen: The smart car camera senses the surrounding environment, triggers virtual reality technology to generate interactive pages on the in-vehicle electronic screen and projection interface, and selects the "clothing image replacement" and "text content synthesis" categories; inputs the size, color, and pattern parameters of the clothing by voice or typing in the in-vehicle system, specifies the copy theme and layout design style, and the system simultaneously collects the passenger's body data or scene images; AI generates the effect of passengers wearing virtual clothing images in real time, embeds customized copy in the landscape picture, and supports touch screen zoom preview; after confirmation, place an order through the in-vehicle payment system, and can choose to purchase high-definition pictures and photos, and physical clothing, to achieve instant shopping in the driving scene;

[0017] D: Smart robot display screen: Enter the interactive interface through the smart robot screen buttons or voice commands, and select the "Clothing Replacement" and "Text Synthesis" categories; input clothing parameters and text requirements through voice and screen typing, and the robot will simultaneously collect user body data or environmental images; AI generates virtual wearing effects or intelligent copywriting design, and supports 360° rotation preview; after confirmation, scan the code on the robot screen to pay, and you can choose to print pictures and photos or purchase physical clothing, and the system will automatically synchronize the order to warehousing and delivery.

[0018] Preferably, the interactive start in S2 refers to the interactive start link that constructs a full-scene access path with multiple carriers: the user clicks the "Replace Product" button on the opening page of the traditional indoor and outdoor electronic screen and the smart robot display screen, and can jump to the mobile phone interactive page after scanning the code, or directly operate on the electronic screen; if VR glasses, smart phones, smart cars and smart robots are used, the environmental images or user images are captured in real time through the device camera, and the virtual reality interactive page is directly triggered without manual scanning of the code. After entering the page, the user can intuitively select the "clothing replacement" or "text generation" core categories - the former supports the input of size, color, and version personalized parameters, and uses multimodal AI technology to accurately "wear" virtual clothing to the user's real-time image, and simultaneously supports 360° preview; the latter can intelligently generate adaptive text content based on the design theme, and embed it into the specified layout template. The entire interactive process is based on the concept of "device as the entrance", breaking the scene restrictions, and realizing seamless connection from the physical world to virtual generation, providing users with a zero-threshold, highly immersive intelligent graphic and text interactive experience.

[0019] Preferably, the command input in S3 refers to the command input link that supports multimodal interaction to meet diverse needs. Users can accurately convey specific requirements for clothing replacement or text generation through voice commands and typing input; for clothing needs, the device relies on cameras and sensors to collect user height, shoulder width, and body contour data in real time, and combines the input color, version, and size parameters to build a personalized virtual try-on model; if it is a text demand, the system will simultaneously analyze the scene image, and the user will specify the theme keywords to generate an adaptive copy style and content direction; the entire set of input logic is centered on "precise data matching needs", and through intelligent collection and semantic understanding, it can achieve efficient conversion from vague ideas to clear commands, providing precise underlying parameters for subsequent AI generation.

[0020] Preferably, the specific steps of generating AI in S4 are as follows:

[0021] Step 1: Multimodal data fusion modeling: Computer vision (CV) is used to extract user body data (height H, shoulder width S, waist circumference W) and clothing parameters (color C, version T, size M), and natural language processing (NLP) is used to parse text requirements to build a multi-dimensional input vector: input vector = [H, S, W, C, T, M, K]. The AI ​​model integrates image features and text semantics to generate a basic model for virtual clothing wear and a text semantic matrix.

[0022] Step 2: Real-time Rendering and Interaction Design: Based on a virtual reality (VR) rendering engine, the clothing model is fitted to the user's body size data to generate a wearing effect. The design template is automatically matched according to the semantic matrix of the copywriting, and the formula is expressed as: Virtual Image = VR Rendering (Clothing Model * Body Size Matrix) Copy Layout = Template Matching (K × D × F × P). The user triggers zoom (Scale), rotation (Rotate), and color change (Recolor) operations through touch screen and voice commands, and the system provides real-time feedback on the effects after parameter adjustment;

[0023] Step 3: Intelligent Optimization and Confirmation Output: An internal reinforcement learning (RL) algorithm is used to optimize the generation result according to the user's interaction behavior (duration of stay, number of adjustments). The calculation formula is: Optimization Weight = ∑(Interaction Behavior i × Preference Coefficient i) (i = 1, 2,..., n). After the user confirms satisfaction, the system outputs a high-definition graphic file (resolution ≥ 300 DPI), generates a production instruction for physical goods, and synchronously stores the user's preference data for subsequent recommendations.

[0024] Preferably, the confirmation of payment in S5 means that the user can directly purchase high-definition graphic photos, support local printing or generate electronic files to meet the needs of commemoration, sharing, or commercial design. The payment link supports full-scenario payment methods - payment is confirmed through scanning codes on the electronic screen, robot screen, in-vehicle system's contactless payment, and eye movement tracking of VR glasses, realizing "one-click ordering"; after the transaction is completed, the user can track the logistics information in real time. Clothing products are delivered by express or picked up at a designated smart cabinet, and the graphic photos are immediately generated and pushed to the user's terminal. The entire process takes "what you see is what you get" as the core, breaks down the barriers between virtual experience and physical consumption, and constructs an efficient and intelligent new retail closed-loop.

[0025] Preferably, the fulfillment and delivery in S6 means that the offline fulfillment link realizes an efficient closed-loop of "instant availability": after the user confirms the purchase of graphic photos, the high-definition printer installed on the offline device can immediately output physical photos, supporting multiple size options such as A4 and A3-sized books, magazines, and postcards to meet the needs of instant commemoration or scene display. For physical clothing products, the system automatically synchronizes the order to the intelligent warehousing center based on the user's virtual fitting parameters, matches the inventory and triggers flexible production through the AI sorting system. The products can be delivered to the user's designated address through the intelligent logistics network and pushed to the self-service cabinet of nearby partner merchants; the user can quickly pick up the physical goods with the pick-up code. The whole process relies on the Internet of Things technology to realize real-time synchronization of order, warehousing, and logistics data, truly realizing seamless connection of the entire link from "virtual fitting, design - online ordering - offline fulfillment", and making the consumption experience of "what you see is what you get" become a tangible real scenario.

[0026] The beneficial effects of the present invention are as follows:

[0027] 1. Under the dual challenges of e-commerce impact and changing consumption habits, traditional offline merchants are facing the dilemmas of a sharp decline in customer flow and rigid business formats. The multi-modal AI vending machine, in the mode of "full-scenario consumption entrance", injects liquidity vitality into offline commerce. Through multiple terminals such as electronic screens, VR glasses, and smart cars, it extends the shopping scenario to shopping malls, transportation hubs, communities, and even moving vehicles, breaking the geographical limitations of stores and building a "mobile consumption network". This innovation not only opens up fragmented traffic conversion channels for merchants, but also reshapes the logic of product display through AI technology. The static display of traditional shelves is upgraded to dynamic virtual try-on and intelligent copywriting generation, making offline media such as display windows and billboards become "interactive consumption touchpoints". Data shows that the offline conversion rate of merchants introducing this device has increased by more than 30%, effectively alleviating the pressure of the shrinkage of the physical business format, promoting the transformation of offline commerce from "rent dependence" to "technology-driven", and regaining the initiative in the consumer market.

[0028] 2. The "people looking for goods" mode of traditional retail has gradually lost its attraction due to the single experience, while the multi-modal AI vending machine builds an unprecedented sense of consumer participation with the dual engines of "technology + creativity": consumers can use voice or gesture commands to let the AI "wear" virtual clothing on their own images or generate exclusive literary copy for travel photos. This instant creation process from "idea to finished product" transforms shopping into an interesting "digital experience game". Especially for the personalized and interactive needs of Generation Z, the device supports 360° rotation to preview the clothing effect, real-time adjustment of the copy style, and even generation of printable souvenir photos, elevating the consumption behavior from "material transaction" to "emotional creation". Pilot data shows that the average interaction time of users exceeds 8 minutes, and the replay rates of parent-child families and young groups reach 65%. This strong-stickiness experience mode not only increases the single consumption amount, but also forms a fission effect of "experience is marketing" through sharing on social platforms, creating long-term brand value for merchants.

[0029] 3. In the present invention, the "delayed feeling" in consumer decision-making often leads to the loss of demand. However, the multi-modal AI vending machine solves this pain point with the logic of "technology shortening the link". After the user is satisfied with the virtual try-on clothing or the intelligently designed graphics and texts, they can immediately choose to print photos or place an order for physical goods. The high-definition printer built into the offline device outputs physical photos within 10 seconds, and the clothing goods achieve extremely fast fulfillment of "placing an order online and arriving in as fast as 1 hour" through the intelligent warehousing system. This zero-delay closed-loop of "experience - payment - delivery" precisely captures the instantaneous interest of consumers and directly converts the "heartbeat" during browsing into "action". Data shows that the impulse consumption ratio of device users is 40% higher than that of traditional retail, especially prominent in the fields of festival gifts and scene-based clothing. More importantly, the "exclusive content" generated by AI has scarce emotional value, further amplifying the consumption willingness and promoting the evolution of the business form from "function satisfaction" to "emotion awakening", redefining the core driving force of offline consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a flowchart of the method steps of the vending machine for replacing multi-modal AI text with clothing product images according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0032] As Figure 1 shown, the embodiment of the present invention provides a vending machine for replacing multi-modal AI text with clothing product images. This vending machine is a vending machine for replacing multi-modal AI text with clothing product images and is an innovative retail device integrating cutting-edge multi-modal AI technology. By using AI to quickly and accurately match and generate corresponding clothing product images and display them on the screen, it intuitively presents the appearance of the products and helps consumers efficiently lock in their favorite clothing.

[0033] Among them, the specific steps of the vending machine for replacing multi-modal AI text with clothing product images and the method are as follows:

[0034] S1: Entrance selection: Trigger the device camera and scanning function through five types of terminal entrances, namely traditional indoor and outdoor electronic displays, VR glasses and mobile phones in virtual reality, intelligent car screens, and intelligent robot displays.

[0035] S2 interactive start: click the button to replace the product on the electronic screen and smart robot screen to directly display the operation on the screen, or scan the QR code with your mobile phone to enter the mobile page for operation; there is also scanning and perception of physical objects, through VR glasses, cars, mobile phones and smart robots as carrier paths to directly enter the interactive page, select clothing replacement or text generation category.

[0036] S3: Command input: Input requirements by voice or typing, such as "black slim dress size L" or "retro travel copywriting", and the device will simultaneously collect the user's body shape data, height, shoulder width, and the scene required by the copywriting you like to generate images and text;

[0037] S4: AI generation: Multimodal AI combined with virtual reality technology generates images of wearing virtual clothing in real time, embeds customized copy in the design template, and supports dynamic preview of zoom and rotation;

[0038] S5: Confirm payment: After being satisfied with the generated results, users can choose high-definition pictures and photos, support local printing or generate electronic files, meet the needs of multiple scenarios, and payment covers all scenarios such as scanning codes, in-vehicle contactless payment, VR eye tracking, etc., to achieve "one-click ordering". After the transaction, the clothing products are delivered by express or smart cabinets, and the pictures and photos are pushed to the terminal in real time. The whole process is centered on "what you see is what you get", breaking through the barriers between virtual and physical consumption and building an efficient and intelligent new retail closed loop;

[0039] S6: Fulfillment and delivery: Offline devices print photos instantly, and clothing products are shipped through smart warehousing and picked up by designated merchants, realizing a full-link closed loop of "what you see is what you get".

[0040] Through five types of terminal entrance trigger functions, you can enter the interactive page to select categories, input requirements by voice or typing and collect data, generate content through multimodal AI, and after confirmation, choose to purchase pictures or physical objects and pay, and then instantly print offline or ship from smart warehousing to achieve a full-link closed loop.

[0041] The specific steps of the entry selection in S1 are as follows:

[0042] A: Traditional indoor and outdoor electronic display screens: electronic screen operation or mobile phone scanning code activation and category selection, click the electronic screen opening page to replace the product button, scan the code to enter the mobile page and select the "clothing image replacement" or "text content synthesis" category on the electronic screen; input clothing size, color, number parameters, or design layout copy theme and style requirements through voice or typing; AI generates virtual clothing wearing effects or customized copy in real time, and places an order after satisfactory preview, and can choose to purchase the generated pictures and photos or the actual clothing and complete the payment;

[0043] B: VR glasses and mobile phones for virtual reality: Use VR glasses or mobile phone cameras to shoot real objects around you, such as your own image or scene images, to automatically trigger the virtual reality page, and select the "clothing image replacement" or "text content synthesis" category; enter clothing size, color, number parameters, or specify copy theme, style and layout design requirements by voice or typing; AI combines real-time images to generate virtual clothing wearing effects or intelligent copy embedded in the layout. After previewing and adjusting to your satisfaction, place an order to purchase the generated pictures and photos or clothing and complete the payment;

[0044] C: Smart car screen: The smart car camera senses the surrounding environment (such as the image of passengers in the car or the scenery outside the car), triggers virtual reality technology to generate interactive pages on the car's electronic screen and projection interface, and selects the "clothing image replacement" or "text content synthesis" category; inputs the size, color, and pattern parameters of the clothing (clothing category) by voice (such as "recommended blue sports jacket size M") or typing in the car system, or specifies the copy theme (such as "self-driving tour circle of friends copy") and layout design style (text category), and the system simultaneously collects the passenger's body data or scene image; AI generates the image effect of passengers wearing virtual clothing in real time, or embeds customized copy in the landscape picture, and supports touch screen zoom preview; after confirmation, place an order through the car payment system, and choose to purchase high-definition pictures and photos (printed directly or sent to mobile phones) or physical clothing (delivered to a designated address), realizing instant shopping in the driving scene;

[0045] D: Smart robot display screen: Enter the interactive interface through the smart robot screen button or voice command (such as "open the dressing function"), and select the "clothing replacement" or "text synthesis" category; enter clothing parameters (size, color, version) or text requirements (copy theme, style) through voice (such as "I want a white loose shirt") or screen typing, and the robot will simultaneously collect user body data or environmental images; AI generates virtual wearing effects or smart copy design, and supports 360° rotation preview; after confirmation, scan the code on the robot screen to pay, and you can choose to print pictures and photos or purchase physical clothing, and the system will automatically synchronize the order to warehousing and delivery.

[0046] Electronic screen, scan the code to enter the mobile phone page, select the category and input the demand by voice / typing, and you can purchase pictures or physical objects after AI generates the effect; VR / mobile phone, take a picture of the physical object to trigger the page, input the parameters and AI combines the image to generate the effect, and place an order after adjustment. Smart car, the camera perceives the environment, and the in-vehicle interaction inputs the demand, AI generates the effect or text, and supports in-vehicle payment and multi-form delivery. Smart robot, button / voice trigger interaction, input the parameters and AI generates the effect, and the order is synchronized to the warehouse after scanning the code to pay.

[0047] Among them, the interactive start in S2 refers to the interactive start link that constructs a full-scene access path with multiple carriers: the user clicks the "Replace Product" button on the opening page on the traditional indoor and outdoor electronic screen and the smart robot display screen, and can jump to the mobile phone interactive page after scanning the code, or directly operate on the electronic screen; if VR glasses, smart phones, smart cars and smart robots are used, the environmental images or user images are captured in real time through the device camera, and the virtual reality interactive page is directly triggered without manual scanning. After entering the page, the user can intuitively select the "clothing replacement" or "text generation" core categories - the former supports the input of size, color, and version personalized parameters, and uses multimodal AI technology to accurately "wear" virtual clothing to the user's real-time image, and simultaneously supports 360° preview; the latter can intelligently generate adaptive text content based on design themes (such as posters, copywriting), and embed it into the specified layout template. The entire interactive process is based on the concept of "device as the entrance", breaking the scene restrictions, and realizing seamless connection from the physical world to virtual generation, providing users with a zero-threshold, highly immersive intelligent graphic and text interactive experience.

[0048] Traditional electronic screens require scanning codes to jump to the mobile phone side, while VR glasses, mobile phones, smart cars, smart robots and other devices use cameras to capture the environment or user images and directly trigger the virtual reality interaction page; after entering, users can choose "clothing replacement" or "text generation". The former uses AI to achieve virtual try-on and 360° preview, and the latter intelligently generates text based on the theme and embeds it into templates, achieving seamless interaction between the physical world and virtual generation.

[0049] Among them, the command input in S3 refers to the command input link that supports multimodal interaction to meet diverse needs. Users can use voice commands (such as "recommended gray hooded sweatshirt size M") or typing input (such as "business-style conference poster copy") to accurately convey specific requirements for clothing replacement or text generation; for clothing needs, the device relies on cameras and sensors to collect user height, shoulder width, body shape and other data in real time, and combines the input color, version (such as slim, loose), and size parameters to build a personalized virtual try-on model; if it is a text demand, the system will simultaneously analyze scene images (such as scenery, portraits) or user-specified theme keywords (such as "retro travel", "workplace wear") to generate an adaptive copy style and content direction; the entire set of input logic is centered on "precise data matching needs", and through intelligent collection and semantic understanding, it realizes efficient conversion from vague ideas to clear commands, providing precise underlying parameters for subsequent AI generation.

[0050] Users can convey their needs through voice (such as "find a pair of blue jeans size L") or typing (such as "generate camping theme poster copy"); for clothing, the device uses cameras and sensors to collect height, shoulder width, and body shape data in real time, and builds a virtual try-on model based on parameters such as color, version, and size; text-based needs analyze scene images or theme keywords (such as "minimalist style" and "spring picnic") to match the copy style and content direction. The entire logic is centered on "data-driven needs". Through intelligent perception and semantic analysis, it converts vague ideas into structured instructions, providing accurate data support for AI generation and achieving efficient interaction of "what you think is what you input, and what you input is what you produce".

[0051] The specific steps of AI generation in S4 are as follows:

[0052] Step 1: Multimodal data fusion modeling: Computer vision (CV) is used to extract user body data (height H, shoulder width S, waist circumference W) and clothing parameters (color C, version T, size M), and natural language processing (NLP) is used to parse text requirements (such as the keyword K of "retro-style copywriting") to construct a multi-dimensional input vector: input vector = [H, S, W, C, T, M, K]. The AI ​​model (such as CLIP+StyleGAN3) fuses image features and text semantics to generate a basic model of virtual clothing or a copywriting semantic matrix.

[0053] Step 2: Real-time rendering and interactive design: Based on the virtual reality (VR) rendering engine, the clothing model is fitted to the user's body shape data to generate the wearing effect, or the design template (such as poster size D, font F, color matching P) is automatically matched according to the copy semantic matrix. The formula is expressed as: virtual image = VR rendering (clothing model * body shape matrix copy layout = template matching (K×D×F×P). The user triggers the scaling (Scale), rotation (Rotate), and color change (Recolor) operations through the touch screen and voice commands, and the system provides real-time feedback on the effect of parameter adjustment;

[0054] Step 3: Intelligent optimization and confirmation output: The built-in reinforcement learning (RL) algorithm optimizes the generated results according to the user's interactive behavior (stay duration, number of adjustments). The calculation formula is: optimization weight = ∑ (interaction behavior i × preference coefficient i) (i = 1, 2, ..., n). After the user confirms that he is satisfied, the system outputs high-definition graphic files (resolution ≥ 300DPI) or generates physical product production instructions, and simultaneously stores user preference data for subsequent recommendations.

[0055] Multimodal Data Fusion Modeling: Extract user body shape data and clothing parameters through computer vision, parse text requirements through natural language processing, and construct a multi-dimensional input vector; the AI model fuses image and text semantics to generate a basic model for virtual clothing wear or a semantic matrix of copywriting; Step 2, Real-time Rendering and Interaction Design: Use a VR rendering engine to generate wearing effects or match design templates, and users can operate through touch screens and voice, and the system provides real-time feedback to adjust the effects; Step 3, Intelligent Optimization and Confirmation Output: Optimize the generation results based on reinforcement learning algorithms, output high-definition graphics and text or production instructions after user confirmation, and store preference data to achieve personalized recommendations.

[0056] Among them, the confirmed payment in S5 means that users can directly purchase high-definition graphic photos, support local printing or generating electronic files (JPG / PNG format, with the option of attaching design source files), meeting the needs of commemoration, sharing or commercial design. For example, the dressing effect diagram can be printed as a physical postcard, and the copy layout can be used for making social media posters. If clothing items are selected, the system automatically synchronizes parameters such as size and color of virtual try-on to the supply chain, triggering intelligent warehouse sorting or flexible production processes; the payment link supports full-scenario payment methods - scanning codes (WeChat / Alipay) through electronic screens and robot screens, contactless payment through in-vehicle systems (binding in-vehicle accounts), and confirmation of payment through eye movement tracking of VR glasses, etc., to achieve "one-click ordering"; after the transaction is completed, users can track logistics information in real time. Clothing products are delivered by express or picked up at designated smart cabinets, while graphic photos are generated and pushed to the user terminal immediately. The whole process takes "what you see is what you get" as the core, breaks through the barriers between virtual experience and physical consumption, and constructs an efficient and intelligent new retail closed-loop.

[0057] Users can purchase high-definition graphic photos, support local printing or generating electronic files, meeting diverse needs; if clothing items are selected, the system automatically synchronizes parameters to the supply chain and initiates intelligent warehouse processes. Payment methods cover code scanning, in-vehicle contactless payment, and VR eye movement tracking. After the transaction, users can track logistics, and products are delivered by express or picked up, while graphics are pushed immediately, achieving an efficient closed-loop from virtual experience to physical consumption.

[0058] Among them, the fulfillment and delivery in S6 refers to the efficient closed-loop of "instant availability upon demand" in the offline fulfillment link: after the user confirms the purchase of graphic photos, the high-definition printers installed on offline devices (such as electronic screen terminals, intelligent robots) can immediately output physical photos, supporting multi-size options for books, magazines, postcards in A4 and A3 sizes, meeting the needs of instant commemoration or scenario-based display. For physical clothing products, the system automatically synchronizes the order to the intelligent warehousing center based on the user's virtual fitting parameters, matches the inventory through the AI sorting system and triggers flexible production, and the products can be delivered to the user's designated address through the intelligent logistics network or pushed to the self-service lockers of nearby partner merchants (such as shopping mall convenience stores, community stations); the user can quickly pick up the physical objects with the pick-up code, and the whole process relies on the Internet of Things technology to realize the real-time synchronization of order, warehousing, and logistics data, truly realizing the seamless connection of the entire link from "virtual fitting, design - online ordering - offline fulfillment", and making the consumption experience of "seeing is believing" become a tangible real scenario.

[0059] After the user selects graphic photos, the high-definition printers of offline devices can immediately output physical photos, supporting multi-size options to meet instant needs; for physical clothing, the order is synchronized to the intelligent warehouse according to the virtual fitting parameters, and after being sorted by AI to match the inventory or through flexible production, it is delivered to the designated address or nearby self-service locker through intelligent logistics; the user picks up the item with the code, and the whole process relies on the Internet of Things to achieve real-time data synchronization, connecting the entire "virtual - online - offline" link, making the consumption experience truly tangible.

[0060] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprises", "comprising", or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device.

[0061] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A vending machine for replacing text of multi-modal AI with images of clothing products, characterized in that: This vending machine is a vending machine that replaces multimodal AI text and clothing product images. It is an innovative retail device that integrates cutting-edge multimodal AI technology. It uses AI to quickly and accurately match and generate corresponding clothing product images, display them on the screen, and intuitively present the appearance of the products, helping consumers to efficiently lock in their favorite clothing.

2. Vending machine and method for replacing text of multimodal AI with clothing product images, characterized in that: The specific steps of the vending machine and method for replacing multimodal AI text and clothing product images are as follows: S1: Entrance selection: Trigger the device camera and code scanning function through five types of terminal entrances: traditional indoor and outdoor electronic display screens, virtual reality VR glasses and mobile phones, smart car screens, and smart robot display screens; S2 interactive start: click the button to replace the product on the electronic screen and smart robot screen to directly display the operation on the screen, or scan the QR code with your mobile phone to enter the mobile page for operation; there is also scanning and perception of physical objects, through VR glasses, cars, mobile phones and smart robots as carrier paths to directly enter the interactive page, select clothing replacement or text generation category. S3: Command input: Input requirements by voice or typing, such as "black slim dress size L" or "retro travel copywriting", and the device will simultaneously collect the user's body shape data, height, shoulder width, and the scene required by the copywriting you like to generate images and text; S4: AI generation: Multimodal AI combined with virtual reality technology generates images of wearing virtual clothing in real time, embeds customized copy in the design template, and supports zooming, rotating dynamic preview; S5: Confirm payment: After being satisfied with the generated results, users can purchase high-definition pictures and photos, support local printing or generate electronic files to meet the needs of multiple scenarios. Payment covers all scenarios such as scanning codes, in-vehicle contactless payment, VR eye tracking, etc., to achieve "one-click ordering". After the transaction, the clothing products are delivered by express or smart cabinets, and the pictures and photos are pushed to the terminal in real time. The whole process is centered on "what you see is what you get", breaking down the barriers between virtual and physical consumption and building an efficient and intelligent new retail closed loop. S6: Fulfillment and delivery: Offline devices print photos instantly, and clothing products are shipped through smart warehousing and picked up by designated merchants, realizing a full-link closed loop of "what you see is what you get".

3. The vending machine and method for replacing text with clothing product images in multi-modal AI according to claim 2, characterized in that: The specific steps of the entry selection in S1 are as follows: A: Traditional indoor and outdoor electronic display screens: electronic screen operation or mobile phone scanning code activation and category selection, click the electronic screen opening page to replace the product button, scan the code to enter the mobile page and on the electronic screen, select "clothing image replacement", "text content synthesis" category; input clothing size, color, number parameters through voice or typing, design the copy theme and style requirements of the layout; AI generates virtual clothing wearing effects in real time, customized copy, and place an order after satisfactory preview, and can choose to purchase the generated pictures and photos or the actual clothing and complete the payment; B: VR glasses and mobile phones for virtual reality: Use VR glasses or mobile phone cameras to shoot real objects around you, such as your own image or scene images, to automatically trigger the virtual reality page, select "clothing image replacement" and "text content synthesis" categories; input clothing size, color, number parameters by voice or typing, and specify the copy theme, style and layout design requirements; AI combines real-time images to generate virtual clothing wearing effects, and intelligent copy is embedded in the layout. After previewing and adjusting to your satisfaction, place an order to purchase the generated pictures, photos, and clothing and complete the payment; C: Smart car screen: Use the smart car camera to sense the surrounding environment, trigger virtual reality technology to generate interactive pages on the car electronic screen and projection interface, and select the "clothing image replacement" and "text content synthesis" categories; Use voice or typing on the in-vehicle system to input clothing size, color, and pattern parameters, specify the copy theme and layout design style, and the system will simultaneously collect passenger body data or scene images; AI generates real-time image effects of passengers wearing virtual clothing, embeds customized copy in the landscape image, and supports touch-screen zoom preview; after confirmation, place an order through the in-vehicle payment system, and choose to purchase high-definition pictures and photos, or physical clothing, to achieve instant shopping in the driving scene; D: Smart robot display screen: Enter the interactive interface through the smart robot screen buttons or voice commands, select the "clothing replacement" and "text synthesis" categories; input clothing parameters and text requirements through voice and screen typing, and the robot will simultaneously collect user body data or environmental images; AI generates virtual wearing effects or intelligent copywriting design, and supports 360° rotation preview; after confirmation, scan the code on the robot screen to pay, and you can choose to print pictures and photos or purchase physical clothing, and the system will automatically synchronize the order to warehousing and delivery.

4. The vending machine and method for replacing text with clothing product images in multi-modal AI according to claim 2, characterized in that: The interactive start in S2 refers to the interactive start link that builds a full-scene access path with multiple carriers. Users click the "Replace Product" button on the opening page of traditional indoor and outdoor electronic screens and smart robot display screens, and can jump to the mobile phone interactive page after scanning the code, or they can operate directly on the electronic screen; if VR glasses, smart phones, smart cars and smart robots are used, the environmental images or user images are captured in real time through the device camera, and the virtual reality interactive page is directly triggered without manual scanning of the code. After entering the page, users can intuitively select the core categories of "clothing replacement" or "text generation" - the former supports the input of personalized parameters such as size, color, and style, and uses multimodal AI technology to accurately "wear" virtual clothing to the user's real-time image, and simultaneously supports 360° preview; the latter can intelligently generate adaptive text content based on the design theme, and embed it into the specified layout template. The entire interactive process is based on the concept of "device as the entrance", breaking the scene restrictions, and realizing seamless connection from the physical world to virtual generation, providing users with a zero-threshold, highly immersive intelligent graphic and text interactive experience.

5. The vending machine and method for replacing text with clothing product images in multimodal AI according to claim 2, characterized in that: The command input in S3 refers to the command input link that supports multimodal interaction to meet diverse needs. Users can use voice commands and typing input to accurately convey specific requirements for clothing replacement or text generation. For clothing needs, the device relies on cameras and sensors to collect user height, shoulder width, and body contour data in real time, and combines the input color, version, and size parameters to build a personalized virtual try-on model. If it is a text requirement, the system will simultaneously analyze the scene image, and the user will specify the theme keywords to generate an adaptive copywriting style and content direction. The entire set of input logic is centered on "precise data matching needs". Through intelligent collection and semantic understanding, it can achieve efficient conversion from vague ideas to clear commands, providing precise underlying parameters for subsequent AI generation.

6. The vending machine and method for replacing text with clothing product images in multimodal AI according to claim 2, characterized in that: The specific steps of AI generation in S4 are as follows: Step 1: Multimodal data fusion modeling: Computer vision (CV) is used to extract user body data (height H, shoulder width S, waist circumference W) and clothing parameters (color C, version T, size M), and natural language processing (NLP) is used to parse text requirements to build a multi-dimensional input vector: input vector = [H, S, W, C, T, M, K]. The AI ​​model integrates image features and text semantics to generate a basic model for virtual clothing wear and a text semantic matrix. Step 2: Real-time rendering and interactive design: Based on the virtual reality (VR) rendering engine, the clothing model is fitted to the user's body shape data to generate the wearing effect, and the design template is automatically matched according to the copy semantic matrix. The formula is expressed as: virtual image = VR rendering (clothing model * body shape matrix copy layout = template matching (K×D×F×P). The user triggers the scaling (Scale), rotation (Rotate), and color change (Recolor) operations through the touch screen and voice commands, and the system provides real-time feedback on the effect of parameter adjustment; Step 3: Intelligent optimization and confirmation output: The built-in reinforcement learning (RL) algorithm optimizes the generated results according to the user's interactive behavior (stay duration, number of adjustments). The calculation formula is: optimization weight = ∑ (interaction behavior i × preference coefficient i) (i = 1, 2, ..., n). After the user confirms that he is satisfied, the system outputs high-definition graphic files (resolution ≥ 300DPI), generates physical product production instructions, and simultaneously stores user preference data for subsequent recommendations.

7. The vending machine and method for replacing text with clothing product images in multi-modal AI according to claim 2, characterized in that: The payment confirmation in S5 means that users can directly purchase high-definition pictures and photos, support local printing or generate electronic files to meet commemorative, sharing or commercial design needs, and the payment process supports full-scenario payment methods - through electronic screens, robot screen scanning, vehicle-mounted system contactless payment, VR glasses eye tracking to confirm payment, and realize "one-click ordering"; after the transaction is completed, users can track logistics information in real time, clothing products are delivered by express or picked up at designated smart cabinets, and pictures and photos are generated instantly and pushed to the user terminal. The whole process is centered on "what you see is what you get", breaking down the barriers between virtual experience and physical consumption, and building an efficient and intelligent new retail closed loop.

8. The vending machine and method for multi-modal AI text and clothing product image replacement according to claim 2, characterized in that: The fulfillment and delivery in S6 refers to the efficient closed-loop of "instant availability upon demand" in the offline fulfillment process: after the user confirms the purchase of picture and text photos, the high-definition printer installed on the offline device can immediately output physical photos, supporting multiple size options such as A4 and A3-sized books, newspapers, magazines, and postcards, meeting the needs of instant commemoration or scenario-based display. For physical clothing products, the system automatically synchronizes the order to the intelligent warehousing center based on the user's virtual fitting parameters, triggers flexible production by matching inventory through the AI sorting system, and the products can be delivered to the user's designated address through the intelligent logistics network and pushed to the self-service lockers of nearby partner merchants; the user can quickly pick up the physical objects with the pick-up code. The whole process relies on the Internet of Things technology to achieve real-time synchronization of order, warehousing, and logistics data, truly realizing the seamless connection of the entire link from "virtual fitting, design - online order - offline fulfillment", and turning the consumption experience of "seeing is believing" into a tangible real scenario.

Citation Information

Patent Citations

  • Intelligent fitting method for mobile terminal, mobile terminal and intelligent fitting system

    CN106326511A

  • Intelligent cloud fitting system for realizing clothes sharing

    CN107220886A

  • Online and offline combined virtual fitting method and system capable of carrying out network transaction

    CN110706076A

  • Intelligent clothing vending machine system

    CN111951061A

  • Commodity shopping system and method based on combination of AR and AI

    CN118052613A