Intelligent retail robot multi-mode interaction system and method based on AI large model

By using an AI-based big data model-driven intelligent retail robot multimodal interaction system, which combines touchscreens, gestures, and facial recognition to dynamically adjust the interface, the system solves the problems of single interaction methods and immature multimodal integration in existing technologies, and achieves personalized recommendations and an efficient shopping experience.

CN120848718APending Publication Date: 2025-10-28SHANDONG TIANDI YITONG TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510704358.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing smart retail robots have limited interaction methods and immature multimodal information integration, resulting in inaccurate product recommendations to consumers and reduced interactive experience and purchasing efficiency.

Method used

The system employs a multimodal interaction system for intelligent retail robots based on a large AI model. It combines touchscreen operation, gesture recognition, and facial recognition to provide personalized product recommendations by analyzing consumers' historical purchase records and behavioral patterns. It dynamically adjusts the touchscreen interface and uses data fusion processing modules and computer vision algorithms to identify consumer intentions and behaviors.

Benefits of technology

It enhances the naturalness and convenience of consumer interaction with retail robots, provides a personalized shopping experience, improves shopping efficiency and sales conversion rates, and enables merchants to develop precise marketing strategies and optimize product recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848718A_ABST
    Figure CN120848718A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent retail robot multi-modal interaction system and method based on an AI large model, and belongs to the technical field of intelligent retail, and the system comprises a data statistics module which is used for carrying out the statistics of the historical sales data of each commodity in a retail robot, and analyzing the sales characteristics of each commodity based on the historical sales data; the image information acquisition module responds to the historical sales data and is used for acquiring face images of consumers corresponding to the commodities when the commodities are purchased according to the historical sales data and constructing a database, and the database comprises consumption preferences and purchase habits of the corresponding consumers. By analyzing the historical purchase record and behavior mode of the consumer, the system can provide personalized commodity recommendation for the consumer, and by combining multiple interaction modes such as touch screen operation, gesture recognition and face recognition, interaction between the consumer and the retail robot is more natural, convenient and efficient, and more personalized interaction experience is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent retail technology, and in particular to a multimodal interaction system and method for intelligent retail robots based on AI big data models. Background Technology

[0002] With the development of artificial intelligence technology, intelligent retail robots are being used more and more widely in shopping malls. However, the existing intelligent retail robots have relatively simple interaction methods, mainly relying on voice or simple touch screen operation, which cannot meet the increasingly diverse and complex interaction needs of consumers.

[0003] Regarding this research, application CN202111413269.9 provides an interactive method and system for smart vending machines. This technical solution includes acquiring purchase records of target groups based on different target scenarios, and creating corresponding product sales lists based on these records. The smart vending machine receives user commands in real time. Voice commands are input into a pre-set voice extraction model to obtain key voice information. The key voice information is then matched with product names in a pre-set product name library to obtain the target product name. This technical solution achieves the goal of shopping through voice commands, simplifying the manual purchase process while avoiding operational malfunctions caused by human intervention, thus providing users with a better shopping experience.

[0004] Another application, CN202110570676.4, provides a vending machine interaction method and system based on reliable gesture recognition. This technical solution includes: Step 1, setting different gesture actions to represent different control commands for the vending machine; Step 2, obtaining the judgment result of the gesture action and converting the gesture recognition result into a corresponding control command based on the current command state; Step 3, sending the control command to the vending machine control device to achieve contactless control of the vending machine. This technical solution, based on reliable gesture recognition using graph convolutional networks and hand trajectory prediction, serves as a human-computer interaction system for vending machines. It can solve the safety and hygiene problems caused by current contact-based vending machine interactions and improve the accuracy of algorithms based on voice or traditional gesture recognition, making vending machines more intelligent.

[0005] However, the aforementioned technical solutions are not yet mature enough in terms of multimodal information fusion, which may result in the fusion effect not meeting expectations. For example, the fusion of voice or touchscreen operation and image recognition results may not accurately reflect the consumer's intent, or in complex scenarios, retail robots may not be able to understand the consumer's multimodal commands well, thereby affecting the accuracy of product recommendations to consumers and reducing the interactive experience and purchase efficiency. Summary of the Invention

[0006] In view of the problems existing in the field of smart retail technology, the present invention is proposed.

[0007] Therefore, one of the objectives of this invention is to provide a multimodal interaction system and method for intelligent retail robots based on AI large models. By analyzing consumers' historical purchase records and behavioral patterns, the system can provide consumers with personalized product recommendations. Furthermore, by combining multiple interaction methods such as touch screen operation, gesture recognition, and facial recognition, the interaction between consumers and retail robots becomes more natural, convenient, and efficient, providing a more personalized interactive experience.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution: On the one hand, this invention provides a multimodal interaction system for intelligent retail robots based on AI large-scale models, including: The data statistics module is used to collect historical sales data for each product within the retail robot and analyze the sales characteristics of each product based on the historical sales data. The image information acquisition module responds to the historical sales data and is used to acquire the facial images of consumers corresponding to each product at the time of purchase based on the historical sales data, and construct a database, which includes the consumption preferences and purchasing habits of the corresponding consumers. The data fusion processing module is used to acquire the product information of each remaining product in the retail robot after each product is sold out, and to adjust the graphical interface on the touch screen of the retail robot based on the product information of each remaining product; the data fusion processing module includes an acquisition unit, a sensing unit, an image acquisition unit, and a data retrieval unit. The acquisition unit is used to acquire the layout of the product name buttons for each remaining product on the graphical interface, and acquire the position of each product name button on the graphical interface according to the layout. The sensing unit is used to sense the degree of proximity and dwell time of consumers to the retail robot in order to determine the consumers' intentions and behaviors. The image acquisition unit responds to the sensing unit and is used to acquire the consumer's facial image based on the determined consumer's intention and behavior; wherein, when it is determined that the consumer will touch the touch screen on the retail robot, the consumer's facial image is acquired; otherwise, it is not acquired. The data retrieval unit responds to the image acquisition unit by retrieving the historical purchase records of the consumer corresponding to the face image at the retail robot in the database, and predicting the consumer's current purchase based on the historical purchase records.

[0009] In a preferred embodiment of the present invention, based on the prediction of the consumer's current purchase, the product name buttons corresponding to the consumer's historical purchase records are displayed on the touch screen, including displaying the products purchased by the consumer in the last 3 to 5 times in the upper left or upper right corner of the touch screen; and displaying relevant information about the products corresponding to the product name buttons, including product details, promotional activities, and inventory. In a preferred embodiment of the present invention, the data fusion processing module further includes a gesture acquisition unit. The gesture acquisition unit is used to acquire images of the consumer's gestures and to analyze and process the images using computer vision algorithms to identify the shape, position, and trajectory of the consumer's gestures. Simultaneously, based on the trajectory, it determines whether the consumer will click the product name button on the upper left or upper right of the touchscreen. If it is determined that the consumer will not click the product name button on the upper left or upper right of the touchscreen, the product name button on the touchscreen is adjusted, including adjusting it to a product name button corresponding to a product purchased by the consumer within the past month.

[0010] In a preferred embodiment of the present invention, the data statistics module analyzes the sales characteristics of each product based on the historical sales data, including analyzing the data in a time series manner, arranging the historical sales data in chronological order, and analyzing the trend and seasonal characteristics of the sales data over time; the time series analysis includes moving average method, exponential smoothing method, and ARIMA model. It also includes analysis using regression analysis, which involves establishing a regression model to analyze the impact of independent variables on the dependent variable in historical sales data. The independent variables include price, promotional activities, and season, while the dependent variables include sales revenue and sales volume. The regression analysis includes linear regression and multinomial regression.

[0011] In a preferred embodiment of the present invention: after replacing the product name buttons with those corresponding to products purchased by the consumer within the past month, the number of product name buttons clicked by the consumer is counted, and the time taken by the consumer to click different product name buttons is calculated from the count. Based on the length of time taken, the clicked product name buttons are sorted, and the product finally purchased by the consumer is obtained from the sorted product name buttons. The purchased product is marked as a reference product. When it is determined in a future time period that the consumer will touch the touch screen on the retail robot, the product name button corresponding to the reference product is displayed on the touch screen first.

[0012] In a preferred embodiment of the present invention, the purchasing characteristics of the consumer who has purchased the reference product in the same historical purchasing history as other consumers are calculated based on the reference product, and the calculation is performed according to the purchase frequency, as shown below: ; In the formula, Indicates purchase frequency. Indicates the number of purchases. Indicates the length of time, which is in months and / or years; This formula represents the number of times a consumer purchases a specific product within a given time period.

[0013] In a preferred embodiment of the present invention, the following method is used: calculating the purchasing characteristics of the consumer who also purchased the reference product in other past purchases, based on the reference product, further includes calculating the characteristics based on the purchase conversion rate, as shown below: ; In the formula, Indicates purchase conversion rate. Indicates the actual number of purchases. Indicates the number of times the product was viewed; This formula is used to calculate the percentage of consumers who actually purchase a particular product after browsing it. Alternatively, it can be calculated through basket analysis, as shown below: ; In the formula, Indicates a specific product, Indicates other goods, Indicates purchasing at the same time and Number of transactions Indicates the total number of transactions; This formula calculates the probability that a consumer will purchase other goods at the same time as a specific product.

[0014] In a preferred embodiment of the present invention, the consumption preferences and / or purchasing habits of other consumers relative to the consumer are obtained based on the calculated purchase characteristics; based on the consumption preferences and / or purchasing habits, products purchased by other consumers in past consumption that are different from those of the consumer are obtained; these products are marked as differentiated products; and when it is determined in a future time period that the consumer will touch the touch screen on the retail robot, product name buttons corresponding to the reference product and the differentiated product are preferentially displayed on the touch screen.

[0015] On the other hand, the present invention provides a method for applying to the AI-based large-scale intelligent retail robot multimodal interaction system as described above, comprising the following steps: Collect historical sales data for each product within the retail robot, and analyze the sales characteristics of each product based on the historical sales data; Based on historical sales data, obtain facial images of consumers corresponding to each product at the time of purchase, and construct a database, which includes the consumption preferences and purchasing habits of the corresponding consumers; After each item is sold out, the remaining item information in the retail robot is obtained, and the graphical interface on the retail robot's touchscreen is adjusted based on the remaining item information. Obtain the layout of the product name buttons for each remaining product on the graphical interface, and obtain the position of each product name button on the graphical interface based on the layout; The robot senses how close consumers are to it and how long they stay, in order to determine their intentions and behaviors. The system collects facial images of consumers based on their perceived intentions and behaviors. Specifically, if it is determined that a consumer will touch the touchscreen on the retail robot, then the consumer's facial image is collected; otherwise, it is not collected. The database is used to retrieve the historical purchase records of the consumer corresponding to the facial image at the retail robot, and the current purchase of the consumer is predicted based on the historical purchase records.

[0016] Beneficial effects: 1. By analyzing consumers' historical purchase records and behavioral patterns, the system can provide personalized product recommendations. For example, displaying recently purchased items in a prominent position on the touchscreen reduces the time consumers spend searching for products and improves shopping efficiency; 2. By combining multiple interaction methods such as touch screen operation, gesture recognition, and facial recognition, the interaction between consumers and retail robots becomes more natural, convenient, and efficient. For example, the system can automatically adjust the product layout on the touch screen based on the consumer's intentions and behaviors, providing a more personalized interactive experience. 3. By analyzing historical sales data, the system can predict product sales trends and demand, and can adjust the product layout on the touchscreen in real time. Based on consumers' purchasing habits and preferences, it can dynamically display the products most likely to be purchased, thereby improving sales conversion rates. 4. By analyzing purchase frequency, purchase conversion rate, and purchase basket, the system can gain a deep understanding of consumers' purchasing behavior and preferences, providing merchants with precise marketing strategies. For example, merchants can develop personalized promotional activities based on consumers' purchase frequency and conversion rate, helping them better understand their target customer groups, optimize product recommendations, and improve consumer loyalty. 5. Leveraging the powerful processing capabilities of AI large-scale models, the system can perform complex data analysis and prediction. It can adaptively adjust interaction strategies and product recommendations based on consumers' real-time behavior and historical data, providing a more accurate and intelligent interactive experience. For example, through time series analysis and regression analysis, the system can predict product sales trends and provide decision support for merchants. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a schematic diagram of the modular structure of the AI-based large-scale intelligent retail robot multimodal interaction system according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the method flow according to an embodiment of the present invention; The diagram is labeled as follows: 110 - Data statistics module; 120 - Image information acquisition module; 130 - Data fusion processing module; 1301 - Acquisition unit; 1302 - Sensing unit; 1303 - Image acquisition unit; 1304 - Data retrieval unit; 1305 - Gesture and motion acquisition unit. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.

[0019] Because existing technologies are not yet mature enough in multimodal information fusion, the fusion effect may not be as expected, thus affecting the accuracy of product recommendations to consumers and reducing the interactive experience and purchase efficiency.

[0020] Based on this, the present invention proposes a multimodal interaction system and method for intelligent retail robots based on AI large models. By analyzing consumers' historical purchase records and behavioral patterns, the system can provide consumers with personalized product recommendations. Furthermore, by combining multiple interaction methods such as touch screen operation, gesture recognition, and facial recognition, the interaction between consumers and retail robots becomes more natural, convenient, and efficient, providing a more personalized interactive experience.

[0021] The present solution will be further described in detail below through embodiments and in conjunction with the accompanying drawings.

[0022] Reference Figures 1 to 2 This is one embodiment of the present invention, which provides a multimodal interaction system for intelligent retail robots based on an AI large model, including: The data statistics module 110 is used to collect historical sales data of each product in the retail robot and analyze the sales characteristics of each product based on the historical sales data. In this embodiment, analyzing the sales characteristics of each product based on historical sales data is a very important part of the retail business. It can help merchants better understand market demand, optimize inventory management, and formulate marketing strategies. Image information acquisition module 120 responds to historical sales data and is used to acquire the facial images of consumers corresponding to each product at the time of purchase based on the historical sales data, and build a database, which includes the consumption preferences and purchasing habits of the corresponding consumers. In this embodiment, the image information acquisition module acquires consumers' facial images based on historical sales data and builds a database containing consumption preferences and purchasing habits. This helps merchants to fully understand consumer characteristics, achieve precise marketing, and improve marketing effectiveness and consumer loyalty. The data fusion processing module 130 is used to acquire the product information of each remaining product in the retail robot after each product is sold out, and adjust the graphical interface on the touch screen of the retail robot based on the product information of each remaining product; the data fusion processing module 130 includes an acquisition unit 1301, a sensing unit 1302, an image acquisition unit 1303 and a data retrieval unit 1304. In this embodiment, consumers can browse product information, filter products, and place orders through touch operations; This dynamic adjustment ensures that the interface always displays product information that best reflects current sales conditions and consumer needs, thereby improving sales efficiency and user experience. The acquisition unit 1301 is used to acquire the layout of the product name buttons for each remaining product on the graphical interface, and to acquire the position of each product name button on the graphical interface according to the layout. The sensing unit 1302 is used to sense the proximity and dwell time of consumers to the retail robot in order to determine the consumer's intentions and behaviors; In this embodiment, the sensing unit includes a distance sensor, an infrared sensor, etc. Image acquisition unit 1303 responds to perception unit, used to acquire facial images of consumers based on the determined intentions and behaviors of consumers; wherein, when it is determined that the consumer will touch the touch screen on the retail robot, the consumer's facial image is acquired; otherwise, it is not acquired. The data retrieval unit 1304 responds to the image acquisition unit and is used to retrieve the historical purchase records of consumers at the retail robot corresponding to the facial images in the database, and to predict the current purchase of consumers based on the historical purchase records. In this embodiment, the collaborative work of the acquisition unit, perception unit, image acquisition unit and data retrieval unit enables multimodal perception and interaction of consumer behavior. The system can predict the consumer's purchase intention based on behavioral characteristics such as the consumer's proximity and dwell time, as well as the facial image recognition results, and provide more intelligent and personalized services, thereby enhancing the interactive experience between consumers and retail robots. Based on the above, according to the prediction results of the consumer's current purchase, the product name buttons corresponding to the consumer's historical purchase records are displayed on the touch screen, including displaying the products purchased by the consumer in the last 3 to 5 times in the upper left or upper right corner of the touch screen; and displaying relevant information about the products corresponding to the product name buttons, including product details, promotional activities and inventory; In this embodiment, this personalized display method can quickly guide consumers to find products they may be interested in, saving shopping time and improving shopping efficiency; By showcasing products recently purchased by consumers and related information, the company increases consumers' attention to and willingness to buy these products, thereby improving sales conversion rates and generating more sales for merchants. When consumers discover that retail robots can remember their purchase history and provide personalized recommendations, they will feel cared for and valued, thereby enhancing their stickiness and loyalty to the merchant and promoting repeat purchases. The data fusion processing module 130 also includes a gesture acquisition unit 1305. The gesture acquisition unit is used to acquire images of consumers' gestures and to analyze and process the images using computer vision algorithms. It identifies the shape, position, and trajectory of the consumer's gestures and determines whether the consumer will click the product name button in the upper left or upper right corner of the touch screen based on the trajectory. If it is determined that the consumer will not click the product name button in the upper left or upper right corner of the touch screen, the product name button on the touch screen is adjusted, including adjusting it to the product name button corresponding to the product purchased by the consumer within the last month. In this embodiment, in one feasible implementation, a camera is installed near the retail robot or touchscreen to capture images of the consumer's hand gestures. This includes capturing thermal radiation images of fingers using an infrared thermal imaging camera, and identifying the movement trajectory of gestures by analyzing the temperature distribution and changes in the thermal images; It also includes using electromagnetic waves of a specific frequency, by setting up transmitters and receivers around the touchscreen, emitting electromagnetic waves and detecting their reflected signals, and identifying the movement trajectory of gestures based on the characteristics of the reflected signals; Computer vision algorithms include image segmentation algorithms, which include threshold-based segmentation, region-based segmentation, edge-based segmentation, and graph cut-based segmentation, etc. Based on the gesture trajectory, the system determines the consumer's click intention. If it determines that the consumer will not click the currently displayed product name button, it automatically adjusts the product button layout on the touch screen to better meet the consumer's actual needs. Gesture recognition, as a natural and intuitive interaction method, allows consumers to interact with retail robots in a more natural way, without complicated operation steps, thus lowering the interaction threshold and improving the naturalness and smoothness of the interaction. At the same time, the interface layout is dynamically adjusted based on consumers' real-time gestures, so that the interface can better adapt to the personalized needs and behavioral habits of different consumers, thereby improving the adaptability of the interface and the user experience. Furthermore, in the data statistics module, the sales characteristics of each product are analyzed based on historical sales data, including analysis in the form of time series analysis, which arranges historical sales data in chronological order and analyzes the trend and seasonality of sales data over time; time series analysis includes moving average method, exponential smoothing method and ARIMA model; It also includes analysis using regression analysis, which involves building a regression model to analyze the impact of independent variables on dependent variables in historical sales data. Independent variables include price, promotional activities, and season, while dependent variables include sales revenue and sales volume. Regression analysis includes linear regression and multinomial regression. In this embodiment, time series analysis methods (such as moving average, exponential smoothing and ARIMA model) are used to analyze historical sales data, which can accurately grasp the trend and seasonal characteristics of sales data over time. This helps merchants to predict market demand in advance, arrange inventory reasonably, avoid inventory backlog or stockouts, reduce operating costs and improve operating efficiency. By building models through regression analysis (including linear regression and multinomial regression) to analyze the impact of independent variables such as price, promotional activities and season on dependent variables such as sales revenue and sales volume, this analysis can help businesses gain a deeper understanding of the driving role of different factors in sales, thereby formulating more precise marketing and pricing strategies. It is important to emphasize in this embodiment that after replacing the product name buttons with those corresponding to the products purchased by the consumer within the past month, the number of product name buttons clicked by the consumer is counted, and the time taken by the consumer to click different product name buttons is calculated from the count. Based on the length of time taken, the clicked product name buttons are sorted, and the product finally purchased by the consumer is obtained from the sorted product name buttons. The purchased product is marked as a reference product. When it is determined that the consumer will touch the touch screen on the retail robot in the future, the product name button corresponding to the reference product is displayed on the touch screen first. In this embodiment, the accuracy and relevance of product recommendations are further improved, so that the optimized product recommendation order can better meet the actual needs of consumers, increase consumers' acceptance of recommended products and their willingness to purchase, thereby improving the recommendation effect and sales conversion rate, and bringing more sales opportunities and revenue to merchants; Based on the above, the purchasing characteristics of consumers who have purchased the same reference product in their past consumption history are calculated according to the reference product, and are obtained based on purchase frequency, as shown below: ; In the formula, Indicates purchase frequency. Indicates the number of purchases. Indicates the length of time, which can be in months and / or years; The purchase frequency in this formula is the number of times a consumer purchases a specific product within a certain period of time. In this embodiment, by calculating the purchase frequency formula, the number of times a consumer purchases a specific product within a certain period of time can be quantified, providing merchants with an intuitive indicator to measure the consumer's purchase activity for different products. This embodiment further illustrates that the purchase frequency calculation results can be used to segment consumers into different groups, such as high-frequency buyers, medium-frequency buyers, and low-frequency buyers. For consumer groups with different purchase frequencies, merchants can develop differentiated marketing strategies and service plans to improve marketing effectiveness and customer satisfaction. By understanding how frequently consumers purchase different products, merchants can optimize the product layout on the retail robot touchscreen, placing frequently purchased products in more prominent positions to increase product exposure and sales opportunities. At the same time, it also provides an important reference for personalized recommendations, making the recommendations more in line with consumers' purchasing habits and needs. Furthermore, the purchasing characteristics of consumers who have purchased the same reference product as other consumers in their past consumption are calculated based on the reference product, including those calculated using purchase conversion rates, as shown below: ; In the formula, Indicates purchase conversion rate. Indicates the actual number of purchases. Indicates the number of times the product was viewed; The purchase conversion rate in this formula is the percentage of consumers who actually make a purchase after browsing a specific product. Alternatively, it can be calculated through basket analysis, as shown below: ; In the formula, Indicates a specific product, Indicates other goods, Indicates purchasing at the same time and Number of transactions Indicates the total number of transactions; The Market Basket Analysis formula calculates the probability that a consumer will purchase other items at the same time as a specific item. In this embodiment, in addition to purchase frequency, by calculating purchase conversion rate and purchase basket analysis, consumers' purchasing behavior can be comprehensively measured from multiple perspectives. Purchase conversion rate reflects the proportion of consumers who actually purchase after browsing products, revealing the product's attractiveness to consumers and its impact on purchasing decisions; purchase basket analysis reveals consumers' associated purchasing behavior when purchasing specific products, helping merchants understand consumers' demand combinations and shopping patterns. Furthermore, the results of purchase basket analysis can be used to achieve accurate related recommendations. When a consumer purchases a specific product, the system can recommend other products that are frequently purchased together with it based on the purchase basket analysis results, thereby increasing the consumer's purchase volume and shopping satisfaction. The calculation results of purchase conversion rate and purchase basket analysis provide important basis for merchants to optimize marketing strategies. Merchants can analyze the reasons for products with low purchase conversion rates and take corresponding improvement measures, such as adjusting product display and optimizing promotional activities; based on the purchase basket analysis results, merchants can design marketing strategies such as bundled promotions and sales. Based on the calculated purchase characteristics, obtain the consumption preferences and / or purchasing habits of other consumers relative to the consumer. Based on the consumption preferences and / or purchasing habits, obtain the products that other consumers have purchased in the past that are different from the consumer's. Mark the products as differentiated products. When it is determined in the future that the consumer will touch the touch screen on the retail robot, prioritize displaying the product name buttons corresponding to the reference product and differentiated products on the touch screen. In this embodiment, by calculating purchase characteristics and obtaining the consumption preferences and purchasing habits of other consumers relative to the target consumer, it is possible to discover other consumer groups with similar purchasing behaviors to the target consumer, and further obtain the products purchased by these consumers that are different from those of the target consumer (differentiated products), which helps to explore the potential needs and unmet consumption needs of the target consumer. By highlighting and prioritizing differentiated products on the touchscreen of the retail robot, consumers are given a wider range of choices, encouraged to try new products, and expanded their purchasing options.

[0023] As can be seen from the above, this application, by combining large AI models and multimodal interaction, not only enhances the consumer shopping experience but also improves retail efficiency, demonstrating significant innovation and practicality.

[0024] This embodiment, in conjunction with the aforementioned AI-based large-scale intelligent retail robot multimodal interaction system, also proposes a working method for this system, as follows: The method applied to the multimodal interaction system of intelligent retail robots based on AI large models as described in claim 1 is characterized by comprising the following steps: Collect historical sales data for each product within the retail robot, and analyze the sales characteristics of each product based on the historical sales data; Based on historical sales data, obtain facial images of consumers corresponding to each product at the time of purchase, and construct a database, which includes the consumption preferences and purchasing habits of the corresponding consumers; After each item is sold out, the remaining item information in the retail robot is obtained, and the graphical interface on the retail robot's touchscreen is adjusted based on the remaining item information. Obtain the layout of the product name buttons for each remaining product on the graphical interface, and obtain the position of each product name button on the graphical interface based on the layout; The robot senses how close consumers are to it and how long they stay, in order to determine their intentions and behaviors. The system collects facial images of consumers based on their perceived intentions and behaviors. Specifically, if it is determined that a consumer will touch the touchscreen on the retail robot, then the consumer's facial image is collected; otherwise, it is not collected. The database is used to retrieve the historical purchase records of the consumer corresponding to the facial image at the retail robot, and the current purchase of the consumer is predicted based on the historical purchase records.

[0025] In summary, by analyzing consumers' historical purchase records and behavioral patterns, the system can provide consumers with personalized product recommendations. Furthermore, by combining multiple interaction methods such as touch screen operation, gesture recognition, and facial recognition, the interaction between consumers and retail robots becomes more natural, convenient, and efficient, providing a more personalized interactive experience.

[0026] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multimodal interaction system for intelligent retail robots based on AI large-scale models, characterized in that: include: The data statistics module is used to collect historical sales data for each product within the retail robot and analyze the sales characteristics of each product based on the historical sales data. The image information acquisition module responds to the historical sales data and is used to acquire the facial images of consumers corresponding to each product at the time of purchase based on the historical sales data, and construct a database, which includes the consumption preferences and purchasing habits of the corresponding consumers. The data fusion processing module is used to acquire the product information of each remaining product in the retail robot after each product is sold out, and to adjust the graphical interface on the touch screen of the retail robot based on the product information of each remaining product; the data fusion processing module includes an acquisition unit, a sensing unit, an image acquisition unit, and a data retrieval unit. The acquisition unit is used to acquire the layout of the product name buttons for each remaining product on the graphical interface, and acquire the position of each product name button on the graphical interface according to the layout. The sensing unit is used to sense the degree of proximity and dwell time of consumers to the retail robot in order to determine the consumers' intentions and behaviors. The image acquisition unit responds to the sensing unit and is used to acquire the consumer's facial image based on the determined consumer's intention and behavior; wherein, when it is determined that the consumer will touch the touch screen on the retail robot, the consumer's facial image is acquired; otherwise, it is not acquired. The data retrieval unit responds to the image acquisition unit by retrieving the historical purchase records of the consumer corresponding to the face image at the retail robot in the database, and predicting the consumer's current purchase based on the historical purchase records.

2. The AI-based large-scale model-based intelligent retail robot multimodal interaction system as described in claim 1, characterized in that, Based on the prediction results of the consumer's current purchase, the product name buttons corresponding to the consumer's historical purchase records are displayed on the touch screen, including displaying the products purchased by the consumer in the last 3 to 5 times in the upper left or upper right corner of the touch screen; and displaying relevant information of the products corresponding to the product name buttons, including product details, promotional activities and inventory.

3. The AI-based large-scale model-based intelligent retail robot multimodal interaction system as described in claim 2, characterized in that, The data fusion processing module also includes a gesture acquisition unit, which is used to acquire images of the consumer's gestures and analyze and process them using computer vision algorithms to identify the shape, position, and trajectory of the consumer's gestures. Based on the trajectory, it determines whether the consumer will click the product name button in the upper left or upper right corner of the touchscreen. If it is determined that the consumer will not click the product name button in the upper left or upper right corner of the touchscreen, the product name button on the touchscreen is adjusted, including adjusting it to the product name button corresponding to the product purchased by the consumer within the last month.

4. The multimodal interaction system for intelligent retail robots based on AI large-scale models as described in claim 1, characterized in that, In the data statistics module, the sales characteristics of each product are analyzed based on the historical sales data, including analysis in the form of time series analysis, arranging the historical sales data in chronological order, and analyzing the trend and seasonal characteristics of the sales data over time. The time series analysis includes moving average method, exponential smoothing method and ARIMA model; It also includes analysis using regression analysis, which involves establishing a regression model to analyze the impact of independent variables on the dependent variable in historical sales data. The independent variables include price, promotional activities, and season, while the dependent variables include sales revenue and sales volume. The regression analysis includes linear regression and multinomial regression.

5. The AI-based large-scale model-based intelligent retail robot multimodal interaction system as described in claim 3, characterized in that, After replacing the product name buttons with those corresponding to the products purchased by the consumer within the past month, the number of product name buttons clicked by the consumer is counted. The time taken by the consumer to click different product name buttons is calculated from the count. Based on the time taken, the clicked product name buttons are sorted. At the same time, the product finally purchased by the consumer is obtained from the sorted product name buttons, and the purchased product is marked as a reference product. When it is determined that the consumer will touch the touch screen on the retail robot in the future, the product name button corresponding to the reference product is displayed on the touch screen first.

6. The AI-based large-scale model-based intelligent retail robot multimodal interaction system as described in claim 5, characterized in that, Based on the reference product, the purchasing characteristics of the consumer and other consumers who have purchased the reference product in their past consumption history are calculated, and the results are obtained based on purchase frequency, as shown below: ; In the formula, Indicates purchase frequency. Indicates the number of purchases. Indicates the length of time, which is in months and / or years; This formula represents the number of times a consumer purchases a specific product within a given time period.

7. The AI-based large-scale model-based intelligent retail robot multimodal interaction system as described in claim 6, characterized in that, The purchase characteristics of the consumer, which were also purchased by other consumers in the past, are calculated based on the reference product, including calculations based on purchase conversion rates, as shown below: ; In the formula, Indicates purchase conversion rate. Indicates the actual number of purchases. Indicates the number of times the product was viewed; This formula is used to calculate the percentage of consumers who actually purchase a particular product after browsing it. Alternatively, it can be calculated through basket analysis, as shown below: ; In the formula, Indicates a specific product, Indicates other goods, Indicates purchasing at the same time and Number of transactions Indicates the total number of transactions; This formula calculates the probability that a consumer will purchase other goods at the same time as a specific product.

8. The AI-based large-scale model-based intelligent retail robot multimodal interaction system as described in any one of claims 6 to 7, characterized in that, Based on the calculated purchase characteristics, obtain the consumption preferences and / or purchasing habits of other consumers relative to the consumer. Based on the consumption preferences and / or purchasing habits, obtain the products that other consumers have purchased in the past that are different from the consumer. Mark the products as differentiated products. When it is determined in a future time period that the consumer will touch the touch screen on the retail robot, prioritize displaying the product name buttons corresponding to the reference product and the differentiated products on the touch screen.

9. The method applied to the multimodal interaction system of intelligent retail robots based on AI large models as described in claim 1, characterized in that, Includes the following steps: Collect historical sales data for each product within the retail robot, and analyze the sales characteristics of each product based on the historical sales data; Based on historical sales data, obtain facial images of consumers corresponding to each product at the time of purchase, and construct a database, which includes the consumption preferences and purchasing habits of the corresponding consumers; After each item is sold out, the remaining item information in the retail robot is obtained, and the graphical interface on the retail robot's touchscreen is adjusted based on the remaining item information. Obtain the layout of the product name buttons for each remaining product on the graphical interface, and obtain the position of each product name button on the graphical interface based on the layout; The robot senses how close consumers are to it and how long they stay, in order to determine their intentions and behaviors. The system collects facial images of consumers based on their perceived intentions and behaviors. Specifically, if it is determined that a consumer will touch the touchscreen on the retail robot, then the consumer's facial image is collected; otherwise, it is not collected. The database is used to retrieve the historical purchase records of the consumer corresponding to the facial image at the retail robot, and the current purchase of the consumer is predicted based on the historical purchase records.

Citation Information

Patent Citations

  • Vending machine interaction method and system based on reliable gesture recognition

    CN113377193A

  • Interaction method and system of intelligent vending machine

    CN113936665A

  • Face recognition-based information push method and system

    CN107463608A

  • Method and device for achieving dynamic layout of interface of vending machine

    CN111696256A

  • Data pushing method based on artificial intelligence

    CN119338562A