Multi-modal large language model training method and system for retail scene
By extracting human features and training historical data on user area images of smart vending terminals, a multimodal large language model was established, which solved the problem of lack of real-time perception of user features in smart vending terminals, achieved dynamic and accurate prediction of product demand, and improved the timeliness and accuracy of the prediction.
Patent Information
- Application Number
- CN202510665479.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies lack real-time perception and utilization of user image features in smart vending terminals, resulting in low accuracy in product demand prediction and an inability to dynamically adapt to fluctuations in product demand caused by changes in population and consumer preferences.
By extracting human features from user area images in historical commodity transaction scenarios, updating user group characteristic information, and training the long-short term memory model in combination with historical commodity sales data, fine-tuning training is performed using user group characteristic information to establish a multimodal large language model, thus achieving dynamic and accurate prediction of commodity demand.
It improves the timeliness and accuracy of commodity demand forecasting, solves the problem of untimely response to demand fluctuations caused by static data forecasting in existing technologies, and realizes accurate modeling and dynamic adjustment of actual consumer groups.
Smart Images

Figure CN120672248A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on April 2, 2025, with the invention name “Retail terminal inventory replenishment method, device and system based on multimodal data” and application number 202510405222.X. Technical Field
[0002] The present invention relates to the field of smart retail terminal technology, and in particular to a multimodal large language model training method and system for retail scenarios. Background Art
[0003] With the development of artificial intelligence, the Internet of Things, and computer vision technologies, smart retail terminals have been widely used in the retail industry. As a new type of unmanned retail terminal, smart retail terminals feature automatic identification, intelligent checkout, and remote monitoring, providing users with a convenient self-service shopping experience. To improve product turnover and replenishment efficiency at smart retail terminals and reduce the risk of stockouts or unsold items, accurately predicting future product demand has become a core task in smart vending scenarios. Effectively predicting product demand allows operators to implement more scientific product replenishment strategies and dynamic inventory management, thereby improving operational efficiency and user satisfaction. Existing technologies typically rely on historical sales data, product category information, and time series features, using deep learning methods such as linear regression, support vector machines, or long short-term memory (LSTM) networks to model product sales. However, most of these methods ignore the impact of actual user behavior and regional demographics on product demand, and are unable to dynamically adjust prediction models based on user changes within a region. Furthermore, existing models lack the continuous updating and utilization of user profile information, resulting in insufficient adaptability of prediction results to sudden changes in customer flow or changes in scenarios. Therefore, existing technologies still have obvious technical bottlenecks in predicting personalized and regionalized product demand in dynamic retail scenarios, and there is an urgent need for an adaptive prediction method that combines user image features with historical product transaction data.
[0004] Existing Chinese patent CN114219412A discloses an automatic replenishment method based on sales forecast of a smart commodity system, comprising: obtaining the inventory quantity, inventory cycle and commodity attributes of each of a plurality of inventory commodities; determining whether there is a commodity that needs to be replenished and has insufficient inventory among the plurality of inventory commodities; if so, retrieving the historical sales database of the smart commodity system, and performing the following steps for each commodity that needs to be replenished to calculate the replenishment quantity: screening out commodities that have a matching degree with the commodity that needs to be replenished in terms of commodity attributes exceeding a preset degree from the retrieved historical sales database as sample commodities; obtaining sales information of the sample commodities, identifying sales change data of the sales information, comparing the sales trends of the commodity that needs to be replenished and the sample commodities, and calculating, based on the comparison result of the sales trends of the two, the sales quantity of the sample commodities within the remaining sales time of the current inventory cycle of the commodity that needs to be replenished; using the calculated sales quantity of the sample commodities as the replenishment quantity of the commodity that needs to be replenished; and allocating inventory commodities according to the calculated replenishment quantities of each commodity that needs to be replenished to achieve replenishment. While the above solution can estimate replenishment quantities based on historical sales data and product attribute similarity, improving the rationality of replenishment decisions to a certain extent, it relies solely on static historical sales information and product attribute matching for analysis. It lacks real-time perception and utilization of current user group behavioral characteristics, and cannot dynamically adapt to fluctuations in product demand caused by factors such as demographic changes and shifts in consumer preferences in actual sales scenarios. Therefore, this solution cannot effectively address the low prediction accuracy and delayed response issues inherent in existing technologies, which stem from a lack of modeling and real-time updating of regional user characteristics.
[0005] Therefore, how to combine user image features with historical commodity transaction data to establish a commodity demand prediction model and improve the accuracy of commodity demand prediction in smart vending machines is a technical problem that needs to be solved urgently. Summary of the Invention
[0006] In view of this, the present invention provides a multimodal large language model training method and system for retail scenarios, which is used to solve the problem of low accuracy in commodity demand prediction in smart vending terminals in the prior art.
[0007] The technical solution adopted in the present invention is:
[0008] In a first aspect, the present invention provides a multimodal large language model training method for retail scenarios, the method comprising:
[0009] Extract human features from user area images in historical commodity transaction scenarios and update user group feature information;
[0010] Using the pre-collected historical sales data and historical product feature information corresponding to each product, the long-short-term memory model is trained to determine the initial model for product demand forecasting;
[0011] The user group characteristic information is input as incremental data into the pre-trained commodity demand forecasting initial model for fine-tuning training to obtain a fine-tuned commodity demand forecasting target model.
[0012] In a second aspect, the present invention provides a multimodal large language model training device for retail scenarios, the device comprising:
[0013] The human body feature extraction module is used to extract human body features from user area images in historical commodity transaction scenarios and update user group feature information;
[0014] The initial model training module is used to train the long-short-term memory model using the pre-collected historical product sales data and historical product feature information corresponding to each product to determine the initial model for product demand forecasting;
[0015] The fine-tuning training module is used to input the user group characteristic information as incremental data into the pre-trained commodity demand forecasting initial model for fine-tuning training to obtain a fine-tuned commodity demand forecasting target model.
[0016] In a third aspect, an embodiment of the present invention further provides a multimodal large language model training system for retail scenarios, characterized in that it includes: a first camera and a second camera; the acquisition range of the first camera is inside and outside the unmanned vending machine, and the acquisition range of the second camera is outside the unmanned vending machine; at least one processor, at least one memory, and computer program instructions stored in the memory, and when the computer program instructions are executed by the processor, the above method is implemented.
[0017] In summary, the beneficial effects of the present invention are as follows:
[0018] The present invention provides a multimodal large language model training method and system for retail scenarios. The method includes: extracting human features from user area images in historical commodity transaction scenarios and updating user group feature information; using pre-collected historical commodity sales data and historical commodity feature information corresponding to each commodity to train a long-short-term memory model to determine an initial commodity demand forecast model; and inputting the user group feature information as incremental data into the pre-trained initial commodity demand forecast model for fine-tuning training to obtain a fine-tuned commodity demand forecast target model. By introducing human feature extraction from user area images in historical commodity transaction scenarios, the present invention dynamically obtains and updates user group feature information, thereby establishing an accurate model of the actual consumer population, effectively compensating for the lack of real-time user behavior and feature perception in the existing technology. Combining pre-collected product sales data and product feature information, the long short-term memory model (LSTM) is used to train a basic prediction model. The updated user group characteristics are then used as incremental data to fine-tune the model, enabling the model to adjust the prediction strategy in real time according to the actual population characteristics of the current region, achieving dynamic and accurate prediction of product demand, improving the timeliness and accuracy of the prediction, and solving the problem of untimely response to demand fluctuations caused by static data prediction in existing technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work, and these are all within the scope of protection of the present invention.
[0020] Figure 1 This is a schematic diagram of the overall working process of the multimodal large language model training method for retail scenarios in Example 1 of the present invention;
[0021] Figure 2 Schematic diagram of the process of performing commodity target detection on each frame of the real-time image in Example 1 of the present invention;
[0022] Figure 3 Schematic diagram of the process of performing multimodal feature extraction on the product area image in Example 1 of the present invention;
[0023] Figure 4 Schematic diagram of a process for performing multimodal fusion of the color feature information, contour feature information, texture feature information, and text feature information in Example 1 of the present invention;
[0024] Figure 5 This is a flow chart of extracting human features from a user area image in a previous commodity transaction scenario in Example 1 of the present invention;
[0025] Figure 6 This is a flow chart of inputting the user group characteristic information as incremental data into the pre-trained commodity demand forecasting initial model for fine-tuning training in Example 1 of the present invention;
[0026] Figure 7 This is a schematic diagram of a process for selecting target products from a preset product list for replenishment and shelf placement according to prediction results in Example 1 of the present invention;
[0027] Figure 8 This is a schematic diagram of a process for selecting target products for replenishment and shelf placement from a preset product list based on optimization processing results in Example 1 of the present invention;
[0028] Figure 9 This is a structural block diagram of a multimodal large language model training device for retail scenarios in Example 3 of the present invention;
[0029] Figure 10 This is a schematic diagram of the structure of the smart retail terminal in Example 4 of the present invention;
[0030] Figure 11 Schematic diagram of the structure of an electronic device in Example 4 of the present invention;
[0031] The reference numerals are as follows:
[0032] 1-First camera; 2-Second camera. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the orientation or position relationship indicated by the terms "center", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like is based on the orientation or position relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further limitations, elements defined by the phrase "comprising..." do not preclude the presence of additional identical elements in the process, method, article, or apparatus comprising the elements. The embodiments of the present invention and the features thereof may be combined with each other if there is no conflict, and all are within the scope of protection of the present invention.
[0034] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant location, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0035] Example 1
[0036] See Figure 1 Embodiment 1 of the present invention discloses a multimodal large language model training method for retail scenarios, the method comprising:
[0037] According to the real-time video data of the unmanned vending machine in the commodity transaction scenario, the real-time video data is decomposed into multiple frames of real-time images;
[0038] Specifically, real-time video data from vending machines is acquired, transaction scenes within the vending machines are captured in real time, and the real-time video data is broken down into multiple frames. This means that each second or shorter portion of the video stream is divided into independent frames. Cameras are used to continuously record the commodity transaction process, and then video processing techniques (such as frame extraction algorithms) are used to convert the video data into a series of continuous image frames for subsequent processing of each frame. The real-time images can accurately capture the commodity status, customer behavior, and transaction dynamics at each point in time, providing timely and accurate visual data support for subsequent commodity inspection and demand forecasting, ensuring that the replenishment system can make real-time decisions based on the latest situation, rather than relying on delayed or outdated data.
[0039] Performing commodity target detection on each frame of the real-time image, and determining a commodity area image based on the target detection result;
[0040] Specifically, product target detection technology is applied to each frame of real-time image to identify the products in the image and determine their location. Deep learning models, such as YOLO and Faster R-CNN, are used to perform target detection on each frame of the real-time image. After training, the model can identify the products in the image and extract the location information of each product to generate a product area image. The product area image refers to the part of the image that contains the product, removing other irrelevant background or objects. Target detection identifies the product by analyzing features in the image (such as shape, color, texture, etc.) and selects the product area. Therefore, it can accurately extract product information from the complex vending machine environment, laying the foundation for subsequent feature extraction and demand forecasting. By accurately identifying the product area, focusing on the product itself and avoiding interference from other background factors, the accuracy and real-time performance of product monitoring are improved, ensuring more accurate replenishment decisions.
[0041] In one embodiment, see Figure 2 The performing commodity target detection on each frame of the real-time image and determining the commodity area image according to the target detection result includes:
[0042] resizing and denoising each frame of the real-time image to determine multiple frames of target images after denoising;
[0043] Specifically, the real-time image is first resized to ensure that the image meets the input requirements of the target detection model (such as uniform size). Then, noise in the image is removed through denoising processing (such as Gaussian filtering, mean filtering and other technologies) to reduce image interference caused by ambient light, camera quality or other factors. The denoised image can more clearly display the characteristics of goods and users, and improve the accuracy of target detection. Because denoising processing can remove interference information, subsequent target detection can focus more on the real target object, avoiding misidentification and missed detection.
[0044] Inputting the target image of each frame into a pre-trained target detection model to determine initial target boundary position information, wherein the initial target includes a product target and a user target;
[0045] Specifically, each denoised target image frame is fed into a pre-trained object detection model (such as YOLO or SSD). This model, trained on a large amount of annotated data, can accurately identify both product and user targets within the image. The target detection model outputs the bounding box (position and size) of each product target, representing the initial boundary position of the target. Since the initial targets include both products and users, the model can distinguish between these two types of targets and provide their location information. This step provides precise target location information for subsequent processing, ensuring that products and users can be accurately distinguished and analyzed.
[0046] Acquiring an initial target area image within the initial target boundary according to the initial target boundary position information;
[0047] Specifically, based on the detected initial target bounding box, the image is cropped to identify the area within the bounding box, resulting in an initial target region image containing both the product and the user. Each target region is then individually extracted for subsequent feature processing. This precise initial target region extraction reduces interference from irrelevant information, allowing for more focused subsequent image processing and enabling more effective analysis of product and user details.
[0048] The initial target area image is subjected to commodity label detection, and the commodity area image is determined based on the detected commodity label position information and the initial target boundary position information.
[0049] Specifically, the initial target area image is further inspected for product labels. The primary goal is to identify whether a product has labels, such as barcodes or price tags, and determine their location. By combining the initial target boundary location information, the product labels are accurately extracted and matched with the product target area to determine the final product area image. Label detection allows for precise location of specific product information (such as price and type), further ensuring accurate demarcation of product areas and providing a more accurate basis for subsequent product analysis and demand forecasting.
[0050] Performing multimodal feature extraction on the product area image to determine product feature information;
[0051] Specifically, multimodal feature extraction is performed on the product area image to extract multiple types of feature information from the product area image to comprehensively describe the characteristics of the product. Multimodal feature extraction includes using deep learning models, such as convolutional neural networks (CNNs), to extract visual features of the image, such as color, shape, and texture, while combining other modal information, such as data such as product weight and temperature collected by embedded sensors, or customer interaction information captured by voice recognition. These multiple types of features are fused together to form comprehensive feature information of the product. This multimodal processing can fully utilize the advantages of different data sources to provide a more comprehensive and accurate product description than a single visual feature, thereby better supporting product demand forecasting, inventory management, and user behavior analysis. Through multimodal feature extraction, it is possible to accurately capture the various characteristics of the product and improve the accuracy and intelligence level of replenishment forecasting.
[0052] In one embodiment, see Figure 3 The performing multimodal feature extraction on the product area image to determine product feature information includes:
[0053] Performing color space conversion on the product area image, and obtaining a color histogram of the product area in the converted color space;
[0054] Specifically, the product area image is first converted to a color space, converting the original image from the RGB color space to other color spaces more suitable for feature extraction, such as HSV or Lab. Color space conversion can better highlight the color information related to the product in the image and reduce the impact of changes in ambient light. Then, based on the converted color space, the color histogram of the product area is calculated. The color histogram reflects the distribution of different colors in the image and can be used to distinguish the color characteristics of the product. The beneficial effect of this step is that through accurate color feature extraction, the main color of the product can be identified, which helps to distinguish different products and improves the reliability and accuracy of image analysis.
[0055] Aggregating and optimizing the color histogram to determine color feature information corresponding to the product area;
[0056] Specifically, the color histogram generated in the previous step is aggregated and optimized. Aggregation involves merging similar color regions to reduce noise, while optimization involves removing insignificant color information and enhancing the product's representative color features. This process aims to extract more stable and representative color features, avoiding color deviations caused by factors such as ambient lighting changes. The resulting color feature information corresponding to the product area will more accurately describe the product's color characteristics, improving the stability and recognition of the color features, allowing the product to be accurately identified in different environments.
[0057] Performing edge detection on the product area image, and determining contour feature information of the product area based on the edge detection result;
[0058] Specifically, edge detection algorithms, such as Canny edge detection, are applied to product area images to extract the contour features of the products. Edge detection identifies areas with significant grayscale changes in the image, marks the boundaries of the object, and then outlines the shape of the product. By analyzing the contour, the geometric shape features of the product, such as rectangles and circles, are extracted. This can effectively distinguish the appearance features of the products, especially for products with large shape differences, which helps to improve the accuracy and robustness of recognition.
[0059] Extracting texture features from the product area image to determine texture feature information of the product area;
[0060] Specifically, Local Binary Pattern (LBP) is a technique widely used in image texture analysis. It extracts detailed texture features of the product surface, such as roughness and smoothness, by performing an LBP transformation on an image of a product region. By comparing the grayscale value of a pixel with that of surrounding pixels, the image is converted into a binary pattern, capturing local texture features. This allows for distinguishing the surface textures of different products, for example, determining whether a product has a smooth or rough surface. It effectively processes texture information and exhibits strong noise immunity, enabling the identification of unique texture features in complex environments, thereby improving product recognition accuracy.
[0061] Performing text recognition on the product area image to determine text information in the product area;
[0062] Specifically, optical character recognition (OCR) technology is used to identify text in product area images and extract labels, brand names, prices or other information on the products. OCR technology analyzes the text area in the product area image, identifies the characters in the product area and converts them into machine-readable text information. It can accurately extract the text information of the product and provide more data support for subsequent product identification and demand forecasting. In particular, for products with logos and labels, it can quickly obtain important product information.
[0063] Performing natural language processing on the text information to determine text feature information;
[0064] Specifically, once text information is recognized through OCR, natural language processing (NLP) technology is applied to further process the text, including word segmentation, part-of-speech tagging, entity recognition and other operations to extract key information in the product description, such as brand, model, type, and price. The text feature information obtained after NLP processing is used for deeper analysis, such as product classification and sales trend prediction. It can convert the text information of the product into structured features to support more intelligent data analysis and decision-making.
[0065] Multimodal data fusion is performed on the color feature information, the outline feature information, the texture feature information, and the text feature information to determine the product feature information.
[0066] Specifically, the system performs multimodal fusion on the extracted color, contour, texture, and text features to generate comprehensive product feature information. Multimodal fusion methods include weighted averaging, concatenation, or deep learning models (such as multimodal neural networks) to combine information from these different sources, thereby obtaining a more comprehensive and accurate product description. By fusing visual and text features, product identification no longer relies on a single data source, but instead comprehensively considers multiple aspects of the product, greatly improving recognition accuracy and robustness.
[0067] In one embodiment, see Figure 4 The performing multimodal data fusion on the color feature information, the contour feature information, the texture feature information, and the text feature information to determine the product feature information includes:
[0068] Segmenting the product area image based on the contour feature information to obtain local area position information corresponding to the product body;
[0069] Specifically, based on the extracted product contour feature information, image segmentation technology, such as edge-based segmentation algorithm or region growing algorithm, is used to extract the product area from the background, identify the specific area of the product body, and accurately separate the product from the image by analyzing the contour, providing a clear target area for subsequent feature extraction. By accurately determining the boundary of the product body, accurate local area images are provided for subsequent feature extraction and classification, thereby improving the accuracy and efficiency of target recognition.
[0070] Extracting local area color feature information from the color feature information according to the local area position information;
[0071] Extracting local area contour feature information from the contour feature information according to the local area position information;
[0072] Extracting local area texture feature information from the texture feature information according to the local area position information;
[0073] Extracting local area text feature information from the text feature information based on the local area position information;
[0074] Specifically, based on the local area image obtained by segmentation, the local area color feature information, local area contour feature information, local area texture feature information and local area text feature information of the local area are further extracted, which can more comprehensively describe the visual and textual attributes of the product area and provide more powerful data support for subsequent product identification and classification.
[0075] Inputting the product area image into a pre-trained target classification model to determine the product category;
[0076] Specifically, the extracted product area images are fed into a pre-trained target classification model, such as a deep convolutional neural network (CNN). This model has been trained on a large number of labeled product images and can identify and classify the types of products in the images. The output is the category to which the product belongs, such as beverages, snacks, or daily necessities. This process automatically and accurately determines the product category, improving the automation level of product recognition and reducing the risk of manual intervention and misclassification.
[0077] Determining, based on the product category and a preset mapping relationship table between product categories and weights, a first weight for the local area color feature information, a second weight for the local area contour feature information, a third weight for the local area texture feature information, and a fourth weight for the local area text feature information;
[0078] Specifically, based on the product category, a preset mapping relationship table between product categories and weights is consulted. Different weights are assigned to different features of local areas of different product categories, such as color, outline, texture, and text. This determines a first weight for local area color feature information, a second weight for local area outline feature information, a third weight for local area texture feature information, and a fourth weight for local area text feature information. This mapping table adjusts the relevance of each feature based on the type of product. For example, for some products, color features may be more important than texture features, and vice versa. By flexibly adjusting the weight of each feature in product identification based on the actual characteristics of the product, the accuracy of multi-feature fusion is improved, making product classification and demand forecasting more precise.
[0079] According to the first weight, the second weight, the third weight and the fourth weight, the local area color feature information, the local area contour feature information, the local area texture feature information and the local area text feature information are weightedly fused, and the fused feature information is determined as the product feature information.
[0080] Specifically, based on the previously determined first, second, third, and fourth weights, a weighted averaging method is used to fuse the local area color feature information, local area contour feature information, local area texture feature information, and local area text feature information. After weighted fusion, a fused product feature information is obtained. This product feature information contains a complete description of the product in all dimensions. Through this weighted fusion, the importance of each feature is effectively combined to ensure the rational integration of feature information. Through weighted fusion, a more accurate and comprehensive product feature description can be obtained, thereby improving the intelligence and precision of product identification, demand forecasting, and replenishment management.
[0081] Extract human features from user area images in historical commodity transaction scenarios and update user group feature information;
[0082] Specifically, feature information related to the human body is extracted from user area images captured in historical commodity transaction scenarios. This process usually includes the use of deep learning models, such as human posture estimation, to analyze the human body structure and behavioral characteristics in the image, such as the user's posture, movement, height, body shape, and movement trajectory. Through human feature extraction, the user's physical characteristics are identified and analyzed, and the customer's behavioral changes are understood in real time, thereby more accurately predicting commodity demand, further optimizing replenishment decisions and the service experience of vending machines. Through this extraction of user group feature information, unmanned vending machines can respond to customer needs more intelligently, improve operational efficiency and customer satisfaction.
[0083] In one embodiment, see Figure 5 The extracting of human features from user area images in historical commodity transaction scenarios and updating of user group feature information include:
[0084] Obtain historical images of commodity trading scenarios;
[0085] Specifically, historical images of commodity transaction scenarios are retrieved from historical order data. Historical images are usually captured in real time by cameras or other monitoring equipment of unmanned vending machines during the interaction between customers and unmanned vending machines. The purpose of obtaining historical images is to capture and record customer behavior and interaction information. These data will serve as the basis for subsequent analysis.
[0086] Performing human body key point detection on the historical image to determine the user area image;
[0087] Specifically, historical images are processed using a human keypoint detection algorithm, such as OpenPose or HRNet, to detect key points within the image, such as the head, shoulders, elbows, and knees. This keypoint location information helps accurately identify and locate the customer's position and posture within the image. Using this detected keypoint information, a sub-image containing the customer area, known as the user area image, can be extracted from the entire image. This allows for precise identification of the customer's location and activities in front of the vending machine, making subsequent feature extraction and behavior analysis more accurate.
[0088] Performing image enhancement processing on the user area image to determine an enhanced user target image, wherein the image enhancement processing includes super-resolution reconstruction, image denoising, and contrast adjustment;
[0089] Specifically, in this step, the image extracted from the user area undergoes image enhancement. A super-resolution algorithm is used to increase the image resolution, making details in the user area clearer and facilitating further feature analysis. Denoising techniques are used to remove noise from the image, minimizing image degradation caused by shooting conditions (such as low light or interference). Image contrast is optimized to make the target area (e.g., the customer's physical features) more prominent, facilitating subsequent feature extraction and analysis. By improving the quality of the user area image, clearer and more accurate visual information is provided for subsequent feature extraction, enhancing the system's ability to perceive customer behavior.
[0090] Inputting the user target image into a pre-trained posture detection model and skin color detection model to determine skin color feature information, posture feature information, and body shape feature information;
[0091] Specifically, the user target image is first input into the pre-trained posture detection model and skin color detection model. The posture detection model analyzes the positions of key points of the human body in the user target image, thereby identifying the posture information of the human body, including standing, walking, raising arms, bending over and other actions. By identifying these actions, the customer's behavior status and intention can be understood, such as whether he is selecting goods, picking up goods or paying. The skin color detection model identifies and extracts the customer's skin color features based on the skin color distribution in the image. It can judge the customer's skin color type from visual information. Combining the outputs of the two models, it can further infer the customer's body feature information, such as height and body shape (thin or fat). Body feature information helps to optimize the personalized recommendation system, such as inferring customer preferences or recommending suitable products based on body shape. Through the posture detection model and the skin color detection model, multi-dimensional information about customer behavior, appearance and interaction patterns can be obtained in real-time or historical images. This not only helps to improve the accuracy of customer identification, but also can provide more intelligent services, such as personalized product recommendations, optimized replenishment forecasts, and understanding customer purchasing decisions. Combining these features, the vending machine can dynamically respond to customer behavior, improve customer experience and optimize operational efficiency.
[0092] Performing clothing type recognition on the user target image to determine clothing feature information of the user;
[0093] Specifically, clothing type recognition is performed on the user target image. The goal of clothing type recognition is to accurately identify the type, color, and style of clothing worn by the user. Specifically, deep learning models such as convolutional neural networks (CNNs) are used to extract features from clothing regions. Clothing region feature information includes, but is not limited to, clothing styles such as T-shirts, shirts, pants, and jackets; colors such as red, blue, and black; patterns such as stripes, solid colors, and prints; and materials such as cotton, synthetic fibers, and wool. By extracting these detailed clothing features, the user's dressing style is further analyzed. To ensure accurate clothing type recognition, a pre-trained model trained on a large-scale dataset contains images of various types of clothing, covering different human poses, lighting conditions, and viewing angles. This pre-trained model can cope with the diversity of real-world scenarios and can identify users' clothing in complex environments. In addition, the extracted clothing features are further processed, for example, through color space conversion, such as from RGB to HSV, to enhance the stability of color recognition, or using edge detection technology to improve the accuracy of clothing outline recognition. This process not only identifies the user's clothing type but also infers their activity context or needs. For example, wearing sportswear may indicate a workout, while formal attire may indicate work. These inferences enable personalized product recommendations based on the user's clothing characteristics. For example, if a user is wearing sportswear, it can be inferred that they may need to purchase beverages or sporting goods, allowing the vending machine's inventory or display to be adjusted accordingly. This refined analysis enhances the user experience and personalizes the service, thereby improving the vending machine's operational efficiency and user satisfaction.
[0094] Preliminarily predicting the user's age and gender based on the clothing feature information, skin color feature information, and user posture feature information to determine an initial age range and user gender;
[0095] Specifically, based on the clothing feature information, skin color feature information and posture feature information extracted from the user's image, a multimodal learning model is used to make a preliminary prediction of the user's age and gender. Specifically, the user's dressing style is first judged by information such as clothing type, color, and style, which provides preliminary clues for the prediction of age and gender. For example, certain clothing styles or colors may be associated with specific age groups or gender groups, such as sportswear or casual wear are usually associated with young people, while formal suits are more common in middle-aged and elderly men. Secondly, skin color feature information provides further auxiliary information. The distribution characteristics of skin color can help the model determine the user's gender to a certain extent, because people of certain genders or age groups have different genders or genders. Finally, the user's posture characteristics (such as standing, sitting, walking style, etc.) can also provide important clues to behavioral patterns. Posture characteristics are related to typical behavioral patterns of age or gender. Combining these multiple characteristics, machine learning or deep learning models are used to make preliminary predictions on the user's age stage, such as children, teenagers, adults, the elderly and gender (male or female). Pre-trained classification models, such as support vector machines (SVMs), decision trees, and deep neural networks are used to comprehensively analyze these multi-dimensional features. The prediction results will give an initial age range (such as 18-25 years old, 26-35 years old) and gender (male or female), providing basic data for subsequent personalized recommendations and service adjustments.
[0096] Performing face detection on the user target image to determine whether the user's face exists in the user target image;
[0097] Specifically, facial detection algorithms, such as Haar Cascade, MTCNN, and Dlib, are used to process the user's target image and determine whether a recognizable facial area exists within the image. Facial detection technology determines the face's location and presence within the image by identifying facial features, such as the relative positions and geometric relationships of the eyes, nose, and mouth. If no clear facial features exist in the image, this step will return a result indicating that the face does not exist, skipping the subsequent facial feature extraction process. If a face is detected, the subsequent steps will proceed. By detecting the presence of a face in the image, the system ensures that the more detailed feature extraction process only begins when the user's facial information is clear. This effectively avoids misoperations or invalid calculations, improving overall processing efficiency and accuracy.
[0098] If the user's face is present, facial feature extraction is performed on the facial region image to determine facial feature information, wherein the facial feature information includes eye feature information and facial texture feature information;
[0099] Specifically, after detecting the face, accurate feature extraction is performed on the facial area image. Facial feature extraction typically uses deep learning models such as CNN and FaceNet, or classic computer vision algorithms such as HOG features and LBP texture analysis to identify key facial areas. First, eye feature information is extracted by detecting features such as the position, shape, and size of the eyes. This information is closely related to age and gender. Second, facial texture feature information is obtained by analyzing the texture, wrinkles, and details of the skin surface. This is particularly useful for age judgment. The extracted eye and texture features form a set of vector information containing facial features. By extracting detailed facial features, more effective information that helps in inferring age and gender can be obtained. Accurate extraction of facial feature information makes age judgment more accurate, especially when the age range is more ambiguous. The details of the eye and facial texture can effectively serve as a basis for correction, thereby improving the accuracy of the prediction.
[0100] performing a second correction on the initial age range based on the eye feature information and the facial texture feature information, and determining the corrected age range as the target age range;
[0101] Specifically, based on the eye feature information and the facial texture feature information, the initial age range of the preliminary prediction is secondary corrected by analyzing the eye features and texture features of the facial area. Eye feature information, such as bags under the eyes and wrinkles at the corners of the eyes, and facial texture feature information, such as skin laxity and the appearance of fine lines, can provide more accurate age estimation. Specifically, facial aging features (such as wrinkles at the corners of the eyes and the degree of skin laxity) can effectively help distinguish subtle differences in age. Through a deep learning model, the initial age range is fine-tuned in combination with the eye feature information and the facial texture feature information to ensure that the judgment of the age group is more accurate. Through secondary correction of facial features, the preliminary prediction results can be refined, especially in the absence of obvious age markers (such as skin color or clothing), facial details provide accurate age judgment, further improving accuracy.
[0102] The user group characteristic information is determined according to the user gender, skin color characteristic information and target age range.
[0103] Specifically, after accurately correcting the initial age range, the user group characteristics are finally determined by combining the previously acquired user gender and skin color characteristics with the corrected target age range. By combining the characteristics of gender, skin color, and age, a more comprehensive and accurate user profile is provided, more accurately determining the user's age group, gender, and skin color, which provides strong support for subsequent personalized recommendations and services.
[0104] Inputting the user group characteristic information as incremental data into a pre-trained commodity demand forecasting initial model for fine-tuning training to obtain a fine-tuned commodity demand forecasting target model;
[0105] Specifically, user group feature information is input as incremental data into the pre-trained initial model for product demand forecasting. The initial model extracts new information from user characteristics such as age, gender, and skin color, and uses these features as incremental data. Combined with the existing product demand forecasting model, the model weights are updated to generate a more accurate target model for product demand forecasting. Specifically, the incremental learning method can seamlessly introduce new data into the existing model without retraining the entire model, effectively improving the model's adaptability and predictive capabilities when real-time data is updated. In this way, forecasts of product demand can be dynamically adjusted to improve forecast accuracy.
[0106] In one embodiment, see Figure 6 , the user group characteristic information is input as incremental data into the pre-trained commodity demand forecasting initial model for fine-tuning training, and the fine-tuned commodity demand forecasting target model is obtained, including:
[0107] Using the pre-collected historical sales data and historical product feature information corresponding to each product, the long-short-term memory model is trained to determine the initial model for product demand forecasting;
[0108] Specifically, historical product sales data and historical product feature information are first collected and organized. The historical sales data includes sales quantity, price, and sales time of each product, while the historical product feature information includes the type, brand, and specifications of each product. These data and feature information are then preprocessed, including outlier removal, missing value filling, and normalization to ensure input data quality. Next, a long short-term memory (LSTM) model is trained based on this processed data. Because LSTM models excel at processing time series data and can capture both long-term dependencies and short-term fluctuations in product demand, they are well-suited for product demand forecasting. By using these data as input, the LSTM model learns the temporal patterns of product sales and gradually adjusts its parameters by optimizing the loss function during training, such as the mean squared error (MSE). Ultimately, an initial model capable of predicting future product demand is obtained. After training, the model is evaluated on a validation set to ensure that the trained model can effectively predict product demand and provide an accurate initial forecast benchmark for subsequent incremental training.
[0109] Classifying the target age ranges to determine age characteristic information corresponding to the target age ranges;
[0110] Specifically, the extracted target age ranges of users are first divided into multiple discrete categories, such as young, middle-aged, or elderly. Each discrete category is customized based on market demand or product characteristics to ensure that different age groups can be segmented. This classification helps the model better identify the demand characteristics and consumption preferences of different age groups when purchasing products. To ensure that the processed features can be input into the model, these category intervals are encoded. A common method is to use one-hot encoding (One-Hot Encoding), converting each age range into a binary feature vector. For example, the age range "18-25 years old" might correspond to a vector [1,0,0], while "26-35 years old" might correspond to [0,1,0], and so on. This encoding method allows the model to fully utilize information from different age groups and avoids the problem of sequential dependencies between category information, thereby improving the product demand forecasting model's ability to predict user demand from different age groups.
[0111] Performing color space conversion on the skin color feature information to determine converted color feature information;
[0112] Specifically, skin color is typically represented in image processing and computer vision using the RGB (red, green, blue) or HSV (hue, saturation, value) color spaces. While the RGB color space is directly based on the intensities of the three primary colors, the HSV space is more consistent with human color perception and is particularly advantageous for distinguishing skin tones because it decomposes color features into three components: hue (H), saturation (S), and value (V). During the conversion process, skin color data can first be converted from the image's RGB space to the HSV space. Hue, saturation, and value are then calculated to extract key skin color features. For example, hue (H) helps the model distinguish the basic colors of different skin tones, while saturation (S) helps describe the purity of the skin tone, and value (V) reflects the lightness or darkness of the skin tone. The converted color feature information can be further standardized or normalized for input into subsequent machine learning models. In this way, skin color information not only reduces interference from varying lighting conditions but also more accurately reflects the differences in consumer behavior among people with different skin tones, thereby improving the accuracy of product demand forecasting models for different user groups.
[0113] Converting and encoding the user's gender to determine binary feature information after processing;
[0114] Specifically, user gender information is converted and encoded, transforming it into a numerical feature that is directly input into the machine learning model. Since gender is a binary feature, binary encoding (also known as one-hot encoding) is often used. In binary encoding, gender is converted into two possible numerical labels: one for "male" and the other for "female." To accommodate most machine learning algorithms, this gender information is typically converted into a binary variable of 0 and 1. For example, suppose "male" is represented by 1 and "female" by 0, or vice versa. This encoding allows the model to directly use these numerical features for training and prediction. Furthermore, if the dataset has a highly imbalanced gender ratio, additional processing may be required, such as class balancing techniques (e.g., oversampling or undersampling) to prevent the model from favoring a particular gender. This processing method converts gender information into a numerical feature that effectively helps the prediction model understand the differences in product demand among users of different genders. For specific scenarios, it can also be expanded to include more categories (e.g., multi-class gender encoding) to support more complex demand prediction.
[0115] determining multi-dimensional feature information based on the age feature information, the color feature information, and the binary feature information;
[0116] Specifically, based on the age, color, and binary feature information, these three features can be concatenated or fused to construct a comprehensive multidimensional feature. The key to this step lies in effectively integrating these heterogeneous data types (such as discrete numerical values for age ranges, numerical color information for skin color, and binary features for gender) so that they can be provided to machine learning models for training and prediction. First, age feature information can be discretized to produce a numerical feature representing the age range. For example, "18-25 years old" might correspond to a value of 1, "26-35 years old" might correspond to a value of 2, and so on. After color space conversion, skin color feature information typically obtains numerical values in RGB or HSV format, with each dimension representing the intensity and distribution of color. These color values can be used as continuous feature inputs. Finally, binary gender information is a simple value of 0 or 1, representing male or female. To effectively combine these different types of features, they can be concatenated along the feature dimensions into a single vector, resulting in a comprehensive feature vector containing information from all dimensions. This vector is provided as input to the product demand forecasting model, helping the model consider the multidimensional information of different users during the forecasting process. In this way, the model can take into account the user's basic information (such as gender and age) and incorporate visual features (skin color) when predicting product demand, thereby improving prediction accuracy and personalization effects.
[0117] According to a preset feature weighting mechanism, the multi-dimensional feature information is input into the initial model for commodity demand forecasting to perform incremental training and model fine-tuning to determine the target model for commodity demand forecasting.
[0118] Specifically, the multi-dimensional feature information is first weighted according to a preset feature weighting mechanism, assigning different weights to each feature. The design of the feature weighting mechanism can be determined based on actual needs and prior knowledge. For example, gender may have a greater impact on demand forecasting for certain products, while age and skin color may be more important for other types of products. This weighting mechanism ensures that the model prioritizes learning certain features, thereby improving prediction accuracy. Next, the weighted multi-dimensional feature information is input into the initial product demand forecasting model, and incremental training and model fine-tuning are performed. The goal of incremental training is to enable the initial model to be further optimized based on newly added user features (such as age, skin color, and gender) without requiring the entire model to be retrained. This process can be achieved through methods such as mini-batch gradient descent, ensuring that the model effectively incorporates new feature information during training while retaining its memory of historical data. During incremental training, the model gradually adapts to training data containing new features by updating weights and parameters, thereby improving prediction accuracy. Fine-tuning meticulously adjusts the model parameters to ensure that it maintains efficient prediction capabilities despite the addition of new feature dimensions. Finally, the model that has undergone incremental training and fine-tuning is the product demand prediction target model, which can more accurately predict product demand based on the personalized characteristics of different users (such as age, skin color, and gender).
[0119] The commodity feature information is input into the commodity demand forecasting target model, and target commodities are screened from a preset commodity list for replenishment and shelf placement according to the forecasting results.
[0120] Specifically, the product feature information is input into the product demand forecasting target model, which is then trained using incremental data to perform real-time forecasting. Based on the target model's accurate prediction of product demand, the system automatically selects the products from a preset product list that best meet demand and replenishes them. This intelligently determines which products are in high demand in the current transaction scenario, optimizing replenishment decisions, reducing manual intervention and inventory backlogs, and improving the product flow efficiency of the vending machine.
[0121] In one embodiment, see Figure 7 Inputting the commodity feature information into the commodity demand forecasting target model and selecting target commodities for replenishment and shelf placement from a preset commodity list according to the forecasting result includes:
[0122] Inputting the commodity feature information into the commodity demand forecasting target model and outputting the commodity demand;
[0123] Specifically, product feature information (such as color, label, shape, and texture) extracted from unmanned vending machines is fed into a trained product demand prediction target model. The target model then calculates the demand for each product in the current context based on this feature information. This prediction provides foundational data for subsequent replenishment decisions, helping to identify products with high demand and those that are likely to be slow-moving.
[0124] Based on the business needs of unmanned vending machines, the optimization goals are determined to be product profits and user popularity;
[0125] Specifically, in the actual replenishment process, the goal of unmanned vending machines is not only to meet the demand for goods, but also to ensure that the replenishment plan is the most optimized for the business. This means that replenishment decisions must not only consider how to meet customer needs (by increasing customer favorability), but also how to bring higher profits (by selecting high-profit goods). Based on business needs, it is necessary to balance these two goals: on the one hand, selecting higher-profit goods will directly increase the profitability of unmanned vending machines; on the other hand, increasing user favorability can promote customer repeat purchases and improve customer satisfaction, thereby increasing sales and reputation. By setting these two optimization goals, economic benefits and user experience are comprehensively considered in the replenishment process, ultimately achieving maximized business value.
[0126] Obtaining preset constraints, including the inventory capacity constraint corresponding to the current unmanned vending machine, the inventory capacity constraint corresponding to the product, the replenishment budget constraint, and the sales cycle constraint;
[0127] Specifically, in practice, replenishment decisions cannot be divorced from the actual operating environment and require consideration of multiple constraints. Inventory capacity refers to the limited space in unmanned vending machines, which cannot accommodate an unlimited number of items. Therefore, replenishment must be arranged based on the available space. Product inventory capacity refers to the quantity of each item in the vending machine. Some items may occupy a large amount of space due to size, packaging, and other factors, thus affecting the placement of other items. Replenishment budget constraints refer to the limited funds involved in the replenishment process, making it impossible to replenish large quantities of every item in a short period of time. Sales cycle constraints refer to the need for items to sell out within a certain sales cycle. Therefore, the sales cycle of each item must be predicted to ensure that replenishment can be completed within a reasonable timeframe. By comprehensively considering these constraints, it is possible to make optimal replenishment decisions within limited resources and time.
[0128] Performing multi-objective optimization on the replenishment quantity of the product based on the constraints and the optimization goal, combined with the demand for the product;
[0129] Specifically, multi-objective optimization algorithms are used to solve the decision-making problem of maximizing preset optimization objectives (product profit and user popularity) under limited resources (such as inventory capacity constraints, replenishment budget constraints, and sales cycle constraints). Multi-objective optimization processing must not only consider product demand, but also take into account constraints such as inventory capacity, budget constraints, and sales cycle. These multiple objectives and constraints are combined through mathematical models to ultimately calculate the most appropriate product replenishment quantity. For example, in the mathematical logic of multi-objective optimization, a function must first be defined for each objective. These objectives are usually competing, such as maximizing product profit and maximizing user popularity. Suppose there are n products, and the optimization goal is to determine the replenishment quantity x1, x2, ..., xn based on the characteristic information of each product. Each objective function is a function fi(x1, x2, ..., xn) related to the replenishment quantity. For example, increasing the replenishment quantity of a certain product will lead to a decrease in profit. To comprehensively consider these objectives, a weighted sum method or Pareto optimization method is used to convert them into a comprehensive objective function. Then, combined with constraints such as inventory capacity C, budget B, and sales cycle T, a mathematical constraint optimization model is formed:
[0130]
[0131] Among them, Maximize means maximization, wi is the weight coefficient of each goal, xi is the replenishment quantity of the i-th product, indicating the importance of each goal, and the constraints include the maximum inventory capacity constraint xi≤Ci of the product, Ci is the inventory capacity constraint of the i-th product, indicating the maximum inventory of the product. Budget constraint ∑ i=1 n p i xi≤B, where B is the total budget constraint, representing the maximum amount of funds available for replenishment. Furthermore, the replenishment quantity must not exceed the maximum sales cycle, xi≤Ti. Ti is the sales cycle constraint for the i-th item, indicating the sales cycle of the item, i.e., the time required for replenishment. By solving this optimization problem, we obtain a set of optimal item replenishment quantities x1, x2, ..., xn that both satisfy the constraints and maximize the overall objective function. The specific optimization method used is linear programming, integer programming, or genetic algorithms, depending on the form of the objective function and constraints. Ultimately, through multi-objective optimization, the system can balance different objectives while meeting business needs to find the most appropriate item replenishment plan.
[0132] Based on the optimization results, target products are selected from the preset product list for replenishment.
[0133] Specifically, after performing multi-objective optimization calculations, the system selects items requiring replenishment from a pre-defined list based on the results, ensuring that each item is stocked at the optimized quantity. This not only improves the efficiency of the replenishment process but also ensures that the merchandise in the vending counters meets market demand and business objectives. By tracking and adjusting the replenishment list in real time, the system allows for flexible replenishment based on changing sales data and market trends, further optimizing the operational efficiency of the vending counters.
[0134] In one embodiment, see Figure 8 , the step of selecting target products for replenishment from a preset product list according to the optimization processing result includes:
[0135] Based on the optimization results, determine whether the current inventory capacity of the target product meets the replenishment demand;
[0136] Specifically, by analyzing the optimization results, we first evaluate whether the current inventory capacity of the target product is sufficient to meet the replenishment quantity provided by the demand forecasting model. If the inventory capacity of the target product is less than the replenishment demand, we identify this deficiency and decide whether to replenish it. The degree of satisfaction of inventory capacity is crucial to optimizing the replenishment plan. Only when the inventory capacity is not satisfied will further product substitution and adjustment processes be triggered to ensure the smooth progress of replenishment.
[0137] When it is determined that the inventory capacity does not meet the replenishment demand, feature extraction is performed on each non-target product in the preset product list to determine the feature information of the non-target products;
[0138] Specifically, when it is determined that the inventory capacity of the target product cannot meet the replenishment demand, feature extraction will be performed on all non-target products in the preset product list. This process includes extracting features from the product's attribute information (such as category, price, sales status, user ratings, etc.) and converting these features into a form that can be processed in the data model. The feature information of non-target products will help in subsequent similarity calculations, allowing evaluation of whether these non-target products can replace the target products to meet the replenishment demand without significantly affecting the ultimate business goals.
[0139] Calculating similarity between the product feature information and the non-target product feature information to determine the similarity between each non-target product and the target product, and obtaining the inventory capacity corresponding to each non-target product;
[0140] Specifically, by calculating the similarity between the target product and each non-target product, it is determined which non-target products are more similar to the target product in terms of characteristics. The similarity calculation is usually based on the multi-dimensional characteristics of the product, such as color, brand, category, price, etc., and uses methods such as cosine similarity and Euclidean distance to quantify the degree of similarity between the two. This calculation helps determine which products can be used as alternative options when the target product is out of stock, thereby reducing business losses caused by insufficient inventory.
[0141] Compare the similarities between each non-target product and the target product to determine the maximum similarity;
[0142] Specifically, based on the similarity calculated previously, the actual inventory capacity information is extracted for each non-target product. Inventory capacity is crucial for replenishment decisions, as it allows knowing the available quantity of each product in actual sales. In this step, the existing inventory status of the product is used to evaluate whether more products can be added or the target product can be replaced. By combining this with inventory capacity, the replenishment strategy is further refined to ensure that there is no over-replenishment or excess inventory.
[0143] Comparing the maximum similarity with a preset similarity threshold to determine a comparison result;
[0144] Specifically, the calculated maximum similarity value is compared with a pre-set similarity threshold. The similarity threshold is a predefined standard that indicates whether the similarity between products is high enough to ensure that the characteristics of the substitute product match the target product to a certain extent. Through this comparison, the system determines which non-target products are similar enough to serve as substitutes for the target product. If the maximum similarity is greater than or equal to the similarity threshold, the non-target product is similar to the target product in characteristics and can be selected as a replenishment candidate. Otherwise, the non-target product is discarded as a replacement option.
[0145] According to the comparison result, when the maximum similarity is greater than the similarity threshold, the non-target product inventory capacity corresponding to the maximum similarity is compared with a preset capacity threshold;
[0146] Specifically, once the non-target products corresponding to the maximum similarity are determined, the inventory capacity of these products is checked. Specifically, the inventory capacity of the non-target products is compared with the preset capacity threshold. The capacity threshold refers to the minimum inventory requirement for each product, indicating whether there is sufficient inventory for replenishment. If the inventory capacity of the non-target product is greater than the preset capacity threshold, it means that the product has the potential to meet the replenishment demand and can be used as a replenishment target; otherwise, the replenishment requirements are not met and other replenishment options will be selected.
[0147] When the inventory capacity of the non-target product is greater than the capacity threshold, the target product and the non-target product corresponding to the maximum similarity are screened out from the preset product list for replenishment.
[0148] Specifically, based on the comparison results of the first two steps, it will be decided which products need to be replenished. If the maximum similarity is greater than the similarity threshold and the inventory capacity of the non-target products exceeds the capacity threshold, the target products and the non-target products corresponding to the maximum similarity will be screened out from the preset product list for replenishment. These products will be put on the shelves together to ensure that user needs can be met when inventory is insufficient and to avoid out-of-stock situations. At the same time, by introducing products with the maximum similarity, users' acceptance of and purchase possibility of alternative products will be improved, thereby maximizing commercial benefits.
[0149] Example 2
[0150] In actual operation, unmanned vending machines are widely distributed in different areas such as office buildings, schools, shopping malls, and subway stations. The demand for goods in each vending machine varies depending on the geographical location, customer flow, and consumption habits. Traditional replenishment methods are often based on fixed time or inventory warnings, which leads to some vending machines being sold out prematurely, affecting sales, or excessive replenishment increasing operating costs. Therefore, in order to intelligently calculate a globally optimal replenishment time point, based on Example 1, after inputting the product feature information into the product demand prediction target model and outputting the product demand, the following steps are further included:
[0151] Obtain the target commodity demand corresponding to the unmanned vending machines in each target area and the location information of each target area;
[0152] Specifically, the method of Example 1 is used to obtain the target product demand for unmanned vending machines in each target area. GIS (Geographic Information System) or GPS positioning technology is used to obtain the geographic location of each vending machine and divide it into corresponding regions for subsequent regional replenishment optimization. This regional division allows for more targeted replenishment plans and improves logistics and distribution efficiency.
[0153] Calculate the replenishment demand for each unmanned vending machine based on the demand for each target product, combined with historical sales data and real-time inventory information;
[0154] Specifically, based on historical sales data, time series analysis (such as ARIMA and LSTM neural networks) is used to predict the demand for goods in the future. The current inventory level is calculated. If the inventory is lower than a certain threshold, such as the safety stock, the replenishment demand calculation is triggered. A replenishment model, such as the inventory control model EOQ, is used to calculate the replenishment quantity required for each vending machine to ensure sufficient supply of goods while avoiding excessive inventory. The demand is dynamically adjusted based on external factors such as holidays, weather, and time periods. For example, when the temperature rises, the demand for beverages may increase, and the predicted replenishment quantity needs to be increased accordingly. Through historical data analysis and real-time inventory calculations, the replenishment demand is accurately calculated to ensure that the replenishment quantity can meet sales demand without causing backlogs. The external environmental factors are combined to optimize the forecast, improve the system adaptability, and enhance the service capabilities of the unmanned vending machine. Machine learning is used to optimize inventory management, reduce corporate operating costs, and increase profit margins.
[0155] Obtaining a sales cycle corresponding to each target area according to the location information of each target area;
[0156] Specifically, statistics are collected for each unmanned vending machine in different time periods, such as hours, days, and weeks, to analyze sales trends. A K-means clustering algorithm is used to group vending machines with similar sales patterns into the same category to determine the sales cycle for each region. For example, unmanned vending machines in office areas have brisk sales during the day on weekdays, while unmanned vending machines in residential areas are in higher demand at night. Combined with business district classification (such as commercial areas, schools, and stations), the sales cycle is further refined to adapt to sales patterns in different environments. By analyzing the sales cycle, the rationality of replenishment time is improved, making the replenishment plan more in line with actual sales conditions. A clustering algorithm is used to improve the accuracy of regional sales cycles and optimize the operation and management of unmanned vending machines. Combined with business district classification and historical data, the forecasting ability is improved, making the replenishment plan more flexible.
[0157] Calculate the replenishment frequency based on the replenishment demand and the sales cycle to determine the optimal replenishment cycle for each unmanned vending machine;
[0158] Specifically, based on the replenishment demand and sales cycle, optimization algorithms such as integer linear programming and dynamic programming are used to calculate the optimal replenishment cycle for each unmanned vending machine. Multiple constraints are set, such as inventory thresholds, product shelf life, and transportation costs, to ensure the rationality of the replenishment cycle. The replenishment frequency of different products is calculated based on their consumption rates. For example, the replenishment cycle for high-selling products should be shorter, while the replenishment cycle for low-selling products can be appropriately extended. The replenishment cycle is dynamically adjusted to adapt to actual operating conditions. By calculating the optimal replenishment cycle, unnecessary replenishments are reduced, logistics costs are lowered, and supply chain efficiency is improved. Optimization is carried out by combining multiple factors to improve the inventory turnover rate of the vending machines and avoid product backlogs or out-of-stocks. A dynamic adjustment mechanism is used to increase the flexibility of the system and make the replenishment plan more adaptable.
[0159] Based on the optimal replenishment cycle of each vending machine, the replenishment time of the vending machines in the same area is optimized synchronously to determine the optimal replenishment time window at the regional level;
[0160] Specifically, a distributed computing method is used to comprehensively calculate the replenishment cycles of all vending machines in the same area to determine the optimal replenishment time window; a genetic algorithm or heuristic optimization algorithm is used to find the optimal solution among multiple replenishment plans to ensure that the replenishment time is arranged reasonably; the replenishment time window is further optimized in combination with the operating time of the replenishment vehicle and the work arrangement of the replenishment staff to reduce replenishment costs and time waste; real-time data monitoring (such as inventory changes, traffic conditions, etc.) is used to dynamically adjust the replenishment time window to respond to emergencies (such as temporary out-of-stock, peak demand surge, etc.); through regional-level replenishment optimization, the number of replenishment times is reduced and the logistics distribution efficiency is improved; an intelligent scheduling algorithm is used to improve the rationality of the replenishment plan and avoid resource waste; combined with real-time monitoring, the system adaptability is improved and the operational risks brought by emergencies are reduced.
[0161] Based on the optimal replenishment time window at the regional level, the replenishment paths between different regions are optimized to determine the optimal replenishment strategy, wherein the replenishment strategy includes the replenishment time point and the replenishment route.
[0162] Specifically, a path optimization algorithm is used to calculate the shortest replenishment path to improve transportation efficiency; combined with real-time traffic data, the replenishment route is dynamically adjusted to avoid traffic congestion; the operation plan of the replenishment vehicle is optimized to reduce transportation costs; for example, priority is given to replenishing vending machines in high-sales areas or severely out-of-stock areas to ensure the rational allocation of replenishment resources; through path optimization, replenishment efficiency is improved and vehicle operating costs are reduced; real-time data adjustments are used to increase the flexibility of the replenishment plan and reduce replenishment delays; combined with the intelligent scheduling system, the overall operational efficiency of the vending machines is improved, and the profitability of the enterprise is improved.
[0163] Example 3
[0164] See Figure 9 Embodiment 3 of the present invention further provides a multimodal large language model training device for retail scenarios, the device comprising:
[0165] An image acquisition module is used to decompose the real-time video data of the unmanned vending machine in the commodity transaction scenario into multiple frames of real-time images;
[0166] The target detection module is used to detect commodity targets in each frame of the real-time image and determine the commodity area image based on the target detection result;
[0167] A multimodal feature extraction module, configured to extract multimodal features from the product area image to determine product feature information;
[0168] The human body feature extraction module is used to extract human body features from user area images in historical commodity transaction scenarios and update user group feature information;
[0169] A target model training module is used to input the user group feature information as incremental data into a pre-trained commodity demand forecasting initial model for fine-tuning training to obtain a fine-tuned commodity demand forecasting target model;
[0170] The replenishment and shelving module is used to input the commodity feature information into the commodity demand prediction target model, and filter out target commodities for replenishment and shelving from a preset commodity list according to the prediction results.
[0171] Specifically, the multimodal large language model training device for retail scenarios provided by the embodiment of the present invention can reflect changes in customer behavior in real time, avoid delayed or over-replenishment of goods, and thus achieve accurate and intelligent replenishment management.
[0172] Example 4
[0173] Embodiment 4 of the present invention further provides a multimodal large language model training system for retail scenarios, the system comprising:
[0174] The first camera 1 and the second camera 2; the first camera 1 captures data inside and outside the smart vending machine, while the second camera 2 captures data outside the unmanned vending machine; Figure 10 A schematic diagram of the structure of an intelligent retail terminal provided by Embodiment 4 of the present invention is shown.
[0175] A multimodal large language model training system for retail scenarios may include a processor and a memory storing computer program instructions.
[0176] Specifically, the processor may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits for implementing the embodiments of the present invention.
[0177] The memory may include a large capacity memory for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include a removable or non-removable (or fixed) medium. Where appropriate, the memory may be inside or outside the data processing device. In a specific embodiment, the memory is a non-volatile solid-state memory. In a specific embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0178] The processor reads and executes computer program instructions stored in the memory to implement any one of the multimodal large language model training methods for retail scenarios in the above embodiments.
[0179] In one example, the multimodal large language model training system for retail scenarios may also include a communication interface and a bus. Figure 11 As shown, the processor, memory, and communication interface are connected via a bus and communicate with each other.
[0180] The communication interface is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiments of the present invention.
[0181] Bus comprises hardware, software or both, couples the parts of described equipment together.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus can comprise one or more buses.Although the embodiment of the present invention describes and shows specific bus, the present invention considers any suitable bus or interconnection.
[0182] In summary, the embodiments of the present invention provide a multimodal large language model training method and system for retail scenarios.
[0183] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0184] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0185] It should also be noted that the exemplary embodiments described herein describe methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.
[0186] The above description is only a specific embodiment of the present invention. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention.
Claims
1. A multimodal large language model training method for retail scenarios, characterized by: The method comprises: Extract human features from user area images in historical commodity transaction scenarios and update user group feature information; Using the pre-collected historical sales data and historical product feature information corresponding to each product, the long-short-term memory model is trained to determine the initial model for product demand forecasting; The user group characteristic information is input as incremental data into the pre-trained commodity demand forecasting initial model for fine-tuning training to obtain a fine-tuned commodity demand forecasting target model.
2. The multimodal large language model training method for retail scenarios according to claim 1 is characterized in that The extracting of human features from user area images in historical commodity transaction scenarios and updating user group feature information includes: Obtain historical images of commodity trading scenarios; Performing human body key point detection on the historical image to determine the user area image; Performing image enhancement processing on the user area image to determine an enhanced user target image, wherein the image enhancement processing includes super-resolution reconstruction, image denoising, and contrast adjustment; Inputting the user target image into a pre-trained posture detection model to determine posture feature information; Performing clothing type recognition on the user target image to determine clothing feature information of the user; Preliminarily predicting the user's age and gender based on the clothing feature information and the user's posture feature information to determine an initial age range and user gender; Performing face detection on the user target image to determine whether the user's face exists in the user target image; If the user's face is present, facial feature extraction is performed on the facial region image to determine facial feature information, wherein the facial feature information includes eye feature information and facial texture feature information; performing a second correction on the initial age range based on the eye feature information and the facial texture feature information, and determining the corrected age range as the target age range; Predicting the user's occupation based on the geographic location information of the unmanned vending machine and the clothing feature information to determine the user's occupation; The user group characteristic information is determined according to the user gender, target age range and the user occupation.
3. The multimodal large language model training method for retail scenarios according to claim 2 is characterized in that: The inputting of the user group characteristic information as incremental data into the pre-trained commodity demand forecasting initial model for fine-tuning training to obtain the fine-tuned commodity demand forecasting target model includes: Classifying the target age ranges to determine age characteristic information corresponding to the target age ranges; Classifying and encoding the user's occupation to determine occupational characteristic information; Converting and encoding the user's gender to determine binary feature information after processing; Performing weighted fusion on the age feature information, occupation feature information, and binary feature information according to preset weights to determine multi-dimensional feature information; The multi-dimensional feature information is input into the initial commodity demand forecasting model for incremental training and model fine-tuning to determine the commodity demand forecasting target model.
4. The multimodal large language model training method for retail scenarios according to any one of claims 1 to 3, characterized in that: The method further comprises: According to the real-time video data of the unmanned vending machine in the commodity transaction scenario, the real-time video data is decomposed into multiple frames of real-time images; Performing commodity target detection on each frame of the real-time image, and determining a commodity area image based on the target detection result; Performing multimodal feature extraction on the product area image to determine product feature information; The commodity feature information is input into the commodity demand forecasting target model, and target commodities are screened from a preset commodity list for replenishment and shelf placement according to the forecasting results.
5. The multimodal large language model training method for retail scenarios according to claim 4 is characterized in that: The performing multimodal feature extraction on the product area image to determine product feature information includes: Performing color space conversion on the product area image, and obtaining a color histogram of the product area in the converted color space; Aggregating and optimizing the color histogram to determine color feature information corresponding to the product area; Performing edge detection on the product area image, and determining contour feature information of the product area based on the edge detection result; Extracting texture features from the product area image to determine texture feature information of the product area; Performing text recognition on the product area image to determine text information in the product area; Performing natural language processing on the text information to determine text feature information; Multimodal data fusion is performed on the color feature information, the outline feature information, the texture feature information, and the text feature information to determine the product feature information.
6. The multimodal large language model training method for retail scenarios according to claim 5 is characterized in that: The performing multimodal data fusion on the color feature information, the outline feature information, the texture feature information, and the text feature information to determine the product feature information includes: Segmenting the product area image based on the contour feature information to determine the local area position information corresponding to the product body; Extracting local area color feature information from the color feature information according to the local area position information; Extracting local area contour feature information from the contour feature information according to the local area position information; Extracting local area texture feature information from the texture feature information according to the local area position information; Extracting local area text feature information from the text feature information based on the local area position information; Inputting the product area image into a pre-trained target classification model to determine the product category; Determining, based on the product category and a preset mapping relationship table between product categories and weights, a first weight for the local area color feature information, a second weight for the local area contour feature information, a third weight for the local area texture feature information, and a fourth weight for the local area text feature information; According to the first weight, the second weight, the third weight and the fourth weight, the local area color feature information, the local area contour feature information, the local area texture feature information and the local area text feature information are weightedly fused, and the fused feature information is determined as the product feature information.
7. The multimodal large language model training method for retail scenarios according to claim 4 is characterized in that: Inputting the commodity feature information into the commodity demand forecasting target model, and screening target commodities for replenishment and shelf placement from a preset commodity list according to the forecasting result comprises: Inputting the commodity feature information into the commodity demand forecasting target model and outputting the commodity demand; Obtaining an optimization goal in a commodity transaction scenario, wherein the optimization goal includes commodity profit and user preference; Obtaining preset constraints, including the inventory capacity constraint corresponding to the current unmanned vending machine, the inventory capacity constraint corresponding to the product, the replenishment budget constraint, and the sales cycle constraint; Performing multi-objective optimization on the replenishment quantity of the product based on the constraints and the optimization goal, combined with the demand for the product; Based on the optimization results, target products are selected from the preset product list for replenishment.
8. The multimodal large language model training method for retail scenarios according to claim 7 is characterized in that: The step of selecting target products for replenishment from the preset product list according to the optimization processing result includes: Based on the optimization results, determine whether the current inventory capacity of the target product meets the replenishment demand; When it is determined that the inventory capacity does not meet the replenishment demand, feature extraction is performed on each non-target product in the preset product list to determine the feature information of the non-target products; Calculating similarity between the product feature information and the non-target product feature information to determine the similarity between each non-target product and the target product, and obtaining the inventory capacity corresponding to each non-target product; Compare the similarities between each non-target product and the target product to determine the maximum similarity; Comparing the maximum similarity with a preset similarity threshold to determine a comparison result; According to the comparison result, when the maximum similarity is greater than the similarity threshold, the non-target product inventory capacity corresponding to the maximum similarity is compared with a preset capacity threshold; When the inventory capacity of the non-target product is greater than the capacity threshold, the target product and the non-target product corresponding to the maximum similarity are screened out from the preset product list for replenishment.
9. A multimodal large language model training device for retail scenarios, characterized by: The device comprises: The human body feature extraction module is used to extract human body features from user area images in historical commodity transaction scenarios and update user group feature information; The initial model training module is used to train the long-short-term memory model using the pre-collected historical product sales data and historical product feature information corresponding to each product to determine the initial model for product demand forecasting; The fine-tuning training module is used to input the user group characteristic information as incremental data into the pre-trained commodity demand forecasting initial model for fine-tuning training to obtain a fine-tuned commodity demand forecasting target model.
10. A multimodal large language model training system for retail scenarios, characterized by: include: A first camera and a second camera; the first camera's collection range is inside and outside the unmanned vending machine, and the second camera's collection range is outside the unmanned vending machine; At least one processor, at least one memory, and computer program instructions stored in the memory, which implement the method according to any one of claims 1 to 8 when the computer program instructions are executed by the processor.
Citation Information
Patent Citations
Automatic replenishment method and replenishment system based on sales prediction of intelligent commodity system
CN114219412A