Retail method, system and device based on large model and robot and medium

Through the combination of depth cameras and large models, two-dimensional and three-dimensional information fusion is generated, and precise product recognition and capture by robots in complex retail environments is achieved, the problem of insufficient identification and capture accuracy in the existing technology is solved, and the level of automation of retail business is improved.

CN120510474APending Publication Date: 2025-08-19SHANTOU UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510483459.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing retail technology has shortcomings in product identification and capture accuracy and flexibility, especially in complex environments, and the poor integration of different technologies leads to high costs and long implementation cycles.

Method used

Depth cameras are used to obtain color images and depth information, combine large models to analyze multimodal information, use deep learning to generate two-dimensional candidate boxes and map them to the three-dimensional point cloud coordinate system, perform three-dimensional segmentation and feature analysis, integrate the two-dimensional three-dimensional classification results, and control the robot to accurately capture the target products.

Benefits of technology

It improves the accuracy of product identification and spatial positioning accuracy, enhances the adaptability to complex retail scenarios, reduces manual intervention, and improves the work efficiency and service quality of retail business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510474A_ABST
    Figure CN120510474A_ABST
Patent Text Reader

Abstract

The invention provides a retail method, system and device based on a large model and a robot and a medium, and relates to the technical field of smart retail, and the method comprises the steps: employing a depth camera, and carrying out the photographing to obtain a color image and depth information of a retail scene; analyzing multi-modal information input by a user by adopting a large model to obtain attribute information of a target commodity; generating a two-dimensional candidate frame by using a deep learning model and calculating a classification result, mapping depth information to a three-dimensional coordinate system to obtain a positioning coordinate, performing three-dimensional segmentation and feature analysis to optimize classification precision, finally fusing the two-dimensional classification result and the three-dimensional classification result to determine a target commodity, and according to the positioning coordinate, determining a commodity classification result. And controlling the robot to accurately grab the target commodity to finish settlement. According to the method, the accuracy of commodity identification and the precision of spatial positioning in a retail scene are effectively improved, the problem of inaccurate single-dimension classification is effectively solved, the adaptability to a complex retail scene is enhanced, manual intervention is reduced, and the working efficiency of retail business is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart retail technology, and in particular to the field of smart retail technology based on large models and robots. Background Art

[0002] With the development of technology, consumers are increasingly demanding higher standards for their shopping experiences, and traditional retail models are gradually becoming deficient in terms of efficiency and personalized service. Existing technologies, including vending machines, unmanned stores, and robotic arms, have made some progress in improving automation in the retail industry, but they generally suffer from various flaws: The limited variety of vending machines cannot meet diverse needs; while unmanned stores reduce labor costs, there is still room for improvement in product management and customer experience; robotic arms lack operational accuracy in object identification and grasping, nor flexibility in complex environments, and their low level of intelligence and difficulty integrating different technologies result in high overall solution costs and long implementation cycles. These limitations make it difficult for existing technologies to fully adapt to the dynamically changing retail environment and provide efficient and personalized services. Summary of the Invention

[0003] The present application provides a retail method, system, device and medium based on large models and robots to solve one or more technical problems existing in the prior art and at least provide a beneficial option or create conditions.

[0004] In one aspect, the present application provides a retail method based on a large model and a robot, comprising the following steps: Use a depth camera to capture color images and depth information of retail scenes; Use a large model to parse the multimodal information input by users and obtain the attribute information of the target product; Generate a two-dimensional candidate frame using a deep learning model based on the attribute information and the color image, and calculate and obtain a two-dimensional classification result of the candidate object within the two-dimensional candidate frame; In combination with the depth information, the two-dimensional candidate box is mapped to a three-dimensional point cloud coordinate system to obtain a three-dimensional point cloud set, and the positioning coordinates of the candidate object are calculated; Performing segmentation processing and feature analysis on the three-dimensional point cloud set to obtain a three-dimensional classification result of the candidate object; Fusing the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain a final classification result of the candidate object; According to the final classification result, determining whether the candidate object is the target product, and if so, using the location coordinates of the candidate object as the location coordinates of the target product; According to the positioning coordinates of the target product, the robot is controlled to grab the target product and perform settlement.

[0005] Furthermore, the combining of the depth information, mapping the two-dimensional candidate box to a three-dimensional point cloud coordinate system to obtain a three-dimensional point cloud set, and calculating the positioning coordinates of the candidate object includes: Extracting the two-dimensional coordinates of each data point in the two-dimensional candidate frame; Extracting the depth value corresponding to each data point from the depth information as its Z-axis coordinate; According to the two-dimensional coordinates and Z-axis coordinates of each data point, the two-dimensional candidate box is mapped to a three-dimensional point cloud coordinate system to obtain the three-dimensional coordinates of each data point, and the three-dimensional point cloud set is obtained accordingly; The three-dimensional coordinates of the center point of the two-dimensional candidate frame are used as the positioning coordinates of the candidate object.

[0006] Furthermore, the segmentation processing and feature analysis of the three-dimensional point cloud set to obtain the three-dimensional classification result of the candidate object includes: Performing plane segmentation on the three-dimensional point cloud set using a random sampling consensus algorithm, eliminating background plane point clouds in the three-dimensional point cloud set, and obtaining a first segmented point cloud subset of the candidate object; Performing Euclidean clustering segmentation on the first segmented point cloud subset to separate point cloud clusters of independent objects based on a preset spatial distance threshold to obtain a second segmented point cloud subset of the candidate object; Extracting geometric features of the candidate object from the second segmented point cloud subset to generate a multi-dimensional feature vector corresponding to the candidate object; The multidimensional feature vector is input into a pre-trained machine learning model to obtain a three-dimensional classification result of the candidate object.

[0007] Furthermore, fusing the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain the final classification result of the candidate object includes the following steps: Calculating the probability confidence of the two-dimensional classification result and the geometric feature confidence of the three-dimensional classification result; Based on the weighted fusion of the probability confidence and the geometric feature confidence, a classification fusion result is obtained and used as the final classification result of the candidate object.

[0008] Furthermore, the deep learning model includes a Faster R-CNN model.

[0009] Furthermore, the machine learning model includes a support vector machine model.

[0010] On the other hand, the present application provides a retail system based on a large model and a robot, including a data acquisition module, a two-dimensional classification module, a positioning coordinate module, a three-dimensional classification module, a classification fusion module, and a target product grasping module; The data acquisition module is used to use a depth camera to capture color images and depth information of the retail scene, and use a large model to analyze the multimodal information input by the user to obtain attribute information of the target product; The two-dimensional classification module is configured to generate a two-dimensional candidate frame based on the attribute information and the color image using a deep learning model, and calculate a two-dimensional classification result of the candidate object within the two-dimensional candidate frame; The positioning coordinate module is used to map the two-dimensional candidate box to a three-dimensional point cloud coordinate system in combination with the depth information to obtain a three-dimensional point cloud set, and calculate the positioning coordinates of the candidate object; The three-dimensional classification module is used to perform segmentation processing and feature analysis on the three-dimensional point cloud set to obtain a three-dimensional classification result of the candidate object; The classification fusion module is used to fuse the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain the final classification result of the candidate object; The target product grabbing module is used to determine whether the candidate object is the target product based on the final classification result, and if so, use the positioning coordinates of the candidate object as the positioning coordinates of the target product; and control the robot to grab the target product and perform settlement based on the positioning coordinates of the target product.

[0011] On the other hand, the present application provides a retail device based on a large model and a robot, including a robot; the robot includes a product information acquisition device, a depth camera, a moving device, a grasping device, and a computing device; The commodity information acquisition device is used to acquire multimodal information input by the user; The computing device is used to analyze the multimodal information using a large model to obtain attribute information of the target product; The depth camera is used to capture and obtain color images and depth information of the retail scene; The computing device is further configured to generate a two-dimensional candidate frame based on the attribute information and the color image using a deep learning model, and calculate a two-dimensional classification result of the candidate object within the two-dimensional candidate frame; map the two-dimensional candidate frame to a three-dimensional point cloud coordinate system in combination with the depth information to obtain a three-dimensional point cloud set, and calculate the positioning coordinates of the candidate object; perform segmentation processing and feature analysis on the three-dimensional point cloud set to obtain a three-dimensional classification result of the candidate object; fuse the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain a final classification result of the candidate object; and determine whether the candidate object is the target product based on the final classification result, and if so, use the positioning coordinates of the candidate object as the positioning coordinates of the target product; The moving device is used to drive the robot to move to the grasping range of the target product according to the positioning coordinates of the target product; The grabbing device is used to grab the target commodity and perform settlement according to the positioning coordinates of the target commodity.

[0012] Furthermore, the depth camera includes an RGB-D camera.

[0013] On the other hand, the present application provides a computer medium having a processor-executable program stored therein. When the processor-executable program is executed by the processor, it is used to implement the aforementioned large model and robot-based retail method.

[0014] The beneficial effects of the present application are as follows: the present application provides a retail method based on a large model and a robot, which adopts a depth camera to capture and obtain color images and depth information of the retail scene; adopts a large model to parse the multimodal information input by the user to obtain the attribute information of the target product; utilizes a deep learning model to generate a two-dimensional candidate frame and calculate the classification result, and then combines the depth information to map it to a three-dimensional coordinate system to obtain the positioning coordinates, performs three-dimensional segmentation and feature analysis to optimize the classification accuracy, and finally integrates the two-dimensional and three-dimensional classification results to determine the target product. Based on this positioning coordinate, the robot is controlled to accurately grasp the target product to complete the settlement. The present application effectively improves the accuracy of product recognition and the accuracy of spatial positioning in the retail scene, effectively solves the problem of inaccurate classification in a single dimension, and enhances the adaptability to complex retail scenes, reduces manual intervention, and improves the work efficiency and service quality of the retail business. The present application also provides corresponding devices, systems and media. The beneficial effects of the devices, systems and media are similar to those of the method and will not be repeated here.

[0015] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation to the technical solution of the present invention.

[0017] Figure 1 This application provides a flowchart of a retail method based on a large model and a robot; Figure 2 This is a schematic diagram of the principle of obtaining the two-dimensional candidate box and the two-dimensional classification results of the candidate object provided by this application; Figure 3 This is a schematic diagram of the principle of obtaining the three-dimensional classification results of candidate objects provided by this application; Figure 4 This is a structural diagram of the retail system based on large models and robots provided by this application; Figure 5 It is a structural diagram of the robot provided by this application. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0019] The present application is further described below in conjunction with the accompanying drawings and specific embodiments. The described embodiments should not be considered as limiting the present application. All other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0022] With technological advancements and growing consumer demand, traditional retail models are increasingly lacking in efficiency and personalized service. Smart retail systems, leveraging advanced robotics and artificial intelligence algorithms, aim to automate and intelligently manage merchandise and services, thereby improving overall operational efficiency and service quality. To meet consumer demand for a convenient and fast shopping experience, these systems must not only efficiently handle daily transactions but also adapt to complex environments.

[0023] Currently, some technologies are attempting to introduce automation into the retail industry. For example, vending machines provide 24-hour self-service, but their product variety is limited. Unmanned stores achieve self-service shopping through technologies such as RFID tags. Although this reduces labor costs, there is still room for improvement in product management and customer experience. Robotic arms are used for cargo handling or simple sorting, but their operation is mostly limited to specific environments, and their flexibility and accuracy also need to be further improved.

[0024] While these technologies have brought progress to the retail industry, they still have flaws. First, most existing equipment can only perform preset tasks and is unable to cope with complex or changing retail environments. Second, their low level of intelligence and inability to effectively analyze and respond to customer behavior make it difficult to provide a personalized shopping experience. Furthermore, existing machine vision and robotic arm control technologies have limited operational accuracy when it comes to object identification and grasping, making them particularly prone to errors when handling irregularly shaped or similarly packaged goods. Finally, poor integration between different technologies leads to high solution costs and long implementation cycles.

[0025] In response to the problems existing in the related art, the embodiments of the present application provide a retail method, system, device and medium based on a large model and a robot. First, a depth camera is used to capture and obtain color images and depth information of the retail scene; a large model is used to parse the multimodal information input by the user to obtain the attribute information of the target product; and a deep learning model is used to generate a two-dimensional candidate box and its classification results. This method effectively improves the recognition accuracy of the target product, especially in complex backgrounds or when multiple products are mixed. Secondly, the two-dimensional candidate box is mapped to a three-dimensional point cloud coordinate system to obtain the positioning coordinates of the candidate object, so that the system can not only identify the product, but also accurately locate its position in three-dimensional space, which is crucial for the subsequent precise grasping of the robot. In addition, by segmenting and feature analyzing the three-dimensional point cloud set, a more detailed three-dimensional classification result is obtained, and it is fused with the two-dimensional classification result to obtain a more accurate final classification result, effectively solving the problem of inaccurate single-dimensional classification. Based on the target product positioning coordinates obtained by all the above steps, the robot is controlled to accurately grasp the target product and complete the settlement process, effectively improving the automation level of the retail business, reducing manual intervention, and improving work efficiency and service quality. The entire process design of the embodiment of this application fully considers various situations that may arise in retail scenarios, such as product placement and changes in lighting conditions. By integrating two-dimensional and three-dimensional information (color information and depth information), it enhances adaptability to different retail environments. In summary, through a series of carefully designed steps, the embodiment of this application effectively improves the accuracy of product recognition and the precision of grasping operations, while achieving a high degree of automation in retail operations, providing consumers with a more convenient and efficient service experience.

[0026] First, the retail method based on large models and robots provided by the embodiments of the present application will be described in detail with reference to the accompanying drawings.

[0027] Reference Figure 1 The implementation process of the retail method based on large models and robots provided in the embodiment of the present application includes but is not limited to the following steps.

[0028] Step 101: Use a depth camera to capture a color image and depth information of a retail scene, and use a large model to analyze the multimodal information input by the user to obtain attribute information of the target product. In step 101, a large model is used to parse the multimodal information input by the user to obtain the target product's attribute information (such as name, size, and color). A depth camera is used to capture color images and depth information of the retail scene, providing the necessary input for subsequent processing. Color images provide rich visual information, helping to identify product features such as color and shape. Depth information, on the other hand, provides spatial information about the product and its surroundings, which is crucial for accurate identification and positioning. Together, these data form the prerequisite for the system to understand the retail scene and execute subsequent operations.

[0029] In some embodiments of the present application, the large model includes large models such as a Large Language Model (LLM) and a Vision-Language Model (VLM).

[0030] LLMs are artificial intelligence models capable of understanding and generating natural language. These models are typically trained using large amounts of text data, enabling them to perform tasks such as translation, question answering, and text generation. In the context of smart retail systems, LLMs can help process and understand customer inquiries, comments, or instructions, providing a smoother customer service experience.

[0031] VLMs are models that understand the relationship between images and text. These models analyze images and their corresponding descriptions to learn how to associate visual and verbal information. In smart retail environments, this capability can be used for product recognition and classification, as well as assisting visual navigation. For example, they can identify products on a shelf and classify them based on their appearance.

[0032] By combining the capabilities of these two models, smart retail systems can improve automation and service quality through understanding customer behavior, product identification, and classification. For example, a robot can use the visual language model to identify bottles of different colors (such as blue, green, or green tea) and, with the help of the large language model to understand task instructions, accurately grasp and classify items. Such a system can significantly improve the efficiency and accuracy of retail operations.

[0033] Step 102: Generate a two-dimensional candidate frame based on the attribute information and the color image using a deep learning model, and calculate and obtain a two-dimensional classification result of the candidate object within the two-dimensional candidate frame.

[0034] In step 102, based on the acquired attribute information and color image, a deep learning model is used to generate 2D candidate bounding boxes and calculate 2D classification results. This process relies primarily on advanced computer vision algorithms to identify possible target product locations on a 2D plane and perform a preliminary classification of objects within these locations. This step effectively improves recognition efficiency, enabling rapid identification of potential targets in complex backgrounds or with a mix of products, laying the foundation for further 3D spatial analysis.

[0035] In step 103 , the two-dimensional candidate box is mapped to a three-dimensional point cloud coordinate system in combination with the depth information to obtain a three-dimensional point cloud set, and the positioning coordinates of the candidate object are calculated.

[0036] In step 103, the previously obtained 2D candidate bounding box is mapped into a 3D point cloud coordinate system, combining depth information to determine the specific location of the candidate object in 3D space. This conversion process not only enhances the system's spatial perception capabilities but also provides a key basis for the robot's precise grasping. In this way, the system can accurately determine the target product's position relative to itself, ensuring the accuracy of subsequent operations.

[0037] Step 104 : performing segmentation processing and feature analysis on the three-dimensional point cloud set to obtain a three-dimensional classification result of the candidate object.

[0038] In step 104, to further refine the classification results, the 3D point cloud is segmented and feature analyzed to extract key geometric features of the candidate objects. These features are then analyzed to obtain a more detailed 3D classification result. This approach effectively addresses the inaccuracy of single-dimensional classification and improves classification accuracy.

[0039] Step 105 : Fusing the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain the final classification result of the candidate object.

[0040] In step 105, the 2D and 3D classification results from the previous two steps are fused, taking into account factors such as probability confidence and geometric feature confidence, to produce a more reliable final classification result. This multimodal information fusion strategy not only enhances classification accuracy but also better copes with complex and changing real-world application scenarios, making system decisions more robust and reliable.

[0041] Step 106 : Based on the final classification result, determine whether the candidate object is the target product. If so, use the location coordinates of the candidate object as the location coordinates of the target product.

[0042] In step 106, the final classification results determine whether the candidate object is the target product. If it is, its location coordinates are used as the exact location of the target product. This step acts as a screening process, ensuring that only products that meet the requirements are selected for the next step, avoiding the possibility of misoperation and ensuring the efficiency and accuracy of the entire process.

[0043] Step 107: According to the positioning coordinates of the target product, the robot is controlled to grab the target product and perform settlement.

[0044] In step 107, based on the location coordinates of the target product, the system controls the robot to perform precise grabbing and complete the checkout process. This automated operation not only effectively reduces manual intervention, improves work efficiency and service quality, but also provides consumers with a more convenient and efficient shopping experience.

[0045] In some embodiments of the present application, the deep learning model includes the Faster R-CNN model. This is a highly efficient object detection algorithm widely used in the field of computer vision. Through its unique architectural design, the Faster R-CNN model efficiently processes image data and generates high-quality object candidate regions, providing strong support for product identification and positioning in smart retail systems.

[0046] Specifically, the Faster R-CNN model consists of four key components. First, there's the feature extraction network (Backbone), which typically uses a pre-trained convolutional neural network like ResNet or VGG to extract rich feature representations from the input image. These feature maps not only capture object shape information but also preserve important texture details, providing a solid foundation for subsequent tasks. Second, there's the Region Proposal Network (RPN), a lightweight fully convolutional network that operates directly on the feature map. Using a sliding window, it scans the entire feature map and predicts regions (i.e., candidate boxes) where the target object may reside. For each location, it generates multiple candidate boxes of varying scales and aspect ratios, effectively improving the speed and efficiency of object detection. Next comes the Region of Interest Pooling (ROI) layer, which converts candidate boxes of varying sizes into fixed-size feature vectors for subsequent classifier and regressor processing. This ensures uniform processing of candidate boxes, regardless of their original size, effectively simplifying the design of subsequent steps. Finally, there's the classifier and bounding box regressor. These two modules are responsible for determining the object category within a candidate box and adjusting the box's position and size to more accurately enclose the target object. The classifier outputs a probability distribution that each candidate box belongs to a specific category, while the bounding box regressor fine-tunes the box's position to achieve more accurate positioning. Through the collaborative work of these four components, the Faster R-CNN model can efficiently and accurately complete object detection tasks.

[0047] In the embodiments of the present application, the application of the Faster R-CNN model is mainly reflected in improving the accuracy of commodity recognition. By utilizing its powerful target detection capabilities, the model can quickly and accurately identify target commodities from complex retail scenarios, and can perform well even in the presence of a mixture of multiple commodities. It effectively improves the accuracy of commodity recognition and positioning in smart retail scenarios, and lays a technical foundation for achieving highly automated retail services. In addition, this solution has good scalability and flexibility, and the model structure and parameter settings can be adjusted according to actual needs to adapt to different application scenarios. Through a series of precisely designed steps, the entire system realizes intelligent management of the entire process from commodity information acquisition to automatic crawling, providing consumers with a more convenient and efficient shopping experience.

[0048] In some embodiments of the present application, in step 103, the two-dimensional candidate box is mapped to a three-dimensional point cloud coordinate system in combination with depth information to obtain a three-dimensional point cloud set, and the implementation process of calculating the positioning coordinates of the candidate object includes but is not limited to the following steps.

[0049] Step 201: extract the two-dimensional coordinates of each data point in the two-dimensional candidate frame.

[0050] In step 201, by extracting the 2D coordinates (X, Y) of each data point within the 2D candidate frame, the system can determine the position of each possible object on the image plane. These coordinates provide preliminary spatial positioning information. Although limited to the 2D plane, they serve as an important starting point for subsequent 3D mapping. This process ensures that key information obtained from the image recognition stage is accurately transferred to the 3D processing stage, laying the foundation for further integration of depth information.

[0051] Step 202: extract the depth value corresponding to each data point from the depth information as its Z-axis coordinate.

[0052] In step 202, depth information (Z-axis coordinate) is added to each 2D coordinate point using depth information. The depth value represents the distance of the object from the camera, which is crucial for building a 3D model. By matching each point in the 2D image with the corresponding depth information, the system can obtain the precise position of each data point in 3D space. This not only enriches the data dimension but also renders the originally two-dimensional information three-dimensional, providing the necessary depth information support for subsequent 3D mapping and positioning.

[0053] In step 203 , the two-dimensional candidate box is mapped to a three-dimensional point cloud coordinate system according to the two-dimensional coordinates and the Z-axis coordinates of each data point, and the three-dimensional coordinates of each data point are obtained, thereby obtaining a three-dimensional point cloud set.

[0054] In step 203, based on the 2D coordinates (X, Y) and depth values (Z) obtained in the previous step, the system maps all data points within the 2D candidate box into a 3D point cloud coordinate system, forming a 3D point cloud collection containing all necessary spatial information. This process not only enhances the system's spatial perception capabilities but also provides a key basis for the robot arm's precise grasping operations. In this way, the system can more accurately understand the positional relationship of the target product in the actual environment, ensuring the accuracy of subsequent operations.

[0055] Step 204 : The three-dimensional coordinates of the center point of the two-dimensional candidate box are used as the positioning coordinates of the candidate object.

[0056] In step 204, by analyzing the center point of the 2D candidate box and its corresponding 2D coordinates (X, Y) and depth value (Z), the system calculates the precise position of the candidate object in 3D space. This approach not only simplifies the positioning process but also improves positioning accuracy. The resulting positioning coordinates directly guide the robot's grasping action, ensuring it can accurately locate and grasp the target product. This process is a core step in the entire visual grasping process, ensuring the efficiency and reliability of automated operations.

[0057] In some embodiments of the present application, a color image and depth information of a retail scene are acquired by an RGB-D camera. Assume that the optical center coordinates of the RGB-D camera are , the focal length on the X axis is , the focal length on the Y axis is ; For any data point in the two-dimensional candidate box , extract its two-dimensional coordinates , extracting data points from depth information Corresponding depth value As its Z-axis coordinate. Map to the 3D point cloud coordinate system to get the data points The three-dimensional coordinates of ,in , , .

[0058] In some embodiments of the present application, in step 104 , the process of performing segmentation processing and feature analysis on the three-dimensional point cloud set to obtain a three-dimensional classification result of the candidate object includes but is not limited to the following steps.

[0059] Step 301 : Using a random sampling consensus algorithm, perform plane segmentation on the three-dimensional point cloud set, remove the background plane point cloud in the three-dimensional point cloud set, and obtain a first segmented point cloud subset of the candidate object.

[0060] In step 301, the Random Sample Consensus (RANSAC) algorithm is used to remove background plane point clouds from the 3D point cloud collection, thereby obtaining the first segmented point cloud subset of the candidate object. The RANSAC algorithm effectively identifies and isolates background planes, such as the floor or shelf surfaces, which occupy a large portion of the space. This background information is often irrelevant or even disruptive to the identification and classification of the target object. This step allows the system to focus on processing the point cloud data truly relevant to the target object, effectively improving the efficiency and accuracy of subsequent processing steps.

[0061] Step 302 : Perform Euclidean clustering segmentation on the first segmented point cloud subset, separate point cloud clusters of independent objects based on a preset spatial distance threshold, and obtain a second segmented point cloud subset of candidate objects.

[0062] In step 302, after obtaining the first segmented point cloud subset after background removal, Euclidean clustering is further employed to separate point cloud clusters of independent objects based on a preset spatial distance threshold, forming a second segmented point cloud subset. Euclidean clustering groups points based on their distance, classifying closely connected points into the same category, thereby distinguishing different independent objects. This method is particularly suitable for scenes containing multiple objects that are close to each other but not touching, ensuring that each independent object can be individually identified, providing clear object boundaries for further feature analysis.

[0063] Step 303 : extracting geometric features of the candidate object from the second segmented point cloud subset, and generating a multi-dimensional feature vector corresponding to the candidate object.

[0064] In step 303, key geometric features of the candidate objects are extracted from the second segmented point cloud subset, and corresponding multidimensional feature vectors are generated. Geometric features may include, but are not limited to, physical properties such as volume, surface area, and aspect ratio. These features not only describe the basic shape of the object but also reflect its uniqueness in three-dimensional space. By extracting these features and converting them into multidimensional feature vectors that are easily processed by machine learning models, the system can better understand and distinguish different types of objects, providing strong data support for the final classification decision.

[0065] In step 304 , the multi-dimensional feature vector is input into a pre-trained machine learning model to obtain a three-dimensional classification result of the candidate object.

[0066] In step 304, the multidimensional feature vector generated in the previous step is input into a pre-trained machine learning model to obtain a three-dimensional classification result for the candidate object. This model is typically trained on a large amount of labeled data and has good generalization capabilities and classification accuracy. In this way, the system can accurately determine the category of the candidate object based on its specific features, and can correctly classify even complex or highly similar objects. This process fully utilizes advanced machine learning technology, enabling the entire system to automatically complete three-dimensional classification tasks with high accuracy, effectively improving the intelligence level and work efficiency of the smart retail system.

[0067] In some embodiments of the present application, the machine learning model includes a support vector machine (SVM) model. SVM is a supervised learning method widely used in classification and regression analysis, particularly for data classification tasks in high-dimensional spaces. Its core concept is to accurately classify new data points by finding an optimal hyperplane that maximizes the gap between different classes.

[0068] This hyperplane not only needs to correctly divide the training samples, but also needs to be as far away from the nearest data points (i.e., support vectors) as possible to ensure that the model has good generalization capabilities. In addition, for data sets that cannot be directly separated by a linear hyperplane, SVM can use kernel functions (such as radial basis functions (RBFs) and polynomial kernels) to map the original feature space to a higher-dimensional space and find a suitable segmentation plane in this space.

[0069] In an embodiment of the present application, an SVM model is used to process multidimensional feature vectors extracted from a three-dimensional point cloud set. These feature vectors represent various geometric properties of candidate objects. By leveraging the powerful classification capabilities of SVM, the system can more accurately distinguish between different product categories, and even products with similar shapes or similar packaging can be correctly identified and classified. This method effectively improves classification accuracy and can adapt to complex retail environments, such as lighting changes or partial occlusion. SVM improves the adaptability and reliability of the system by optimizing solutions in high-dimensional space.

[0070] The SVM model also plays a crucial role in optimizing decision-making. By analyzing multidimensional feature vectors, it generates high-quality classification results, providing a solid foundation for subsequent operations. For example, after identifying a target product, the system can use the SVM classification results to guide the robot to perform precise grasping actions, reducing the possibility of misoperation and improving overall work efficiency and service quality. This automated operation based on precise classification results effectively enhances the system's intelligence and user experience.

[0071] In summary, the use of the support vector machine model in the embodiments of this application not only effectively improves the accuracy of product identification and classification, but also enhances the system's adaptability and robustness to complex environments. The efficient performance and good scalability of the SVM model enable the smart retail system to not only handle current task requirements, but also continuously optimize and upgrade as technology advances and market demand changes. By combining advanced computer vision technology and machine learning algorithms, the entire smart retail system achieves intelligent management of the entire process from product information acquisition to automatic crawling, providing consumers with a more convenient and efficient shopping experience.

[0072] In some embodiments of the present application, the implementation process of the random sampling consensus algorithm includes but is not limited to the following steps.

[0073] First, assume that the background plane point cloud is the internal point and the point cloud corresponding to the candidate object is the external point, and build a plane model. A 3D point cloud collection of point clouds Randomly select three non-collinear points 、 and , assuming the plane parameters are 、 、 and , where 、 and are the normal vectors of the plane, which define the direction of the plane; and is a parameter related to the distance from the origin to the plane, which can be understood as the signed distance from the origin of the coordinate system to the plane. This distance is measured in the direction of the plane normal vector. The plane equation satisfies the following formula (1): (1); Then, by vector cross product ,in , , get the plane normal vector , and normalized.

[0074] Secondly, calculate each point in the 3D point cloud set Euclidean distance to the plane , satisfying the following formula (2): (2); like ( is a preset threshold, such as ), then the point Determined as an internal point (i.e., background plane point cloud); Furthermore, calculate the internal point ratio of the current plane model , is the current number of internal points; iterate repeatedly until the maximum number of iterations is found or When the preset threshold is reached, the plane model with the largest number of inliers is retained as the optimal solution, and the background plane point cloud is eliminated to obtain the first segmentation point cloud subset of the candidate object. .

[0075] In some embodiments of the present application, the implementation process of the Euclidean clustering segmentation method includes but is not limited to the following steps.

[0076] First, for the first segmented point cloud subset , build a KD tree (K-dimension Tree) to accelerate the search of neighborhood points.

[0077] Then, calculate Any two points and Euclidean distance between , satisfying the following formula (3): (3); like , is the preset spatial distance threshold (such as ), then and Classified into the same category.

[0078] Finally, traverse All points, according to Separate independent objects and output point cloud cluster sets , where each cluster corresponds to a candidate object, and the second segmented point cloud subset of the candidate object is obtained.

[0079] In some embodiments of the present application, the training process of the SVM model includes but is not limited to the following steps.

[0080] First, obtain the training set; during the training process of the SVM model, the number of candidate objects given for training is , the multidimensional feature vector corresponding to each given candidate object

[0081] Then, SVM finds the optimal classification hyperplane by minimizing the following loss function to ensure high generalization ability for objects in different postures. The following formula (4) is satisfied: (4); In formula (4), is the normal vector of the hyperplane (i.e., weight vector); is the bias term; is the slack variable, , allowing a small amount of misclassification; is a penalty parameter used to control the balance between classification error and interval; for any multidimensional feature vector , loss function The constraints satisfy the following formula (5): (5); In formula (5), is a multidimensional feature vector Corresponding classification labels; is a kernel function used to map data into a high-dimensional space to handle nonlinear separable problems; in the embodiment of the present application, the kernel function adopts a Gaussian kernel, and the multidimensional feature vector and The kernel function value between The following formula (6) is satisfied: (6); In formula (6), Indicates that the width of the kernel function is controlled, that is, it determines the two sample points and How the similarity in feature space changes as the distance between them changes.

[0082] Furthermore, for multi-classification tasks (such as multiple item classifications), a "one-vs-rest" strategy is adopted. Assume that the number of item classifications is , for each category , train a two-class SVM model, and the final classification result is the category with the highest prediction score, and the three-dimensional classification result of the candidate object is obtained , satisfying the following formula (7): (7); In formula (7), Representation category The corresponding bias term, Representation category The corresponding weight vector.

[0083] Finally, the SVM model trained through the above steps is obtained.

[0084] In some embodiments of the present application, given a trained SVM model, for a newly entered candidate object, the second segmented point cloud subset of each candidate object is Extract geometric features such as dimensions (i.e., the length, width, and height of the bounding box), shape (such as surface curvature, main axis direction, sphericity), and color (such as RGB mean) to generate a multidimensional feature vector corresponding to the candidate object , and calculate its distance to the hyperplane And predict the category, satisfying the following formula (8): (8); In formula (8), express The corresponding bias term determines the position of the hyperplane; express The corresponding weight vector determines the direction of the hyperplane; represents a symbolic function, if It means that the candidate object belongs to the positive class (that is, the candidate object belongs to the target product), otherwise it belongs to the negative class (that is, the candidate object does not belong to the target product).

[0085] In some embodiments of the present application, in step 105 , the process of fusing the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain the final classification result of the candidate object includes but is not limited to the following steps.

[0086] Step 401 : Calculate the probability confidence of the two-dimensional classification result and the geometric feature confidence of the three-dimensional classification result.

[0087] In step 401, the probability confidence score of the 2D classification result is first calculated. This is the probability distribution generated when classifying each candidate object in the image based on a deep learning model (such as Faster R-CNN). A high probability confidence score indicates a high degree of confidence in the classification result. Simultaneously, the system also calculates the geometric feature confidence score of the 3D classification result. This score determines the degree of match between the geometric features (such as shape and volume) extracted from the 3D point cloud data and the known categories. The geometric feature confidence score reflects the confidence level in determining that the object belongs to a certain category based on the 3D features. These two confidence scores provide a quantitative basis for the subsequent fusion processing, ensuring the reliability and accuracy of the classification results.

[0088] Step 402 : Based on the weighted fusion of the probability confidence and the geometric feature confidence, a classification fusion result is obtained, and the result is used as the final classification result of the candidate object.

[0089] In step 402, the probability confidence of the two-dimensional classification result and the geometric feature confidence of the three-dimensional classification result are weighted and fused to generate a comprehensive classification fusion result, which is used as the final classification result of the candidate object. Through reasonable weight distribution, the importance of two-dimensional image information and three-dimensional geometric information can be balanced, so that the final classification result takes into account both visual recognition accuracy and the consistency of actual physical features. For example, in some cases, two-dimensional images may be inaccurately classified due to occlusion or lighting problems, while three-dimensional geometric features can provide additional support; and vice versa. This multimodal information fusion strategy not only improves the accuracy of classification, but also better copes with complex and changeable actual application scenarios, making system decisions more robust and reliable. Ultimately, this fusion process ensures that the system can output the most credible classification results, providing a solid foundation for subsequent operations.

[0090] Secondly, an embodiment of the present application provides a retail system based on a large model and a robot, including a data acquisition module, a two-dimensional classification module, a positioning coordinate module, a three-dimensional classification module, a classification fusion module and a target product grabbing module.

[0091] The data acquisition module uses a depth camera to capture color images and depth information of the retail scene, and uses a large model to analyze the multimodal information input by the user to obtain the attribute information of the target product; The two-dimensional classification module is used to generate a two-dimensional candidate box based on attribute information and color images using a deep learning model, and calculate the two-dimensional classification results of the candidate objects within the two-dimensional candidate box.

[0092] The positioning coordinate module is used to combine depth information, map the two-dimensional candidate box to the three-dimensional point cloud coordinate system, obtain a three-dimensional point cloud set, and calculate the positioning coordinates of the candidate object.

[0093] The 3D classification module is used to perform segmentation processing and feature analysis on the 3D point cloud set to obtain the 3D classification results of the candidate objects.

[0094] The classification fusion module is used to fuse the two-dimensional classification results and the three-dimensional classification results of the candidate object to obtain the final classification result of the candidate object.

[0095] The target product grabbing module is used to determine whether a candidate object is the target product based on the final classification results. If so, the candidate object's location coordinates are used as the target product's location coordinates. Based on the target product's location coordinates, the robot is controlled to grab the target product and perform settlement.

[0096] Furthermore, an embodiment of the present application provides a retail device based on a large model and a robot, including a robot comprising a product information acquisition device, a depth camera, a moving device, a grasping device, and a computing device.

[0097] The commodity information acquisition device is used to acquire multimodal information input by the user; The computing device is used to use a large model to analyze multimodal information and obtain attribute information of the target product; The depth camera is used to capture color images and depth information of the retail scene; The computing device is further configured to generate a two-dimensional candidate frame using a deep learning model based on the attribute information and the color image, and calculate a two-dimensional classification result of the candidate object within the two-dimensional candidate frame; map the two-dimensional candidate frame to a three-dimensional point cloud coordinate system in combination with the depth information to obtain a three-dimensional point cloud set, and calculate the positioning coordinates of the candidate object; perform segmentation processing and feature analysis on the three-dimensional point cloud set to obtain a three-dimensional classification result of the candidate object; fuse the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain a final classification result of the candidate object; and determine whether the candidate object is a target product based on the final classification result, and if so, use the positioning coordinates of the candidate object as the positioning coordinates of the target product; The moving device is used to drive the robot to move to the grasping range of the target product according to the positioning coordinates of the target product.

[0098] The grabbing device is used to grab the target product and perform settlement according to the positioning coordinates of the target product.

[0099] In some embodiments of the present application, the product information acquisition device includes a voice interaction device and a text interaction device. The product information acquisition device parses the user's multimodal input data (including voice commands or text input) to extract key semantic features of the target product, performs feature matching queries against a pre-set product database, and generates corresponding structured attribute information (such as product name, category, and specifications).

[0100] Voice interaction devices allow users to communicate directly with the system using natural language, quickly and accurately expressing their needs. For example, customers can convey information about the product they wish to purchase to the system through voice commands such as "I want to buy a bottle of green tea" or "Please help me find a bag of coffee beans." The system's built-in voice recognition technology converts these voice commands into text and further analyzes them to extract key product attributes, such as product name (green tea, coffee beans) and category (beverage, food). Furthermore, advanced natural language processing (NLP) technology can understand user intent and provide more personalized and precise service. For example, if a user mentions "the green tea I bought last time," the system can identify the specific brand and specifications based on the user's purchase history. This type of interaction is particularly suitable for users who want to complete shopping tasks quickly or who are not comfortable manually operating the device.

[0101] Text interaction devices provide another flexible way to input information, especially for users who prefer to communicate through writing or typing. Users can enter product-related information or keywords through mobile applications, in-store terminals, or other digital interfaces. For example, users can enter "organic milk" or "sugar-free cola" in the search bar. After receiving these text inputs, the system can quickly parse and match them to the corresponding product database to obtain detailed attribute information, such as brand, specifications, price, etc. This method is not only convenient and fast, but also allows users to describe their needs more accurately. In addition, text interaction also supports multilingual input, allowing users from different language backgrounds to easily use the system.

[0102] In some embodiments of the present application, the voice interaction device includes a speaker and a microphone array.

[0103] The speaker provides user feedback. It not only plays preset prompts or voice messages but also generates and plays real-time responses based on user queries. High-quality speakers ensure clear, distortion-free sound, allowing users to clearly hear system feedback even in noisy retail environments. Furthermore, the speaker supports multiple languages, making the system more inclusive and international. It can also provide personalized feedback, such as product recommendations or special offers, further enhancing the user experience.

[0104] A microphone array is another key component in voice interaction devices, responsible for capturing user voice commands. Compared to a single microphone, a microphone array offers superior directionality and noise reduction, enabling accurate recognition of user voices in complex environments while filtering out background noise. This design enables the system to operate stably in a variety of practical scenarios, maintaining excellent recognition performance in both crowded supermarkets and quiet specialty stores. By using multiple microphones working together, the microphone array accurately determines the direction of the sound source, focusing on the specific user's voice commands while ignoring distracting sounds from other directions.

[0105] In smart retail scenarios, the combined use of speakers and microphone arrays effectively enhances the user experience and increases the intelligence of the system. For example, upon entering a store, customers can interact with the system through simple voice commands. The system captures these commands through the microphone array, recognizes and processes the voice, and provides instant feedback through the speakers. This seamless interaction not only simplifies the shopping process but also provides customers with a more intuitive and user-friendly shopping experience. Furthermore, voice interaction devices can be integrated with other smart devices to form a multimodal interactive system, further enriching the user's shopping experience and enhancing the automation and intelligence of retail operations.

[0106] In some embodiments of the present application, a computing device has a built-in multimodal large model module. Its core function is to fuse and process input data from different sources and types, including image, text, and voice data. By integrating these different types of data, the system can provide richer and more comprehensive product descriptions, thereby improving the accuracy of recognition and classification. For example, image data captures the appearance characteristics of the product, text data provides detailed descriptions such as brand and specifications, and voice data provides users with a natural and intuitive way to interact.

[0107] The multimodal large model module possesses powerful feature extraction capabilities, capable of extracting high-dimensional feature representations from a variety of data sources. This not only helps improve product recognition accuracy but also enables a better understanding and prediction of user needs. This module maps data from different modalities into a common feature space, allowing information from different modalities to complement and enhance each other. For example, combining color information in an image with brand descriptions in text enables more accurate product identification. Leveraging advanced deep learning algorithms (such as the Transformer architecture), the multimodal large model can be trained on large datasets, learning more complex patterns and relationships, thereby improving overall performance.

[0108] In some embodiments of this application, the depth camera includes an RGB-D camera (Red-Green-Blue and Depth Camera), which can simultaneously capture color images and depth information, providing rich visual data for smart retail devices. RGB-D cameras not only provide high-resolution color images but also generate precise depth maps, enabling the system to accurately identify and locate products in three-dimensional space.

[0109] The RGB portion of an RGB-D camera is used to capture high-quality color images. These images contain rich visual information such as the product's color, texture, and shape. Using high-resolution cameras, RGB-D cameras can produce clear color images, helping systems identify specific product features. For example, in a grocery store, RGB images can help systems distinguish packaging designs from different brands or identify the color and ripeness of fruit. These details are crucial for accurate product identification, especially when a variety of products are mixed.

[0110] The D portion of the RGB-D camera is responsible for capturing depth information in the scene, specifically the distance from each pixel to the camera. This depth perception capability enables the system to construct a three-dimensional model of the scene and accurately calculate the position and posture of objects in three-dimensional space. Depth information is particularly important for understanding the specific placement of products on shelves, determining the relative distances between products, and planning the movement path of the robotic arm. For example, when a specific item needs to be grabbed from a shelf, depth information can help the system determine its exact location, thereby guiding the robotic arm for precise operation.

[0111] A key feature of RGB-D cameras is their ability to simultaneously capture and integrate color images and depth information. This means that the color image and the corresponding depth map captured at the same moment are perfectly matched, greatly facilitating subsequent processing and analysis. The system can leverage this synchronized data for more complex tasks, such as 3D reconstruction, object detection, and tracking. For example, by jointly analyzing the RGB image and depth map of the same scene, the system can more accurately identify and locate target products, maintaining high recognition accuracy even in complex or dynamically changing environments.

[0112] In smart retail scenarios, the use of RGB-D cameras effectively enhances the system's intelligence. By combining color images and depth information, the system can not only identify the appearance of products but also understand their layout and distribution in three-dimensional space. As an advanced depth camera, RGB-D cameras can simultaneously capture color images and depth information, providing powerful visual support for smart retail systems. Their application not only improves the accuracy of product identification and positioning, but also enhances the system's intelligence and user experience.

[0113] In some embodiments of the present application, the mobile device includes an omnidirectional mobile platform, which provides the robot with excellent maneuverability and flexibility. The omnidirectional mobile platform enables the robot to move efficiently and flexibly in complex retail environments, allowing it to more accurately and quickly reach the location of target products for grabbing and checkout operations.

[0114] The omnidirectional chassis design allows the robot to move in any direction without rotating the body. This is achieved primarily through a special wheel layout and drive mechanism, such as Mecanum wheels or omnidirectional wheels. This unique mobility gives the robot extreme flexibility, enabling it to move freely in tight spaces and easily navigate around obstacles. This omnidirectional chassis capability is particularly important in crowded shelves or areas with dense customer traffic.

[0115] Thanks to the omnidirectional mobile chassis, robots can navigate retail environments more efficiently. Based on the calculated coordinates of the target product, they can quickly adjust their direction and move in a straight line to the target location without complex steering maneuvers. This not only reduces the robot's movement time and energy consumption, but also improves the overall system's responsiveness and service efficiency. The advantages of the omnidirectional mobile chassis are particularly evident in dynamic environments where frequent position changes are required, such as during promotional events.

[0116] The omnidirectional mobile chassis is designed with the diverse and complex retail environment in mind. Whether navigating closely packed shelves, unexpected pedestrians, or uneven surfaces, it provides stable and reliable mobility. Furthermore, this type of chassis can be equipped with advanced sensors and obstacle avoidance systems to further enhance safety, ensuring the robot avoids collisions with its surroundings while performing its tasks.

[0117] The use of omnidirectional mobile chassis not only improves robot operational efficiency but also enhances the user shopping experience. For example, in a large supermarket or shopping mall, users can select products through a smart terminal, and then a robot equipped with an omnidirectional mobile chassis quickly and accurately delivers the products to them. This service not only saves users time but also provides customers with a novel and convenient shopping experience, helping to improve customer satisfaction and loyalty.

[0118] In some embodiments of the present application, the grasping device includes a robotic arm equipped with a gripper. Based on the positioning coordinates of the target product, the robotic arm is driven by inverse kinematics solution and path planning to adaptively grasp the target product.

[0119] Inverse Kinematics (IK) is used in smart retail devices to calculate the required angles or displacements of each joint in a robotic arm based on the three-dimensional coordinates of the target item, enabling the end arm to accurately reach the target location. This process takes into account the physical structure and motion limitations of the robotic arm, such as the maximum rotation angles of the joints and the length of the connecting rods. This ensures the accuracy of the grasping action and avoids collisions and other unexpected situations during the movement of the robotic arm. Through this precise calculation, the robotic arm can complete grasping tasks efficiently and safely.

[0120] Path planning further optimizes the robot's motion path based on inverse kinematics, ensuring it takes the optimal route from its initial to its target position. This process comprehensively considers factors such as obstacle avoidance, maximum efficiency, and smooth transitions. For example, the A* search algorithm or the RRT random tree algorithm can be used to generate the optimal path in real time and dynamically adjust to environmental changes. Path planning not only ensures the robot's safe operation but also improves overall efficiency, ensuring a smooth trajectory and minimizing unnecessary stops and sharp turns, thereby extending device life and enhancing the user experience.

[0121] Adaptive grasping means that the robot arm automatically adjusts its grasping strategy based on actual conditions to accommodate different product characteristics and placement methods. This includes aspects such as shape adaptability, force control, and dynamic adjustment. For example, for irregularly shaped products, the robot arm can ensure stable grasping by equipping it with flexible grippers or multi-point contact technology. Built-in force sensors can sense and adjust the grasping force applied to the product to prevent damage or dropping. If the actual position is found to be inconsistent with the expected position during the grasping process, the robot arm can adjust its motion trajectory in real time to ensure successful completion of the task. Adaptive grasping effectively improves the flexibility and reliability of the system, enabling it to operate stably in complex and changing real-world environments.

[0122] By combining inverse kinematics, path planning, and adaptive gripping, the smart retail device automates the entire process, from target product identification to final grasping. These highly integrated technical solutions not only effectively improve the automation level of retail operations and reduce manual intervention, but also provide consumers with a more convenient and efficient shopping experience. Furthermore, they also bring higher operational efficiency and service quality to retailers, helping to achieve true smart retail. The entire system, through precisely designed steps, ensures the accuracy and efficiency of the robotic arm's movements, making retail operations more intelligent and flexible.

[0123] In addition, an embodiment of the present application provides a computer medium storing a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to implement the aforementioned retail method based on large models and robots.

[0124] In summary, the retail method, system, device, and medium based on large models and robots provided in the embodiments of the present application have the following technical effects.

[0125] The embodiments of the present application achieve efficient product identification and positioning by acquiring attribute information, color images, and depth information of the target product, and using a deep learning model to generate a two-dimensional candidate box and its classification results. This method further combines the color image and depth information synchronously acquired by the RGB-D camera to map the two-dimensional candidate box to a three-dimensional point cloud coordinate system, thereby accurately determining the position of the candidate object in three-dimensional space. In addition, by segmenting and analyzing the three-dimensional point cloud set, extracting key geometric features, and using machine learning models such as support vector machines (SVM) for classification, the two-dimensional and three-dimensional classification results are finally integrated to obtain a more reliable classification decision. The application of this series of technologies effectively improves the accuracy of product identification and positioning, and enhances the system's ability to cope with complex scenarios.

[0126] In addition, the robot provided in the embodiment of the present application integrates a product information acquisition device, a depth camera, a mobile device, a grasping device and a computing device. The product information acquisition device collects user needs through voice interaction and text interaction, and uses an RGB-D camera to capture color images and depth information of the retail scene. The computing device is responsible for generating two-dimensional candidate frames from these data, calculating positioning coordinates, performing three-dimensional classification, and fusing classification results. The omnidirectional mobile chassis enables the robot to move flexibly in a complex retail environment, while the adaptive grasping manipulator completes the precise grasping task according to the positioning coordinates of the target product. This highly integrated design not only improves the intelligence level and service efficiency of the system, but also provides consumers with a convenient and efficient shopping experience.

[0127] In some optional embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation schematic diagram. For example, depending on the functions / operations involved, the two boxes shown in succession may actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flow chart of the present application are provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.

[0128] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, a person skilled in the art will be able to implement the present application as set forth in the claims using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.

[0129] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several programs that enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0130] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable programs for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, a program execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can retrieve and execute a program from a program execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, a program execution system, apparatus, or device.

[0131] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting, or processing it in a suitable manner as necessary, and then storing it in a computer memory.

[0132] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0133] In the above description of this specification, reference to the terms "one embodiment / implementation," "another embodiment / implementation," or "certain embodiments / implementations" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in the embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0134] Although the embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

[0135] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A retail method based on large models and robots, characterized by: The following steps are involved: Use a depth camera to capture color images and depth information of retail scenes; Use a large model to parse the multimodal information input by users and obtain the attribute information of the target product; Generate a two-dimensional candidate frame using a deep learning model based on the attribute information and the color image, and calculate and obtain a two-dimensional classification result of the candidate object within the two-dimensional candidate frame; In combination with the depth information, the two-dimensional candidate box is mapped to a three-dimensional point cloud coordinate system to obtain a three-dimensional point cloud set, and the positioning coordinates of the candidate object are calculated; Performing segmentation processing and feature analysis on the three-dimensional point cloud set to obtain a three-dimensional classification result of the candidate object; Fusing the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain a final classification result of the candidate object; According to the final classification result, determining whether the candidate object is the target product, and if so, using the location coordinates of the candidate object as the location coordinates of the target product; According to the positioning coordinates of the target product, the robot is controlled to grab the target product and perform settlement.

2. The retail method based on large models and robots according to claim 1, characterized in that: The combining of the depth information, mapping the two-dimensional candidate box to a three-dimensional point cloud coordinate system to obtain a three-dimensional point cloud set, and calculating the positioning coordinates of the candidate object includes: Extracting the two-dimensional coordinates of each data point in the two-dimensional candidate frame; Extracting the depth value corresponding to each data point from the depth information as its Z-axis coordinate; According to the two-dimensional coordinates and Z-axis coordinates of each data point, the two-dimensional candidate box is mapped to a three-dimensional point cloud coordinate system to obtain the three-dimensional coordinates of each data point, and thereby obtain the three-dimensional point cloud set; The three-dimensional coordinates of the center point of the two-dimensional candidate frame are used as the positioning coordinates of the candidate object.

3. The retail method based on large models and robots according to claim 1, characterized in that: The performing segmentation processing and feature analysis on the three-dimensional point cloud set to obtain the three-dimensional classification result of the candidate object includes: Performing plane segmentation on the three-dimensional point cloud set using a random sampling consensus algorithm, eliminating background plane point clouds in the three-dimensional point cloud set, and obtaining a first segmented point cloud subset of the candidate object; Performing Euclidean clustering segmentation on the first segmented point cloud subset to separate point cloud clusters of independent objects based on a preset spatial distance threshold to obtain a second segmented point cloud subset of the candidate object; Extracting geometric features of the candidate object from the second segmented point cloud subset to generate a multi-dimensional feature vector corresponding to the candidate object; The multidimensional feature vector is input into a pre-trained machine learning model to obtain a three-dimensional classification result of the candidate object.

4. The retail method based on large models and robots according to claim 1, characterized in that: The fusing of the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain the final classification result of the candidate object includes the following steps: Calculating the probability confidence of the two-dimensional classification result and the geometric feature confidence of the three-dimensional classification result; Based on the weighted fusion of the probability confidence and the geometric feature confidence, a classification fusion result is obtained and used as the final classification result of the candidate object.

5. The retail method based on large models and robots according to claim 1, characterized in that: The deep learning model includes a Faster R-CNN model.

6. The retail method based on large models and robots according to claim 3, characterized in that: The machine learning model includes a support vector machine model.

7. Retail system based on large models and robots, characterized by: It includes data acquisition module, two-dimensional classification module, positioning coordinate module, three-dimensional classification module, classification fusion module and target product grabbing module; The data acquisition module is used to use a depth camera to capture color images and depth information of the retail scene, and use a large model to analyze the multimodal information input by the user to obtain attribute information of the target product; The two-dimensional classification module is configured to generate a two-dimensional candidate frame based on the attribute information and the color image using a deep learning model, and calculate a two-dimensional classification result of the candidate object within the two-dimensional candidate frame; The positioning coordinate module is used to map the two-dimensional candidate box to a three-dimensional point cloud coordinate system in combination with the depth information to obtain a three-dimensional point cloud set, and calculate the positioning coordinates of the candidate object; The three-dimensional classification module is used to perform segmentation processing and feature analysis on the three-dimensional point cloud set to obtain a three-dimensional classification result of the candidate object; The classification fusion module is used to fuse the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain the final classification result of the candidate object; The target product grabbing module is used to determine whether the candidate object is the target product based on the final classification result, and if so, use the positioning coordinates of the candidate object as the positioning coordinates of the target product; and control the robot to grab the target product and perform settlement based on the positioning coordinates of the target product.

8. Retail installation based on large models and robots, characterized by: The robot comprises a product information acquisition device, a depth camera, a moving device, a grasping device and a computing device; The commodity information acquisition device is used to acquire multimodal information input by the user; The computing device is used to analyze the multimodal information using a large model to obtain attribute information of the target product; The depth camera is used to capture and obtain color images and depth information of the retail scene; The computing device is further configured to generate a two-dimensional candidate frame using a deep learning model based on the attribute information and the color image, and calculate a two-dimensional classification result of the candidate object within the two-dimensional candidate frame; map the two-dimensional candidate frame to a three-dimensional point cloud coordinate system in combination with the depth information to obtain a three-dimensional point cloud set, and calculate the positioning coordinates of the candidate object; and perform segmentation processing and feature analysis on the three-dimensional point cloud set to obtain a three-dimensional classification result of the candidate object; Fusing the two-dimensional classification result and the three-dimensional classification result of the candidate object to obtain a final classification result of the candidate object; According to the final classification result, determining whether the candidate object is the target product, and if so, using the location coordinates of the candidate object as the location coordinates of the target product; The moving device is used to drive the robot to move to the grasping range of the target product according to the positioning coordinates of the target product; The grabbing device is used to grab the target commodity and perform settlement according to the positioning coordinates of the target commodity.

9. The large model and robot based retail device according to claim 8, characterized in that The depth camera includes an RGB-D camera.

10. A computer medium storing a program executable by a processor, characterized in that: The processor-executable program is used to implement the large model and robot-based retail method according to any one of claims 1 to 6 when executed by the processor.

Citation Information

Patent Citations

  • Visual capture method and device based on depth image and readable storage medium

    CN107748890A

  • Object recognition and positioning method and device and terminal equipment

    CN111178250A

  • Irregular object pose estimation method and device based on depth camera

    CN113450408A

  • Robot grabbing detection method based on multi-mode visual information fusion

    CN115861999A

  • Target detection method and system based on environment self-adaptive robot vision system

    CN116486287A