A retail guide robot and an interaction control method of the retail guide robot

By combining multimodal interaction and cascaded visual models with an emotion strategy library and mobile grasping control, the shopping guide robot achieves accurate product recognition and grasping in complex retail environments. This solves the problem of insufficient service refinement of existing shopping guide robots in complex environments and improves user experience and service intelligence.

CN122086251BActive Publication Date: 2026-07-21SHANGHAI FOURIER INTELLIGENCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI FOURIER INTELLIGENCE CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing shopping guide robots struggle to accurately locate target products and provide refined services in complex retail environments, resulting in delayed responses, inconsistent instructions, or execution failures, which negatively impacts user experience.

Method used

By employing a multimodal interaction module and a cascaded visual model, combined with an emotion strategy library and a mobile grasping control strategy, the system generates precise interactive actions and grasping control commands by collecting image and voice information, thereby enabling the identification and grasping of target products.

Benefits of technology

It improves the accuracy and efficiency of the shopping guide robot in complex retail scenarios, enhances the naturalness and responsiveness of user interaction, and improves the overall intelligence level and user experience of the service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122086251B_ABST
    Figure CN122086251B_ABST
Patent Text Reader

Abstract

The application provides a retail guide robot and an interaction control method of the retail guide robot. The method comprises the following steps: collecting a first image and a first interaction voice of a user, inputting the first interaction voice into a guide large language model to obtain purchase intention information of the user and determine a guide service type of the user; if the guide service type is a product guide, generating an interaction action strategy according to the first interaction voice through an emotion strategy library, determining target product information according to the purchase intention information, inputting the first image into a cascade visual model to obtain a product arrangement state and a product arrangement position, generating a movement and grabbing control strategy of the retail guide robot according to the product arrangement state and the product arrangement position, generating an action interaction instruction for grabbing a target product according to the movement and grabbing control strategy and the interaction action strategy, and controlling the retail guide robot to interact with the user according to the action interaction instruction. In this way, the experience of the user during purchase can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics technology, and in particular to a retail shopping guide robot and an interactive control method for the retail shopping guide robot. Background Technology

[0002] Currently, with the development of smart retail and unmanned services, shopping guide robots are gradually being applied in shopping malls, brand stores, and large-scale warehouse retail scenarios to replace or assist human sales guides in completing product recommendations, path guidance, and basic interactive services. However, in practical applications, existing shopping guide robots mostly rely on preset rules or single-modal interaction methods, such as keyword-based voice recognition or simple path navigation, lacking a deep understanding of the user's true purchasing intent. Furthermore, in complex retail environments, with a wide variety of products and diverse display methods (such as stacking, hanging, and zoned displays), as well as real-time inventory changes and dynamic pedestrian flow interference, robots struggle to accurately locate target products and provide refined services. This results in problems such as slow response, inconsistent instructions, or execution failures in product inquiries, product guidance, and product retrieval, thus affecting the overall service quality.

[0003] Therefore, how to improve the interaction quality and enhance the user experience of shopping guide robots when providing services to users is an urgent issue to be addressed. Summary of the Invention

[0004] The purpose of this application is to provide a retail shopping guide robot and an interactive control method for the retail shopping guide robot, in order to solve the problem that most existing robots cannot effectively provide accurate interactive services and shopping guide services to shoppers, resulting in a poor shopping experience for shoppers.

[0005] To achieve the objectives of this application, the following technical solution is provided:

[0006] In a first aspect, this application provides an interactive control method for a retail shopping guide robot, the method comprising:

[0007] Capture the first image and the user's first interactive voice;

[0008] The first interactive voice is input into the shopping guide language model to obtain the user's purchase intention information; and the shopping guide service type for the user is determined by the shopping guide robot based on the purchase intention information; if the shopping guide service type is detected to be product guidance, an interactive action strategy is generated based on the first interactive voice through the emotion strategy library; the target product information is determined based on the purchase intention information; the first image is input into the cascaded vision model to obtain the target product placement state and product placement position; a movement grasping control strategy for the shopping guide robot is generated based on the product placement state and product placement position; an action interaction command for grasping the target product is generated based on the movement grasping control strategy and the interactive action strategy, and the shopping guide robot is controlled to interact with the user according to the action interaction command.

[0009] Secondly, this application provides a retail shopping guide robot, which includes a vision module, a voice module, a mobility mechanism, and a control module, wherein:

[0010] The vision module is used to acquire a first image of the area surrounding the retail shopping guide robot through a camera module;

[0011] The voice module is used to collect the user's first interactive voice through the voice module;

[0012] The mobile mechanism is used to drive the retail shopping guide robot to move;

[0013] The control module is configured to input the first interactive voice into the shopping guide language model to obtain the user's purchase intent information; and determine the type of shopping guide service provided by the retail shopping guide robot to the user based on the purchase intent information; if the shopping guide service type is detected to be product guidance, then generate an interactive action strategy based on the first interactive voice through the emotion strategy library; determine the target product information based on the purchase intent information; input the first image into the cascaded visual model to obtain the target product placement state and product placement position; generate a movement grasping control strategy for the retail shopping guide robot based on the product placement state and product placement position; generate an action interaction command to grasp the target product based on the movement grasping control strategy and the interactive action strategy, and control the retail shopping guide robot to interact with the user according to the action interaction command.

[0014] Thirdly, this application provides a control device, including a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for performing steps in any of the methods of the first aspect of the embodiments of this application.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in any method of the first aspect of this application.

[0016] Fifthly, this application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in any method of the first aspect of the embodiments of this application. The computer program product may be a software installation package.

[0017] By implementing the embodiments of this application, the following beneficial effects are achieved:

[0018] This application provides a retail shopping guide robot and an interactive control method for the retail shopping guide robot, applied to the control device of a retail robot. The method includes: acquiring a first image and a user's first interactive voice; inputting the first interactive voice into a shopping guide language model to obtain the user's purchase intent information; determining the type of shopping guide service offered by the retail shopping guide robot to the user based on the purchase intent information; if the shopping guide service type is detected to be product guidance, generating an interactive action strategy based on the first interactive voice using an emotion strategy library; determining the target product information based on the purchase intent information; inputting the first image into a cascaded visual model to obtain the target product's placement state and position; generating a movement and grasping control strategy for the retail shopping guide robot based on the product's placement state and position; generating an action interaction command to grasp the target product based on the movement and grasping control strategy and the interactive action strategy; and controlling the retail shopping guide robot to interact with the user according to the action interaction command. This enables the retail shopping guide robot to accurately identify and efficiently grasp target products in complex retail scenarios. Meanwhile, by introducing an emotional strategy library to optimize interactive action strategies, the naturalness and responsiveness of interactions with shoppers are effectively improved. This enhances the overall intelligence level and user experience of the shopping guide service while ensuring accuracy and stability in data capture. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1This is an architecture diagram of an interactive control system for a retail shopping guide robot provided in an embodiment of this application;

[0021] Figure 2 This is a schematic diagram of the structure of a control device provided in an embodiment of this application;

[0022] Figure 3 This is a flowchart illustrating an interactive control method for a retail shopping guide robot provided in an embodiment of this application;

[0023] Figure 4 This is a schematic diagram illustrating a scenario where a retail shopping guide robot grasps folded goods, as provided in an embodiment of this application.

[0024] Figure 5 This is a schematic diagram illustrating a scenario where a retail shopping guide robot grasps suspended goods, as provided in an embodiment of this application.

[0025] Figure 6 This is a flowchart illustrating another interactive control method for a retail shopping guide robot provided in an embodiment of this application;

[0026] Figure 7 This is a schematic diagram of the interactive control action response of a retail robot provided in an embodiment of this application;

[0027] Figure 8 This is a schematic diagram of the structure of a retail shopping guide robot provided in an embodiment of this application. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0029] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0030] It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document indicates that the preceding and following related objects are in an "or" relationship. In the embodiments of this application, "multiple" refers to two or more.

[0031] In the embodiments of this application, "at least one item" or its similar expression refers to any combination of these items, including any combination of a single item or a plurality of items. "One or more" means one or more, while "multiple" means two or more. For example, "at least one item" of a, b, or c can represent the following seven cases: a, b, c; a and b; a and c; b and c; a, b, and c. Each of a, b, and c can be an element or a set containing one or more elements.

[0032] In this application embodiment, "connection" refers to various connection methods such as direct connection or indirect connection to realize communication between devices. This application embodiment does not limit this in any way.

[0033] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0034] Please see Figure 1 , Figure 1 This is an architecture diagram of an interactive control system for a retail shopping guide robot provided in an embodiment of this application. The interactive control system 100 of the retail shopping guide robot includes a robot 110 and a retail shopping guide cloud platform 120.

[0035] The robot 110, serving as the terminal execution carrier for the shopping guide service, includes a multimodal interaction module 11, a cascaded visual model 12, and a motion control module 13. It is used to complete user interaction perception, product visual recognition, and terminal action execution in retail scenarios. The multimodal interaction module 11 integrates a camera module 111 and a voice module 112. The camera module 111 collects visual image information of the robot's surrounding environment, target products, and shelf areas, providing a visual data source for product recognition, scene localization, and motion planning. The voice module 112 collects user voice commands, interactive dialogue, and environmental audio, enabling front-end perception and noise preprocessing for human-machine voice interaction, ensuring the accuracy of interactive commands. The cascaded visual model 12 is a lightweight visual processing unit on the robot's end. It receives image data transmitted from the multimodal interaction module 11 and completes visual tasks such as product feature extraction, shelf area localization, and clothing status recognition (e.g., vertical hanging, folded placement). Simultaneously, it reduces cloud computing power consumption through edge-cloud collaboration. The motion control module 13 includes a motion drive module 131 and a gripping actuator 132. The motion drive module 131 is used to control the trajectory planning, speed regulation and power output of the robot's moving mechanism, robotic arm and other moving parts, adapting to the motion needs of different shopping guide scenarios such as shelf guidance and merchandise retrieval. The gripping actuator 132, as the end effector of the humanoid robot, is used to complete fine operations such as clothing gripping and merchandise retrieval, realizing the terminal actions of shopping guide services and ensuring the feasibility of non-standard merchandise services.

[0036] Among them, the retail shopping guide cloud platform 120 serves as a cloud computing power and data platform for shopping guide services, including a shopping mall resource database 121, a retail vertical action visual model 122, a retail knowledge base 123, and a robot action library 124, which are used to provide cloud data support, model reasoning empowerment, and action library scheduling for the robot. The shopping mall resource database 121 stores real-time product data, shelf layout information, tasting time information, and pedestrian flow heatmaps, among other scene resources. The retail vertical action vision model 122, as a cloud-based inference model, achieves standardized data interaction with the robot, retail knowledge base 123, and robot action library 124 through a model context protocol. It completes core tasks such as user intent parsing, service decision generation, and action instruction planning. This model is vertically optimized for retail shopping guide scenarios and can adapt to the shopping guide needs of different retail scenarios such as supermarkets and membership-based warehouse stores. The retail knowledge base 123 and robot action library 124 communicate with the retail vertical action vision model 122 through the model context protocol. The retail knowledge base 123 stores shopping guide domain knowledge such as product attributes, promotional rules, shopping guide language, and emotional tags, providing knowledge support for intent parsing and service generation. The robot action library 124 stores standardized action sequences for humanoid robots, including action templates such as clothing grabbing, shelf guidance, and interactive gestures, providing callable action resources for the motion control module 13 and ensuring the standardization and stability of action execution. In one possible embodiment, the model context protocol is used to standardize the data interaction format and communication logic between the retail vertical action visual model 122 and the retail knowledge base 123 and robot action library 124. This ensures that the model can call the shopping guide knowledge in the knowledge base and the action templates in the action library in real time during the model inference process, avoiding the disconnect between model inference and scene knowledge, and improving the scene adaptability and decision accuracy of shopping guide services. The cascaded visual model 12 and the retail vertical action visual model 122 form an edge-cloud collaborative visual processing architecture, which not only ensures the real-time response capability of the robot end, but also leverages the complex inference capability of the cloud model, realizing efficient and accurate visual processing and service decision-making.

[0037] In some possible embodiments, robot 110 enters a low-power standby state at a preset service point in the shopping mall. Multimodal interaction module 11 detects the user's shopping guide instructions. Upon receiving a shopping guide request initiated by the user via voice or touch, camera module 111 captures the first image of the surrounding environment, and voice module 112 interacts with the user, capturing the first voice message, thus completing the acquisition and preprocessing of multimodal perception data. Then, cascaded visual model 12 performs feature extraction and product positioning on the image data. Simultaneously, the voice data is transmitted to the retail shopping guide cloud platform 120 via a communication link. The retail vertical action visual model 122, combined with the retail knowledge base 123, completes the analysis of the user's shopping guide intent and the determination of the service type. If the service type is product recommendation, the cloud platform combines the shopping mall resource database... The real-time shelf layout and tasting time information in 121 generate recommended content, which is sent back to the robot and displayed to the user by the multimodal interaction module 11. If the service type is product consultation or shelf guidance, the cloud platform calls the robot motion library 124 through the model context protocol to generate motion execution instructions containing movement paths and grasping trajectories, and sends them to the robot. The motion control module 13 controls the movement mechanism to guide the user to the shelf sub-area or tasting area of ​​the target shelf, or controls the grasping actuator 132 to complete the picking and placing of vertically hung or folded clothing. Finally, after the robot completes the shopping guide service, it sends the interaction log and user preference data back to the cloud platform, updates the mall resource database 121 and the retail knowledge base 123, and returns to the preset service point to enter standby mode.

[0038] As can be seen, the interactive control system 100 of this retail shopping guide robot integrates multimodal interaction, cascaded visual recognition, vertical model reasoning, knowledge base scheduling, and motion control through edge-cloud collaboration. Addressing pain points in retail shopping guide scenarios such as insufficient refinement of shelf guidance, low integration of tasting scenarios, and poor precision in serving non-standard products, it achieves an integrated shopping guide service encompassing product recommendation, product consultation, and refined shelf guidance. This effectively solves the problems of insufficient scenario adaptability, low service refinement, and poor user experience of existing retail shopping guide robots, improving the service capabilities of humanoid shopping guide robots in retail scenarios such as supermarkets and membership-based warehouse stores.

[0039] The following is combined with Figure 2 The control device in the embodiments of this application will be described. Figure 2 This is a schematic diagram of the structure of a control device provided in an embodiment of this application, such as... Figure 2 As shown, the control device 200 includes a processor 210, a memory 220, a communication interface 230, and one or more programs 221. The processor 210 is communicatively connected to the memory 220 and the communication interface 230 via an internal communication bus.

[0040] The one or more programs 221 are stored in the memory 220 and configured to be executed by the processor 210. The one or more programs 221 include instructions for performing any step in the above method embodiments.

[0041] The processor 210 can be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, units, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication unit can be a communication interface, transceiver, transceiver circuit, etc., and the storage unit can be a memory.

[0042] The memory 220 can be volatile memory or non-volatile memory, or it can include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0043] It is understood that the control device 200 may include more or fewer structural elements than those shown in the block diagram above, such as a power module, physical buttons, a Wi-Fi module, a speaker, a Bluetooth module, sensors, a display module, etc., without limitation herein. It is understood that the control device may be equipped with... Figure 1 The architecture of the interactive control system of a retail shopping guide robot.

[0044] After understanding the software and hardware architecture of this application, the following will be combined with... Figure 3 This application describes an interactive control method for a retail shopping guide robot. Figure 3 This is a flowchart illustrating an interactive control method for a retail shopping guide robot provided in an embodiment of this application, specifically including the following steps:

[0045] Step S310: Acquire the first image and the user's first interactive voice.

[0046] The first image is environmental image data obtained in real time by the camera module integrated into the retail guide robot to perceive its surrounding environment. The first image is used to represent the spatial environment information and target object distribution in the current retail scene, including shelf area, product display status, user location and interactive behavior, etc. The first interactive voice is the voice command or voice request signal issued by the user in the shopping guide scene collected by the voice module. The first interactive voice is used to represent the user's explicit interactive input information for subsequent semantic understanding and shopping guide intent parsing.

[0047] Specifically, when the retail shopping guide robot is in standby or service mode, its camera module continuously scans the surrounding environment at a preset frame rate, acquiring continuous video stream data as the first image. The first image can include the overall layout information of the shelves where the goods are located and the placement characteristics of the goods, such as the folding or hanging status of clothing. Simultaneously, the voice module collects the user's voice through a built-in microphone array and preprocesses the raw voice signal using voice front-end processing algorithms (such as noise reduction, echo cancellation, and voice enhancement) to obtain a clear first interactive voice signal. Furthermore, the voice module can trigger the voice acquisition process through a wake-up word detection mechanism. For example, when a preset wake-up command is detected, the voice recognition process is initiated, thereby improving the accuracy and real-time performance of the system response. After acquiring the first image and the first interactive voice, the first image will serve as input to the subsequent cascaded visual model to identify the category, placement location, and placement status of the target goods; the first interactive voice will be input into the shopping guide's large language model to extract the user's purchase intent information and the type of shopping guide needs. Through the synchronous acquisition of image and voice information, multimodal data fusion input is achieved.

[0048] It should be noted that the acquisition processes of the first image and the first interactive voice can be executed in parallel. That is, the camera module and the voice module complete environmental perception and user voice acquisition respectively within the same time window, thereby avoiding the information loss problem caused by single-modal acquisition. Furthermore, the acquisition range of the first image can be dynamically adjusted according to the robot's location. For example, the acquisition range can be reduced when the user approaches to improve target recognition accuracy, and expanded when the user moves away to enhance environmental perception. The acquisition of the first interactive voice can also be combined with sound source localization technology, prioritizing the extraction of voice signals from the direction of the target user, thereby reducing environmental noise interference and improving the robustness of voice recognition. Furthermore, the first image and the first interactive voice can be aligned in the time dimension to ensure the consistency of semantic and visual information in the subsequent multimodal fusion process.

[0049] Step S320: Input the first interactive voice into the shopping guide language model to obtain the user's purchase intention information; and determine the type of shopping guide service provided by the retail shopping guide robot to the user based on the purchase intention information.

[0050] Among them, the shopping guide language model is a vertical language understanding model deployed in the retail shopping guide robot system. It is pre-trained or fine-tuned based on retail shopping guide scenarios and can combine a domain knowledge base to perform semantic parsing and intent recognition of user speech. Purchase intent information is used to characterize the user's purchase needs during the current interaction, including product category needs (e.g., T-shirts, shirts), attribute needs (e.g., white, size), behavioral needs (e.g., browse, pick one up), and scenario needs (e.g., recommend, take me there). The shopping guide service type is the category of retail robot service execution determined based on the purchase intent information, used to guide subsequent calls to different functional modules for responses, such as product recommendations, product consultations, or product guidance.

[0051] Specifically, after acquiring the first interactive voice, the voice signal is first converted into a text sequence by the speech recognition module, and the text sequence is then input into the shopping guide language model. Based on its pre-trained semantic understanding capabilities and a loaded retail shopping guide knowledge base, the shopping guide language model performs multi-layer semantic parsing on the text sequence, extracting keywords, contextual semantic relationships, and potential intentions. For example, when a user says, "Find me a white T-shirt one size larger," the shopping guide language model can parse out multi-dimensional semantic features such as "product category = T-shirt," "attribute = white," "demand = one size larger," and "behavior = search / acquire," thereby generating structured purchase intent information. Furthermore, the shopping guide language model can also combine contextual historical interaction information to complete or correct user intent; for example, when the user does not explicitly specify the product category, it infers the range of their needs based on previous dialogue. Then, based on preset intent-service type mapping rules or through model inference mechanisms, the corresponding shopping guide service type is determined. For example, when the purchase intent information contains keywords such as "recommend" or "what's suitable," it is determined to be a product recommendation service type; when it contains keywords such as "where" or "take me there," it is determined to be a product guidance service type; when it contains query expressions such as "do you have a size up?" or "are there other colors?", it is determined to be a product consultation service type. This mapping process can be implemented through a rule engine, or it can be automatically determined by calling predefined service classification functions through the shopping guide language model combined with the Function Calling mechanism, thereby improving classification accuracy and system scalability. In addition, in the process of determining the shopping guide service type, real-time data in the retail shopping guide knowledge base can also be used for auxiliary judgment. For example, when a user asks whether a certain product is in stock, the inventory information can be queried at the same time to determine whether to trigger alternative recommendations or guidance services, thereby achieving a more accurate service type decision.

[0052] It is evident that by inputting the first interactive voice into the shopping guide's large language model for semantic parsing and combining it with a preset mapping mechanism to determine the shopping guide service type, an automatic conversion process from "natural language input" to "structured service decision-making" is achieved. This not only accurately understands the diverse expressions of users but also enables intelligent switching between multiple service modes in complex retail scenarios, thereby significantly improving the system's interactive flexibility and response efficiency, and further enhancing the user interaction experience and service efficiency.

[0053] In one possible embodiment, the purchase intent information includes: product attribute information, product status information, and user demand information. Determining the type of shopping guide service provided by the retail shopping guide robot to the user based on the purchase intent information specifically includes the following steps:

[0054] 321. Determine the user's purchase intention value based on the product attribute information and the product status information, and obtain multiple purchase intention values;

[0055] 322. When the first purchase intention value is greater than a preset sales guidance intention threshold, determine the sales guidance scenario corresponding to the first purchase intention value; the first purchase intention value is any one of the plurality of purchase intention values;

[0056] 323. Obtain the scenario constraint information for the retail shopping guide robot to execute the shopping guide scenario;

[0057] 324. Based on the scenario constraint information and the user demand information, determine the type of shopping guide service that satisfies the user.

[0058] Among them, product attribute information is used to characterize the basic characteristics of the products that users are interested in, such as product category, color, size, and material, to reflect users' preferences for the products themselves; product status information is used to characterize the dynamic status of the products in the current retail environment, including whether the inventory is sufficient and whether the products are on sale; user demand information is used to characterize the user's behavioral goals in the current interaction, such as querying product information, obtaining recommendations, being guided to the shelf, or directly picking up the products; scenario constraint information may include the shelf layout in the current environment, the spatial location of the target product, the accessibility of the robot's movement path, the robotic arm's grasping range, and the real-time inventory status.

[0059] Specifically, firstly, after acquiring the product attribute information and product status information, the user's potential purchase tendency is quantitatively assessed based on the matching relationship between the two, generating a corresponding purchase intention value. In this process, the user's currently expressed needs can be matched with product information in the knowledge base. Combining the product's inventory status, promotional status, and the degree of fit with the user's needs, different candidate products or intentions are scored, resulting in multiple purchase intention values ​​to represent the priority of different candidate shopping guide directions. Next, these multiple purchase intention values ​​are filtered. When a purchase intention value exceeds a preset shopping guide intention threshold, the corresponding candidate intention is considered to have high execution value, thus determining its corresponding shopping guide scenario as the current target shopping guide scenario. This shopping guide scenario can specifically manifest as a product recommendation scenario, a product guidance scenario, or a product retrieval scenario, thereby achieving preliminary classification and positioning of user needs. Then, after determining the shopping guide scenario, the scenario constraint information encountered by the retail shopping guide robot during the execution of this shopping guide scenario is further acquired. By introducing scenario constraint information, the executability of the shopping guide scenario can be verified, avoiding service behaviors that cannot be completed in a real-world environment. For example, if the target product is located on a high shelf and beyond the reach of the robotic arm, the existing shopping guide scenario needs to be adjusted. Finally, based on the joint analysis of scenario constraint information and user demand information, and through a preset service type mapping relationship or model inference mechanism, a target shopping guide service type that meets the current constraints is selected from multiple candidate shopping guide service types. This target shopping guide service type not only reflects the user's real needs but can also be effectively executed under the current environment and equipment conditions, thereby achieving a closed-loop process from intent recognition to service decision-making. The preset service type mapping relationship can be pre-set or predicted by machine learning methods.

[0060] It is evident that by refining users' purchase intent information into multi-dimensional semantic features and combining them with purchase intention assessment and scenario constraint screening mechanisms, the dynamic determination of the type of shopping guide service is achieved. This enables the shopping guide robot to accurately respond to user needs in complex retail environments, thereby improving the decision-making accuracy and service reliability of the shopping guide robot in complex retail scenarios.

[0061] Step S330: If the shopping guide service type is detected to be product guidance, then an interactive action strategy is generated based on the first interactive voice through the emotion strategy library; the target product information is determined based on the purchase intent information; the first image is input into the cascaded visual model to obtain the product placement state and product placement position; a movement grasping control strategy for the retail shopping guide robot is generated based on the product placement state and product placement position; an action interaction instruction for grasping the target product is generated based on the movement grasping control strategy and the interactive action strategy, and the retail shopping guide robot is controlled to interact with the user based on the action interaction instruction.

[0062] The emotional strategy library stores predefined interactive behavior templates and their corresponding emotional tag mapping relationships. Different emotional tags correspond to different voice tones, body movements, and interaction rhythms to enhance the naturalness and friendliness of human-computer interaction. The target product information is the specific product entity corresponding to the user's current needs, including product category, model, size, and target shelf area in the retail environment. The cascaded vision model is used to perform layered recognition and reasoning on environmental images to obtain the placement status of the target product (such as stacked or hung) and its specific spatial location on the shelf. This cascaded vision model is integrated into the retail shopping guide robot to improve its response efficiency. The mobile grasping control strategy is the overall control scheme for the retail shopping guide robot to move from its current position to the location of the target product and complete the grasping action.

[0063] Specifically, firstly, upon detecting that the current shopping guide service type is product guidance, the system matches corresponding emotional tags from the emotional strategy library based on the user's initial voice interaction and generates corresponding interactive action strategies. These strategies can include voice feedback content, tone of voice, and physical actions performed by the robot during the guidance process, such as turning around to guide or waving, thus establishing a positive interactive atmosphere before the shopping guide task is executed. Next, the system performs structured analysis of user needs based on purchase intent information, extracting product-related attribute features and combining this with a retail shopping guide knowledge base to determine target product information, including the shelf area where the target product is located and its possible placement range, thereby providing target constraints for subsequent visual perception. Then, the first image is input into the cascaded vision model. By performing target detection and feature matching on the environmental image, the actual placement and specific spatial location of the target product in the current scene are identified. During this process, the cascaded vision model can compare the appearance features of the product with a pre-loaded product feature library and determine whether it is stacked or suspended. At the same time, it outputs the product's hierarchical position and relative coordinate information on the shelf. After obtaining the visual perception results, a mobile grasping control strategy is further generated based on the product's placement and location. This control strategy comprehensively considers factors such as the robot's current pose, the target shelf position, path accessibility, and the robotic arm's grasping range to plan the robot's movement path and grasping method. For example, when the product is stacked, a forward approach path is prioritized and a pinch grasping method is used; when the product is suspended, a lateral approach path is planned and an unslip grasping method is used. Finally, the motion grasping control strategy is integrated with the interactive action strategy to generate unified action interaction commands. These commands include not only the robot's movement in space and the robotic arm's grasping actions, but also voice and body movements that interact synchronously with the user during execution, thus enabling parallel task execution and human-computer interaction. The retail guide robot performs motion control based on these action interaction commands, simultaneously outputting voice prompts while guiding the user to the target shelf, and completing the grasping and display of the target product upon reaching the target location, thereby achieving complete product guidance and service.

[0064] It is evident that by integrating emotional interaction strategies with mobile grasping control strategies, the shopping guide robot achieves coordinated behavior execution and human-computer interaction during the product guidance task, thereby significantly enhancing user experience and service intelligence while improving task execution efficiency.

[0065] In one possible embodiment, generating the movement and grasping control strategy for the retail guide robot based on the product placement state and the product placement position specifically includes the following steps:

[0066] 331. Input the first image and the purchase intention information into the visual language action model to obtain the first action strategy of the retail guide robot and the first visual strategy for guiding the visual perception of the retail guide robot.

[0067] 332. Determine the target shelf location corresponding to the target product according to the first action strategy;

[0068] 333. According to the first visual strategy, a second image is acquired by the camera module within the product position range corresponding to the product placement position;

[0069] 334. Generate a first movement action command for the retail guide robot based on the second image; and generate a first grasping action command for the robotic arm of the retail guide robot based on the product placement status;

[0070] 335. Control the retail guide robot to move according to the first action strategy; when the retail guide robot reaches the target shelf position, generate the movement and grasping control strategy of the retail guide robot according to the first movement action command and the first grasping action command.

[0071] The visual-language-action model is used to fuse visual and semantic information to jointly model and generate strategies for robot behavior. Its inputs include a first image and purchase intent information, and its outputs include a first action strategy to guide the overall behavior of the robot and a first visual strategy to guide the visual perception process. The first action strategy is used to represent the robot's global execution intent in the current shopping guide task, including the movement target, path selection, and task priority. The first visual strategy is used to limit the area of ​​focus and acquisition method of visual perception to improve the accuracy and efficiency of target product recognition. The first movement action command is used to represent the motion control information for the robot to make fine adjustments to its position near the target shelf. The first grasping action command is used to represent the grasping execution method generated by the robotic arm for different product placement states.

[0072] Specifically, firstly, the first image is fused with the purchase intent information and input into the visual language action model. Through the model's joint understanding of environmental visual information and user semantic needs, a first action strategy and a first visual strategy are generated. In this process, the visual language action model can not only identify the shelf distribution and candidate product areas in the current environment, but also semantically constrain the target product based on the user's purchase intent information, thereby outputting a goal-oriented action planning strategy. Next, based on the first action strategy, the target shelf location corresponding to the target product is determined from a retail guide knowledge base or environmental map. This location may include a specific area number and shelf hierarchy information, thus providing a clear spatial target for the robot's subsequent movement. Then, according to the first visual strategy, the camera module directionally collects data from the product's location range to obtain a second image. In this process, the first visual strategy limits the camera module's acquisition range and viewing angle, ensuring that the collected image data is concentrated on the target product area, thereby reducing interference from irrelevant information and improving the efficiency of visual recognition. Based on the second image, the spatial location of the target product is further refined, and a first movement action command is generated for the retail guide robot. This command controls the robot to adjust its position near the target shelf, enabling the robotic arm to accurately align with the target product. Simultaneously, based on the product's placement (e.g., stacked or suspended), the corresponding robotic arm skill model or grasping strategy library is invoked to generate a first grasping action command for the robotic arm, thus determining the specific grasping method and action parameters. Finally, the retail guide robot is controlled to move globally according to the first action strategy, gradually approaching the target shelf. After reaching the target shelf, the robot's movement behavior in the local space and its grasping behavior are coordinated and planned in conjunction with the first movement action command and the first grasping action command, generating a unified movement and grasping control strategy. This control strategy includes not only path adjustment information but also the robotic arm's grasping trajectory and execution sequence, achieving a smooth transition from global navigation to precise local operation, ensuring the robot can accurately and efficiently complete the task of retrieving the target product.

[0073] It is evident that by integrating visual and semantic information to generate hierarchical action strategies, and combining visual guidance with grasping strategies for collaborative planning, continuous control of the robot from global navigation to precise local grasping is achieved, thereby improving the accuracy and efficiency of the retail shopping guide robot in locating and grasping products.

[0074] In one possible embodiment, the product placement state includes a stacked state and a suspended state. The step of generating a first movement command for the retail guide robot based on the second image, and generating a first grasping command for the robotic arm of the retail guide robot based on the product placement state, specifically includes the following steps:

[0075] 3341. Based on the second image, identify the spatial position of the target product to obtain the first pose of the target product relative to the retail guide robot;

[0076] 3342. Determine the pose adjustment parameters of the retail shopping guide robot based on the first pose;

[0077] 3343. Generate a first movement action command based on the pose adjustment parameters using the robot movement skill library, and control the movement mechanism of the retail guide robot to move according to the first movement action command, so that the end effector of the robotic arm is aligned with the second pose of the target product;

[0078] 3344. Determine the gripping mode corresponding to the robotic arm based on the product placement status;

[0079] 3345. If the product placement state is the stacked state, the first grasping action parameters for the robotic arm to perform pinching, lifting and translation operations are generated based on the first pose, the second pose and the robotic arm skill library; and the first grasping action command of the robotic arm is determined based on the grasping mode and the first grasping action parameters.

[0080] 3346. If the product placement state is the suspended state, the second grasping action parameters for the robotic arm to perform the unhanging and extraction operation are generated by the robotic arm skill library according to the first pose and the second pose; and the first grasping action command of the robotic arm is determined according to the grasping mode and the first grasping action parameters.

[0081] The second image contains visual information about the local area where the target product is located. By analyzing the second image, the spatial position of the target product and its position and posture relative to the retail guide robot can be obtained. The first pose describes the position and posture information of the target product in the robot coordinate system. The second pose describes the target alignment posture that the end effector of the robotic arm needs to achieve when performing the grasping action. The pose adjustment parameters describe the amount of pose correction that the robot body needs to make in order to achieve precise alignment between the robotic arm and the target product. The grasping mode is used to distinguish the grasping strategy type under different product placement states.

[0082] Specifically, firstly, after acquiring the second image, a visual recognition algorithm is used to locate and estimate the pose of the target product in the image, extracting its position coordinates and pose information in three-dimensional space, thus obtaining the first pose of the target product relative to the retail guide robot. Based on this, the deviation between the current pose of the retail guide robot and the target product is calculated according to the first pose, generating pose adjustment parameters for correcting the robot's spatial position and orientation. These parameters may include information such as translation and rotation angles. Next, by calling the robot's movement skill library, a first movement command is generated based on the pose adjustment parameters, and the robot's movement mechanism is controlled to perform corresponding fine-tuning movements, causing the robotic arm to gradually approach the target product and ultimately align its end effector with the second pose of the target product, thus providing a precise spatial basis for subsequent grasping operations. Then, after completing the pose adjustment of the robot body, the corresponding grasping mode is determined according to the product's placement. When the product is stacked, it indicates that the target product is usually laid flat or piled up, in which case a pinch-grabbing mode is preferred; when the product is suspended, it indicates that the target product is usually suspended by a hanger or hook, in which case an unhanging grasping mode is required. After determining the grasping mode, the robotic arm's skill library is invoked to generate corresponding grasping action parameters for different placement states. Specifically, when the product is stacked, first grasping action parameters are generated based on the first and second poses to allow the robotic arm to perform pinching, lifting, and translation operations. These parameters may include information such as the opening angle of the robotic gripper, the grasping path, and the lifting height, thereby ensuring that the clothing is not squeezed or deformed during the grasping process. When the product is suspended, second grasping action parameters are generated based on the first and second poses to allow the robotic arm to perform unhooking and extraction operations. These parameters may include information such as the rotation angle of the robotic arm, the unhooking path, and the extraction direction, thereby preventing the clothing from snagging on adjacent items during the operation. Finally, based on the determined grasping mode and the corresponding grasping action parameters, a first grasping action command for the robotic arm is generated to control the robotic arm to perform grasping actions according to a predetermined trajectory.

[0083] In this way, complete control is achieved from target product pose recognition and robot pose adjustment to grasping mode selection and motion parameter generation, thereby ensuring that the robot can complete accurate and stable grasping operations in different product placement states, thus improving the grasping success rate and operational stability.

[0084] For easier understanding, please refer to Figure 4 , Figure 4This is a schematic diagram illustrating a scenario where a retail guide robot grasps folded goods, as provided in an embodiment of this application. As can be seen, folded goods (clothes) are placed on a shelf in a stacked state, forming a regular stacked structure. The retail guide robot uses a camera module to visually perceive the shelf area, identifying the specific position and posture information of the clothes within the stacked area. The robotic arm is located at the robot's front end, with its end effector (claw) facing the area where the clothes are located, and is in an initial posture ready to grasp. Specifically, firstly, the retail guide robot uses the camera module to acquire images of the folded clothes on the shelf, obtaining a second image of the clothes. This second image is then analyzed using a cascaded vision model to identify the spatial position and stacking state of the clothes on the shelf. Based on the stacking state, the robot further extracts the first pose information of the clothes, including its position coordinates in the shelf coordinate system and its surface orientation. Then, based on this first pose, the deviation between the robot's current pose and the target product is calculated, generating pose adjustment parameters. The robot's body is then controlled by a robot movement skill library to perform fine-tuning movements, causing the robotic arm to gradually approach the location of the target product. After the robot completes its pose adjustment, the robotic arm enters a pinching gripping mode based on the folded state of the clothing. At this point, it calls upon the robotic arm's skill library to generate corresponding gripping motion parameters based on the initial pose and the target gripping pose. These parameters include the opening angle of the gripper, the gripping contact position, the clamping force, and the lifting path. The opening angle of the gripper is adaptively adjusted according to the width of the folded clothing to ensure stable gripping of the target item while avoiding compression or deformation. During the gripping process, the robotic arm executes operations according to a "pinch-lift-translate" motion trajectory: first, the gripper clamps at the edge of the clothing; then, it lifts vertically to remove the clothing from the folded area; finally, it translates along a preset path to move the clothing to the designated location (such as the user's location). Simultaneously, pressure sensors monitor and adjust the gripping force in real time throughout the entire gripping process to ensure stability and safety, preventing deformation of the clothing due to excessive force or failure to grip due to insufficient force. If slippage or deviation is detected during the grasping process, the system can dynamically correct the robotic arm's movements based on feedback information, thereby further improving the grasping success rate.

[0085] With the above Figure 4 For implementation examples that are consistent with the above, please refer to [link / reference]. Figure 5 , Figure 5This is a schematic diagram of a retail guide robot grasping suspended goods, provided in an embodiment of this application. As can be seen, the target product is clothing, which is suspended on a hanging rod or hanger structure, forming a vertically suspended display. The retail guide robot is located in the front area of ​​the clothing and uses a camera module to visually perceive the suspended area, identifying the specific position and posture information of the clothing in the suspended space. A robotic arm is located at the front end of the robot body, with its end effector (mechanical claw) facing the suspension position of the clothing and in an initial posture ready to perform a grasping or unsustaining operation. Specifically, firstly, the retail guide robot uses the camera module to acquire images of the suspended clothing, obtaining a second image of the clothing. This second image is then input into a cascaded vision model for analysis and processing, thereby identifying the spatial position of the clothing in the suspension structure, the hanger position, and the orientation information of the clothing. Since the clothing is suspended, the retail guide robot extracts the first posture information of the clothing, including the spatial position of the clothing relative to the robot's coordinate system and the position of the hanger point. Then, based on the first pose information, the deviation between the robot's current pose and the target garment is calculated, pose adjustment parameters are generated, and the robot body is fine-tuned in position and orientation using the robot's movement skill library. This allows the robotic arm to gradually approach the hanging position of the garment and align the end effector with the hanger's hook, thus forming a second pose suitable for performing the unhooking operation. After completing the pose adjustment, the robotic arm enters the unhooking grasping mode based on the garment's hanging state. At this time, the robotic arm skill library is invoked to generate corresponding grasping action parameters based on the first pose and the target grasping pose, including the robotic arm's rotation angle, approach path, unhooking direction, and extraction trajectory. During execution, the robotic arm first positions and grips the hanger hook or the upper structure of the garment using the end effector, and then performs the "unhooking-extraction" action along a preset direction: first, the hanger is detached from the hanging rod by rotation or lifting, and then the garment is removed as a whole in a direction away from the hanging rod and moved to the target location (such as delivering it to the user or placing it in a designated area). During this process, the movement of the robotic arm is adjusted in real time through force control or position feedback mechanisms to avoid snagging or interfering with adjacent clothing during the unhooking process.

[0086] In one possible embodiment, generating the movement and grasping control strategy for the retail shopping guide robot based on the first movement action command and the first grasping action command specifically includes the following steps:

[0087] 3351. Obtain the first motion constraint parameters of the mobile mechanism of the retail shopping guide robot; and obtain the second motion constraint parameters of the robotic arm;

[0088] 3352. Control the retail guide robot to move based on the first moving action command, and obtain the first moving parameters of the moving mechanism;

[0089] 3353. Control the robotic arm to perform grasping based on the first grasping action command, and obtain the first grasping trajectory parameters of the robotic arm;

[0090] 3354. Determine the movement control strategy of the retail guide robot based on the first action constraint parameters and the first movement parameters;

[0091] 3355. Determine the grasping control strategy of the robotic arm based on the second motion constraint parameters and the first grasping trajectory parameters;

[0092] 3356. The motion control strategy and the grasping control strategy are aligned using a preset robot motion library to obtain the motion grasping control strategy.

[0093] The first motion constraint parameter characterizes the kinematic and environmental constraints experienced by the retail shopping guide robot's mobile mechanism during movement, including speed, turning radius, obstacle avoidance safety distance, and path traversability. The second motion constraint parameter characterizes the structural and control constraints experienced by the robotic arm during grasping, including joint range of motion, grasping force limits, end effector posture constraints, and pressure thresholds to prevent clothing deformation. The first movement parameter describes the actual motion state information of the robot during the execution of the first movement command, such as displacement path, speed changes, and posture adjustment process. The first grasping trajectory parameter describes the motion trajectory information of the robotic arm during the execution of the grasping action, including grasping path, joint angle changes, and end effector motion trajectory.

[0094] Specifically, firstly, after obtaining the first movement action command and the first grasping action command, the first motion constraint parameters of the moving mechanism and the second motion constraint parameters of the robotic arm are obtained respectively, thus clarifying the constraints that the retail guide robot needs to satisfy during the movement and grasping process. Next, the retail guide robot is controlled to execute the first movement action command for motion control, gradually approaching the target product location along a predetermined path, and its actual movement state is recorded during the movement to obtain the corresponding first movement parameters. Simultaneously, the robotic arm is controlled to execute the first grasping action command for grasping action planning or pre-execution, thereby obtaining the corresponding first grasping trajectory parameters. These first grasping trajectory parameters characterize the movement path and posture changes of the robotic arm from its initial position to the target grasping position. Then, after obtaining the first movement parameters and the first grasping trajectory parameters, strategies are generated by combining them with the corresponding constraint parameters. Specifically, the first motion parameters are constrained and analyzed based on the first motion constraint parameters. The feasibility of the motion path is verified and optimized to generate a motion control strategy that meets environmental and motion constraints. For example, sharp turns in the path are smoothed out or detours are planned for potential collision areas. Simultaneously, the first grasping trajectory parameters are constrained and analyzed based on the second motion constraint parameters. The robotic arm's grasping trajectory is corrected to generate a grasping control strategy that meets mechanical structure and grasping safety requirements. For example, the grasping force is adjusted or the grasping path is optimized to avoid damage to clothing. Finally, the motion control strategy and the grasping control strategy are unified, integrated, and aligned using a pre-set robot motion library to obtain the motion-grasping control strategy. In this process, the motion library pre-stores various typical motion and grasping combination strategies. By matching and integrating the currently generated motion control strategy and the grasping control strategy, they can achieve coordination and consistency in the time and space dimensions, thereby avoiding conflicts between motion and grasping actions. For example, the robotic arm's posture can be adjusted synchronously as the robot approaches the target position, so that it can directly execute the grasping action upon reaching the target position, thereby improving overall execution efficiency.

[0095] It is evident that by performing constraint analysis and strategy generation on movement and grasping actions respectively, and by achieving their collaborative integration through an action library, a conversion mechanism from single action instructions to composite control strategies is constructed, enabling the retail shopping guide robot to achieve stable and efficient collaborative control of movement and grasping in complex retail environments.

[0096] With the above Figure 3 For embodiments consistent with those described, please refer to [link / reference]. Figure 6 , Figure 6 This is a flowchart illustrating another interactive control method for a retail shopping guide robot provided in this application embodiment. The method is applied to a control device and specifically includes the following steps:

[0097] S601. If the service type is detected to be the product recommendation, then the user's purchase preference information and purchase demand information are determined based on the purchase intent information.

[0098] S602. Generate recommended product information based on the purchase preference information and the purchase demand information using a retail shopping guide knowledge base;

[0099] S603. Control the voice module to display the recommended product information to the user.

[0100] Among them, purchase intent information refers to information expressed by users through the first interactive voice regarding product attributes, functional requirements, usage scenarios, brand preferences, etc., which can reflect users' potential purchase interests and needs; purchase preference information is the user's personalized preferences for product characteristics, styles, brands, or combinations extracted from the purchase intent information; purchase demand information refers to the user's requirements for product functions, uses, quantities, price ranges, etc. during the purchase process; the retail guide knowledge base refers to a database or knowledge management system that stores product information, inventory information, sales strategies, promotional information, and user historical interaction information in the retail scenario; recommended product information refers to a list of products that meet the user's needs, selected from the retail guide knowledge base based on the user's purchase preference information and purchase demand information, including product name, model, attributes, price, and related promotional information.

[0101] Specifically, firstly, the system analyzes users' purchase intent information to further uncover their purchase preferences and needs. Based on this information, it identifies users' potential interests and purchasing tendencies, such as preferences for the color, size, or brand of a particular type of product, as well as their specific needs regarding price ranges, functional combinations, or bundled purchases. Next, the retail shopping guide robot inputs the acquired user preference and purchase need information into a retail shopping guide knowledge base for matching and processing. This knowledge base includes data such as product attribute data, inventory information, sales strategies, promotional information, and user interaction history. Through the knowledge base's rule engine or recommendation algorithm, user preferences are matched with product attributes to generate a list of recommended products that match the user's interests and needs. During the recommendation process, not only is product matching considered, but also inventory availability, product popularity, combination effects, and other marketing strategies are taken into account to ensure that recommended products both meet user needs and optimize the use of sales resources. For example, for users with specific functional requirements, products that meet those requirements are prioritized, and a personalized recommendation order is generated based on the user's past preferences to improve recommendation accuracy and user satisfaction. Then, the generated recommended product information is displayed to the user via a voice module. In this process, the voice module not only converts textual product information into speech output, but also adjusts tone, rhythm, and wording using emotional or vocal expression strategies to enhance the interactive experience. For example, for highlighted products, an emphasis tone or additional promotional information can be used to make the recommendations easier for users to accept and understand. Furthermore, the voice module's output can be combined with screen displays or the graphical interface of mobile devices to achieve multimodal display, improving the intuitiveness and interactivity of information delivery. It can also adjust recommendation strategies in real time based on user feedback, such as replacing products that users are not interested in or adjusting the recommendation order, achieving dynamic optimization.

[0102] It is evident that by extracting user preferences and needs from purchase intent information, generating personalized recommendations by combining them with a retail shopping guide knowledge base, and then providing intelligent displays through a voice module, this approach not only enhances the user interaction experience and shopping convenience but also optimizes sales strategies and inventory management through data-driven methods, thereby increasing the service intelligence level and commercial value of retail shopping guide robots.

[0103] In one possible embodiment, after the control voice module displays the recommended product information to the user, the method further includes the following steps:

[0104] 61. Control the voice module to acquire the user's second interactive voice;

[0105] 62. Input the second interactive voice into the shopping guide's large language model to obtain user feedback information;

[0106] 63. Determine the attribute information of the first product based on the user feedback information to obtain the attribute information of the first product;

[0107] 64. Determine the product model in the first product attribute information to obtain the first product model;

[0108] 65. Filter out the product information corresponding to the first product model from the preset shopping guide knowledge base to obtain the second product information, and output the second product information to the user through the voice module, and send the second product information to the shopping mall resource cloud platform.

[0109] The second interactive voice refers to the user's feedback on the recommended products, expressed through voice after receiving the product information. This feedback includes the user's level of interest, functional requirements, color, model, quantity, or other specific preferences, as well as any questions or suggestions for modification. The shopping guide's big language model is a pre-trained natural language processing model capable of understanding the semantic information of the user's voice input. It combines product information from the retail knowledge base, historical interaction data, and contextual information to generate structured user feedback information, thereby achieving a deep analysis of the user's purchase intentions and preferences. The user feedback information is the structured data output by the big language model, including the user's level of approval of the recommended products, modification suggestions, changes in preferences, and updated information on potential purchase intentions, used to guide subsequent product recommendations or guidance operations. The first product refers to the specific product that the user explicitly expresses interest in or selects in the user feedback information. The first product attribute information is a detailed description of the product's characteristics in terms of appearance, function, model, specifications, price, etc., used for precise retrieval in the knowledge base. The pre-set shopping guide knowledge base is a database system that integrates product information, inventory information, promotional strategies, price information, and historical user interaction data, supporting rapid retrieval, intelligent matching, and product recommendation services. The second product information consists of complete product information selected from the shopping guide knowledge base and corresponding to user feedback. This includes product name, model, attribute parameters, inventory status, price, and related promotional information. The shopping mall resource cloud platform is a cloud-based system that centrally manages product information, inventory data, and sales data within the shopping mall. It supports data uploads and status synchronization for the shopping guide robot, enabling unified product information management and sales data statistics.

[0110] Specifically, firstly, the voice module collects the user's second interactive voice and transmits this voice data in real time to the large-scale language model of the retail guide robot for semantic analysis. The large-scale language model extracts semantic features using natural language understanding technology, including the user's functional needs, appearance preferences, quantity requirements, and potential purchase intentions for the product, and outputs structured user feedback information. Next, based on the user feedback information, the attributes of the first product are identified and matched to generate first product attribute information, thereby determining the specific product and its model that the user is interested in. After obtaining the first product model, a precise search is performed in a preset guide knowledge base to filter out the complete product information corresponding to that model, i.e., the second product information. At the same time, the voice module outputs the second product information to the user in voice format so that the user can quickly confirm the desired product. Finally, to ensure the accuracy of information synchronization and inventory management in the retail system, the second product information is also sent to the mall resource cloud platform to achieve real-time synchronization of product information between the local guide robot and the cloud management system.

[0111] As can be seen, through this embodiment, the retail shopping guide robot can quickly identify user intent and accurately match specific products based on real-time user voice feedback, realizing dynamic updates and interactive responses to product recommendations. This not only improves the efficiency and accuracy of user interaction with the shopping guide robot but also enhances the personalization and intelligence of the recommendation service. Simultaneously, it supports cloud-based synchronous management of mall inventory and product information, providing an efficient, intelligent, and scalable technical solution for retail services.

[0112] In one possible embodiment, after controlling the retail shopping guide robot to interact with the user according to the action interaction instructions, the method further includes the following steps:

[0113] A1. Control the camera module to acquire a third image;

[0114] A2. Input the third image into the cascaded visual model to obtain visual information;

[0115] A3. Determine whether a second product exists based on the visual information;

[0116] A4. When the second product is not present in the visual information, generate fault-tolerant action instructions for the robotic arm of the retail guide robot based on the preset fault-tolerant algorithm and the visual information.

[0117] A5. Control the robotic arm to move according to the fault-tolerant action command to grab the second product; and / or control the voice module to output abnormal information to the user to notify the mall operation and maintenance personnel to handle the corresponding abnormality.

[0118] The second product refers to a specific product identified during the shopping guide process through a knowledge base or user feedback, indicating clear user intent. The third image is image data obtained by a camera module controlled by the retail shopping guide robot after interaction, capturing the location of the second product. This provides accurate visual information to determine the product's actual presence and status. The visual information is the output of the cascaded vision model processing the third image, including the product's presence, its position coordinates on the shelf or display area, its posture angle, the distribution of surrounding obstacles, and lighting or occlusion conditions, providing a basis for subsequent grasping actions. The fault-tolerant algorithm is a pre-defined robotic arm motion control strategy used to dynamically adjust and compensate for grasping actions when product recognition is abnormal or the product is not placed in the expected position. This includes path replanning, end effector adjustment, and switching of grasping modes to ensure the robot can complete grasping tasks under different environmental conditions. The fault-tolerant action command is a specific control command generated based on visual information and the fault-tolerant algorithm, used to guide the robotic arm in posture adjustment, movement path correction, and grasping operations to achieve safe and reliable grasping of the second product. Anomaly information refers to the prompt message output by the voice module to the user or mall maintenance personnel when the target product is not detected in the visual information or the capture fails. It is used to inform that there is an anomaly in the current operation or that manual intervention is required.

[0119] Specifically, firstly, after the retail guide robot completes the interaction command with the user, it captures an image of the second product through the camera module, generating a third image. Next, the third image is input into a cascaded vision model for processing. The model first performs object detection to identify whether the second product is in the expected position, and further analyzes the product's posture and placement, while assessing the potential impact of the surrounding environment on grasping, thus generating complete visual information. Then, based on the visual information, it determines whether the second product exists. If the model outputs confirmation of the product's existence, the retail guide robot can complete the operation according to the originally set grasping strategy; if the model determines that the second product is not detected in the visual information, the retail guide robot will initiate a fault-tolerant processing procedure. This fault-tolerant processing first generates fault-tolerant action commands for the robotic arm based on the visual information and a preset fault-tolerant algorithm, including path replanning, end effector position adjustment, and dynamic switching of the grasping mode, to ensure operational robustness under different environmental conditions. Simultaneously, the robotic arm is controlled to move according to the fault-tolerant action commands to achieve compensatory grasping of the second product. In addition, to ensure user experience and system security, the voice module can output abnormal prompts to users, informing them that the current product has not been successfully detected or has failed to be captured, and can simultaneously notify mall maintenance personnel for intervention.

[0120] For easier understanding, please refer to Figure 7 , Figure 7This is a schematic diagram of the interactive control action response of a retail robot provided in an embodiment of this application. As can be seen, the interactive control action response process starts with "head interaction and voice reception", performs voice parsing of user input through "large model intent recognition (ASR, TTS, etc.)", and performs hierarchical screening and structured processing of semantic information through "two-layer information data filtering (through custom scenarios and general actions)". Then, the processing results are input into "database quick query" and combined with "cloud action library data table (providing query services for large models)" to perform action retrieval and matching, thereby realizing the linkage output of "filtering action commands" and "action library discovery commands", and further enters the "data inspection and execution" stage. After execution, the result feedback is provided through "execution completion return status or execution failure return data code". Finally, "the caller obtains the action status, completes the interaction loop, and starts the next round according to the status" to realize the complete closed-loop control process.

[0121] Specifically, firstly, the user's voice is collected through the head-interaction voice acquisition module. This voice data, as the raw input signal, is transmitted to the large-scale model's intent recognition module. This module converts the speech into text using ASR (Automatic Speech Recognition) technology and combines it with TTS (Text-to-Speech) and natural language understanding capabilities to parse the user's intent, thereby obtaining a structured semantic expression result. Next, this semantic result is input to a two-layer information data filtering module. In this module, the semantic information undergoes two-layer filtering processing through preset custom scenario rules and general action rules. On one hand, the semantics are filtered based on scenario constraints according to retail shopping guidance scenarios (such as product recommendation, product guidance, or product retrieval); on the other hand, basic action semantics are matched using a general action library to obtain standardized action request data. After completing the information filtering, the processing result is input to the database fast query module. This module achieves efficient retrieval of action commands by calling the cloud-based action library data table. Specifically, the cloud-based action library data table provides the large-scale model with standard action templates, execution parameters, and action mapping relationships, enabling the system to quickly match the corresponding candidate action set based on the current semantic request. Then, the candidate actions are further filtered by the action command filtering module to determine the optimal action command, which is then passed to the action library discovery module for final confirmation of the action instruction. After the action command is determined, the data checking and execution phase begins. In this phase, the action command is first validated for legality and execution conditions, including the completeness of action parameters, feasibility of the execution environment, and compatibility with device status. After successful validation, the retail robot is controlled to execute the corresponding action instruction, such as navigation, robotic arm grasping, or interactive feedback. After the action is executed, the execution result (including success status, failure status, or exception code) is fed back to the upper-level caller through the execution completion return status or execution code data feedback module. Finally, the caller obtains the action status, completing the entire interaction loop, and decides whether to start the next round of interaction based on the returned status information. For example, when the action is executed successfully, the subsequent service process can begin; when the execution fails or an exception occurs, a fault tolerance mechanism can be triggered or the action path can be replanned to ensure the continuity and stability of the system operation. It can be seen that by introducing large-scale model intent recognition, a two-layer information filtering mechanism, and a cloud-based action library query mechanism, control from user voice input to action execution output is achieved. This not only improves the accuracy and response speed of action command matching, but also enhances the system's adaptability to complex retail scenarios, thereby effectively improving the intelligence level and interactive experience of the retail shopping guide robot.

[0122] As can be seen, through this embodiment, the retail shopping guide robot can perceive the existence and status of goods in real time during actual operation. When an anomaly is detected or the goods are not placed in the expected position, the robot can dynamically adjust the grasping strategy of the robotic arm through a fault-tolerant algorithm to achieve autonomous correction of the grasping action. This not only improves the success rate and robustness of the robot's grasping operation, but also enhances the reliability and continuity of user interaction, while ensuring the accuracy of product information and the safety of operation in the retail scenario.

[0123] It can be seen that, through the above Figure 3 The illustrated embodiment provides an interactive control method for a retail shopping guide robot. By acquiring a first image and a user's first interactive voice, the first interactive voice is input into a large-scale language model to obtain the user's purchase intent information. Based on the purchase intent information, the type of shopping guide service provided by the retail shopping guide robot to the user is determined. If the shopping guide service type is detected to be product guidance, an interactive action strategy is generated based on the first interactive voice using an emotion strategy library. The target product information is determined based on the purchase intent information, and the first image is input into a cascaded visual model to obtain the target product's placement state and position. Based on the product placement state and position, a movement and grasping control strategy for the retail shopping guide robot is generated. Based on the movement and grasping control strategy and the interactive action strategy, action interaction instructions for grasping the target product are generated, and the retail shopping guide robot is controlled to interact with the user according to the action interaction instructions. This method improves the robot's scenario applicability, flexibility, and accuracy, and enhances the user's shopping experience.

[0124] The above mainly describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the control device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0125] This application embodiment can divide the control device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0126] When dividing each function into modules according to its corresponding function. Figure 8 This is a schematic diagram of the structure of a retail shopping guide robot provided in an embodiment of this application. The retail shopping guide robot 800 includes: a voice module 801, a vision module 802, a movement mechanism 803, and a control module 804, wherein:

[0127] The voice module 801 is used to collect the user's first interactive voice through the voice module;

[0128] The vision module 802 is used to acquire a first image of the area surrounding the retail shopping guide robot through a camera module;

[0129] The mobile mechanism 803 is used to drive the retail guide robot to move.

[0130] The control module 804 is used to input the first interactive voice into the shopping guide language model to obtain the user's purchase intention information; and determine the type of shopping guide service provided by the retail shopping guide robot to the user based on the purchase intention information; if the shopping guide service type is detected to be product guidance, then generate an interactive action strategy based on the first interactive voice through the emotion strategy library; determine the target product information based on the purchase intention information; input the first image into the cascaded visual model to obtain the product placement status and product placement position; generate a movement grasping control strategy for the retail shopping guide robot based on the product placement status and product placement position; generate an action interaction command for grasping the target product based on the movement grasping control strategy and the interactive action strategy, and control the retail shopping guide robot to interact with the user according to the action interaction command.

[0131] In one possible embodiment, the control module 804 is specifically used for generating the movement and grasping control strategy for the retail guide robot based on the product placement state and the product placement position:

[0132] The first image and the purchase intent information are input into the visual language action model to obtain the first action strategy of the retail shopping guide robot and the first visual strategy for guiding the visual perception of the retail shopping guide robot.

[0133] The target shelf location corresponding to the target product is determined according to the first action strategy;

[0134] According to the first visual strategy, a second image is captured by the camera module within the product location range corresponding to the product placement location;

[0135] The retail guide robot generates a first movement action command based on the second image; and generates a first grasping action command for the robotic arm of the retail guide robot based on the product placement status.

[0136] The retail guide robot is controlled to move according to the first action strategy; when the retail guide robot reaches the target shelf position, the movement and grasping control strategy of the retail guide robot is generated according to the first movement action command and the first grasping action command.

[0137] In one possible embodiment, the product placement state includes a stacked state and a suspended state. Specifically, the control module 804 is used to generate a first movement command for the retail guide robot based on the second image, and to generate a first grasping command for the robotic arm of the retail guide robot based on the product placement state, in the following aspects:

[0138] Based on the second image, the spatial position of the target product is identified to obtain the first pose of the target product relative to the retail guide robot;

[0139] The pose adjustment parameters of the retail shopping guide robot are determined based on the first pose.

[0140] The robot generates a first movement command based on the pose adjustment parameters using the robot's movement skill library, and controls the movement mechanism of the retail guide robot to move according to the first movement command, so that the end effector of the robotic arm is aligned with the second pose of the target product.

[0141] The gripping mode of the robotic arm is determined based on the product placement status.

[0142] If the product is placed in the stacked state, the robotic arm skill library is used to generate the first grasping action parameters for pinching, lifting and translating operations by the robotic arm, based on the first pose, the second pose and the first grasping action parameters; and the first grasping action command of the robotic arm is determined based on the grasping mode and the first grasping action parameters.

[0143] If the product is in the suspended state, the robotic arm skill library generates second grasping action parameters for the robotic arm to perform unsustaining and extraction operations based on the first pose and the second pose; and determines the first grasping action command of the robotic arm based on the grasping mode and the first grasping action parameters.

[0144] In one possible embodiment, the control module 804, in generating the movement and grasping control strategy for the retail guide robot based on the first movement action command and the first grasping action command, is specifically configured to:

[0145] Obtain the first motion constraint parameters of the mobile mechanism of the retail shopping guide robot; and obtain the second motion constraint parameters of the robotic arm;

[0146] The retail guide robot is controlled to move based on the first movement action command, and the first movement parameters of the movement mechanism are obtained.

[0147] The robotic arm is controlled to grasp based on the first grasping action command, and the first grasping trajectory parameters of the robotic arm are obtained.

[0148] The movement control strategy of the retail guide robot is determined based on the first action constraint parameters and the first movement parameters;

[0149] The grasping control strategy of the robotic arm is determined based on the second motion constraint parameters and the first grasping trajectory parameters.

[0150] The motion-grasping control strategy is obtained by aligning the motion control strategy and the grasping control strategy using a preset robot motion library.

[0151] In one possible embodiment, the purchase intent information includes: product attribute information, product status information, and user demand information. Specifically, the control module 804, in determining the type of shopping guide service provided by the retail shopping guide robot to the user based on the purchase intent information, is used for:

[0152] Based on the product attribute information and the product status information, the user's purchase intention value is determined, resulting in multiple purchase intention values;

[0153] When the first purchase intention value is greater than a preset sales guidance intention threshold, the sales guidance scenario corresponding to the first purchase intention value is determined; the first purchase intention value is any one of the plurality of purchase intention values;

[0154] Obtain the scenario constraint information for the retail shopping guide robot to execute the shopping guide scenario;

[0155] Based on the scenario constraint information and the user demand information, the type of shopping guide service that satisfies the user is determined.

[0156] In one possible embodiment, the shopping guide service type further includes product recommendations, and the control module 804 is specifically used for:

[0157] If the service type is detected to be the product recommendation, then the user's purchase preference information and purchase demand information are determined based on the purchase intent information;

[0158] Based on the purchase preference information and purchase demand information, recommended product information is generated using a retail shopping guide knowledge base;

[0159] The voice control module displays the recommended product information to the user.

[0160] In one possible embodiment, after the control voice module displays the recommended product information to the user, the control module 804 is further configured to:

[0161] Control the voice module to acquire the user's second interactive voice;

[0162] The second interactive voice is input into the shopping guide's large language model to obtain user feedback information;

[0163] The attribute information of the first product is determined based on the user feedback information, and the attribute information of the first product is obtained;

[0164] Determine the product model from the first product attribute information to obtain the first product model;

[0165] The system filters out the product information corresponding to the first product model from the preset shopping guide knowledge base, obtains the second product information, outputs the second product information to the user through the voice module, and sends the second product information to the shopping mall resource cloud platform.

[0166] In one possible embodiment, after controlling the retail guide robot to interact with the user according to the action interaction command, the control module 804 is further configured to:

[0167] Control the camera module to acquire a third image;

[0168] The third image is input into the cascaded visual model to obtain visual information;

[0169] Determine whether the second product exists based on the visual information;

[0170] When the second product is not present in the visual information, fault-tolerant action instructions for the robotic arm of the retail guide robot are generated based on the preset fault-tolerant algorithm and the visual information.

[0171] The robotic arm is controlled to move according to the fault-tolerant action command to grasp the second product; and / or, the voice module is controlled to output abnormal information to the user to notify the mall's operation and maintenance personnel to handle the abnormality accordingly.

[0172] It should be noted that the specific implementation of each operation can adopt the methods described above. Figure 3 The description of an embodiment of the interactive control method for a retail shopping guide robot is shown below. The retail shopping guide robot 800 can be used to execute the above-described method embodiment of this application, and will not be repeated here.

[0173] As can be seen, the embodiments of this application provide a retail shopping guide robot. By deploying a visual language action model and a large language model in the retail shopping guide knowledge domain, the visual action model is equipped with a knowledge base, and the skill base and other information that needs to be obtained in real time are input into the large language model. Through multimodal interaction, the robot performs retail shopping guide tasks, thereby improving the robot's scenario applicability, flexibility and accuracy, as well as enhancing the user's purchase experience.

[0174] This application also provides a computer-readable storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes a control device.

[0175] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer includes a control device.

[0176] It should be noted that, for the sake of simplicity, the above embodiments are all described as a series of actions. Those skilled in the art should understand that this application is not limited to the described order of actions, as some steps in the embodiments of this application can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions, steps, modules, or units involved are not necessarily essential to the embodiments of this application.

[0177] In the above embodiments, the descriptions of each embodiment in this application have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0178] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0179] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in the embodiments of this application can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0180] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the embodiments of this application. It should be understood that the above descriptions are merely specific embodiments of the embodiments of this application and are not intended to limit the protection scope of the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solutions of the embodiments of this application should be included within the protection scope of the embodiments of this application.

Claims

1. An interactive control method for a retail shopping guide robot, characterized in that, include: The system acquires a first image and the user's first interactive voice; the first image is environmental image data perceived in real time by the retail guide robot. The first interactive voice is input into the shopping guide's large language model to obtain the user's purchase intention information; The purchase intent information includes: product attribute information, product status information, and user demand information; based on the product attribute information and the product status information, the user's purchase intention value is determined, resulting in multiple purchase intention values; when a first purchase intention value is greater than a preset shopping guide intention threshold, a shopping guide scenario corresponding to the first purchase intention value is determined; the first purchase intention value is any one of the multiple purchase intention values; scenario constraint information for the retail shopping guide robot to execute the shopping guide scenario is obtained; based on the scenario constraint information and the user demand information, the type of shopping guide service that satisfies the user is determined; If the shopping guide service type is detected to be product guidance, an interactive action strategy is generated based on the first interactive voice through the emotion strategy library; target product information is determined based on the purchase intent information; the first image is input into the cascaded visual model to obtain the product placement state and product placement position; the product placement state includes stacked state and hanging state; the first image and the purchase intent information are input into the visual language action model to obtain the first action strategy of the retail shopping guide robot, and a first visual strategy for guiding the visual perception of the retail shopping guide robot; the target shelf position corresponding to the target product is determined based on the first action strategy; and the product position range corresponding to the product placement position is collected by the camera module based on the first visual strategy. The system generates a second image within the product display area; generates a first movement action command for the retail guide robot based on the second image; generates a first grasping action command for the robotic arm of the retail guide robot based on the product display state; determines the grasping mode corresponding to the robotic arm based on the product display state; controls the retail guide robot to move according to the first action strategy; when the retail guide robot reaches the target shelf position, generates a movement grasping control strategy for the retail guide robot based on the first movement action command and the first grasping action command; generates an action interaction command for grasping the target product based on the movement grasping control strategy and the interaction action strategy, and controls the retail guide robot to interact with the user according to the action interaction command.

2. The method as described in claim 1, characterized in that, The step of generating a first movement command for the retail guide robot based on the second image; and generating a first grasping command for the robotic arm of the retail guide robot based on the product placement status, includes: Based on the second image, the spatial position of the target product is identified to obtain the first pose of the target product relative to the retail guide robot; The pose adjustment parameters of the retail shopping guide robot are determined based on the first pose. The robot generates a first movement command based on the pose adjustment parameters using the robot's movement skill library, and controls the movement mechanism of the retail guide robot to move according to the first movement command, so that the end effector of the robotic arm is aligned with the second pose of the target product. If the product is placed in the stacked state, the robotic arm skill library is used to generate the first grasping action parameters for pinching, lifting and translating operations by the robotic arm, based on the first pose, the second pose and the first grasping action parameters; and the first grasping action command of the robotic arm is determined based on the grasping mode and the first grasping action parameters. If the product is in the suspended state, the robotic arm skill library generates second grasping action parameters for the robotic arm to perform unsustaining and extraction operations based on the first pose and the second pose; and determines the first grasping action command of the robotic arm based on the grasping mode and the first grasping action parameters.

3. The method as described in claim 1, characterized in that, The step of generating the movement and grasping control strategy for the retail shopping guide robot based on the first movement action command and the first grasping action command includes: Obtain the first motion constraint parameters of the mobile mechanism of the retail shopping guide robot; and obtain the second motion constraint parameters of the robotic arm; The retail guide robot is controlled to move based on the first movement action command, and the first movement parameters of the movement mechanism are obtained. The robotic arm is controlled to grasp based on the first grasping action command, and the first grasping trajectory parameters of the robotic arm are obtained. The movement control strategy of the retail guide robot is determined based on the first action constraint parameters and the first movement parameters; The grasping control strategy of the robotic arm is determined based on the second motion constraint parameters and the first grasping trajectory parameters. The motion-grasping control strategy is obtained by aligning the motion control strategy and the grasping control strategy using a preset robot motion library.

4. The method according to any one of claims 1-3, characterized in that, The shopping guide service type also includes product recommendations, and the method further includes: If the service type is detected to be the product recommendation, then the user's purchase preference information and purchase demand information are determined based on the purchase intent information; Based on the purchase preference information and purchase demand information, recommended product information is generated using a retail shopping guide knowledge base; The voice control module displays the recommended product information to the user.

5. The method as described in claim 4, characterized in that, After the control voice module displays the recommended product information to the user, the method further includes: Control the voice module to acquire the user's second interactive voice; The second interactive voice is input into the shopping guide's large language model to obtain user feedback information; The attribute information of the first product is determined based on the user feedback information, and the attribute information of the first product is obtained; Determine the product model from the first product attribute information to obtain the first product model; The system filters out the product information corresponding to the first product model from the preset shopping guide knowledge base, obtains the second product information, outputs the second product information to the user through the voice module, and sends the second product information to the shopping mall resource cloud platform.

6. The method according to any one of claims 1-3, characterized in that, After controlling the retail shopping guide robot to interact with the user according to the action interaction instructions, the method further includes: Control the camera module to acquire a third image; The third image is input into the cascaded visual model to obtain visual information; Determine whether a second product exists based on the visual information; When the second product is not present in the visual information, fault-tolerant action instructions for the robotic arm of the retail guide robot are generated based on the preset fault-tolerant algorithm and the visual information. The robotic arm is controlled to move according to the fault-tolerant action command to grasp the second product; and / or, the voice module is controlled to output abnormal information to the user to notify the mall's operation and maintenance personnel to handle the abnormality accordingly.

7. A retail shopping guide robot, characterized in that, The retail shopping guide robot includes a vision module, a voice module, a mobility mechanism, and a control module, wherein: The vision module is used to acquire a first image of the area around the retail guide robot through a camera module; the first image is environmental image data perceived in real time by the retail guide robot. The voice module is used to collect the user's first interactive voice through the voice module; The mobile mechanism is used to drive the retail shopping guide robot to move; The control module is used to input the first interactive voice into the shopping guide language model to obtain the user's purchase intention information; the purchase intention information includes: product attribute information, product status information, and user demand information; determine the user's purchase intention value based on the product attribute information and the product status information to obtain multiple purchase intention values; when the first purchase intention value is greater than a preset shopping guide intention threshold, determine the shopping guide scenario corresponding to the first purchase intention value; the first purchase intention value is any one of the multiple purchase intention values; obtain the scenario constraint information for the retail shopping guide robot to execute the shopping guide scenario; determine the shopping guide service type that satisfies the user based on the scenario constraint information and the user demand information; if the shopping guide service type is detected to be product guidance, generate an interactive action strategy based on the first interactive voice through the emotion strategy library; determine the target product information based on the purchase intention information; input the first image into the cascaded visual model to obtain the product placement state and product placement position; the product placement state includes stacked state and hanging state; and input the first image and the purchase intention information into the cascaded visual model to obtain the product placement state and product placement position. Information is input into the visual language action model to obtain a first action strategy for the retail guide robot and a first visual strategy for guiding the visual perception of the retail guide robot; the target shelf position corresponding to the target product is determined according to the first action strategy; a second image is acquired by a camera module within the product position range corresponding to the product placement position according to the first visual strategy; a first movement action command for the retail guide robot is generated based on the second image; and a first grasping action command for the robotic arm of the retail guide robot is generated based on the product placement state; the grasping mode corresponding to the robotic arm is determined based on the product placement state; the retail guide robot is controlled to move according to the first action strategy; when the retail guide robot reaches the target shelf position, a movement grasping control strategy for the retail guide robot is generated based on the first movement action command and the first grasping action command; an action interaction command for grasping the target product is generated based on the movement grasping control strategy and the interaction action strategy, and the retail guide robot is controlled to interact with the user according to the action interaction command.

8. A control device, characterized in that, include: Processor, memory, communication interface, and one or more programs; The one or more programs are stored in the memory and configured to be executed by the processor, the programs including instructions for performing the steps of the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Robot multi-mode interaction control method and related device

    CN121179445A

  • Control method and artificial intelligence experiment system

    CN121354556A