Robot multi-mode interaction control method and related device
By analyzing user interaction information to generate target action sequences of joint and voice actions, the problem of insufficient action interaction in humanoid shopping guide robots in vehicle sales scenarios is solved, thereby improving their functionality and intelligence.
Patent Information
- Application Number
- CN202511745789.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-11-26
AI Technical Summary
Existing humanoid shopping guide robots lack action-interaction tasks based on their own abilities in vehicle sales scenarios, resulting in insufficient functionality and intelligence.
By acquiring user interaction information, analyzing the type of shopping guide service and the target vehicle parts, and using a pre-set vehicle knowledge base and motion library, target motion sequences of joint movements and voice movements are generated to collaboratively achieve multimodal interaction.
It improves the personalized action interaction capabilities of humanoid robots in sales scenarios, enhancing the functionality and intelligence of intelligent shopping guides.
Smart Images

Figure CN121179445A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, specifically to a multimodal interactive control method and related apparatus for robots. Background Technology
[0002] To improve service efficiency and reduce the cost of human sales guides, humanoid robots are gradually being introduced into sales scenarios such as vehicle dealerships, serving as intelligent sales guide devices to assist in completing basic service tasks. However, the core capabilities of humanoid sales guide robots currently used in vehicle sales scenarios are limited to basic information queries and simple question answering. They lack the ability to perform action-interaction tasks based on their own abilities in the sales scenario, and still have shortcomings in terms of functionality and intelligence. Summary of the Invention
[0003] This application provides a robot multimodal interaction control method and related apparatus to improve the functionality and intelligence of humanoid robots as intelligent shopping guide robots.
[0004] In a first aspect, embodiments of this application provide a robot multimodal interactive control method, applied to a robot controller in a robot control system, the robot control system including the robot controller and a target robot, the method comprising: Acquire user interaction information, which includes user voice information and / or user action information; Analyze the user interaction information to determine the type of shopping guide service and the target vehicle parts; Based on the type of shopping guide service, obtain the first component information of the target vehicle component from a preset vehicle knowledge base; Based on the shopping guide service type and the first component information, a target action sequence is determined from a preset action library. The target action sequence includes joint actions and voice actions. The voice actions are used to output voice shopping guide information, and the voice shopping guide information is determined based on the first component information. A first control command is sent to the target robot, the first control command being used to instruct the target robot to execute the target action sequence.
[0005] Secondly, embodiments of this application provide a robot multimodal interactive control device, applied to a robot controller in a robot control system. The robot control system includes the robot controller and a target robot. The device includes: The first acquisition unit is used to acquire user interaction information, which includes user voice information and / or user action information. The parsing unit is used to parse the user interaction information and determine the type of shopping guide service and the target vehicle parts; The second acquisition unit is used to acquire the first component information of the target vehicle component from a preset vehicle knowledge base according to the type of the shopping guide service. The determining unit is configured to determine a target action sequence from a preset action library based on the shopping guide service type and the first component information. The target action sequence includes joint actions and voice actions. The voice actions are used to output voice shopping guide information. The voice shopping guide information is determined based on the first component information. The sending unit is used to send a first control command to the target robot, the first control command being used to instruct the target robot to execute the target action sequence.
[0006] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, a communication interface, and one or more programs, the one or more programs being stored in the memory and configured to be executed by the processor, the programs including instructions for performing the steps in the first aspect of embodiments of this application.
[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in the first aspect of this embodiment.
[0008] Fifthly, this application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of this application. The computer program product may be a software installation package.
[0009] As can be seen, in this embodiment, the robot controller first acquires user interaction information containing user voice information and / or user action information, then parses the user interaction information to determine the type of shopping guide service and the target vehicle component to accurately obtain the user's actual needs. Next, based on the type of shopping guide service, it obtains the first component information of the target vehicle component from a preset vehicle knowledge base to improve the adaptability of the acquired information to the user's actual needs and service scenarios. Subsequently, based on the type of shopping guide service and the first component information, it determines the target action sequence containing joint actions and voice actions from a preset action library. The voice actions are used to output the voice shopping guide information determined based on the first component information, realizing the coordination of joint actions and voice shopping guide information. Finally, it sends a first control command to the target robot to instruct it to execute the target action sequence. Ultimately, it can respond to user interaction information and realize the robot's multimodal action output, supporting the realization of personalized action interaction tasks in sales scenarios, thereby improving the functionality and intelligence of the humanoid robot as an intelligent shopping guide robot. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of the architecture of a robot control system provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a robot multimodal interactive control method provided in an embodiment of this application; Figure 3 This is a schematic diagram of a vehicle selection page provided in an embodiment of this application; Figure 4 This is a schematic diagram of a reservation information confirmation page provided in an embodiment of this application; Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of this application; Figure 6 This is a functional unit block diagram of a robot multimodal interactive control device provided in an embodiment of this application; Figure 7 This is a block diagram of the functional units of another robot multimodal interactive control device provided in the embodiments of this application. Detailed Implementation
[0012] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0013] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0015] The embodiments of this application will now be described with reference to the accompanying drawings.
[0016] The technical solution of this application can be applied to, for example... Figure 1 The robot control system shown is, specifically, as follows: Figure 1 As shown, the robot control system may include a robot controller 100 and a target robot 200, which are communicatively connected.
[0017] Specifically, the robot controller 100 can be integrated into the target robot 200. For example, the robot controller 100 can be a processor system in the target robot 200. Alternatively, in other embodiments, the robot controller 100 can be set independently of the target robot 200. No specific restrictions are imposed here.
[0018] Specifically, the target robot 200 can be a humanoid robot. In particular, the humanoid robot can be an intelligent device that imitates the basic appearance and interaction methods of humans. For example, it can have multiple sets of joints distributed in the mechanical arm, torso, head and other parts, and can flexibly realize joint movements similar to human activities such as raising hands, pointing, and turning around. It can also support the output of voice information to interact with users, and can simulate facial expressions such as smiling through facial screens or lighting components.
[0019] In a specific implementation, the target robot 200 may also be configured with acquisition devices such as image acquisition devices (e.g., cameras) and audio acquisition devices. When obtaining user authorization requests (e.g., prompting the user via voice or page display content that the acquisition device will collect user images or voice to support subsequent shopping guide services, and detecting the user's voice reply confirmation information or click operation on the confirmation control, etc.), the target robot 200 can obtain user interaction information such as actions and / or voice through the acquisition devices, and then send the user interaction information collected by the acquisition devices to the robot controller, so that the robot controller 100 can implement shopping guide services for the user based on the user interaction information.
[0020] Specifically, the robot controller 100 can parse the user interaction information to determine the type of shopping guide service and the target vehicle component, thereby accurately acquiring the user's actual needs. Then, based on the shopping guide service type, it retrieves the first component information of the target vehicle component from a preset vehicle knowledge base to improve the adaptability of the acquired information to the user's actual needs and service scenario. Subsequently, based on the shopping guide service type and the first component information, it determines a target action sequence including joint movements and voice movements from a preset action library. The voice movement is used to output voice-guided shopping information determined based on the first component information, achieving coordination between joint movements and voice-guided shopping information. Finally, it sends a first control command to the target robot 200, instructing it to execute the target action sequence. The target robot 200 is mainly responsible for controlling its own joint movements and voice movements according to the first control command to complete the target action sequence and achieve interaction with the user.
[0021] Ultimately, the robot controller 100 and the target robot 200 in the robot control system work together to respond to user interaction information, realize the robot's multimodal action output, support the realization of personalized action interaction tasks in the sales scenario, and thus help improve the functionality and intelligence of the humanoid robot as an intelligent shopping guide robot.
[0022] In addition, such as Figure 1 As shown, the robot control system may also include a server 300 and a terminal device 400, and the server 300 may be communicatively connected to the terminal device 400 and the robot controller 100 respectively.
[0023] Specifically, the terminal device 400 can be various types of terminal devices such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals. The server 300 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The server 300, terminal device 400, and robot controller 100 can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.
[0024] In specific implementation, server 300 can receive configuration information from terminal device 400, enabling the configuration and updating of information in the preset vehicle knowledge base and the configuration and updating of action sequences in the preset action library. It can also synchronize updated information in the vehicle knowledge base and action library to robot controller 100, updating information stored locally on robot controller 100. Furthermore, in practical applications, robot controller 100 can also utilize Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and large language models. Model (LLM) and other technologies support user interaction and user intent recognition (i.e., deploying an LLM model for the vertical vehicle sales domain, pre-loading a vehicle knowledge base specific to the vehicle sales domain, and integrating dynamically updated action libraries and real-time information from the vehicle knowledge base into the LLM model or Vector Lookup Array (VLA). VLA can convert text, documents, and other information in the knowledge base into computable vector forms for storage, and quickly retrieve the knowledge fragments that best match the user's needs during interaction, ensuring that the LLM calls the latest domain content. When interacting with the user, the robot controller triggers the interaction process based on the user input through the main control program MCP. First, the LLM parses the intent, and then calls the VLA through the Function Calling mechanism to obtain knowledge such as vehicle component information stored in the vehicle knowledge base, or processes multimodal user interaction information through cascaded small models such as speech recognition models and image recognition models. Finally, it integrates the results of all modules to generate a response or action sequence instruction that meets the domain requirements, i.e., generates the target action sequence, and controls the target robot to complete the interaction with the user through the first control instruction). The updates of such technical parameters and model parameters can also be configured locally on the robot controller 100 by the server 300.
[0025] In practice, when users configure and update motion sequences in the motion library through the terminal device 400, they can fill in the configuration file according to the preset format template to set timing information such as the start and end times of joint movements, as well as parameter information such as the direction and angle of joint movement. The server can receive and save the motion sequence in the motion library after the motion sequence has passed the test. Users can also upload category tags for each motion through the terminal device 400 so that the robot controller can query and use it as needed later.
[0026] Understandable Figure 1 This is just an example of a robot control system. In practical applications, the types and number of devices included in a robot control system can be much larger. Figure 1The number of target robots and robot controllers can be different. For example, the number of target robots and robot controllers can be an integer greater than 1, such as 2. Alternatively, the target robot can be equipped with other displays besides the face display screen to support the display of page information to enrich interactive content, etc. No specific restrictions are imposed here.
[0027] Please see Figure 2 , Figure 2 This is a flowchart illustrating a robot multimodal interactive control method provided in an embodiment of this application. This method can be applied to, for example... Figure 1 The robot control system shown includes a robot controller and a target robot, such as... Figure 2 As shown, the multimodal interactive control method for this robot includes: Step S201: The robot controller acquires user interaction information.
[0028] The user interaction information includes user voice information and / or user action information.
[0029] In practice, after obtaining user authorization through voice inquiry or page display, the target robot can collect user voice information through a voice acquisition device or acquire user images through an image acquisition device, and determine user action information by recognizing user images.
[0030] In step S202, the robot controller parses the user interaction information to determine the type of shopping guide service and the target vehicle parts.
[0031] Specifically, when user interaction information includes user voice information, the robot controller can convert the user voice information into text using speech-to-text technology, and then determine the type of shopping guide service and the target vehicle component by keyword extraction and tag matching, or by using a pre-trained intent recognition model. The intent recognition model can be, for example, a large language model (LLM).
[0032] Specifically, when determining the service type through the intent recognition model, the robot controller can input text into the intent recognition model, which will then output a service type label. This label can include, for example, one of the following: vehicle information explanation, vehicle comparison, operation demonstration, operation guidance, vehicle location guidance, or process guidance. Correspondingly, each action in the preset action library can have a corresponding service type label, facilitating the subsequent determination of the target action sequence based on these labels. For instance, if the intent recognition model outputs a service type label of "vehicle information explanation," the target action sequence can be determined based on the action in the action library corresponding to this label.
[0033] Specifically, the intent recognition model can also output the weight probability value of the shopping guide type label. When the weight probability value is greater than a preset value, the output shopping guide service type label is directly used as the shopping guide service type. When the weight probability value is less than or equal to the preset value, the target robot is controlled to ask the user whether to select the shopping guide service type label as the shopping guide service type through voice inquiry or page display. If a confirmation operation is received, the shopping guide type label is used as the shopping guide service type. Otherwise, the robot controller can re-acquire user interaction information and redetermine the shopping guide service type.
[0034] Furthermore, the intent recognition model can output a comprehensive tag, which includes not only a shopping guide type tag but also an explanation style tag and its corresponding weighted probability value. The explanation style tag can include, for example, one of the following: general explanation, professional explanation, or everyday explanation. Correspondingly, the component information for the same vehicle part in the vehicle knowledge base can also include explanation text corresponding to different explanation style tags, facilitating the subsequent determination of the explanation text and the voice shopping guide information to be output based on the explanation style tags.
[0035] In practical applications, the target action sequence can also include facial expressions, and the comprehensive label can also include other types of labels such as emotion labels. Based on the combination of information of multiple types of labels and their corresponding weight probability values, the target action sequence is determined together to further improve the flexibility and efficiency of action sequence determination.
[0036] Similarly, when determining the type of shopping guide service through keyword extraction and tag matching, the robot controller can also convert the user's voice information into text, extract keywords from the text and compare them with preset tags to determine the shopping guide service type tag, explanation style tag, and emotional tag.
[0037] Regarding the determination of target vehicle parts, after the robot controller converts the user's voice information into text, it can also determine the target vehicle parts by extracting information such as vehicle model and part name directly contained in the text through keyword extraction. Alternatively, it can determine the functional tags corresponding to the user's interaction information through an intent recognition model or keyword extraction. Correspondingly, the functional tags corresponding to each vehicle and part can be pre-stored in the preset vehicle knowledge base. The robot controller can then determine the corresponding target vehicle parts from the preset vehicle knowledge base based on the functional tags corresponding to the user's interaction information.
[0038] When user interaction information includes user action information, the robot controller can specifically identify user image information to determine user orientation, hand movements, facial expressions, and other user action information. Based on the determined user action information, it can determine the type of shopping guide service and the target vehicle part. For example, when the target robot is providing operation guidance for a certain vehicle part and detects that the user's facial expression indicates confusion, it can continue to identify that vehicle part as the target vehicle part, and determine the shopping guide service type as a function demonstration or operation demonstration. It can also adjust the subsequent target action sequence and the output voice shopping guide information accordingly, changing from guiding the user to operate independently to having the target robot demonstrate the operation first. Another example is that the robot controller can identify the vehicle part that the user's hand gestures are pointing at as the target vehicle part.
[0039] Furthermore, when user interaction information includes user action information and user voice information, the type of shopping guide service and the target vehicle part can be determined primarily based on the user's voice information. If the type of shopping guide service cannot be determined based on the user's voice information, then the user's action information can be used to determine it. For example, if the user's voice information includes the keyword "this" but lacks other descriptive information about the vehicle or part, image recognition can be used to determine that the user's hand gesture is pointing to a certain vehicle part, and thus that vehicle part can be identified as the target vehicle part. Specifically, the robot controller can use an image recognition model cascaded with the intent recognition model to determine the user action information and vehicle part in the image captured by the target robot.
[0040] In step S203, the robot controller obtains the first component information of the target vehicle component from a preset vehicle knowledge base according to the type of shopping guide service.
[0041] The vehicle knowledge base stores vehicle component information, which may include component explanation text and component operation parameters. When determining the first component information based on the type of sales guide service, the first component information can be differentiated based on whether the sales guide service type requires the target robot to directly operate the vehicle. For example, for sales guide service types that do not require the target robot to directly operate the vehicle, such as vehicle information explanation, vehicle comparison, operation guidance, vehicle location guidance, and process guidance, the target robot can complete the service through general interactive actions such as pointing combined with voice guidance information. Therefore, the first component information may not include component operation parameters and may consist of the component explanation text of the target vehicle component. However, for sales guide service types that require the target robot to directly operate the vehicle, such as operation demonstrations, the determined first component information must include not only the explanation text of the target vehicle component but also the component operation parameters of the target vehicle component. This allows the target robot to accurately execute the specific demonstration operation actions for the target vehicle component based on the target action sequence determined by the first component information, thereby accurately achieving the operation demonstration.
[0042] In step S204, the robot controller determines the target action sequence from a preset action library based on the shopping guide service type and the information of the first component.
[0043] The target action sequence includes joint movements and voice movements. The voice movements are used to output voice-guided shopping information, which is determined based on the first component information.
[0044] In practical implementation, the target action sequence may include multiple actions, as well as the execution order of each action or time information such as the start and end timestamps corresponding to each action, so as to complete the actions in an orderly manner. In cases where some joint actions and voice actions need to be executed synchronously, the explanatory text in the first component information obtained by different tasks may differ, thus the voice-guided shopping information determined based on the first component information may also differ, and the duration of the voice-guided shopping information output by the voice action may also vary. The robot controller can fine-tune the duration of the joint action according to the required duration of the voice action. For example, if the execution time of the joint action is preset to an adjustable range, and the duration of the voice action is within this range, the duration of the joint action can be directly adjusted according to the duration of the voice action. Alternatively, if the duration of the voice action is outside this adjustable range, the number of words in the voice-guided shopping information or the output speed (speech rate) of the voice-guided shopping information corresponding to the voice action can be adjusted according to the duration range of the joint action, thereby adjusting the duration of the voice action to match the duration of the joint action.
[0045] In step S205, the robot controller sends a first control command to the target robot.
[0046] The first control command is used to instruct the target robot to execute the target action sequence.
[0047] In practice, when the target robot executes the target action sequence, it can execute joint actions in an orderly manner and output voice shopping guide information according to the action execution order and other information set in the sequence, so as to realize multimodal interaction with the user.
[0048] As can be seen, in this embodiment, the robot controller first acquires user interaction information containing user voice information and / or user action information, then parses the user interaction information to determine the type of shopping guide service and the target vehicle component to accurately obtain the user's actual needs. Next, based on the type of shopping guide service, it obtains the first component information of the target vehicle component from a preset vehicle knowledge base to improve the adaptability of the acquired information to the user's actual needs and service scenarios. Subsequently, based on the type of shopping guide service and the first component information, it determines the target action sequence containing joint actions and voice actions from a preset action library. The voice actions are used to output the voice shopping guide information determined based on the first component information, realizing the coordination of joint actions and voice shopping guide information. Finally, it sends a first control command to the target robot to instruct it to execute the target action sequence. Ultimately, it can respond to user interaction information and realize the robot's multimodal action output, supporting the realization of personalized action interaction tasks in sales scenarios, thereby improving the functionality and intelligence of the humanoid robot as an intelligent shopping guide robot.
[0049] In one possible example, the target vehicle component is determined through the following steps: parsing the user interaction information to determine whether the user interaction information indicates a first component; if it indicates a first component, identifying the first component as the target vehicle component; if it does not indicate a first component, identifying a first vehicle based on the user interaction information, the first vehicle including a first number of preset tags matching the user interaction information; identifying a second component from the vehicle components corresponding to the first vehicle based on the first number of preset tags; obtaining first priority information of the second component corresponding to the first vehicle; and identifying the target vehicle component from the second component based on the first priority information.
[0050] In this context, the user interaction information indicating the first component can be, for example, explicitly indicating that the user interaction information wants to know about a specific component of a specific vehicle, in which case the specific component of that specific vehicle (i.e., the first component) can be identified as the target vehicle component.
[0051] Among them, the preset tags can be, for example, the aforementioned function tags. That is, when the user does not have a specific vehicle or component they want to know about, the function tags of the vehicle functions that the user is interested in can be determined based on the user's interaction information. Then, the robot controller can query the function tags corresponding to each vehicle stored in the vehicle knowledge base, and then determine the first vehicle whose function tag matches the function tag corresponding to the user's interaction information.
[0052] The value of the first quantity can be set as needed. For example, a ratio can be preset, and the first quantity is determined based on the ratio and the number of function tags corresponding to the user interaction information. Assuming the ratio is 0.8 and the number of function tags corresponding to the user interaction information is 10, the first quantity can be 8. That is, the function tags corresponding to the vehicle include 8 out of the 10 function tags corresponding to the user interaction information, and it can be identified as the first vehicle. Of course, if there is only one function tag corresponding to the user interaction information, then the first quantity is 1 regardless of the ratio. That is, the function tags corresponding to the vehicle include only one function tag corresponding to the user interaction information, and it can be identified as the first vehicle.
[0053] Additionally, if the first vehicle comprises multiple vehicles, the robot controller can also control the target robot to display, for example... Figure 3 The vehicle selection page displays images of multiple vehicles, along with vehicle descriptions and location information, allowing users to select a vehicle of interest. Specifically, the displayed vehicle images can include the vehicle's exterior or a localized area of a vehicle component associated with a function tag. The vehicle description can also be descriptive information corresponding to the function tag. Furthermore, the vehicle images and descriptions can be retrieved from a vehicle knowledge base, and the vehicle location can specifically be the vehicle's position within the current sales showroom.
[0054] After identifying the first vehicle, a list of vehicle components corresponding to the first vehicle can be obtained, and components that match the first number of functional tags can be selected as the second components. At this time, if a single component matches at least one of the first number of functional tags, it can be identified as the second component. For example, if the first number of functional tags includes tag 1, tag 2, tag 3 and tag 4, component 1 corresponds to tag 1 and can be identified as the second component. Component 2 corresponds to tags 2 and 4 and can also be identified as the second component.
[0055] The first priority information can be pre-associated with the first vehicle and stored in the vehicle knowledge base. In other words, for example, the marketing manager of the first vehicle can pre-configure the first priority information of each component in the first vehicle according to marketing needs. For example, the priority of the component corresponding to the vehicle's selling points can be set higher. When the robot controller determines the target vehicle component from the second component, it can give priority to the component corresponding to the vehicle's selling points, thereby improving the flexibility and intelligence of the determination of the target vehicle component.
[0056] In specific implementation, the target vehicle component is determined from the second components. For example, it can be sorted according to the first priority information, and a preset number of second components that rank first are identified as target vehicle components. Each target vehicle component is then explained or demonstrated in sequence according to its priority. At this time, the robot controller can determine the movement route to the first vehicle based on the robot's position, the pre-stored exhibition hall map, the position information of the first vehicle in the exhibition hall, and the priority of each component in the target vehicle component. In the case of multiple target vehicle components, it can further determine the path around the first vehicle when explaining or demonstrating it. For example, the starting point of this path can be the location of the component with the highest first priority among the target vehicle components, and the route to the first vehicle is the movement route in the preset safe movement path on the exhibition hall map, starting from the robot's position and ending at the location of the component with the highest first priority among the first vehicle components.
[0057] When multiple target vehicle components are identified at once, the robot controller can determine the corresponding target action sequence for each component. Following the explanation or demonstration order of the target vehicle components, the controller sets the execution order of these action sequences. Furthermore, after the execution of the previous target vehicle component's action sequence ends and before the start of the next, it connects the target action sequences of different components using pointing and moving actions. In addition to the pre-determined target vehicle components, the robot controller can also add target vehicle components in real-time based on the latest user interaction information and adjust the explanation or demonstration order of subsequent target vehicle components. For example, after explaining component 1, if the user lingers at component 3 for a considerable time while moving to component 2, component 3 can be identified as the target vehicle component. The robot controller can then explain or demonstrate component 3 before proceeding to explain or demonstrate component 2.
[0058] In particular, if the order of explanation of the target vehicle parts is strictly determined according to the order of the first priority information, and the corresponding path around the vehicle is determined, there may be a problem that the target robot leads the user to repeatedly turn back and forth, which will affect the user experience. The path around the vehicle can be directly taken as the starting point of the target vehicle part with the highest first priority, and the direction of the path around the vehicle can be taken as the direction of the target vehicle part that is closer to the second highest first priority target vehicle part. The robot will always move along this direction to explain around the first vehicle. At this time, the target robot can collect vehicle images and identify whether the vehicle part at the current location is the target vehicle part by recognizing the vehicle images through cascaded small models, and then determine whether to explain or demonstrate the part.
[0059] Furthermore, in practical applications, if it is determined from user interaction information that a user has a specific vehicle they want to know about, but no specific component they want to know about, that specific vehicle can be directly identified as the first vehicle. Similarly, the function tag can be determined through user interaction information, and then the second component in the first vehicle can be directly matched through the function tag. In this way, the second component can be determined based on the first priority information corresponding to the first vehicle.
[0060] Alternatively, in practical applications, if user interaction information indicates that a user clearly wants to know about a specific vehicle and component, but the user cannot accurately provide the vehicle model and component name, the robot controller can still determine the appearance and function labels based on the user interaction information. Furthermore, based on these labels, it can identify possible vehicles and components. If the vehicle's location in the showroom is less than a preset distance from the target robot's location, the target robot can directly instruct the user to view and confirm via a pointing gesture. If the vehicle's location in the showroom is greater than or equal to the preset distance from the target robot's location, the target robot can still, for example... Figure 3 The vehicle selection page displays vehicle images and information for users to view and confirm. Figure 3 The example uses five images, from vehicle image 1 to vehicle image 5. All five images can be of the same vehicle, and each image can be a partial image of the vehicle corresponding to an appearance tag or function tag in the user interaction information. The vehicle description information can be the descriptive information associated with the vehicle's appearance tag or function tag. Alternatively, the target robot can match multiple possible vehicles, with each of the five images corresponding to a different vehicle. Users can click on an image to zoom in on the location of vehicle image 1, and the vehicle information below will specifically be the information of that zoomed-in vehicle; no specific restrictions are imposed here. Based on this, the robot controller can quickly determine the vehicle and component that the user actually wants to learn about.
[0061] In particular, when determining the target vehicle component based on the aforementioned user interaction information, the user interaction information may include not only the latest interaction information obtained in real time, but also historical interaction information authorized by the user within the most recent period.
[0062] As can be seen, in this example, when the user interaction information indicates a specific first component, the robot controller directly identifies the first component as the target vehicle component to accurately match the user's needs. However, when the user interaction information does not explicitly indicate a specific component, it first identifies a first vehicle that includes a first number of preset tags that match the user interaction information. Then, it further identifies a second component that matches the first number of tags from the vehicle components of the first vehicle. This process gradually filters out the second component that matches the user's needs based on the user interaction information. Finally, it obtains the first priority information of the second component corresponding to the first vehicle and further identifies the target vehicle component from the second component based on the first priority information. This improves the flexibility of determining the first vehicle while matching the user's needs.
[0063] In one possible example, determining the target vehicle component from the second component based on the first priority information includes: determining the user's level of attention to each preset tag among the first number of preset tags based on the user interaction information; determining the second priority information of the second component based on the level of attention to each preset tag; determining the third priority information of the second component based on the first priority information and the second priority information; and determining the target vehicle component from the second component based on the third priority information.
[0064] Specifically, based on the attention level of each preset tag, the second priority information of the second component is determined. For example, for each second component, the attention level of the tag with the highest attention level among its corresponding preset tags can be used as the attention level of the second component, and the second priority of the second component can be further determined based on the attention levels of each of the multiple second components.
[0065] In a specific implementation, the third priority information of the second component is determined based on the first priority information and the second priority information. For example, for each second component, the average or weighted average of its first priority and second priority is calculated to obtain the third priority information of the second component.
[0066] As can be seen, in this example, the robot controller determines the user's level of attention to each of the first number of preset tags based on user interaction information, and determines the second priority information of the second component based on the level of attention information. Then, it combines the second priority information and the first priority information to comprehensively determine the third priority information of each second component, and finally determines the target vehicle component from the second components based on the third priority information. This helps to further improve the flexibility and intelligence of vehicle component determination.
[0067] In one possible example, obtaining first component information of the target vehicle component from a preset vehicle knowledge base according to the shopping guide service type includes: if the shopping guide service type is a preset type, obtaining first component information of the target vehicle component from the vehicle knowledge base, the first component information including first explanatory text and first operation parameters; determining a target action sequence from a preset action library according to the shopping guide service type and the first component information includes: if the first operation parameters indicate that the target vehicle component supports demonstration operation, determining whether a first action set in the action library includes a first action sequence matching the first component identifier of the target vehicle component; if it includes, determining the target action sequence based on the first action sequence and the first explanatory text; if it does not include, obtaining second component information of the target vehicle component from the vehicle knowledge base, the second component information including second explanatory text; determining a second action sequence from the second action set in the action library based on the second explanatory text; and determining the target action sequence based on the second action sequence and the second explanatory text.
[0068] The preset type could be, for example, an operation demonstration.
[0069] The action library can be divided into a first action set and a second action set. The first action set can save a sequence of exclusive operation actions for a specific component, so that it can be directly called to demonstrate the operation of the specific component. The second action set can save general actions that are not directly related to the component, such as pointing, nodding, and shaking the head.
[0070] In specific implementation, the first operation parameter can include two parts. One part identifies whether the component supports demonstration operation. For components that support demonstration operation, the first operation parameter can include specific operation parameters related to the component, so that the specific operation action sequence (i.e., the first action sequence) for the target vehicle component can be adjusted as needed. For components that do not support demonstration operation, the first operation parameter can include parameter information such as the component's function triggering method and its state after triggering. For example, some component functions in a vehicle, such as airbag deployment, cannot be safely demonstrated by the target robot in the showroom or may damage the vehicle. However, user interaction information may indicate that an operation demonstration should be performed for such components that cannot be demonstrated. In this case, although the robot control will not match the first action sequence according to the first component identifier to perform the operation demonstration, it can generate a simulated operation scenario explanation text according to the first operation parameter, and further combine it with the first explanation text to generate the final explanation text for the target vehicle component. Based on the final explanation text, it can obtain general actions from the second action set to generate the final target action sequence to achieve supplementary explanation.
[0071] In scenarios where the first action set does not include the first action sequence specific to the target vehicle component, if the target robot generates and executes the action sequence on its own without simulation, it may damage the vehicle or affect user safety. Therefore, the robot controller can switch the shopping guide service type from the original operation demonstration to explanation or guidance. That is, the second component information is no longer included in the operation parameters, but the target action sequence is determined directly based on the second explanation text to better help users understand the operation method or guide users to complete the operation on their own.
[0072] Accordingly, for each component, the pre-stored explanatory text in the vehicle knowledge base can be categorized into several types, such as general explanatory sub-text, explanatory sub-text corresponding to non-preset types of sales guide services, explanatory sub-text corresponding to preset types, and explanatory sub-text converted from preset types to non-preset types. For example, explanatory sub-text corresponding to preset types typically includes more detailed descriptions of operational actions, while explanatory sub-text corresponding to non-preset types typically includes more descriptions of component appearance and functional application scenarios. Explanatory sub-text converted from preset types typically includes more descriptions of function triggering conditions and detailed descriptions of simulated operational actions. Based on this, the first explanatory text can include both general explanatory sub-text and explanatory sub-text corresponding to preset types, while the second explanatory text can include both general explanatory sub-text and explanatory sub-text converted from preset types to non-preset types. This allows for the rapid acquisition of explanatory text adapted to different scenario requirements based on the sales guide service type.
[0073] Furthermore, if the shopping guide service type is not the preset type, i.e., operation demonstration, the robot controller can directly obtain the third explanatory text as the first component information. The third explanatory text may specifically include general explanatory sub-text and explanatory sub-text corresponding to the non-preset type of shopping guide service. The robot controller can determine the third action sequence from the second action set based on the third explanatory text, and determine the target action sequence based on the second action sequence and the second explanatory text.
[0074] Specifically, when determining the second and third action sequences, the robot controller can match the corresponding general action sequences based on the keywords in the second or third explanatory text. When determining the target action sequence based on the action sequence and the explanatory text (for example, determining the target action sequence based on the first action sequence and the first explanatory text, or determining the second action sequence based on the second action sequence and the second explanatory text, or determining the target action sequence based on the third action sequence and the third explanatory text), the controller can refer to the aforementioned output duration of the voice shopping guide information corresponding to the explanatory text and the duration of the joint movements in the action sequence to fine-tune the duration of the voice movements or joint movements to obtain the final target action sequence.
[0075] In other embodiments, for scenarios where the first action set does not include a first action sequence specific to the target vehicle component, and the robot controller replaces the shopping guide service type from operation demonstration to explanation or operation guidance, the robot controller can also determine a specific operation action sequence of the same component type from similar models in the first action set based on the component type and vehicle model of the target vehicle component, and modify its operation parameters according to the operation parameters of the target vehicle component, thereby obtaining the operation action sequence to be verified, and uploading it to the server for simulation verification. The server also sends a verification notification to the terminal device of the corresponding staff. After the staff confirms that the simulation verification result is correct, the verified operation action sequence is associated with the first component identifier of the target vehicle component and stored in the first action set as the specific operation action sequence of the target vehicle component.
[0076] In addition, after submitting the simulation verification, the robot controller can also confirm with the user, through voice interaction or page display, whether they need to schedule a next demonstration. This includes confirming vehicle information such as the vehicle model, the service item, the time, and the showroom, as well as user information such as the user's name and contact details. This information is used to notify the user to attend the scheduled service. If the reservation information is confirmed via page display, the robot controller can fill in the vehicle and user information obtained from the voice interaction or page operation into a display window. Figure 4 The appointment information confirmation page shown can save the appointment information when a selection action is detected on the confirmation control or when the user confirms via voice.
[0077] As can be seen in this example, when the shopping guide service type is a preset type, the robot controller first obtains the first explanatory text and the first operation parameters of the target vehicle component. When the first operation parameters indicate that demonstration operation is supported, it further determines whether there is a unique first action sequence in the first action set that matches the target vehicle component. If there is, the target action sequence is directly determined based on the unique first action sequence and the first explanatory text, improving the precision of the target action sequence determination. If the target vehicle component does not have a unique first action sequence, the robot controller can re-obtain the second explanatory text of the target vehicle component and determine a general second action sequence from the second action set based on the second explanatory text. Then, it determines a target action sequence based on the second action sequence and the second explanatory text. When the demonstration operation cannot be directly performed, the shopping guide service type is changed to explanation, and the explanation text and action sequence of the explanation scenario are re-obtained to determine the final target action sequence, which helps to further improve the flexibility and intelligence of the target action sequence determination.
[0078] In one possible example, determining the target action sequence based on the first action sequence and the first explanatory text includes: obtaining a first update time of the first action sequence and a second update time of the first operation parameters; if the first update time is earlier than the second update time, determining an update parameter item from the first operation parameters; if the update parameter item is a preset parameter type, obtaining a preset threshold range corresponding to the update parameter item; if the parameter value corresponding to the update parameter item matches the preset threshold range, adjusting the first action sequence according to the update parameter item to obtain a third action sequence; and determining the target action sequence based on the third action sequence and the first explanatory text.
[0079] Among them, the parameter value corresponding to the updated parameter item matches the preset threshold range, that is, the updated parameter value of the parameter item is within the preset threshold range.
[0080] In practice, preset parameter types can be set as needed. For example, for parameters such as the pressing force of vehicle buttons, if relevant personnel make minor adjustments to the values in the vehicle knowledge base without timely updating the parameters in the first action sequence, and such minor adjustments are within the preset threshold range of the parameter, there will be no safety or damage risk. The robot controller can then directly update the first action sequence based on the latest operating parameters to obtain the third action sequence without requiring the user to operate again, thus improving the efficiency of determining subsequent target action sequences. However, for parameters such as the opening and closing angle of vehicle doors, changes in such parameters may cause collisions with other facilities in the actual site. Therefore, these parameters are not set as preset parameter types. For such changes, the user must manually confirm the corresponding action sequence adjustment to ensure safety.
[0081] As can be seen, in this example, when the update time of the first action sequence is earlier than the update time of the first operation parameter, the robot controller directly adjusts the first action sequence according to the updated parameter to obtain a third action sequence that matches the latest operation parameter when the updated parameter item is a preset parameter type and the updated parameter value matches the preset threshold range of the updated parameter item. Furthermore, the robot controller determines the target action sequence based on the third action sequence and the first explanatory text, which helps to further improve the flexibility, intelligence and efficiency of the target action sequence determination.
[0082] In one possible example, the first explanatory text is determined by the following steps: obtaining a first sub-text corresponding to the target vehicle component; determining an explanatory style tag based on the user interaction information; obtaining a second sub-text corresponding to the target vehicle component based on the explanatory style tag; generating the first explanatory text based on the first sub-text and the second sub-text; the second explanatory text is determined by the following steps: obtaining a third sub-text corresponding to the target vehicle component; generating the second explanatory text based on the first sub-text, the second sub-text, and the third sub-text.
[0083] The explanation style tags can include, for example, one of the tags such as: general explanation, professional explanation, and everyday explanation. The first sub-text can be a general explanation text for the component, while different explanation style tags (such as professional explanation and everyday explanation) can correspond to different supplementary text, i.e., the second sub-text. For example, if the target vehicle component is the electric tailgate, the first sub-text can include descriptive information such as "This car's electric tailgate has an anti-pinch protection function. When it encounters an obstacle when closing, it will automatically stop and spring back," covering the core function of the component. The second sub-text corresponding to the professional explanation can include professional descriptive information such as "The anti-pinch system uses millimeter-wave radar and pressure sensor dual-mode detection," while the second sub-text corresponding to the everyday explanation can include usage scenario description information such as "If it accidentally touches a child's hand or shopping bag when closing, it will immediately stop and spring back, as if someone is gently supporting it, so there is no need to worry about being pinched."
[0084] The first explanatory text can be generated by combining the first and second sub-texts. For the second explanatory text, in addition to the first and second sub-texts, a third sub-text specific to the scenario where the service type changes from a preset type to a non-preset type can also be obtained. For example, the third sub-text may include descriptions of function trigger conditions and detailed descriptions of simulated operation actions. These are then combined with the aforementioned general first sub-text, the second sub-text determined based on the explanation style tag, and the third sub-text to generate the second explanatory text. The second sub-text may contain detailed descriptions of operation actions that are synchronously output with the specific operation action sequence, such as descriptions like "Please look at the position of my finger pressing." Since no operation action targeting the vehicle component is performed at this time, such sub-texts can be deleted.
[0085] As can be seen in this example, when the shopping guide type is a preset type and demonstration operation is supported, the robot controller distinguishes whether the target vehicle part has a unique first action sequence in the first action set and sets a differentiated explanation text determination process. When there is a unique first action sequence, the first explanation text is determined based on the first sub-text and the second sub-text corresponding to the explanation style tag determined by parsing user interaction information. When there is no unique first action sequence, the second explanation text is generated by combining the third sub-text specially set for this scenario. This helps to avoid the problem that setting only a single explanation text cannot adapt to changes in the shopping guide service scenario, and further improves the flexibility and intelligence of explanation text determination.
[0086] In one possible example, the first action sequence includes multiple joint action sub-sequences, the multiple joint action sub-sequences including a first sequence, and the step of determining the target action sequence based on the first action sequence and the first explanatory text includes: obtaining the operation type label and duration constraint information corresponding to the first sequence; determining a fourth sub-text from the first explanatory text based on the operation type label; determining first speech rate information based on the user interaction information; determining a first duration based on the fourth sub-text and the first speech rate information; and determining a second duration of the first voice action corresponding to the first sequence and the first voice shopping guide information corresponding to the first voice action based on the first duration, the duration constraint information, and the fourth sub-text.
[0087] In practical implementation, the division of different joint action sub-sequences can be based on the operation type. For example, taking the first action sequence as a demonstration operation sequence for the power tailgate, the first action sequence can include a joint action sub-sequence labeled "unlocking the power tailgate," where the joint action is used to press the unlock button, and a joint action sub-sequence labeled "anti-pinch simulation blocking," where the joint action is used to prevent the tailgate from closing. Both of these joint action sub-sequences have voice guidance information that needs to be output synchronously, meaning that the voice action needs to be performed simultaneously with the joint action. Taking the operation type label corresponding to the first sequence as "unlock button pressing" as an example, the explanatory text related to pressing the unlock button can be determined from the first explanatory text as the fourth sub-text based on its operation type label.
[0088] Specifically, the duration constraint information can be the maximum execution duration corresponding to the first sequence. After the maximum execution duration ends, the joint action will reset to allow for the execution of subsequent joint actions. For example, if the maximum execution duration corresponding to the first sequence being the unlock button press is 8 seconds, then the target robot must complete the process from starting the action to pressing the unlock button to resetting the joint within 8 seconds. Therefore, the second duration of the first voice action corresponding to the first sequence must be completed within 8 seconds to ensure synchronization between the joint action and the voice action.
[0089] The first speech rate information could be, for example, the average speech rate of the user within a recent preset time period, determined based on user interaction information. The first duration could be the duration required to output the voice-guided shopping information corresponding to the fourth sub-text based on the first speech rate information.
[0090] Specifically, the second duration and the content of the first voice shopping guide information are determined based on the first duration, duration constraint information, and the fourth sub-text. Specifically, if the first duration is greater than the duration constraint information, the first duration corresponding to the first voice action needs to be adjusted to the second duration, and part of the text in the fourth sub-text is deleted to obtain the first voice shopping guide information, so that the output duration required for the adjusted voice shopping guide information under the first speech rate information is the second duration; if the first duration is not greater than the duration constraint information, the first duration can be directly determined as the second duration, and the fourth sub-text can be determined as the first voice shopping guide information.
[0091] For example, taking the first sequence as the unlock press and the duration constraint as 8 seconds, if the fourth text contains 27 words and the first speech rate is 3 words per second, then the first duration is 9 seconds. Since the first duration is greater than 8 seconds, some information in the fourth text needs to be deleted to obtain the first voice shopping guide information. For example, after deletion, the first voice shopping guide information will have 15 words remaining, and the second duration will be 5 seconds, which is less than the duration constraint.
[0092] In particular, when the first duration exceeds the duration constraint information and some text in the fourth sub-text needs to be deleted, the text type of the fourth sub-text can be further subdivided to determine the text to be deleted. For example, the aforementioned general explanatory sub-text or supplementary text outside the first sub-text can be deleted first, while retaining the most core functional description. Alternatively, compressible or incompressible tags can be set in advance for different sub-texts corresponding to the target component, and the sub-texts with compressible tags can be deleted first.
[0093] Specifically, if only the general explanation subtext, the first subtext, or the subtext of the uncompressible tag in the fourth subtext is retained, and the duration required for outputting the information at the first speech rate is still greater than the duration constraint information, the speech rate can be increased to reduce the duration required for the voice action and adapt to the duration constraint information of the joint actions in the first sequence. It is understandable that the speech rate adjustment can also be set to a range, and can only be adjusted within the set range. If the speech rate still cannot adapt to the duration constraint information within the set range, the robot controller can add a voice action and a corresponding joint action sequence after the first sequence to ensure that the voice-guided shopping information at least covers the core functional description of the component. At this time, the joint action sequence can be used to control the target robot to face the user for explanation, ensuring a good user interaction experience. For example, after the target robot finishes pressing the unlock button, it faces the user and maintains the voice action output, and then enters the next joint action subsequence and executes the corresponding voice action.
[0094] As can be seen in this example, for each joint action subsequence in the first action sequence, the robot controller first obtains the operation type label and duration constraint information corresponding to the first sequence, then determines the fourth sub-text corresponding to the first sequence from the first explanatory text based on the operation type label, and determines the first speech rate information based on the user interaction information. Then, it determines the first duration based on the fourth sub-text and the first speech rate information. Finally, by combining the first duration required to output the fourth sub-text at the first speech rate, the duration constraint information of the first sequence, and the fourth sub-text, the fourth sub-text is adjusted to determine the voice shopping guide information that needs to be output synchronously with the first sequence, and the duration of the first voice action for outputting the first voice shopping guide information is determined. This helps to further ensure the synchronization of voice actions and joint actions, and improves the flexibility and intelligence of the target action sequence determination.
[0095] Please see Figure 5 , Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of this application. Specifically, the electronic device 30 may be the aforementioned robot controller, which can be used to execute the above-described method. Specifically, the electronic device 30 may include a processor 310, a memory 320, a communication interface 330, and one or more programs 321. The one or more programs 321 are stored in the memory 320 and configured to be executed by the processor 310. The one or more programs 321 include instructions for performing any step executed by the robot controller in the above-described method embodiment.
[0096] The communication interface 330 is used to support communication between the electronic device 30 and other devices. The processor 310 may be, for example, a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, units, and circuits described in conjunction with the embodiments of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0097] The memory 320 can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SynchLinkDRAM, SLDRAM), and direct memory bus RAM (DRRAM).
[0098] In a specific implementation, the processor 310 is used to execute any step in the above method embodiments, and when performing data transmission such as sending, it can choose to call the communication interface 330 to complete the corresponding operation.
[0099] It should be noted that the above schematic diagram of the electronic device 30 is only an example, and the actual number of components included may be more or less, and no single limitation is made here.
[0100] This application can divide the device into functional units based on the above method examples. For example, each function can be divided into its own functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0101] Figure 6 This is a functional unit block diagram of a robot multimodal interaction control device provided in an embodiment of this application. This robot multimodal interaction control device can be applied to, for example... Figure 1 The robot control system shown includes a robot controller and a target robot. The robot multimodal interaction control device includes: The first acquisition unit 401 is used to acquire user interaction information, which includes user voice information and / or user action information. The parsing unit 402 is used to parse the user interaction information and determine the type of shopping guide service and the target vehicle parts; The second acquisition unit 403 is used to acquire the first component information of the target vehicle component from a preset vehicle knowledge base according to the type of the shopping guide service. The determining unit 404 is used to determine a target action sequence from a preset action library based on the shopping guide service type and the first component information. The target action sequence includes joint actions and voice actions. The voice actions are used to output voice shopping guide information. The voice shopping guide information is determined based on the first component information. The sending unit 405 is used to send a first control command to the target robot, the first control command being used to instruct the target robot to execute the target action sequence.
[0102] In one possible example, the parsing unit 402 is specifically configured to determine the target vehicle component through the following steps: parsing the user interaction information to determine whether the user interaction information indicates a first component; if it indicates a first component, determining the first component as the target vehicle component; if it does not indicate a first component, determining a first vehicle based on the user interaction information, the first vehicle including a first number of preset tags matching the user interaction information; determining a second component from the vehicle components corresponding to the first vehicle based on the first number of preset tags; obtaining first priority information of the second component corresponding to the first vehicle; and determining the target vehicle component from the second component based on the first priority information.
[0103] In one possible example, regarding the determination of the target vehicle component from the second component based on the first priority information, the parsing unit 402 is specifically configured to: determine the user's level of attention to each preset tag among the first number of preset tags based on the user interaction information; determine the second priority information of the second component based on the level of attention to each preset tag; determine the third priority information of the second component based on the first priority information and the second priority information; and determine the target vehicle component from the second component based on the third priority information.
[0104] In one possible example, the second acquisition unit 403 is specifically configured to: if the shopping guide service type is a preset type, acquire first component information of the target vehicle component from the vehicle knowledge base, the first component information including first explanatory text and first operation parameters; the determination unit 404 is specifically configured to: if the first operation parameters indicate that the target vehicle component supports demonstration operation, determine whether the first action set of the action library includes a first action sequence matching the first component identifier of the target vehicle component; if it includes, determine the target action sequence based on the first action sequence and the first explanatory text; if it does not include, acquire second component information of the target vehicle component from the vehicle knowledge base, the second component information including second explanatory text; determine a second action sequence from the second action set of the action library based on the second explanatory text; determine the target action sequence based on the second action sequence and the second explanatory text.
[0105] In one possible example, regarding the determination of the target action sequence based on the first action sequence and the first explanatory text, the determining unit 404 is specifically configured to: obtain a first update time of the first action sequence and a second update time of the first operation parameters; if the first update time is earlier than the second update time, determine an update parameter item from the first operation parameters; if the update parameter item is a preset parameter type, obtain a preset threshold range corresponding to the update parameter item; if the parameter value corresponding to the update parameter item matches the preset threshold range, adjust the first action sequence according to the update parameter item to obtain a third action sequence; and determine the target action sequence based on the third action sequence and the first explanatory text.
[0106] In one possible example, the second acquisition unit 403 is specifically used to determine the first explanatory text through the following steps: acquiring a first sub-text corresponding to the target vehicle component; determining an explanation style tag based on the user interaction information; acquiring a second sub-text corresponding to the target vehicle component based on the explanation style tag; and generating the first explanatory text based on the first sub-text and the second sub-text. The determining unit 404 is specifically used to determine the second explanatory text through the following steps: acquiring a third sub-text corresponding to the target vehicle component; and generating the second explanatory text based on the first sub-text, the second sub-text, and the third sub-text.
[0107] In one possible example, the first action sequence includes multiple joint action sub-sequences, the multiple joint action sub-sequences including a first sequence. In determining the target action sequence based on the first action sequence and the first explanatory text, the determining unit 404 is specifically configured to: obtain the operation type label and duration constraint information corresponding to the first sequence; determine a fourth sub-text from the first explanatory text based on the operation type label; determine first speech rate information based on the user interaction information; determine a first duration based on the fourth sub-text and the first speech rate information; and determine a second duration of the first voice action corresponding to the first sequence, and first voice shopping guide information corresponding to the first voice action, based on the first duration, the duration constraint information, and the fourth sub-text.
[0108] In the case of using integrated units, the functional unit composition block diagram of another robot multimodal interactive control device provided in this application embodiment is as follows: Figure 7 As shown. In Figure 7The robot multimodal interaction control device includes a processing module 520 and a communication module 510. The processing module 520 is used to control and manage the actions of the robot multimodal interaction control device, for example, the steps executed by the first acquisition unit 401, the parsing unit 402, the second acquisition unit 403, the determining unit 404, and the sending unit 405, and / or other processes used to execute the techniques described herein. The communication module 510 is used to support the interaction between the robot multimodal interaction control device and other devices. Figure 7 As shown, the robot multimodal interaction control device may also include a storage module 530, which is used to store the program code and data of the robot multimodal interaction control device.
[0109] The processing module 520 can be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the embodiments of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication module 510 can be a transceiver, RF circuitry, or a communication interface, etc. The storage module 530 can be a memory.
[0110] All relevant content in each scenario involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here. The above-mentioned robot multimodal interaction control devices can all execute the above-mentioned... Figure 2 The steps performed by the robot controller in the robot multimodal interactive control method shown are illustrated.
[0111] This application also provides a computer-readable storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes the aforementioned electronic device (robot controller).
[0112] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. This computer program product can be a software installation package.
[0113] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0114] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0115] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0116] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0117] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0118] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0119] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include a flash drive, ROM, RAM, disk, or optical disk, etc.
[0120] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A multimodal interactive control method for a robot, characterized in that, A robot controller used in a robot control system, the robot control system including the robot controller and a target robot, the method comprising: Acquire user interaction information, which includes user voice information and / or user action information; Analyze the user interaction information to determine the type of shopping guide service and the target vehicle parts; Based on the type of shopping guide service, obtain the first component information of the target vehicle component from a preset vehicle knowledge base; Based on the shopping guide service type and the first component information, a target action sequence is determined from a preset action library. The target action sequence includes joint actions and voice actions. The voice actions are used to output voice shopping guide information, and the voice shopping guide information is determined based on the first component information. A first control command is sent to the target robot, the first control command being used to instruct the target robot to execute the target action sequence.
2. The method according to claim 1, characterized in that, The target vehicle component is determined through the following steps: Analyze the user interaction information to determine whether the user interaction information instructs the first component; If indicated, the first component will be identified as the target vehicle component; If no instruction is given, a first vehicle is determined based on the user interaction information, wherein the first vehicle includes a first number of preset tags that match the user interaction information; Based on the first number of preset tags, the second component is determined from the vehicle components corresponding to the first vehicle; Obtain the first priority information of the second component corresponding to the first vehicle; The target vehicle component is determined from the second component based on the first priority information.
3. The method according to claim 2, characterized in that, The step of determining the target vehicle component from the second component based on the first priority information includes: Based on the user interaction information, determine the user's level of attention to each preset tag among the first number of preset tags; Based on the attention level information of each preset label, the second priority information of the second component is determined; Based on the first priority information and the second priority information, the third priority information of the second component is determined; The target vehicle component is determined from the second component based on the third priority information.
4. The method according to claim 1, characterized in that, The step of obtaining the first component information of the target vehicle component from a preset vehicle knowledge base according to the type of shopping guide service includes: If the shopping guide service type is a preset type, the first component information of the target vehicle component is obtained from the vehicle knowledge base. The first component information includes a first explanatory text and a first operation parameter. The step of determining the target action sequence from a preset action library based on the shopping guide service type and the first component information includes: If the first operation parameter indicates that the target vehicle component supports demonstration operation, determine whether the first action set of the action library includes a first action sequence that matches the first component identifier of the target vehicle component; If included, the target action sequence is determined based on the first action sequence and the first explanatory text; If not included, obtain the second component information of the target vehicle component from the vehicle knowledge base, the second component information including the second explanatory text; Based on the second explanatory text, a second action sequence is determined from the second action set of the action library; The target action sequence is determined based on the second action sequence and the second explanatory text.
5. The method according to claim 4, characterized in that, Determining the target action sequence based on the first action sequence and the first explanatory text includes: Obtain the first update time of the first action sequence and the second update time of the first operation parameter; If the first update time is earlier than the second update time, the update parameter item is determined from the first operation parameter; If the updated parameter item is a preset parameter type, obtain the preset threshold range corresponding to the updated parameter item; If the parameter value corresponding to the updated parameter item matches the preset threshold range, the first action sequence is adjusted according to the updated parameter item to obtain the third action sequence; The target action sequence is determined based on the third action sequence and the first explanatory text.
6. The method according to claim 4, characterized in that, The first explanatory text was determined through the following steps: Obtain the first sub-text corresponding to the target vehicle component; Determine the explanation style tags based on the user interaction information; Based on the explanation style tags, obtain the second sub-text corresponding to the target vehicle component; The first explanatory text is generated based on the first subtext and the second subtext; The second explanatory text was determined through the following steps: Obtain the third sub-text corresponding to the target vehicle component; The second explanatory text is generated based on the first subtext, the second subtext, and the third subtext.
7. The method according to claim 6, characterized in that, The first action sequence includes multiple joint action sub-sequences, the multiple joint action sub-sequences include a first sequence, and the step of determining the target action sequence based on the first action sequence and the first explanatory text includes: Obtain the operation type label and duration constraint information corresponding to the first sequence; The fourth sub-text is determined from the first explanatory text based on the operation type label; The first speech rate information is determined based on the user interaction information; The first duration is determined based on the fourth sub-text and the first speech rate information; Based on the first duration, the duration constraint information, and the fourth sub-text, determine the second duration of the first voice action corresponding to the first sequence, and the first voice shopping guide information corresponding to the first voice action.
8. A robot multimodal interactive control device, characterized in that, A robot controller used in a robot control system, the robot control system including the robot controller and a target robot, the device comprising: The first acquisition unit is used to acquire user interaction information, which includes user voice information and / or user action information. The parsing unit is used to parse the user interaction information and determine the type of shopping guide service and the target vehicle parts; The second acquisition unit is used to acquire the first component information of the target vehicle component from a preset vehicle knowledge base according to the type of the shopping guide service. The determining unit is configured to determine a target action sequence from a preset action library based on the shopping guide service type and the first component information. The target action sequence includes joint actions and voice actions. The voice actions are used to output voice shopping guide information. The voice shopping guide information is determined based on the first component information. The sending unit is used to send a first control command to the target robot, the first control command being used to instruct the target robot to execute the target action sequence.
9. An electronic device, characterized in that, The method includes a processor, a memory, a communication interface, and one or more programs, said one or more programs being stored in the memory and configured to be executed by the processor, said programs including instructions for performing the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program for storing electronic data interchange is provided, wherein the computer program causes a computer to perform the steps of the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Humanoid robot self-service cash register and unattended convenience store operation system
CN107610376A
Novel intelligent retail shopping guide robot and method based on machine vision and AR technology
CN108748218A
Information pushing method, information pushing device and voice interaction equipment
CN109545232A
Information recommendation method, device and system, storage medium and intelligent interaction equipment
CN110097400A
Intelligent shopping guide scheme generation method and device, electronic equipment and storage medium
CN116823299A
Cited By
Retail shopping guide robot and interaction control method thereof
CN122086251A
A retail guide robot and an interaction control method of the retail guide robot
CN122086251B