A method for voice calling navigation of a domestic robot
By employing a voice-guided navigation method for home robots, utilizing sound source localization, speech-to-text, voiceprint recognition, and event recognition technologies, the problem of inaccurate robot localization in noisy home environments and multi-person environments has been solved, achieving navigation with high success rate and high accuracy.
Patent Information
- Application Number
- CN202310694124.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-12
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-06-12
AI Technical Summary
In noisy home environments, home robots struggle to accurately identify the location of sound sources through voice commands, leading to navigation deviations or inability to determine the location of the person giving the command, especially when multiple family members are gathered together.
By collecting the voice call information of the person giving the instructions, analyzing and predicting the initial location, and combining event recognition technology to navigate to the actual location, including sound source localization, speech-to-text, voiceprint recognition, video obstacle avoidance monitoring and home map modeling, the system ensures accurate location of the person giving the instructions in noisy environments.
It improves the success rate and accuracy of voice-guided navigation for home robots, enabling them to accurately identify and navigate to the location of the person giving the command in noisy and crowded environments, thus solving the problem of inaccurate positioning caused by loss of voice call information.
Smart Images

Figure CN116839580B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of household robots, and more particularly relates to a household robot voice calling navigation method. BACKGROUND
[0002] A household robot is a service robot used in a home environment or a similar environment for the purpose of meeting the life needs of a user, and is mainly a robot for home service. The household robot can be classified into a household robot, an educational robot, an entertainment robot, a robot for the aged and the disabled, a household safety monitoring robot, a chef robot, and a carrying robot.
[0003] When a family member needs service, the simplest way is to call the robot through voice, and the robot is positioned in front of the person through voice recognition. For example, CN112017661A discloses a voice control system and method of a sweeping robot, which needs to recognize the voice of a sound source and the position of the sound source, generate a planned walking path for walking to the sound source according to the voice recognition signal and the sound source position signal, and control the sweeping robot to walk according to the planned walking path.
[0004] However, in a household environment with high environmental noise, the voice calling information collection usually has a loss, and it is difficult to simultaneously recognize the voice of the sound source and other information such as the position of the sound source, which leads to inaccurate positioning, and the robot will deviate from the direction. If multiple family members gather at the same time, the robot cannot determine the person who gives the voice calling information and the actual position of the person. SUMMARY
[0005] The purpose of the embodiment of the application is to provide a household robot voice calling navigation method to solve the technical problems of insufficient success rate and accuracy in the voice calling navigation process of the household robot in the prior art.
[0006] To achieve the above purpose, the technical solution adopted by the application is to provide a household robot voice calling navigation method, which comprises:
[0007] Collecting voice calling information of a person giving an instruction;
[0008] Analyzing the voice calling information to predict a preliminary position of the person giving an instruction;
[0009] Navigating to the preliminary position of the person giving an instruction;
[0010] Predicting the person giving an instruction through event recognition;
[0011] Navigating to the actual position of the person giving an instruction.
[0012] Preferably, the method for analyzing the voice calling information to predict the preliminary position of the person giving an instruction comprises:
[0013] Sound source localization is performed on the voice call information to obtain a preliminary position of the commander.
[0014] Preferably, the method for analyzing the voice call information to predict the preliminary position of the commander comprises:
[0015] Command recognition is performed on the voice call information, the voice is converted into text, and the text is analyzed to obtain the preliminary position of the commander.
[0016] Preferably, the method for analyzing the voice call information to predict the preliminary position of the commander comprises:
[0017] Voiceprint recognition is performed on the voice call information to output the commander.
[0018] The position of the commander is obtained through human body recognition or face recognition.
[0019] Preferably, the method for navigating to the preliminary position of the commander comprises:
[0020] A route to the preliminary position is planned;
[0021] The video obstacle avoidance monitoring and event recognition are simultaneously started during the movement.
[0022] Preferably, the method for predicting the commander through event recognition comprises:
[0023] Video tracking is performed on all the personnel near the preliminary position;
[0024] An action service event is obtained based on the tracking video;
[0025] The personnel who make the action service event are taken as the commander.
[0026] Preferably, the method for predicting the commander through event recognition comprises:
[0027] The environment near the preliminary position is scanned;
[0028] An environmental service event in the environmental information is obtained;
[0029] The personnel who are closest to the environmental service event are taken as the commander.
[0030] Preferably, after the navigation to the preliminary position of the commander, the method further comprises the step of:
[0031] If the commander is not predicted after reaching the preliminary position, the commander is requested to issue the order again, and the voice call information of the commander is re-collected.
[0032] Preferably, before the voice call information of the commander is collected, the method further comprises the step of:
[0033] Modeling a family map;
[0034] Setting a map name by area.
[0035] Preferably, before collecting the voice call information of the commander, the method further comprises the steps of:
[0036] Collecting the biometric information and voice features of each family member;
[0037] Associating the family members with the corresponding areas.
[0038] The home robot voice call navigation method provided by the present application can accurately identify the commander from multiple family members and navigate to the actual location of the commander even if there is loss in the collection of voice call information, thereby improving the success rate and accuracy of voice call navigation of the home robot. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0040] Fig. 1 The flowchart of the home robot voice call navigation method provided by the present application is shown in the figure.
[0041] Fig. 2 Another flowchart of the home robot voice call navigation method provided by the present application is shown in the figure. DETAILED DESCRIPTION
[0042] In order to make the technical problems, technical solutions and beneficial effects of the present application more clearly understood, the following will further describe the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0043] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.
[0044] It should be understood that the terms "length", "width", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.
[0045] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0046] Please refer to Figs. 1-2 The voice calling navigation method for household robots provided by the embodiments of the present application will be described. The voice calling navigation method for household robots comprises:
[0047] Step S1, collecting voice calling information of an instruction person;
[0048] Step S2, analyzing the voice calling information to predict the preliminary position of the instruction person;
[0049] Step S3, navigating to the preliminary position of the instruction person;
[0050] Step S4, predicting the instruction person through event recognition;
[0051] Step S5, navigating to the actual position of the instruction person.
[0052] It can be understood that in a home environment, because multiple family members exist at the same time, the robot usually cannot determine the instruction person who sends the voice calling information and the actual position of the instruction person at the first time.
[0053] Therefore, in the present application, a two-step approach is adopted. After collecting the voice calling information of the instruction person in step S1, the voice calling information is analyzed in steps S2 to S3 to predict the preliminary position of the instruction person, and the robot navigates to the preliminary position of the instruction person. Then in steps S3 to S4, the instruction person is predicted through event recognition during the movement or after reaching the preliminary position, and the robot navigates to the actual position of the instruction person.
[0054] Therefore, the household robot voice calling navigation method provided by the application can solve the technical problems of unclear calling voice, loss of voice calling information collection due to large environmental noise, and inability to accurately locate, can solve the technical problem of lack of field of view and long distance to find the instruction person, and can solve the technical problem of multiple family members gathering to find the wrong instruction person.
[0055] Compared with the prior art, the household robot voice calling navigation method provided by the application can still accurately identify the instruction person from multiple family members and navigate to the actual position of the instruction person in the case of loss of voice calling information collection, improve the success rate and accuracy of voice calling navigation of the household robot, and predict the preliminary position of the instruction person by analyzing the voice calling information, and then predict the instruction person by event recognition to navigate to the actual position of the instruction person.
[0056] In another embodiment of the application, the method for predicting the preliminary position of the instruction person by analyzing the voice calling information comprises:
[0057] The voice calling information is subjected to sound source positioning to locate the preliminary position of the instruction person.
[0058] It can be understood that there are many interference of reflected waves in the room, so when the distance is far, there are many fluctuations in the distance measurement, which reflect many standing wave interferences in the space. However, in the present embodiment, only the preliminary position of the instruction person needs to be located, and then the instruction person is predicted by event recognition after reaching the preliminary position to navigate to the actual position of the instruction person, and the prior art can meet the technical requirement.
[0059] In another embodiment of the application, the method for predicting the preliminary position of the instruction person by analyzing the voice calling information comprises:
[0060] The voice calling information is subjected to instruction recognition to convert the voice into text, and the text is analyzed to obtain the preliminary position of the instruction person.
[0061] It can be understood that after converting the voice into text, for example, searching for instruction keywords, negative words and interrogative words in the text, for example, the voice conversion obtains "Xiaozhi, come to the kitchen.", searching for the instruction keyword "kitchen" obtains the preliminary position of the instruction person in the kitchen. Of course, the preliminary position of the instruction person can also be predicted by comprehensively analyzing the whole semantics, for example, the voice conversion obtains "Xiaozhi, I want the TV remote control.", and the preliminary position of the instruction person is predicted in the living room by comprehensively analyzing the whole semantics.
[0062] In another embodiment of the application, the method for predicting the preliminary position of the instruction person by analyzing the voice calling information comprises:
[0063] The voice calling information is subjected to voiceprint recognition to output the instruction person;
[0064] The position of the instruction person is obtained through human body recognition or face recognition.
[0065] It can be understood that the position of the instruction person obtained through voiceprint recognition is relatively accurate, but voiceprint recognition is limited to recognizing a human body or a face.
[0066] In another embodiment of the present application, a method for navigating to a preliminary position of an instruction person includes:
[0067] planning a route to reach the preliminary position;
[0068] starting video obstacle avoidance monitoring and event recognition simultaneously during the movement.
[0069] It can be understood that since the video obstacle avoidance monitoring and the event recognition are both based on video streams, the video obstacle avoidance monitoring and the event recognition can be performed based on pictures obtained by the same camera.
[0070] In another embodiment of the present application, a method for predicting an instruction person through event recognition includes:
[0071] tracking all persons near the preliminary position through video;
[0072] obtaining a motion service event based on the tracking video;
[0073] taking the person who makes the motion service event as the instruction person.
[0074] It can be understood that the motion service event can be a garbage throwing action, a waving action, a finger pointing action, a tea cup holding action, etc., which is pre-recorded in the program of the robot. When the corresponding motion service event is recognized, the robot can determine the instruction person. After reaching the position of the instruction person, the robot can also directly provide corresponding services to the instruction person.
[0075] In another embodiment of the present application, a method for predicting an instruction person through event recognition includes:
[0076] scanning the environment near the preliminary position;
[0077] obtaining an environmental service event in the environmental information;
[0078] taking the person closest to the location of the environmental service event as the instruction person.
[0079] It can be understood that the environmental service event can be a banana peel still on the ground, water left on a table, a stool fallen on the ground, etc. The environmental service event is pre-recorded in the program of the robot. When the corresponding environmental service event is recognized, the robot takes the person closest to the location of the environmental service event as the instruction person.
[0080] In another embodiment of the present application, the navigation to the preliminary position of the commander further comprises the steps of:
[0081] If the commander is not predicted after reaching the preliminary position, the commander is requested to issue an order again, and the voice calling information of the commander is collected again.
[0082] It can be understood that the voice calling information of the commander is collected again after reaching the preliminary position, and the voice calling information can be obtained at a close distance, and the voice calling information is more complete. Preferably, the positions of the commanders can be determined by combining the voice calling information collected twice.
[0083] In another embodiment of the present application, before collecting the voice calling information of the commander, the method further comprises the steps of:
[0084] Modeling the home map;
[0085] Setting the map name by area.
[0086] It can be understood that the home map is modeled including a room, a living room, a kitchen, a bathroom, etc. The user can set the map name for the room, the living room, the kitchen, the bathroom, etc. according to the habits of the user, so that the robot can determine the user's order according to voice recognition.
[0087] Further, before collecting the voice calling information of the commander, the method further comprises the steps of:
[0088] Collecting the biometric information and the voice features of each family member;
[0089] Associating the family members with the corresponding areas.
[0090] It can be understood that after the family members are associated with the corresponding areas, it is beneficial for the robot to quickly determine the position of the commander according to the voice information. For example, after the voice calling information of “Xiaozhi, come to my room” is collected, the voiceprint of the commander is recognized, and the accurate position of the commander can be found based on the biometric information through the area associated with the commander.
[0091] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A voice-activated navigation method for a home robot, characterized in that, Including the following steps: Collect the voice call information of the person giving the instruction; Analyzing voice call information predicts the initial location of the person giving the instruction; Navigate to the initial location of the person giving the instructions; The person giving the instructions can be predicted through event identification; Methods for predicting the person giving instructions through event recognition include: Scan the environment near the initial location; acquire environmental service events from the environmental information; designate the person closest to the location of the environmental service event as the giver; Navigate to the actual location of the person giving the instructions.
2. The voice-activated navigation method for a home robot as described in claim 1, characterized in that, Methods for analyzing voice call information to predict the initial location of the person giving the command include: The voice call information is used to locate the sound source and obtain the initial location of the person giving the instruction.
3. The voice-activated navigation method for a home robot as described in claim 1, characterized in that, Methods for analyzing voice call information to predict the initial location of the person giving the command include: The system performs command recognition on voice call messages, converts the speech into text, and parses the text to obtain the initial location of the person giving the command.
4. The voice-activated navigation method for a home robot as described in claim 1, characterized in that, Methods for analyzing voice call information to predict the initial location of the person giving the command include: Voiceprint recognition is performed on the voice call information to output the command recipient; The location of the person giving the instruction can be obtained through human body recognition or facial recognition.
5. The voice-activated navigation method for a home robot as described in any one of claims 1 to 4, characterized in that, Methods for navigating to the initial location of the person giving the instructions include: Plan the route to the initial location; During the movement, video obstacle avoidance monitoring and event recognition are activated simultaneously.
6. The voice-activated navigation method for a home robot as described in any one of claims 1 to 4, characterized in that, Methods for predicting the person giving instructions through event recognition include: Video tracking was conducted on all personnel near the initial location; Action service events are obtained based on tracking video; The person who performs the action to serve the event is designated as the instructor.
7. The voice-activated navigation method for a home robot as described in claim 1, characterized in that, After navigating to the initial location of the person giving the instructions, the following steps are also included: If the commander is not predicted within the time limit after reaching the initial position, the commander is requested to issue another command, and the commander's voice call information is collected again.
8. The voice-activated navigation method for a home robot as described in claim 1, characterized in that, Before collecting the voice call information from the person giving the instruction, the following steps are also included: Modeling the family map; Set map names for different regions.
9. The voice-activated navigation method for a home robot as described in claim 8, characterized in that, Before collecting the voice call information from the person giving the instruction, the following steps are also included: Collect biometric and voice characteristics of each family member; Associate family members with their corresponding regions.
Citation Information
Patent Citations
Voice control system and method of sweeping robot and sweeping robot
CN112017661A
Voice control action system of robot
CN105856261A
Voice interaction control method and device for intelligent equipment
CN106328132A
Sweeper control method and device
CN110946518A
Robot control method and device, electronic equipment, robot and server
CN113601511A