A positioning method, device, storage medium and robot

By using voice commands to determine the robot's initial target position in home or small space scenarios, and then adjusting it to the final target position by combining image information, the problem of inaccurate sound positioning is solved, achieving more accurate user location positioning and improving user experience.

CN116048089BActive Publication Date: 2026-01-02YANTAI IRAY TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310111000.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2026-01-02
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

In home or small space scenarios, the accuracy of determining the user's location based on sound is low, and location determination errors are prone to occur.

Method used

By acquiring the user's voice commands to the robot, the initial target position is determined, and image information around the initial target position is acquired using an image acquisition device. The initial target position is then adjusted to determine a more precise final target position, and the robot is controlled to move within a preset range of the final target position.

Benefits of technology

It improved the accuracy of user location positioning and optimized the user experience of using the robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116048089B_ABST
    Figure CN116048089B_ABST
Patent Text Reader

Abstract

The application discloses a positioning method and device, a storage medium and a robot, and is applied to the field of smart homes. In the scheme, the robot acquires a voice instruction issued by a user to the robot, and determines a primary target position of the user according to the voice instruction; acquires image information around the primary target position, and determines a final target position of the user according to the image information; and controls the robot to move to a preset range of the final target position. In the application, the target position of the user is not directly determined based on the recognized sound, but the primary target position of the user is first determined according to the sound, the primary target position is adjusted according to the image information around the primary target position, so that a more accurate final target position is determined, and then the robot is controlled to move to the preset range of the final target position, the positioning of the user position is more accurate, and therefore the experience of the user using the robot is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of smart home, in particular to a positioning method and device, a storage medium and a robot. BACKGROUND

[0002] With the development of technology, more and more robots are applied in industrial scenarios, but there are still few robots in home scenarios. In the field of home robots, it is particularly important to realize voice interaction through sound recognition. The specific implementation of the current sound recognition method is as follows: a sound collecting device is used to collect the voice output by a user, and the position of the user is determined according to the time and sound intensity of the collected voice.

[0003] However, when the robot is applied in a home scenario or other scenarios with many walls, the sound wave will be reflected when it encounters a wall. The accuracy of determining the position of the user according to the sound wave is low, and the position determination is often wrong. SUMMARY

[0004] The purpose of the present application is to provide a positioning method, device, storage medium and robot. Instead of directly determining the target position of the user based on the recognized sound, the preliminary target position of the user is first determined according to the sound, and then the preliminary target position is adjusted according to the image information around the preliminary target position, so as to determine a more accurate final target position. The robot is controlled to move to the preset range of the final target position, so that the positioning of the user's position is more accurate, thereby optimizing the user's experience of using the robot.

[0005] To solve the above technical problems, the present application provides a positioning method applied to a robot, comprising:

[0006] obtaining a voice instruction sent by a user to the robot, and determining a preliminary target position of the user according to the voice instruction;

[0007] obtaining image information around the preliminary target position, and determining a final target position of the user according to the image information;

[0008] controlling the robot to move to a preset range of the final target position.

[0009] Preferably, a plurality of sound collecting devices are arranged on the robot, and the obtaining of the voice instruction sent by the user to the robot and the determination of the preliminary target position of the user according to the voice instruction comprise:

[0010] collecting the voice instruction sent by the user to the robot through each sound collecting device;

[0011] determining the preliminary target position of the user according to the time and sound intensity of the voice instruction collected by each sound collecting device.

[0012] Preferably, the controlling the robot to move into the preset range of the final target position further comprises:

[0013] performing voice recognition on the voice instruction, and determining whether the voice instruction is a moving instruction related to the current position of the user according to a voice recognition result;

[0014] If yes, the step of controlling the robot to move into the preset range of the final target position is entered.

[0015] Preferably, the method further comprises:

[0016] when receiving a control instruction sent by the host computer, determining whether the robot is currently performing an action corresponding to the voice instruction;

[0017] If yes, sending prompt information to the host computer;

[0018] Otherwise, controlling the robot to perform an action corresponding to the control instruction.

[0019] Preferably, the robot is provided with an image acquisition device for acquiring the image information; and the acquiring the image information around the primary target position and determining the final target position of the user according to the image information comprises:

[0020] controlling the robot to rotate so that the image acquisition device on the robot faces the direction of the primary target position;

[0021] acquiring image information around the primary target position by using the image acquisition device, and performing target recognition based on the image information to determine the final target position of the user according to a target recognition result.

[0022] Preferably, the performing target recognition based on the image information to determine the final target position of the user according to a target recognition result comprises:

[0023] when the image information acquired by the image acquisition device includes a human body image, determining the position of the human body image as the final target position of the user.

[0024] Preferably, the performing target recognition based on the image information to determine the final target position of the user according to a target recognition result comprises:

[0025] when the image information acquired by the image acquisition device includes multiple human body images, determining the position of the human body image closest to the center point of the image information as the first target position;

[0026] controlling the robot to move to the first target position and to perform voice interaction with a user at the first target position during the movement to determine whether the user at the first target position is the user who sends the voice instruction;

[0027] if yes, determining that the first target position is the final target position.

[0028] Preferably, when determining that the user at the first target position is not the user who sends the voice instruction, the method further comprises:

[0029] sequentially taking a position of a human body image in the image information other than the human body image of the first target position as a second target position;

[0030] controlling the robot to move to the second target position and to perform voice interaction with a user at the second target position during the movement to determine whether the user at the second target position is the user who sends the voice instruction, until determining a final target position of the user who sends the voice instruction or determining that there is no user who sends the voice instruction in the image information.

[0031] To solve the above technical problems, the application further provides a positioning device, comprising:

[0032] a memory for storing a computer program;

[0033] a processor for implementing the steps of the positioning method as described above when storing the computer program.

[0034] To solve the above technical problems, the application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the positioning method as described above.

[0035] To solve the above technical problems, the application further provides a robot, comprising the positioning device as described above.

[0036] The application provides a positioning method, device, storage medium and robot, and is applied to the field of smart home. The robot in the scheme acquires a voice instruction issued by a user to the robot, and determines a primary target position of the user according to the voice instruction; acquires image information around the primary target position, and determines a final target position of the user according to the image information; and controls the robot to move to a preset range of the final target position. In the application, the target position of the user is not directly determined based on the recognized sound, but the primary target position of the user is first determined according to the sound, the primary target position is adjusted according to the image information around the primary target position, so as to determine a more accurate final target position, and then the robot is controlled to move to the preset range of the final target position, so that the positioning of the user position is more accurate, thereby optimizing the experience of the user using the robot. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0038] Figure 1 A flowchart of a positioning method provided by the application is shown in the figure.

[0039] Figure 2 A structure block diagram of a positioning device provided by the application is shown in the figure. DETAILED DESCRIPTION

[0040] The core of the application is to provide a positioning method, device, storage medium and robot, and the target position of the user is not directly determined based on the recognized sound, but the primary target position of the user is first determined according to the sound, the primary target position is adjusted according to the image information around the primary target position, so as to determine a more accurate final target position, and then the robot is controlled to move to the preset range of the final target position, so that the positioning of the user position is more accurate, thereby optimizing the experience of the user using the robot.

[0041] In order to make the purpose, technical scheme and advantages of the embodiments of the application more clear, the technical scheme in the embodiments of the application will be clearly and completely described in the following with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are some of the embodiments of the application, but not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0042] Please refer to Figure 1 , Figure 1A flowchart of a positioning method provided in the present application is shown in the figure, which is applied to a robot and includes the following steps:

[0043] S11: obtaining a voice instruction sent by a user to the robot, and determining a primary target position of the user according to the voice instruction;

[0044] Specifically, the application scenario of the positioning method in the present application is mainly that the user sends a voice instruction to the robot, and the robot positions the user according to the voice instruction. Therefore, first, the voice instruction sent by the user to the robot is obtained, and the user can send the voice instruction to the robot by directly speaking to wake up the function of the robot to recognize the voice instruction. After the robot obtains the voice instruction, the primary target position of the user is determined according to the voice instruction, which is a relatively rough position determined by the robot recognizing the sound.

[0045] As a preferred embodiment, a plurality of sound collecting devices are arranged on the robot, and the step of obtaining the voice instruction sent by the user to the robot and determining the primary target position of the user according to the voice instruction includes the following steps:

[0046] collecting the voice instruction sent by the user to the robot through the sound collecting devices;

[0047] determining the primary target position of the user according to the time and sound intensity of the voice instruction collected by the sound collecting devices.

[0048] Specifically, the specific implementation of determining the primary target position of the user according to the voice instruction can be as follows: first, the voice instruction sent by the user is collected through the plurality of sound collecting devices arranged on the robot, and then the processor in the robot obtains the voice information (such as the time and sound intensity of the voice instruction) related to the voice instruction output by the sound collecting devices, and calculates the primary target position of the user according to the time and sound intensity of the voice instruction collected by the sound collecting devices. The specific calculation method can be as follows: the source of the sound is calculated according to the time difference and sound intensity of the voice instruction collected by the sound collecting devices, and then the primary target position is determined.

[0049] As a preferred embodiment, the plurality of sound collecting devices are uniformly distributed around the robot.

[0050] In order to ensure the accuracy of the primary target position, in the present application, the sound collecting devices are evenly arranged around the robot, that is, each sound collecting device is located at a different position of the robot, so that when each sound collecting device collects the voice command, there is a difference in the time and sound intensity of collecting the voice command, and then the primary target position can be calculated according to the time difference and the sound intensity difference. For example, four sound collecting devices can be arranged, and the four sound collecting devices surround the robot by 360 degrees, which can be arranged at 3 o'clock, 6 o'clock, 9 o'clock and 12 o'clock. Similarly, six sound collecting devices can also be used, and the principle is similar. In addition, the above is only one specific implementation manner of the embodiment, and each sound collecting device can also be arranged at different positions on the robot, which can not be evenly distributed, and the present application is not limited herein.

[0051] Further, the robot in the present application can also be provided with a voice wake-up function, and the voice collecting function of the robot can be activated by a wake-up word to perform voice recognition. For example, the wake-up word can be set as the name of the robot, and the present application is not limited herein.

[0052] S12: acquiring image information around the primary target position, and determining the ultimate target position of the user according to the image information;

[0053] Specifically, considering that when there are many walls or obstacles in the place where the robot is located, the sound wave will be reflected when encountering the wall / obstacle, and the accuracy of the method of simply determining the position of the user according to the sound wave is relatively low. Therefore, in the present application, after the primary target position of the user is determined by the sound wave, the primary target position is adjusted according to the image information around the primary target position to obtain the ultimate target position. The image information can be, but is not limited to, the portrait information collected near the primary target position, and then the primary target position is adjusted according to the portrait information (because usually only a person can issue a voice command), thereby improving the accuracy of positioning the position of the user.

[0054] As a preferred embodiment, the robot is provided with an image collecting device for collecting the image information, for example, a visible light camera and / or an infrared camera can be used, which is not limited herein; and the acquiring the image information around the primary target position and determining the ultimate target position of the user according to the image information comprises:

[0055] controlling the robot to rotate so that the image collecting device on the robot is directed to the direction of the primary target position;

[0056] The image acquisition device is used to acquire image information around the primary target position, and target recognition is performed based on the image information, and the final target position of the user is determined according to the target recognition result.

[0057] Specifically, the embodiment aims to provide a specific implementation manner for adjusting the primary target position, which can but is not limited to acquiring image information around the primary target position by using an image acquisition device arranged on the robot. When the primary target position is determined by using sound, the primary target position is only a rough position, and then the robot controls the image acquisition device (such as a camera) to intervene to finely adjust the determined primary target position to obtain the final target position. Specifically, the robot is controlled to rotate according to the approximate positioning of the sound source, and is rotated to the approximate direction of the sound source, that is, the robot is controlled to rotate so that the image acquisition device faces the direction of the primary target position. In a specific embodiment, the primary target position can be in front of the image acquisition device. After the robot completes the rotation, the image acquisition device intervenes to acquire image information around the primary target position by using the image acquisition device (in a specific embodiment, a target recognition algorithm can but is not limited to be used to capture a person in the field of view of the image acquisition device), so as to adjust the primary target position according to the image information acquired by the image acquisition device.

[0058] As a preferred embodiment, the target recognition based on the image information and the determination of the final target position of the user according to the target recognition result include:

[0059] When the image information acquired by the image acquisition device includes a human body image, the position of the human body image is determined as the final target position of the user.

[0060] Specifically, when there is only one human body image in the image information acquired by the image acquisition device, the person is determined to be the user who issues the voice instruction, at this time, distance information corresponding to the human body image and coordinate information relative to the center point of the field of view of the camera can be fed back to the robot, and the position of the human body image is determined as the final target position of the user, and then the robot is controlled to adjust the direction of the robot according to the final target position and to approach the final target position.

[0061] As a preferred embodiment, the target recognition based on the image information and the determination of the final target position of the user according to the target recognition result include:

[0062] When the image information acquired by the image acquisition device includes multiple human body images, the position of the human body image closest to the center point of the image information is determined as the first target position.

[0063] controlling the robot to move to the first target position and to voice interact with the user at the first target position during the movement to determine whether the user at the first target position is the user who sent the voice instruction;

[0064] If yes, the first target position is determined as the final target position.

[0065] Specifically, when the image information collected by the image collection device includes multiple human body images, the robot cannot directly determine who is the user who sent the voice instruction. At this time, the human body image closest to the center point of the image information is determined as the user who sent the voice instruction, i.e., the position closest to the center point of the image information is determined as the first target position, and the robot moves to the first target position. During the movement, the user is asked (voice interacts with the user at the first target position), and the response of the user is used to determine whether the user at the first target position is the user who sent the voice instruction, i.e., the response of the user is used to determine whether the first target position is the final target position. If yes, the first target position is directly determined as the final target position, and the robot is controlled to move to the final target position. Otherwise, the adjustment is continued.

[0066] As a preferred embodiment, when it is determined that the user at the first target position is not the user who sent the voice instruction, the method further comprises:

[0067] the positions of the human body images in the image information other than the human body image at the first target position are sequentially determined as second target positions;

[0068] the robot is controlled to move to the second target positions and to voice interact with the users at the second target positions during the movement to determine whether the users at the second target positions are the user who sent the voice instruction, until the final target position of the user who sent the voice instruction is determined or it is determined that there is no user who sent the voice instruction in the image information.

[0069] When the human closest to the center point of the image information is not the user who sent the voice instruction, the positions of the human body images in the image other than the human body image at the first target position are sequentially determined as second target positions, which can but not limited to be the positions of the human body images in the image in the order from small to large distance between each human body image and the center point of the image. After the second target position is determined, the robot is controlled to move to the second target position, and during the movement to the second target position, the user at the second target position is voice interacted with to determine whether the user at the second target position is the user who sent the voice instruction, until the final target position of the user who sent the voice instruction is determined or it is determined that there is no user who sent the voice instruction in the image information.

[0070] Further, when the above determination does not find the user sending the voice instruction, the robot issues an inquiry, re-performs the sound source positioning according to the user response, and then re-enters the step of determining the final target position.

[0071] S13: Control the robot to move into a preset range of the final target position.

[0072] As a preferred embodiment, before the step of controlling the robot to move into a preset range of the final target position, the method further comprises:

[0073] voice recognizing the voice instruction, and determining whether the voice instruction is a moving instruction related to the current position of the user according to the voice recognition result.

[0074] If yes, the step of controlling the robot to move into a preset range of the final target position is entered.

[0075] Specifically, it is considered that not all voice instructions need to control the robot to move into a preset range around the final target position of the user, such as some voice control instructions that do not need to follow. Therefore, before the step of controlling the robot to move into a preset range of the final target position, the voice instruction is recognized in the present application to determine whether the voice instruction is a moving instruction related to the current position of the user, for example, the moving instruction related to the current position of the user can be "come to me" or "follow me" which is a moving instruction related to the current position of the user. If it is not a moving instruction related to the current position of the user, the action corresponding to the voice instruction can be directly executed, and if it is, the robot is controlled to move into a preset range of the final target position, and then can interact with the user.

[0076] As a preferred embodiment, the method further comprises:

[0077] When receiving a control instruction sent by the user through the upper computer, it is determined whether the robot is currently performing an action corresponding to the voice instruction;

[0078] If yes, a prompt information is sent to the upper computer;

[0079] Otherwise, the robot is controlled to perform an action corresponding to the control instruction.

[0080] Further, the user can also control the robot through the host computer (such as APP), and the priority of the APP remote control is set to be lower than that of the voice instruction control in the application, that is, if the robot is currently executing an action corresponding to the voice instruction, a prompt information (for example, displaying the words "someone is voice controlling, please try later" on the user APP interface) is sent to the host computer, and the action corresponding to the control instruction sent by the host computer is not executed; otherwise, the robot is controlled to execute the action corresponding to the control instruction.

[0081] In addition, if the robot has an automatic inspection or other automatic triggered task, the priority can be set as: the priority of the voice instruction control > the priority of the remote control > the priority of the automatic inspection or the automatic triggered task.

[0082] In summary, the positioning method provided by the application is not to directly determine the target position of the user based on the recognized sound, but to first determine the primary target position of the user according to the sound, then adjust the primary target position according to the image information around the primary target position, thereby determining a more accurate final target position, and then control the robot to move to the preset range of the final target position, so that the positioning of the user position is more accurate, thereby optimizing the experience of the user using the robot.

[0083] Please refer to Figure 2 , Figure 2 The structure block diagram of a positioning device provided by the application is shown in the figure, and the device comprises:

[0084] The memory 21 is used for storing the computer program.

[0085] The processor 22 is used for realizing the steps of the positioning method as described above when storing the computer program.

[0086] For the introduction of the positioning device, please refer to the above embodiments, which will not be repeated here.

[0087] To solve the above technical problems, the application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the steps of the positioning method as described above. For the introduction of the computer readable storage medium, please refer to the above embodiments, which will not be repeated here.

[0088] To solve the above technical problems, the application further provides a robot comprising the positioning device as described above. For the introduction of the robot, please refer to the above embodiments, which will not be repeated here.

[0089] It is also noted that, in this disclosure, relational terms such as first and second, and the like, can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises," "comprising," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0090] The above description of disclosed embodiments provides enabling concepts for practicing or using the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A positioning method, characterized by, The application is applied to a robot, comprising: acquiring a voice instruction sent by a user to the robot, and determining a primary target position of the user according to the voice instruction; acquiring image information around the primary target position, and determining a final target position of the user according to the image information; wherein, when the image information collected by an image collection device comprises multiple human body images, a position of a human body image closest to a center point of the image information is determined as a first target position; the robot is controlled to move to the first target position, and voice interaction is performed between the robot and a user at the first target position during the movement, so as to determine whether the user at the first target position is the user sending the voice instruction; if yes, the first target position is determined as the final target position; if no, positions of human body images in the image information except the human body image at the first target position are sequentially determined as second target positions, the robot is controlled to move to the second target positions, and voice interaction is performed between the robot and a user at the second target position during the movement, so as to determine whether the user at the second target position is the user sending the voice instruction, until the final target position of the user sending the voice instruction is determined or it is determined that the user sending the voice instruction does not exist in the image information; the robot is controlled to move to a preset range of the final target position.

2. The positioning method of claim 1, wherein, The robot is provided with a plurality of sound collection devices, and the acquisition of the voice instruction sent by the user to the robot and the determination of the primary target position of the user according to the voice instruction comprise: collecting the voice instruction sent by the user to the robot through each sound collection device; determining the primary target position of the user according to the time and sound intensity at which the voice instruction is collected by each sound collection device.

3. The positioning method of claim 1, wherein, Before the control of the robot to move to the preset range of the final target position, the method further comprises: performing voice recognition on the voice instruction, and determining whether the voice instruction is a moving instruction related to the current position of the user according to the voice recognition result; if yes, the step of controlling the robot to move to the preset range of the final target position is entered.

4. The positioning method of claim 1, wherein, The method further comprises: when a control instruction sent by the user through a host computer is received, it is determined whether the robot is currently performing an action corresponding to the voice instruction; if yes, prompt information is sent to the host computer; otherwise, the robot is controlled to perform an action corresponding to the control instruction.

5. The positioning method according to any one of claims 1 to 4, characterized in that, The robot is provided with an image collection device for collecting the image information; the acquisition of the image information around the primary target position and the determination of the final target position of the user according to the image information comprise: controlling the robot to rotate so that the image collection device on the robot faces the direction of the primary target position; collecting image information around the primary target position by using the image collection device, and performing target recognition based on the image information, and determining the final target position of the user according to the target recognition result.

6. The positioning method of claim 5, wherein, The target recognition based on the image information and the determination of the final target position of the user according to the target recognition result comprise: When the image information collected by the image collection device includes a human body image, a position where the human body image is located is determined as the ultimate target position of the user.

7. A positioning device, characterized in that Comprise: a memory for storing a computer program; a processor for implementing the steps of the positioning method according to any one of claims 1-6 when storing the computer program.

8. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium, and the computer program is executed by the processor to implement the steps of the positioning method according to any one of claims 1-6.

9. A robot, characterized in that Comprise the positioning device according to claim 7.

Citation Information

Patent Citations

  • Mobile robot and positioning method thereof

    CN105929827A

  • Mobile robot and patrol path setting method therefor

    CN106292657A

  • Motion control device, mobile robot, and method for moving to best interaction point

    CN106548231A

  • Control method for on call robot, storage medium and robot

    CN111055288A