Robot and method for controlling the same

By generating ultrasonic waves and receiving reflected sounds using a robot, and prioritizing them based on reflectivity information, the usability problem of user voice location estimation in indoor environments was solved, achieving more accurate user voice recognition.

CN121335784APending Publication Date: 2026-01-13SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480037193.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-30
Filing Date
2024-06-10
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing technologies suffer from usability issues due to reflected sound when estimating user voice location in indoor environments, and require additional operations to identify the location via LiDAR or camera.

Method used

The robot generates ultrasonic waves using sensors, receives reflected sound, combines this with microphone information to generate a reflectivity map, identifies the direction of the user's voice, prioritizes reflectivity information, and identifies the location of the user's voice.

Benefits of technology

It improves the accuracy of estimating user voice location in indoor environments, reduces reliance on additional sensors, and simplifies the location recognition process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121335784A_ABST
    Figure CN121335784A_ABST
Patent Text Reader

Abstract

A robot and a method for controlling the robot are provided. The robot comprises: at least one sensor; a speaker; a microphone; a driving unit; at least one memory to store one or more instructions; and at least one processor configured to execute one or more instructions. The one or more instructions, when executed by the at least one processor, generate a map including information about a plurality of objects based on sensing information acquired by the at least one sensor; generating ultrasonic waves with respect to the plurality of objects, respectively, by a speaker; acquiring reflectance information on the plurality of objects based on reflected sounds that have been respectively reflected from the plurality of objects and received through the microphone; and storing the reflectivity information. The reflected sound reflected from each object is at least a portion of the ultrasonic wave reflected from each object. If a user voice is received through a microphone, information on an intensity of the user voice is acquired for each of a plurality of directions. On the basis of information on the intensity of the user speech for each of the plurality of directions, information on a plurality of candidate directions in which the user speech is received is acquired. Priority information about the plurality of candidate directions is acquired based on the position of the robot and the stored reflectivity information, and information about the direction in which the user speech is emitted among the plurality of candidate directions is acquired based on the priority information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to a robot and a method for controlling the same, and more particularly, to a robot for estimating a location from which a user utters a voice and a method for controlling the same. BACKGROUND

[0002] In recent years, a technology regarding a sound source direction estimation for estimating a sound source location of a user's voice or another audio signal is being developed. The sound source direction estimation can refer to a technology for estimating a user direction based on a microphone by measuring an intensity of a voice signal for each direction of a sound including user voice information.

[0003] However, this method lacks usability in an indoor environment in which a sound is easily reflected from a wall or the like, and is used together with another sensor such as a camera to compensate for the above-described problem. However, if this method uses another sensor such as a laser radar sensor or a camera, it is possible to identify that an object exists in a relevant location, but there is a problem in that an additional operation is required to repeatedly move the robot or rotate the robot in order to identify the relevant location through the laser radar sensor or the camera. SUMMARY

[0004] TECHNICAL SOLUTION According to an aspect of the disclosure, a robot includes at least one sensor, a speaker, a microphone, a driver, at least one memory storing one or more instructions, and at least one processor configured to execute the one or more instructions, wherein the one or more instructions, when executed by the at least one processor, cause the robot to: generate a map including information about a plurality of objects based on sensing information obtained through the at least one sensor; generate, through the speaker, an ultrasonic wave toward each of the plurality of objects; obtain reflectivity information about the plurality of objects based on a reflected sound reflected from each of the plurality of objects and received through the microphone, and store the reflectivity information, the reflected sound reflected from each of the plurality of objects being at least a part of the ultrasonic wave reflected from each of the plurality of objects; obtain information about an intensity of a user's voice for each of a plurality of directions based on the user's voice being received through the microphone; obtain information about a plurality of candidate directions from which the user's voice is received based on the information about the intensity of the user's voice for each of the plurality of directions; obtain priority order information for the plurality of candidate directions based on a location of the robot and the stored reflectivity information; and obtain information about a direction from which the user's voice is uttered from the plurality of candidate directions based on the priority order information.

[0005] The one or more instructions, when executed by the at least one processor, can further cause the robot to: generate ultrasonic waves at a preset distance interval with respect to a wall object among the plurality of objects; and obtain reflectivity information about the wall object based on reflected sound reflected from the wall object and received through the microphone, wherein the reflected sound reflected from the wall object is at least a portion of the generated ultrasonic waves at the preset interval that is reflected from the wall object.

[0006] The one or more instructions, when executed by the at least one processor, can further cause the robot to: generate ultrasonic waves from two or more directions toward an object other than the wall object among the plurality of objects; and obtain reflectivity information about the object by obtaining an average of reflected sound reflected from the object and received through the microphone, the reflected sound reflected from the object being at least a portion of the generated ultrasonic waves from the two or more directions that is reflected from the object.

[0007] The one or more instructions, when executed by the at least one processor, can further cause the robot to: obtain information about intensity of a user voice with respect to each of a plurality of directions based on a position of the robot; and identify a preset number of directions in which a respective intensity of the user voice exceeds a predetermined threshold among the plurality of directions as a plurality of candidate directions.

[0008] The one or more instructions, when executed by the at least one processor, can further cause the robot to: identify an object among the plurality of objects that is located in the plurality of candidate directions with respect to a position of the robot; and identify a priority order with respect to the plurality of candidate directions based on respective information about intensity of the user voice corresponding to each of the plurality of candidate directions and reflectivity information corresponding to the identified object.

[0009] The one or more instructions, when executed by the at least one processor, can further cause the robot to: obtain a corrected intensity of the user voice with respect to each of the plurality of candidate directions by multiplying a weight value by a respective intensity of the user voice corresponding to each of the plurality of candidate directions; and identify the priority order based on the corrected intensity with respect to the plurality of candidate directions, wherein the weight value corresponds to the reflectivity information corresponding to the identified object.

[0010] The weight value and the reflectivity information corresponding to the identified object can be inversely proportional to each other.

[0011] For each of the plurality of candidate directions, the priority order can be proportional to the corrected intensity, and the one or more instructions, when executed by the at least one processor, can further cause the robot to: identify a candidate direction having a highest priority order from among the plurality of candidate directions as a direction from which the user voice is emitted.

[0012] The one or more instructions, when executed by the at least one processor, can further cause the robot to: perform speech recognition on the user speech by performing beamforming in a direction from which the user speech is emitted.

[0013] According to an aspect of the disclosure, a method for controlling a robot includes generating a map including information about a plurality of objects based on sensing information obtained through at least one sensor of the robot; generating, by a speaker of the robot, an ultrasonic wave toward each of the plurality of objects; obtaining reflectivity information about the plurality of objects based on reflected sound reflected from each of the plurality of objects and received through a microphone of the robot, and storing the reflectivity information, the reflected sound reflected from each of the plurality of objects being at least a portion of the ultrasonic wave reflected from each of the plurality of objects; obtaining information about intensity of user speech for each of a plurality of directions based on the user speech being received through the microphone; obtaining information about a plurality of candidate directions from which the user speech is received based on the information about the intensity of the user speech for each of the plurality of directions; obtaining priority order information for the plurality of candidate directions based on a location of the robot and the stored reflectivity information; and obtaining information about a direction from which the user speech is emitted from the plurality of candidate directions based on the priority order information.

[0014] The operation of generating the ultrasonic wave toward each of the plurality of objects can include generating the ultrasonic wave at a preset distance interval for a wall object among the plurality of objects, and the operation of obtaining the reflectivity information can include obtaining reflectivity information about the wall object based on reflected sound reflected from the wall object and received through the microphone, wherein the reflected sound reflected from the wall object is at least a portion of the ultrasonic wave generated at the preset interval and reflected from the wall object.

[0015] The operation of generating the ultrasonic wave toward each of the plurality of objects can include generating the ultrasonic wave toward objects other than the wall object among the plurality of objects from two or more directions, and obtaining reflectivity information for the objects by obtaining an average value of reflected sound reflected from the objects and received through the microphone, wherein the reflected sound reflected from the objects is at least a portion of the ultrasonic wave generated from the two or more directions and reflected from the objects.

[0016] The operation of obtaining the information about the plurality of candidate directions can include identifying a preset number of directions in which respective intensities of the user speech among the plurality of directions exceed a predetermined threshold as the plurality of candidate directions.

[0017] The operation of obtaining the priority order information can include identifying objects of the plurality of objects located in a plurality of candidate directions with respect to a position of the robot, and identifying a priority order for the plurality of candidate directions based on respective information on intensity of a user voice corresponding to each of the plurality of candidate directions and reflectance information corresponding to the identified objects.

[0018] The operation of obtaining the priority order information can include obtaining a corrected intensity of the user voice for each of the plurality of candidate directions by multiplying a weight value by a respective intensity of the user voice corresponding to each of the plurality of candidate directions, and identifying the priority order based on the corrected intensity for the plurality of candidate directions, and wherein the weight value corresponds to the reflectance information corresponding to the identified objects.

[0019] According to an aspect of the disclosure, there is provided a non-transitory computer-readable medium storing instructions which, when executed by at least one processor, cause the at least one processor to perform a method of controlling a robot, wherein the method includes generating a map including information on a plurality of objects based on sensing information obtained by at least one sensor of the robot, generating, by a speaker of the robot, ultrasonic waves toward each of the plurality of objects, obtaining reflectance information on the plurality of objects based on reflected sound reflected from each of the plurality of objects and received by a microphone of the robot, and storing the reflectance information, wherein the reflected sound reflected from each of the plurality of objects is at least a portion of the ultrasonic waves reflected from each of the plurality of objects, obtaining information on intensity of a user voice for each of a plurality of directions based on the user voice being received by the microphone, obtaining information on a plurality of candidate directions from which the user voice is received based on the information on the intensity of the user voice for each of the plurality of directions, obtaining priority order information for the plurality of candidate directions based on a position of the robot and the stored reflectance information, and obtaining information on a direction from which the user voice is emitted from the plurality of candidate directions based on the priority order information.

[0020] For the non-transitory computer-readable medium, the operation of generating the ultrasonic waves toward each of the plurality of objects can include generating the ultrasonic waves at a preset distance interval for a wall object of the plurality of objects, and the operation of obtaining the reflectance information can include obtaining reflectance information on the wall object based on reflected sound reflected from the wall object and received by the microphone, wherein the reflected sound reflected from the wall object is at least a portion of the ultrasonic waves generated at the preset interval and reflected from the wall object.

[0021] For the non-transitory computer-readable medium, the generating of the ultrasonic waves toward each of the plurality of objects can include generating the ultrasonic waves toward the objects other than the wall object from two or more directions, and the obtaining of the reflectivity information can include obtaining the reflectivity information for the objects by obtaining an average of reflected sound reflected from the objects and received through the microphone, wherein the reflected sound reflected from the objects is at least a portion of the ultrasonic waves generated from the two or more directions that is reflected from the objects.

[0022] For the non-transitory computer-readable medium, the obtaining of the information about the plurality of candidate directions can include identifying a preset number of directions in which respective intensities of the user voice exceed a predetermined threshold among the plurality of directions as the plurality of candidate directions.

[0023] For the non-transitory computer-readable medium, the obtaining of the priority order information can include identifying objects among the plurality of objects that are located in the plurality of candidate directions with respect to a position of the robot, and identifying a priority order for the plurality of candidate directions based on respective information about intensities of the user voice corresponding to each of the plurality of candidate directions and reflectivity information corresponding to the identified objects. BRIEF DESCRIPTION OF DRAWINGS

[0024] The above and other aspects, features, and advantages of certain embodiments of the disclosure will be more apparent from the following description taken in conjunction with the accompanying drawings, in which: Figure 1 is a block diagram illustrating a configuration of a robot according to an embodiment of the disclosure; Figure 2 is a block diagram illustrating a plurality of configurations for generating a map including reflectivity information and identifying a direction from which a user voice is emitted using the generated map according to an embodiment of the disclosure; Figure 3a , Figure 3b and Figure 3c are diagrams illustrating a method of generating a map including information about objects according to an embodiment of the disclosure; Figure 4a , Figure 4b and Figure 4c are diagrams illustrating a method of obtaining reflectivity information for objects and matching the same with a stored map according to an embodiment of the disclosure; Figure 1 Figure 5a , Figure 5b , Figure 5c and Figure 5d are diagrams illustrating a method of identifying a direction from which a user voice is emitted using reflectivity information according to an embodiment of the disclosure; and Figure 6 is a flowchart illustrating a control method of a robot according to an embodiment of the disclosure.​ DETAILED DESCRIPTION

[0025] Various modifications can be made to the embodiments of the present disclosure, and various types of embodiments can exist. Thus, specific embodiments will be shown in the drawings and described in detail in the detailed description. However, it should be noted that the above is not intended to limit the scope of the present disclosure to the specific embodiments, but should be interpreted to include all modifications, equivalents, or alternatives of the embodiments included in the scope of the ideas and technologies disclosed herein. With regard to the description of the drawings, the same reference numbers can be used to indicate the same elements.

[0026] In describing the present disclosure, detailed descriptions of related known technologies can be omitted if it is determined that such detailed descriptions can unnecessarily obscure the gist of the present disclosure.

[0027] In addition, the following embodiments can be modified in various different forms, and it should be understood that the scope of the technical spirit of the present disclosure is not limited to the following embodiments. Rather, the embodiments are provided so that the present disclosure will be thorough and complete, and fully convey the technical spirit of the present disclosure to those skilled in the art.

[0028] The terms used in the present disclosure have been used simply to describe specific embodiments, and are not intended to limit the scope of protection. Unless otherwise specified, the singular expression also includes the plural expression.

[0029] In the present disclosure, expressions such as "have," "may have," "include," and "may include" are used to designate the presence of corresponding features (for example, elements such as numerical values, functions, operations, or components), and do not exclude the presence or possession of additional features.

[0030] In the present disclosure, expressions such as "A or B," "at least one of A and / or B," or "one or more of A and / or B" can include all possible combinations of the items listed together. For example, "A or B," "at least one of A and B," or "at least one of A or B" can refer to all cases including: (1) only A, (2) only B, or (3) both A and B.

[0031] Expressions such as "1st," "2nd," "first," or "second" used in the present disclosure can limit various elements regardless of order and / or importance, and can be used only to distinguish one element from another element, without limiting the relevant elements.

[0032] When a certain element (e.g., a first element) is indicated as being "in combination with" / "in operative or communicative combination with" another element (e.g., a second element) or "connected to" another element (e.g., a second element), it can be understood that the certain element is directly combined with / directly combined to or connected to the other element or combined through yet another element (e.g., a third element).

[0033] On the other hand, when a certain element (e.g., a first element) is indicated as being "directly combined with" / "directly combined to" another element (e.g., a second element) or "directly connected to" another element (e.g., a second element), it can be understood that there is no yet another element (e.g., a third element) between the certain element and the other element.

[0034] The expression "configured to" (or "set to") used in the disclosure can be used interchangeably with, for example, "adapted to", "have the capability of", "designed to", "suitable for", "manufactured to", or "capable of", based on the situation. The term "configured to" can not necessarily mean "designed as" in terms of hardware.

[0035] On the contrary, in a certain case, the expression "an apparatus configured to" can mean that the apparatus "can perform" together with another apparatus or component. For example, the phrase "a processor configured to (or set to) perform A, B, and C" can mean a dedicated processor (e.g., an embedded processor) for performing the relevant operations or a general-purpose processor (e.g., a central processing unit (CPU) or an application processor) capable of performing the relevant operations by running one or more software programs stored in a memory apparatus.

[0036] The term "module" or "component" in the embodiments herein performs at least one function or operation and can be implemented in hardware or software, or implemented in a combination of hardware and software. Also, except for the "module" or "component" that needs to be implemented to a specific hardware, a plurality of "modules" or a plurality of "components" can be integrated in at least one module and implemented as at least one processor.

[0037] The various elements and regions of the drawings have been schematically shown. Accordingly, the technical spirit of the disclosure is not limited by the relative sizes and distances shown in the drawings.

[0038] Hereinafter, embodiments according to the disclosure will be described in detail with reference to the accompanying drawings to help a person of ordinary skill in the art to understand.

[0039] Figure 1 is a block diagram showing a configuration of a robot according to an embodiment of the disclosure. As shown in FIG. 1, the robot 100 according to an embodiment of the disclosure can include a body 110, a head 120, a neck 130, a waist 140, a left arm 150, a right arm 160, a left leg 170, and a right leg 180. Figure 1As illustrated, the robot 100 can include at least one sensor 110, a speaker 120, a microphone 130, a driver 140, at least one memory 150, and at least one processor 160. The configuration of the robot 100 is not limited to Figure 1 the configuration illustrated in the middle, and configurations apparent to one of ordinary skill in the art can be added or omitted. Furthermore, the robot 100 according to the present disclosure can be a cleaning robot, but this is only one embodiment, and the robot 100 can be a robot that provides various services in a home or a service robot (e.g., a guide robot, etc.) that provides services in various places such as, for example, and without limitation, an airport, a hotel, a supermarket, a clothing store, a warehouse, a hospital, etc.

[0040] The at least one sensor 110 can obtain various information about a state of the robot 100 or a surrounding environment of the robot 100. Specifically, the at least one sensor 110 can include a laser radar sensor and an inertial measurement unit (IMU) sensor. The laser radar sensor can project a ray (e.g., laser, near-infrared light, visible light, ultraviolet, etc.) to an object, and obtain sensing information for obtaining information about a distance from the object by detecting light reflected by the object. The IMU sensor can be a sensor for sensing a movement of the robot 100, and can include at least one of a geomagnetic sensor, an acceleration sensor, and a gyro sensor. However, as described above, using the laser radar sensor to obtain information about a distance from an object is only one embodiment, and various sensors such as a depth sensor can be used to obtain information about a distance from an object.

[0041] Specifically, the at least one processor 160 can obtain information about distances from a plurality of objects and movement information of the robot 100 based on sensing information obtained through the laser radar sensor and the IMU sensor. However, the above is only one embodiment, and information about distances from a plurality of objects and movement information of the robot 100 can be obtained based on sensing information obtained by the at least one sensor 110.

[0042] Furthermore, the at least one sensor 110 can include a camera for capturing an image. Specifically, the camera can obtain an image of an object by capturing a surrounding environment of the robot 100. The robot 100 can obtain information about the object (e.g., a type of the object, a size of the object, a shape of the object, etc.) by inputting the image of the object into a trained neural network model (e.g., an object recognition model).

[0043] The speaker 120 can be a configuration for outputting an audio signal. Specifically, the speaker 120 can output an ultrasonic wave to an object according to a control of the at least one processor 160. The speaker 120 can also generate an ultrasonic wave at a preset distance interval with respect to a wall object among the plurality of objects. In addition, the speaker 120 can generate an ultrasonic wave from two or more directions toward the remaining objects except for the wall object among the plurality of objects.

[0044] The microphone 130 can be configured to receive various audio signals. Specifically, the microphone 130 can receive a reflection sound in which an ultrasonic wave output through the speaker 120 is reflected by an object. In addition, the microphone 130 can receive a user voice. The microphone 130 can be disposed as a plurality, and can receive a user voice received from a plurality of directions through the plurality of microphones.

[0045] The driver 140 can make the robot 100 travel or move in other ways according to a control of the at least one processor 160. The traveling can include an operation in which the robot 100 moves using power. According to one or more embodiments of the disclosure, the traveling can include an operation in which the robot 100 moves in a random direction using power. Alternatively, the traveling can include an operation in which the robot 100 moves along a preset line or path using power.

[0046] Specifically, the driver 140 can include a wheel that allows the robot 100 to travel and a wheel driving motor that rotates the wheel. Specifically, the driver 140 can move the robot 100 within a home space. In addition, when a direction in which a user voice is emitted is identified, the driver 140 can move the robot 100 in the direction in which the user voice is emitted.

[0047] The at least one memory 150 can store instructions or data related to an operating system (OS) and elements of the robot 100 to control overall operations of the elements of the robot 100. Specifically, the at least one memory 150 can include a plurality of modules for generating a map including reflectance information and identifying a direction in which a user voice is emitted using the generated map. For example, as shown in FIG. 2, the at least one memory 150 can include a map generation module 210 including a terrain information obtaining module 211, an object information obtaining module 212, and a reflectance information obtaining module 213, and an identification module 220 including a user voice obtaining module 221, a preprocessing module 222, an intensity measurement module 223, a candidate direction obtaining module 224, a priority order information obtaining module 225, a direction identification module 226, and a user voice identification module 227. Figure 2

[0048] ​Specifically, if a map including reflectivity information is generated and a function for identifying a direction from which a user's voice is emitted is executed using the generated map, the robot 100 can load data for various modules stored in the non-transitory memory to perform various operations in the volatile memory. Here, loading can denote an operation of calling data stored in the non-volatile memory and storing it in the volatile memory to be accessed by the at least one processor 160.

[0049] In addition, the at least one memory 150 can store information on a trained neural network model to obtain information on an object by inputting an image obtained by a camera.

[0050] The at least one memory 150 can be implemented as a non-volatile memory (e.g., a hard disk, a solid state drive (SSD), a flash memory), a volatile memory (which can include a memory within the at least one processor 160), or the like.

[0051] The at least one processor 160 can control the robot 100 according to at least one instruction stored in the memory 150.

[0052] Specifically, the at least one processor 160 can include one or more processors. Particularly, the one or more processors can include one or more of a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), an integrated many-core (MIC), a digital signal processor (DSP), a neural processing unit (NPU), a hardware accelerator, or a machine learning accelerator. The one or more processors can control one or a random combination of the other elements of the robot 100 and perform operations associated with communication or data processing. The one or more processors can execute one or more programs or instructions stored in the memory. For example, the one or more processors can perform a method according to an embodiment of the disclosure by executing one or more instructions stored in the memory.

[0053] When a method according to an embodiment of the disclosure includes a plurality of operations, the plurality of operations can be performed by one processor or by a plurality of processors. That is, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation can all be performed by a first processor, or the first operation and the second operation can be performed by a first processor (e.g., a general-purpose processor) and the third operation can be performed by a second processor (e.g., an artificial intelligence dedicated processor). For example, the robot 100 can perform an operation of generating a map, etc. using a general-purpose processor, and perform an operation of identifying an object, an operation of identifying a user's voice, etc. using an artificial intelligence dedicated processor.

[0054] The one or more processors can be implemented as a single core processor including one core, or as one or more multi-core processors including a plurality of cores (e.g., homogeneous multi-cores or heterogeneous multi-cores). If the one or more processors are implemented as a multi-core processor, each of the plurality of cores included in the multi-core processor can include a memory (such as a cache memory and an on-chip memory) internal to the processor, and a common cache shared by the plurality of cores can be included in the multi-core processor. Furthermore, each of the plurality of cores included in the multi-core processor (or a part of the plurality of cores) can independently read and execute program commands for implementing a method according to an embodiment of the disclosure, or read and execute program commands for implementing a method according to an embodiment of the disclosure due to interconnection of all (or a part of) the plurality of cores.

[0055] When a method according to an embodiment of the disclosure includes a plurality of operations, the plurality of operations can be executed by one core of the plurality of cores, or by the plurality of cores included in the multi-core processor. For example, when a first operation, a second operation, and a third operation are executed by a method according to an embodiment, the first operation, the second operation, and the third operation can all be executed by a first core included in the multi-core processor, or the first operation and the second operation can be executed by the first core included in the multi-core processor, and the third operation can be executed by a second core included in the multi-core processor.

[0056] In one or more embodiments of the disclosure, the at least one processor 160 can generate a map including information about a plurality of objects based on sensing information obtained through the at least one sensor 110. The plurality of objects can include articles, walls, windows, etc. included in a house. Furthermore, the map can include a drawing representing a floor plan by reducing a state of a space within the house to a certain ratio. According to an embodiment of the disclosure, the map can include a drawing in which a floor structure inside a house is reduced to a certain ratio and represented with a predetermined symbol. Furthermore, the map can include a drawing in which a floor structure inside a house is represented with a line. However, the above is not limited thereto, and the map can include locations of main objects inside a house. The map can be implemented as at least one of a two-dimensional grid map, a two-dimensional linear vector map, a three-dimensional grid map, or a three-dimensional map.

[0057] The at least one processor 160 can generate, through the speaker 120, an ultrasonic wave toward each of the plurality of objects, receive, through the microphone 130, a reflected sound reflected from each of the plurality of objects, obtain reflectance information for each of the plurality of objects, and store the reflectance information in the at least one memory 150. The reflectance information related to each of the plurality of objects can include information on a ratio of an intensity of the ultrasonic wave output through the speaker 120 to an intensity of the reflected sound received through the microphone. Information on the reflectance information can be obtained for each object, but this is only one embodiment, and information on the reflectance information can be obtained according to a direction even for the same object.

[0058] Specifically, the at least one processor 160 can obtain information on reflectance in different methods according to a type of an object.

[0059] In an embodiment, the at least one processor 160 can generate an ultrasonic wave at a preset distance interval for a wall object among the plurality of objects, and obtain reflectance information for the wall object by receiving, through the microphone 130, a reflected sound reflected from the wall object. In an embodiment, the at least one processor 160 can generate an ultrasonic wave toward the remaining objects except for the wall object among the plurality of objects from two or more directions, receive, through the microphone 130, a sound reflected from the remaining objects based on the ultrasonic wave generated from the two or more directions, and obtain reflectance information for the remaining objects by obtaining an average of the received sounds reflected from the two or more directions.

[0060] With the method as described above, the at least one processor 160 can store the reflectance information for the plurality of objects and the generated map.

[0061] When a user voice is received through the microphone 130, the at least one processor 160 can obtain information on an intensity of the user voice for each of the plurality of directions. The intensity of the user voice can represent a sound wave energy transmitted per unit time and unit area, and can be replaced with different terms such as, for example, and without limitation, a magnitude of the user voice, a volume of the user voice, an energy of the user voice, etc. Further, the at least one processor 160 can obtain information on the intensity of the user voice for each of the plurality of directions for each preset interval (for example, 1 degree) based on the robot 100.

[0062] The at least one processor 160 can obtain information about a plurality of candidate directions from which the user's voice is received, based on information about the intensity of the user's voice. The plurality of candidate directions can be directions from which the user's voice is likely to be emitted among the plurality of directions. Specifically, the at least one processor 160 can obtain information about the intensity of the user's voice received from the plurality of directions, based on the location of the robot 100. Furthermore, the at least one processor 160 can identify a preset number (e.g., five) of directions from the plurality of directions in which the intensity of the user's voice is high as the plurality of candidate directions. As used herein, the intensity can be considered to be "high" if the intensity exceeds a predetermined threshold. Alternatively, the intensity can be considered to be "high" if the intensity is in a subset of the highest intensity values among a total set of obtained intensity values (e.g., the five highest intensity values in a larger set of intensity values).

[0063] The at least one processor 160 can obtain priority order information for the plurality of candidate directions, based on the location of the robot 100 and the stored reflectance information. The priority order information can include information about directions from which the user's voice is likely to be emitted among the plurality of directions.

[0064] Specifically, the at least one processor 160 can identify objects located in the plurality of candidate directions, based on the location of the robot 100. Then, the at least one processor 160 can identify a priority order for the plurality of candidate directions, based on the intensity of the user's voice measured from each of the plurality of candidate directions and the reflectance information for the identified objects. The at least one processor 160 can correct the intensity of the user's voice obtained from the plurality of candidate directions by multiplying a weight value corresponding to the respective reflectance of the objects located in the plurality of candidate directions by the intensity of the user's voice obtained from the plurality of candidate directions. The weight value can be set to be low in the case where the reflectance of the object located in the candidate direction is high, and the weight value can be set to be high in the case where the reflectance of the object located in the candidate direction is low. Then, the at least one processor 160 can identify the priority order based on the corrected intensity of the user's voice in each of the plurality of candidate directions.

[0065] The at least one processor 160 can obtain information about a direction from which the user's voice is emitted from the plurality of candidate directions, based on the priority order information. Specifically, when the corrected intensity of the user's voice for each of the plurality of candidate directions is large, the at least one processor 160 can identify the priority order to be high, and obtain information about a direction from which the user's voice is emitted by identifying a candidate direction in which the priority order is the highest among the plurality of candidate directions.

[0066] The at least one processor 160 can perform various functions based on the information about the direction from which the user's voice is emitted. In an embodiment, the at least one processor 160 can perform voice recognition on the user's voice by performing beamforming in the direction from which the user's voice is emitted. The operation of performing voice recognition on the user's voice can include at least one of an operation of obtaining text corresponding to the user's voice, an operation of performing natural language understanding using the text corresponding to the user's voice, and an operation of providing a response or controlling the robot 100 based on a result of the natural language understanding. In an embodiment, the at least one processor 160 can rotate or move the robot 100 so that the robot 100 faces the direction from which the user's voice is emitted.

[0067] As described above, by identifying the direction from which the user's voice is emitted using the reflectance information, the robot 100 can provide various functions by more accurately identifying the direction in which the user is located.

[0068] Hereinafter, a method of generating a map including reflectance information and identifying a direction from which a user's voice is emitted using the generated map will be described with reference to Figures 2 to 5d The detailed description generates a map including reflectance information and identifies a direction from which a user's voice is emitted using the generated map.

[0069] Figure 2 is a block diagram illustrating a plurality of configurations for generating a map including reflectance information and identifying a direction from which a user's voice is emitted using the generated map according to an embodiment of the disclosure. As Figure 2 shown, the robot 100 can include a map generation module 210 and an identification module 220.

[0070] The map generation module 210 can generate a map including various information about objects based on sensing information received from the at least one sensor 110 and reflected sound obtained through the microphone 130. Specifically, the map generation module 210 can include a topographic information obtaining module 211, an object information obtaining module 212, and a reflectance information obtaining module 213 as Figure 2 shown.

[0071] The topographic information obtaining module 211 can obtain topographic information of a home space based on sensing information obtained through a laser radar sensor and an IMU sensor while the robot 100 is traveling. For example, as Figure 3a shown, if the robot 100 in the home space 310 generates a map according to a user command, the topographic information obtaining module 211 can control the robot 100 to travel in a preset pattern. The topographic information obtaining module 211 can obtain sensing information through the laser radar sensor and the IMU sensor while the robot 100 is traveling. The topographic information obtaining module 211 can obtain a map 320 of topographic information as Figure 3b shown, based on the obtained sensing information, the map 320 including an outermost line 321 of the home space 310.

[0072] The object information obtaining module 212 can obtain information about objects through an image obtained via the camera. Specifically, the object information obtaining module 212 can obtain information about a plurality of objects included in the indoor space by inputting an image obtained through the camera into a trained neural network model (e.g., an object recognition model). For example, as Figure 3a indicated, if the robot 100 in the indoor space 310 generates a map according to a user command, the object information obtaining module 212 can obtain information about the first object 331, information about the second object 332, and information about the third object 333 by inputting an image obtained through the camera into a trained neural network model. The information about the objects can include at least one of shape information, type information, size information, and location information of the objects. In addition, the object information obtaining module 212 can store the obtained information about the objects in the generated map 320 in the form of metadata.

[0073] The reflectivity information obtaining module 213 can generate, after the generation of the map 320 or while the map 320 is being generated, ultrasonic waves directed to each of the plurality of objects through the speaker 120, receive reflected sound reflected from each of the objects through the microphone 130, obtain reflectivity information for the plurality of objects, and store the reflectivity information. The reflectivity information obtaining module 213 can obtain information about reflectivity in various methods according to the type of the object.

[0074] In an embodiment, the reflectivity information obtaining module 213 can generate ultrasonic waves at a wall at a certain distance (e.g., 2 m) apart for the outermost object (e.g., a wall object, a door object, etc.) and obtain information about reflectivity by measuring the amount of reflection of the reflected sound. For example, as Figure 4a indicated, the reflectivity information obtaining module 213 can generate ultrasonic waves at a wall from the first position 410-1, measure a first reflectivity by measuring the amount of reflection of the reflected sound, generate ultrasonic waves at a wall from the second position 410-2 after having moved a certain distance (e.g., 2 m) from the first position 410-1, and measure a second reflectivity by measuring the amount of reflection of the reflected sound. The reflectivity information obtaining module 213 can obtain reflectivity information for the outermost object by moving in the indoor space 310 in the method as described above. In addition, the reflectivity information obtaining module 213 can obtain a representative value (e.g., an average value or a mode value, etc.) of the reflectivity obtained in the method as described above and store it in the generated map 320. However, the above description relates to only one embodiment, and the reflectivity information obtaining module 213 can store reflectivity information for the outermost object obtained from a plurality of positions in the generated map 320 for each position.

[0075] In an embodiment, the reflectivity information obtaining module 213 can generate an ultrasound wave from two or more directions pointing to the rest of the objects (e.g., home appliances, furniture, etc.) except for the outermost object among the plurality of objects, receive reflected sound reflected from the two or more directions toward the rest of the objects, and obtain reflectivity information for the rest of the objects by obtaining an average of the received sound reflected from the two or more directions toward each of the rest of the objects. For example, as shown in Figure 4b the reflectivity information obtaining module 213 can generate an ultrasound wave at the second object 332 from a first direction 420-1 based on the second object 332, measure a first reflectivity by measuring a reflection amount of reflected sound, generate an ultrasound wave at the second object 332 from a second direction 420-2, measure a second reflectivity by measuring a reflection amount of reflected sound, and generate an ultrasound wave at the second object 332 from a third direction 420-3. The reflectivity information obtaining module 213 can obtain reflectivity information for the second object 332 based on the sound reflected from the plurality of directions based on the above-described method. In addition, the reflectivity information obtaining module 213 can store representative values of the first reflectivity to the third reflectivity obtained in the method as described above in the generated map 320. However, the above describes only one embodiment, and the reflectivity information obtaining module 213 can store reflectivity information for the rest of the objects obtained from the plurality of directions in the generated map 320 for each direction.

[0076] The reflectivity information obtaining module 213 can store the reflectivity information obtained in the method as described above in the form of metadata of the map 320. In an example, as shown in Figure 4c the reflectivity information obtaining module 213 can store reflectivity information for the first object 331 as 0.2, store reflectivity information for the second object 332 as 0.7, store reflectivity information for the third object 333 as 0.4, and store reflectivity information for the wall object as 0.8.

[0077] The recognition module 220 can identify a sound emission location of a user voice using the map obtained by the map generation module 210 and identify the user voice by using the identified sound emission location. As shown in Figure 2 the recognition module 220 can include a user voice obtaining module 221, a preprocessing module 222, an intensity measurement module 223, a candidate direction obtaining module 224, a priority order information obtaining module 225, a direction recognition module 226, and a user voice recognition module 227.

[0078] The user voice obtaining module 221 can obtain a user voice through the microphone 130. The user voice can be a voice uttered by a user and can be distinguished from noise or the like. Specifically, the user voice obtaining module 221 can identify whether a user voice is included in an audio signal obtained by inputting the audio signal into a trained neural network model. For example, as shown in FIG. 10, the user voice obtaining module 221 can obtain a user voice "Robot nano, can you look at me?" uttered by the user 10. Figure 5a

[0079] The preprocessing module 222 can perform a preprocessing operation on the obtained user voice. Specifically, the preprocessing module 222 can obtain an audio signal for a part of a bandwidth from a user voice obtained through a band-pass filter. Then, the preprocessing module 222 can perform a preprocessing operation of removing noise, echo, or the like.

[0080] The intensity measuring module 223 can measure intensities of user voices obtained from a plurality of directions. Then, the intensity measuring module 223 can perform sampling on user voices obtained from a plurality of directions. Specifically, the intensity measuring module 223 can measure intensities of user voices obtained from a plurality of directions by using a steered response power phase transform (SRP-PHAT) algorithm. The intensity measuring module 223 can measure intensities of user voices for each preset interval (e.g., 1 degree) based on the robot 100 from a plurality of directions (e.g., 360 directions).

[0081] The candidate direction obtaining module 224 can identify a preset number of directions in which intensities of user voices are considered to be high from among a plurality of directions as a plurality of candidate directions. For example, as shown in FIG. 11, the candidate direction obtaining module 224 can obtain five directions in which intensities of user voices are high from among directions of a plurality of directions (e.g., 360 directions). That is, an intensity of a user voice obtained in a first candidate direction 510-1 can be 1.2, an intensity of a user voice obtained in a second candidate direction 510-2 can be 0.8, an intensity of a user voice obtained in a third candidate direction 510-3 can be 0.8, an intensity of a user voice obtained in a fourth candidate direction 510-4 can be 0.7, and an intensity of a user voice obtained in a fifth candidate direction 510-5 can be 1.1. Figure 5b Figure 5b The intensities shown in FIG. 11 can be relative values. The candidate direction obtaining module 224 can identify the highest intensity information from among intensity information of user voices received from adjacent directions (e.g., directions within a range of 10 degrees) among a plurality of directions as a candidate direction, and exclude the remaining adjacent directions from the candidate directions even if intensities of user voices in the remaining adjacent directions are greater than those of other directions.

[0082] ​​The priority order information obtaining module 225 can obtain priority order information for the plurality of candidate directions based on the position of the robot 100 and the stored reflectivity information.

[0083] Specifically, the priority order information obtaining module 225 can identify objects located in the plurality of candidate directions based on the position of the robot. For example, the priority order information obtaining module 225 can identify first to fifth objects located in first to fifth candidate directions based on the position of the robot 100. The objects can not exist in a part of the plurality of candidate directions.

[0084] Further, the priority order information obtaining module 225 can identify weight values corresponding to reflectivity for the objects located in the plurality of candidate directions. In a case where the reflectivity of the object located in the candidate direction is high, the weight value can be set to be low, and in a case where the reflectivity of the object located in the candidate direction is low, the weight value can be set to be high. For example, the weight value and the reflectivity can have a relationship as in Equation 1 below.

[0085] Equation 1: Weight value = 1 - Reflectivity However, the above Equation 1 is only one embodiment, and the relationship between the weight value and the reflectivity can be expressed in an equation in which the weight value and the reflectivity have a different inverse relationship.

[0086] The priority order information obtaining module 225 can correct the intensity of the user voice obtained from the plurality of candidate directions by multiplying the weight value corresponding to the reflectivity for the objects located in the plurality of candidate directions by the intensity of the user voice obtained from the plurality of candidate directions. For example, the priority order information obtaining module 225 can correct the intensity of the user voice obtained from the first candidate direction to 0.36 by multiplying the intensity information of the user voice obtained from the first candidate direction by 0.3, which is the weight value corresponding to the reflectivity for the first object located in the first candidate direction among the plurality of candidate directions, and correct the intensity of the user voice obtained from the third candidate direction to 0.56 by multiplying the intensity information of the user voice obtained from the third candidate direction by 0.7, which is the weight value corresponding to the reflectivity for the third object located in the third candidate direction among the plurality of candidate directions.

[0087] The priority order information obtaining module 225 can identify priority order information for the plurality of candidate directions based on the corrected intensity of the user voice for each of the plurality of candidate directions. That is, when the corrected intensity of the user voice for each of the plurality of candidate directions is high, the priority order information obtaining module 225 can identify the priority order of a given candidate direction to be high, and when the intensity of the user voice is low, identify the priority order of the given candidate direction to be low. For example, as in Equation 2 below, the priority order information obtaining module 225 can identify the priority order of the first candidate direction to be high, the priority order of the second candidate direction to be low, the priority order of the third candidate direction to be high, the priority order of the fourth candidate direction to be low, and the priority order of the fifth candidate direction to be high.Figure 5c As shown, the priority order information obtaining module 225 can identify the intensity of the user voice for the first candidate direction 510-1 as 0.36, the intensity of the user voice for the second candidate direction 510-2 as 0.8, the intensity of the user voice for the third candidate direction 510-3 as 0.56, the intensity of the user voice for the fourth candidate direction 510-4 as 0.7, and the intensity of the user voice for the fifth candidate direction 510-5 as 1.1. Then, the priority order information obtaining module 225 can identify the fifth candidate direction 510-5 as the first priority order, the second candidate direction 510-2 as the second priority order, the fourth candidate direction 510-4 as the third priority order, the third candidate direction 510-3 as the fourth priority order, and the first candidate direction 510-1 as the fifth priority order.

[0088] The direction identifying module 226 can identify the direction from which the user voice is emitted based on the priority order information. Specifically, the direction identifying module 226 can identify the candidate direction of the first priority order as the direction from which the user voice is emitted. For example, as shown, the direction identifying module 226 can identify the fifth candidate direction 510-5, which is identified as the first priority order, as the direction from which the user voice is emitted. Figure 5d As shown, the direction identifying module 226 can identify the fifth candidate direction 510-5, which is identified as the first priority order, as the direction from which the user voice is emitted. In another example, the direction identifying module 226 can identify a candidate area in which the user is located by capturing the candidate directions with the camera in the order of the high priority order, and identify the identified candidate area as the direction from which the user voice is emitted.

[0089] The user voice recognizing module 227 can perform speech recognition on the user voice. The user voice recognizing module 227 can perform speech recognition on the (additional) user voice by performing beamforming in the direction from which the user voice is emitted. The user voice recognizing module 227 can convert the additional user voice into text, perform natural language understanding on the converted text, and provide a response to the user 10 or control the robot 100 according to the user voice based on the result of the natural language understanding. For example, the user voice recognizing module 227 can identify the user voice "Robot, can you look at me?" as the user voice and make the robot 100 rotate toward the direction from which the user voice is emitted.

[0090] Figure 6 is a flowchart illustrating a control method of a robot according to an embodiment of the disclosure.

[0091] First, the robot 100 can obtain sensing information through at least one sensor (S605). Specifically, if an event for map creation (for example, an event in which a cleaning robot is initially installed and a user command for map creation is received, etc.) occurs, the robot 100 can travel in the in-home space by the pushing force of the driver 140. Then, the robot 100 can obtain sensing information for map generation while traveling in the in-home space by using at least one sensor 110. In an example, the robot 100 can obtain a sensing value for obtaining distance information between the robot 100 and an object using a laser radar sensor. In an example, the robot 100 can obtain a sensing value for obtaining movement information of the robot 100 by using an IMU sensor. In an example, the robot 100 can obtain an image for obtaining information about an object around the robot 100 by using a camera.

[0092] The robot 100 can generate a map including information about a plurality of objects based on the sensing information (S610). Specifically, the robot 100 can obtain terrain information of the in-home space using an IMU sensor and a laser radar sensor, and obtain information about a plurality of objects included in the in-home space by inputting an image obtained through a camera into a trained neural network model. Various information such as, for example, and without limitation, type information of each of the plurality of objects, size information of each of the plurality of objects, shape information of each of the plurality of objects, position information of each of the plurality of objects, etc. can be included in the information about the plurality of objects. In addition, the robot 100 can store the information about the objects in the form of metadata, as well as a map including information about the in-home space.

[0093] The robot 100 can generate an ultrasonic wave at each of the plurality of objects (S615). The robot 100 can generate an ultrasonic wave using different methods based on the type of the plurality of objects.

[0094] In an example, the robot 100 can generate an ultrasonic wave at a wall at a specific distance interval (for example, 3 m) for a wall object among the plurality of objects. For an outermost object (such as a window or a door) other than the wall object, an ultrasonic wave can be generated at a specific distance interval as described above. In an example, the robot 100 can determine a number of two or more samples (for example, five) within an angle range in which the robot 100 can see for the remaining objects (for example, furniture, articles, etc. in the home) other than the outermost object among the plurality of objects, and generate an ultrasonic wave based on the object from two or more directions by equally dividing the angle according to the number of samples.

[0095] The robot 100 can generate ultrasound waves at each of the plurality of objects while traveling in the indoor space to generate a map, but this is only one embodiment, and the robot 100 can generate ultrasound waves at each of the plurality of objects after having generated a map.

[0096] The robot 100 can obtain reflectivity information for the object by receiving reflected sound for each of the plurality of objects, and store the obtained reflectivity information (S620). Specifically, the robot 100 can obtain information about the intensity of the reflected sound by receiving reflected sound for each of the plurality of objects. Then, the robot 100 can obtain information about the reflectivity of each of the plurality of objects based on the ratio of the intensity of the output ultrasound wave to the intensity of the received reflected sound.

[0097] The robot 100 can obtain information about the reflectivity for each of the specific distance intervals by receiving reflected sound about ultrasound waves generated at specific distance intervals for the outermost object. The robot 100 can store information about the reflectivity for each of the specific distance intervals, and store a representative value (e.g., an average value, a mode value, etc.) of the reflectivity obtained at each of the specific distance intervals as information about the reflectivity of the object. In addition, the robot 100 can obtain information about the reflectivity according to at least two directions by receiving reflected sound obtained from at least two directions for the remaining objects other than the outermost object. The robot 100 can store information about the reflectivity for each of the at least two directions, and store a representative value (e.g., an average value, a mode value, etc.) of the reflectivity obtained in each of the at least two directions as information about the reflectivity of the object.

[0098] The robot 100 can store information about the reflectivity of the object as well as the map. As described above, the robot 100 can store information about the reflectivity of the object on the map in the form of metadata.

[0099] The robot 100 can generate a map including reflectivity information for the object by operating S605 to operation S620, and store it in the robot 100. The robot 100 can not only store the map in the robot 100, but also transmit the map to an external server to store the map in the external server. In addition, operation S605 to operation S620 can be performed by another electronic device other than the robot 100.

[0100] After operation S620 (i.e., after generating the map), the robot 100 can identify whether a user voice is received through the microphone 130 (S625). Specifically, the robot 100 can identify whether a user voice is received using a trained neural network model to identify whether the voice is a human voice, but this is only one embodiment, and whether a user voice is received can be identified in another method.

[0101] If the user voice is recognized to be received (S625 - Yes), the robot 100 can pre-process the user voice (S630). Specifically, the robot 100 can obtain the user voice of a specific frequency range using a preset filter (e.g., a band-pass filter, etc.), and perform pre-processing (such as noise removal) of the user voice.

[0102] The robot 100 can obtain information about the intensity of the user voice for each of the plurality of directions (S635). Specifically, the robot 100 can obtain information about the intensity of the user voice for each of the plurality of directions at a preset interval (e.g., 1 degree) based on the robot 100. In an example, the robot 100 can obtain information about the intensity of the user voice for the first direction to the three hundred sixty directions.

[0103] The robot 100 can obtain information about a plurality of candidate directions in which the user voice is received (S640). Specifically, the robot 100 can identify a preset number (e.g., five) of directions in which the intensity of the user voice is high as the plurality of candidate directions based on the information about the intensity of the user voice obtained for each of the plurality of directions. The robot 100 can identify the maximum intensity information among the intensity information of the user voice received from adjacent directions (e.g., directions within a 10-degree range) among the plurality of directions as a candidate direction, and exclude the remaining adjacent directions from the candidate directions even if the intensity of the user voice in the remaining adjacent directions is greater than that of the other directions.

[0104] The robot 100 can obtain priority order information for the plurality of candidate directions (S645). Specifically, the robot 100 can obtain the priority order information for the plurality of candidate directions based on the corrected intensity of the user voice after having corrected the intensity of the user voice received from the candidate directions based on the position and reflectance information of the robot 100.

[0105] Specifically, the robot 100 can identify objects located in the plurality of candidate directions based on the position of the robot. For example, the robot 100 can identify a first object located in a first candidate direction and a second object located in a second candidate direction based on the position of the robot 100.

[0106] Then, the robot 100 can identify weight values corresponding to reflectivity with respect to objects located in the plurality of candidate directions. In a case where reflectivity with respect to objects located in a candidate direction is high, the weight value can be set to be low, and in a case where reflectivity with respect to objects located in a candidate direction is low, the weight value can be set to be high. That is, since it can be a reflected sound of user voice reflected from an object rather than actual user voice when reflectivity with respect to objects located in a candidate direction is high, the weight value can be set to be low. The value of the weight value can be a value between 0 and 1, but the disclosure is not limited thereto.

[0107] The robot 100 can correct the intensity of user voice obtained from the plurality of candidate directions by multiplying the weight values corresponding to reflectivity with respect to objects located in the plurality of candidate directions by the intensity of user voice obtained from the plurality of candidate directions. That is, the robot 100 can multiply the intensity of user voice received from a direction in which an object having high reflectivity is located by a weight value set to be low, and multiply the intensity of user voice received from a direction in which an object having low reflectivity is located by a weight value set to be high.

[0108] Then, the robot 100 can identify priority order information with respect to the plurality of candidate directions based on the corrected intensity of user voice with respect to each of the plurality of candidate directions. That is, when the corrected intensity of user voice with respect to each of the plurality of candidate directions is great, the robot 100 can identify the priority order to be high, and when the intensity of user voice is small, can identify the priority order to be low.

[0109] The robot 100 can obtain information about a direction from which user voice is emitted based on the priority order information (S650). Specifically, the robot 100 can identify a candidate direction having the highest priority order as a direction from which user voice is emitted. Alternatively, the robot 100 can identify a candidate area in which a user is located by capturing the candidate direction with the camera in order of high priority order, and identify the identified candidate area as a direction from which user voice is emitted.

[0110] Then, the robot 100 can identify user voice (S655). Specifically, the robot 100 can drive the robot 100 to face a direction in which a user is located based on the direction from which user voice is emitted. In addition, the robot 100 can perform voice recognition on (additional) user voice by performing beamforming in the direction from which user voice is emitted.

[0111] The functions associated with artificial intelligence according to the disclosure can be operated by a processor and a memory of the robot 100.

[0112] The processor can be formed of one or more processors. The one or more processors can include at least one of a CPU, a GPU, and an NPU, but are not limited to the examples of the processors described above.

[0113] The CPU can be a general-purpose processor that can not only perform general operations but also perform artificial intelligence operations, and can efficiently run complex programs through a multi-tier cache structure. The CPU can be advantageous in a serial processing method that allows organic connection between previous calculation results and subsequent calculation results through successive calculations. The general-purpose processor is not limited to the examples described above, except when designated as the CPU described above.

[0114] The GPU can be a processor for large-scale operations, such as floating-point operations used in graphics operations, and performs large-scale operations in parallel by integrating a large number of cores. Specifically, the GPU can be advantageous in a parallel processing method, such as a convolution operation, compared to the CPU. In addition, the GPU can function as a co-processor for supplementing the functions of the CPU. The processor for large-scale operations is not limited to the examples described above, except when designated as the GPU described above.

[0115] The NPU can be a processor dedicated to artificial intelligence operations using an artificial neural network, and can implement each layer forming an artificial neural network as hardware (e.g., silicon). Since the NPU can be designed specifically according to the requirements of a company, the NPU can have a lower degree of freedom compared to the CPU or the GPU, but can efficiently process artificial intelligence operations required by the company. As a processor dedicated to artificial intelligence operations, the NPU can be implemented in various forms such as, for example, and without limitation, a tensor processing unit (TPU), an intelligent processing unit (IPU), a visual processing unit (VPU), etc. The artificial intelligence processor is not limited to the examples described above, except when designated as the NPU described above.

[0116] In addition, the one or more processors can be implemented as a system on chip (SoC). The SoC can include a memory and a network interface, such as a bus for data communication between the processor and the memory, in addition to the one or more processors.

[0117] If a plurality of processors are included in a system on chip (SoC) included in a robot, the robot can use a part of the plurality of processors to perform operations associated with artificial intelligence (e.g., operations associated with learning or inference of an artificial intelligence model). For example, the robot can use at least one of a GPU, an NPU, a VPU, a TPU, and a hardware accelerator dedicated to artificial intelligence operations, such as convolution operations and matrix multiplication operations, among the plurality of processors, to perform operations associated with artificial intelligence. However, the above is only one embodiment, and a general-purpose processor such as a CPU can be used to process operations associated with artificial intelligence.

[0118] Further, the robot can perform an operation regarding a function associated with artificial intelligence by using a multi-core (e.g., dual-core, quad-core, etc.) included in one processor. Specifically, the robot can perform an artificial intelligence operation (such as, for example, but not limited to, a convolution operation, a matrix multiplication operation, etc.) in parallel using the multi-core included in the processor.

[0119] The one or more processors can be controlled to process input data according to a predefined operation rule or an artificial intelligence model stored in the memory. The predefined operation rule or the artificial intelligence model is characterized by being created through learning.

[0120] Here, the creation through learning can refer to forming a predefined operation rule or an artificial intelligence model of a desired characteristic by applying a learning algorithm to a plurality of learning data. The learning can be performed in the machine itself that performs artificial intelligence according to the disclosure, or by a separate server / system.

[0121] The artificial intelligence model can be formed of a plurality of neural network layers. At least one layer can have at least one weight value, and an operation of the layer is performed by an operation result of a previous layer and at least one defined operation. Examples of the neural network can include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, and a transformer, and the neural network of the disclosure is not limited to the above examples unless otherwise specified.

[0122] The learning algorithm can be a method for training a predetermined target machine (e.g., a robot) to make decisions or predictions on its own using a plurality of learning data. Examples of the learning algorithm can include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, and the learning algorithm of the disclosure is not limited to the above examples unless otherwise specified.

[0123] The method according to various embodiments of the disclosure can be provided and included in a computer program product. The computer program product can be traded between a seller and a buyer as goods. The computer program product can be distributed in the form of a machine-readable storage medium (e.g., a compact disc read only memory (CD-ROM)) or online through an application store (e.g., PLAYSTORE™) (e.g., downloaded or uploaded) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product (e.g., an app) can be stored or temporarily generated in a storage medium such as a memory of a manufacturer's server, an application store's server, or a relay server.

[0124] Methods according to various embodiments of the disclosure can be implemented by software including instructions stored in a machine-readable storage medium (e.g., computer). The machine can invoke the stored instructions from the storage medium, and as an apparatus operable according to the invoked instructions, can include a robot according to the above-described embodiments.

[0125] The machine-readable storage medium can be provided in the form of a non-transitory storage medium. Here, the "non-transitory storage medium" only means that the storage medium is a tangible device, and does not include a signal (e.g., an electromagnetic wave), and the term does not distinguish between data semi-permanently stored in the storage medium and data temporarily stored in the storage medium. In an example, the "non-transitory storage medium" can include a buffer that temporarily stores data.

[0126] Based on the instructions executed by the processor, the processor can directly or under the control of the processor using other elements to perform functions corresponding to the instructions. The instructions can include code generated by a compiler or run by an interpreter.

[0127] While the disclosure has been shown and described with reference to various example embodiments thereof, it will be understood that the various example embodiments are intended to be illustrative only and are not limiting. Those skilled in the art will understand that various changes in form and detail can be made without departing from the true spirit and full scope of the disclosure, including the appended claims and their equivalents.

Claims

1. A robot comprising: at least one sensor; a speaker; a microphone; a driver; at least one memory storing one or more instructions; and at least one processor configured to execute the one or more instructions, wherein the one or more instructions, when executed by the at least one processor, cause the robot to: generate a map including information about a plurality of objects based on sensing information obtained through the at least one sensor, generate, through the speaker, ultrasonic waves toward each of the plurality of objects, obtain reflectivity information about the plurality of objects based on reflected sound reflected from each of the plurality of objects and received through the microphone, the reflected sound reflected from each of the plurality of objects being at least a portion of the ultrasonic waves reflected from each of the plurality of objects, and store the reflectivity information, obtain information about intensity of a user voice for each of a plurality of directions based on the user voice being received through the microphone, obtain information about a plurality of candidate directions from which the user voice is received based on the information about the intensity of the user voice for each of the plurality of directions, obtain priority order information for the plurality of candidate directions based on a position of the robot and the stored reflectivity information, and obtain information about a direction from which the user voice is emitted from the plurality of candidate directions based on the priority order information. the one or more instructions, when executed by the at least one processor, further cause the robot to:

2. The robot of claim 1, wherein, generate ultrasonic waves at a preset distance interval for a wall object among the plurality of objects, and obtain reflectivity information about the wall object based on reflected sound reflected from the wall object and received through the microphone, the reflected sound reflected from the wall object being at least a portion of the ultrasonic waves generated at the preset interval and reflected from the wall object. the one or more instructions, when executed by the at least one processor, further cause the robot to:

3. The robot of claim 2, wherein, generate ultrasonic waves toward objects other than the wall object among the plurality of objects from two or more directions, and obtain reflectivity information about the objects by obtaining an average of reflected sound reflected from the objects and received through the microphone, the reflected sound reflected from the objects being at least a portion of the ultrasonic waves generated from the two or more directions and reflected from the objects. the one or more instructions, when executed by the at least one processor, further cause the robot to:

4. The robot of claim 1, wherein, obtain information about intensity of the user voice for each of the plurality of directions based on a position of the robot, and identify a preset number of directions in which a respective intensity of the user voice among the plurality of directions exceeds a predetermined threshold as the plurality of candidate directions. the one or more instructions, when executed by the at least one processor, further cause the robot to:

5. The robot of claim 1, wherein, ​ identifying objects among the plurality of objects located in the plurality of candidate directions with respect to a position of the robot, and identifying a priority order for the plurality of candidate directions based on respective information on intensities of the user voice corresponding to each of the plurality of candidate directions and reflectivity information corresponding to the identified objects.

6. The robot of claim 5, wherein, The one or more instructions, when executed by the at least one processor, further cause the robot to: obtain a corrected intensity of the user voice for each of the plurality of candidate directions by multiplying a weight value by a respective intensity of the user voice corresponding to each of the plurality of candidate directions, and identify the priority order based on the corrected intensities for the plurality of candidate directions, and wherein the weight value corresponds to the reflectivity information corresponding to the identified objects.

7. The robot of claim 6, wherein, The weight value and the reflectivity information corresponding to the identified objects are inversely proportional to each other.

8. The robot of claim 6, wherein, For each of the plurality of candidate directions, the priority order is proportional to the corrected intensity, and wherein the one or more instructions, when executed by the at least one processor, further cause the robot to: identify a candidate direction having a highest priority order from among the plurality of candidate directions as a direction from which the user voice is emitted.

9. The robot of claim 1, wherein, The one or more instructions, when executed by the at least one processor, further cause the robot to: perform voice recognition on the user voice by performing beamforming in the direction from which the user voice is emitted. 10.A method for controlling a robot, the method comprising: generating a map including information on a plurality of objects based on sensing information obtained through at least one sensor of the robot; generating, through a speaker of the robot, ultrasonic waves toward each of the plurality of objects; obtaining reflectivity information on the plurality of objects based on reflected sounds reflected from each of the plurality of objects and received through a microphone of the robot, and storing the reflectivity information, the reflected sounds reflected from each of the plurality of objects being at least a portion of the ultrasonic waves reflected from each of the plurality of objects; obtaining information on intensities of a user voice for each of a plurality of directions based on the user voice being received through the microphone; obtaining information on a plurality of candidate directions from which the user voice is received from the plurality of directions based on the information on the intensities of the user voice for each of the plurality of directions; obtaining priority order information for the plurality of candidate directions based on a position of the robot and the stored reflectivity information; and obtaining information on a direction from which the user voice is emitted from the plurality of candidate directions based on the priority order information. The operation of generating the ultrasonic waves toward each of the plurality of objects includes generating the ultrasonic waves at a preset distance interval for a wall object among the plurality of objects, 11. The method of claim 10, wherein, ​ The operation of obtaining the reflectivity information includes obtaining reflectivity information about the wall object based on reflected sound reflected from the wall object and received through the microphone, wherein the reflected sound reflected from the wall object is at least a portion of the ultrasonic wave generated at the preset interval and reflected from the wall object.

12. The method of claim 11, wherein, The operation of generating ultrasonic waves toward each of the plurality of objects includes generating ultrasonic waves toward objects other than the wall object among the plurality of objects from two or more directions, and The operation of obtaining the reflectivity information further includes obtaining reflectivity information for the object by obtaining an average of reflected sound reflected from the object and received through the microphone, wherein the reflected sound reflected from the object is at least a portion of the ultrasonic wave generated from the two or more directions and reflected from the object.

13. The method of claim 10, wherein, The operation of obtaining information about the plurality of candidate directions includes: A preset number of directions in which the respective intensity of the user voice exceeds a predetermined threshold among the plurality of directions are identified as the plurality of candidate directions.

14. The method of claim 10, wherein, The operation of obtaining the priority order information includes: identifying objects among the plurality of objects whose positions relative to the robot are in the plurality of candidate directions; and identifying a priority order for the plurality of candidate directions based on respective information about the intensity of the user voice corresponding to each of the plurality of candidate directions and reflectivity information corresponding to the identified objects.

15. The method of claim 14, wherein, The operation of obtaining priority order information includes: obtaining a corrected intensity of the user voice for each of the plurality of candidate directions by multiplying a weight value by the respective intensity of the user voice corresponding to each of the plurality of candidate directions; and identifying a priority order based on the corrected intensities for the plurality of candidate directions, and wherein the weight value corresponds to reflectivity information corresponding to the identified objects.