Robot, learning data collection device, learning data collection method, and program

The robot efficiently collects and processes voice data through stimulus detection and adaptive collection conditions, improving speaker identification and interaction capabilities.

JP2026023899APending Publication Date: 2026-02-13CASIO COMPUTER CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024126218
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing robots face challenges in efficiently collecting training data for accurate speaker identification from voice inputs.

Method used

A robot equipped with external stimulus detection means, data collection means, and condition change means to selectively store sound data and adjust collection conditions based on detected stimuli, facilitating efficient data collection and identification.

Benefits of technology

Enables efficient collection and utilization of training data for speaker identification, enhancing the robot's ability to mimic human interaction and improve lifelikeness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023899000001_ABST
    Figure 2026023899000001_ABST
Patent Text Reader

Abstract

To efficiently collect learning data for identifying a speaker from a voice.SOLUTION: In the robot 200, the sensor unit 210 detects an external stimulus. When the sensor unit 210 detects a voice satisfying a predetermined collection condition as an external stimulus, the data collection unit 113 stores voice data indicating the detected voice in the storage unit 120 as the learning data 122. When an external stimulus satisfying a specific condition is detected by the sensor unit 210, the condition changing unit 116 changes the collection condition so that the sound detected by the sensor unit 210 more easily satisfies the collection condition until a predetermined time elapses after the external stimulus satisfying the specific condition is detected.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a robot, a learning data collection device, a learning data collection method, and a program. [Background technology]

[0002] Robots that mimic living creatures are known. For example, Patent Document 1 discloses a technology in which a pet-type robot learns information about a user's face and voice in advance, and the robot identifies the user, thereby behaving differently for each user. In particular, Patent Document 1 discloses that, in order to efficiently collect information used for user identification, the robot performs an action to collect voice data from a user who does not have enough voice data. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2019 / 181144 Summary of the Invention [Problem to be solved by the invention]

[0004] For a robot capable of identifying the speaker from the above-mentioned voice, there is a demand for more efficient collection of training data to enable accurate identification of the speaker.

[0005] The present invention is intended to solve the above-mentioned problems, and aims to provide a robot, a training data collection device, a training data collection method, and a program that can efficiently collect training data for identifying a speaker from a voice. [Means for solving the problem]

[0006] In order to achieve the above-mentioned object, one aspect of the robot of the present invention is a robot capable of collecting learning data for identifying a speaker from a sound, and is characterized by comprising: an external stimulus detection means for detecting external stimuli; a data collection means for, when the external stimulus detection means detects a sound that satisfies predetermined collection conditions as the external stimulus, storing sound data representing the detected sound in a storage means as the learning data; and a condition change means for, when the external stimulus detection means detects an external stimulus that satisfies specific conditions, changing the collection conditions so that the sound detected by the external stimulus detection means is more likely to satisfy the collection conditions until a predetermined time has elapsed since the external stimulus that satisfies the specific conditions was detected. [Effects of the Invention]

[0007] According to the present invention, it is possible to efficiently collect training data for identifying a speaker from a voice. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram showing the appearance of a robot according to a first embodiment. [Figure 2] 1 is a cross-sectional side view of a robot according to a first embodiment. [Figure 3] 1 is a block diagram showing the functional configuration of a robot according to a first embodiment. FIG. [Figure 4] FIG. 4 is a diagram illustrating an example of an event table according to the first embodiment. [Figure 5] FIG. 3 is a diagram illustrating an example of training data according to the first embodiment. [Figure 6] FIG. 4 is a diagram showing an example of classifying audio data into a plurality of groups in the first embodiment. [Figure 7] 4 is a flowchart showing the flow of a robot control process according to the first embodiment. [Figure 8] 1 is a flowchart showing the flow of data collection processing according to the first embodiment. [Figure 9] FIG. 10 is a block diagram showing the functional configuration of a data collection device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that identical or corresponding parts in the drawings are denoted by the same reference numerals. A robot 200 according to embodiment 1 is a device that simulates a living creature and is capable of simulating various states of a living creature. In particular, the robot 200 according to embodiment 1 is a pet-type robot that can recognize the voice of a specific user, such as a simulated owner. As an example, as shown in FIG. 1 , the robot 200 according to embodiment 1 is a pet robot that simulates a small animal. The robot 200 includes an exterior 201 having decorative parts 202 that resemble eyes and fluffy fur 203. As shown in FIG. 2 , the robot 200 includes a housing 207. The housing 207 is covered by the housing 201 and stored inside the housing 201. The housing 207 includes a head 204, a connecting part 205, and a body 206. The connecting part 205 connects the head 204 and the body 206.

[0010] The exterior 201 is an example of an exterior member, and has a bag-like shape that is long in the front-to-rear direction and can accommodate the housing 207 inside. The exterior 201 is formed in a cylindrical shape from the head 204 to the body 206, and integrally covers the body 206 and the head 204. With the exterior 201 having such a shape, the robot 200 is formed in a prone position. The outer surface of the exterior 201 is made of an artificial pile fabric that resembles the fur 203 of a small animal, in order to simulate the feel of the skin of a small animal. The lining of the exterior 201 is made of a flexible material such as leather, resin, or rubber. Because it is made of a flexible material, the exterior 201 follows the movement of the housing 207. Specifically, the exterior 201 follows the rotation of the head 204 relative to the body 206.

[0011] The body 206 extends in the front-to-rear direction, and comes into contact with a support surface such as a floor or a table on which the robot 200 is placed, via the exterior 201. The body 206 is provided with a twist motor 221 at its front end. The head 204 is connected to the front end of the body 206 via a connecting unit 205. The connecting unit 205 is provided with a vertical motor 222. Note that although the twist motor 221 is provided in the body 206 in FIG. 2, it may be provided in the connecting unit 205. The twist motor 221 and the vertical motor 222 connect the head 204 to the body 206 so as to be rotatable about axes in the left-right direction (X-axis direction) and the front-to-back direction (Y-axis direction) of the robot 200.

[0012] The connecting portion 205 connects the body portion 206 and the head portion 204 to be rotatable about a first rotation axis that passes through the connecting portion 205 and extends in the front-to-rear direction (Y-axis direction) of the body portion 206. The twist motor 221 is a servo motor for rotating the head portion 204 clockwise (right-hand rotation) and counterclockwise (left-hand rotation) relative to the body portion 206 about the first rotation axis. The connecting portion 205 also connects the body portion 206 and the head portion 204 to be rotatable about a second rotation axis that passes through the connecting portion 205 and extends in the left-to-right direction (X-axis direction) of the body portion 206. The up-down motor 222 is a servo motor for rotating the head portion 204 upward (forward rotation) and downward (reverse rotation) about the second rotation axis.

[0013] The robot 200 is equipped with touch sensors 211 on the head 204 and the torso 206. The robot 200 is also equipped with an acceleration sensor 212, a microphone 213, a gyro sensor 214, an illuminance sensor 215, a speaker 231, a battery 250, and a communication unit 260 on the torso 206. Note that at least some of the acceleration sensor 212, the microphone 213, the gyro sensor 214, the illuminance sensor 215, and the speaker 231 may be provided not only in the torso 206 but also in the head 204, or may be provided in both the torso 206 and the head 204.

[0014] Next, the functional configuration of the robot 200 will be described with reference to Fig. 3. As shown in Fig. 3, the robot 200 includes a control device 100, a sensor unit 210, a drive unit 220, an output unit 230, and an operation unit 240. These units are connected via a bus line BL, for example. Note that instead of the bus line BL, a wired interface such as a USB (Universal Serial Bus) cable or a wireless interface such as Bluetooth (registered trademark) may be used.

[0015] The control device 100 includes a control unit 110, which is an example of a control means, and a memory unit 120, which is an example of a memory means. The control device 100 controls the operation of the robot 200 using the control unit 110 and the memory unit 120. The control unit 110 includes a CPU (Central Processing Unit). The CPU is, for example, a microprocessor, and is a central processing unit that executes various processes and calculations. In the control unit 110, the CPU reads out a control program stored in ROM and controls the operation of the entire device (robot 200) while using RAM as a work memory. Furthermore, although not shown, the control unit 110 includes a clock function, a timer function, etc., and can measure the date and time, etc. The control unit 110 may also be called a "processor."

[0016] The storage unit 120 includes a ROM (Read Only Memory), a RAM (Random Access Memory), a flash memory, etc. The storage unit 120 stores programs and data used by the control unit 110 to perform various processes, including an OS (Operating System) and application programs. The storage unit 120 also stores data generated or acquired by the control unit 110 performing various processes. Specifically, the storage unit 120 stores an event table 121, training data 122, and a trained model 123. These will be described in detail later.

[0017] The sensor unit 210 includes the touch sensor 211, acceleration sensor 212, microphone 213, gyro sensor 214, and illuminance sensor 215 described above. The control unit 110 acquires, via the bus line BL, detection values ​​detected by the various sensors included in the sensor unit 210. Note that the sensor unit 210 may also include sensors other than these. Increasing the types of sensors included in the sensor unit 210 can increase the types of external stimuli that the control unit 110 can acquire.

[0018] The touch sensor 211 includes, for example, a pressure sensor or a capacitance sensor, and detects whether or not there is contact with some object, and the strength of the contact. Based on the detection value of the touch sensor 211, the control unit 110 can detect that the user has stroked or hit the head 204 or the body 206.

[0019] The acceleration sensor 212 detects acceleration applied to the body 206 of the robot 200. The gyro sensor 214 detects angular velocity applied to the body 206 of the robot 200. The control unit 110 can detect the current posture and changes in posture of the robot 200 using the acceleration sensor 212 and the gyro sensor 214. Furthermore, the control unit 110 can detect using the acceleration sensor 212 and the gyro sensor 214 that the user has lifted the robot 200, changed the direction of the robot 200, or thrown the robot 200.

[0020] The microphone 213 detects sounds around the robot 200. For example, the control unit 110 detects human voice, such as a user speaking to the robot 200, based on the sound components detected by the microphone 213. The control unit 110 also detects sounds other than human voice, based on the sound components detected by the microphone 213. Examples of sounds other than human voice include the sound of a user clapping their hands, environmental sounds generated around the robot 200, and sudden sounds. The touch sensor 211, acceleration sensor 212, microphone 213, and gyro sensor 214 of the sensor unit 210 are examples of external stimulus detection means that detect external stimuli.

[0021] The illuminance sensor 215 detects the illuminance around the robot 200. Based on the illuminance detected by the illuminance sensor 215, the control unit 110 can detect whether the area around the robot 200 has become brighter or darker.

[0022] The drive unit 220 includes the twist motor 221 and the up-down motor 222 described above, and is driven by the control unit 110. The robot 200 can express the action of twisting the head 204 sideways by the twist motor 221, and the action of raising and lowering the head 204 by the up-down motor 222. The output unit 230 includes a speaker 231, and when the control unit 110 inputs sound data to the output unit 230, sound is output from the speaker 231. For example, when the control unit 110 inputs data of the cry of the robot 200 to the output unit 230, the robot 200 emits a pseudo cry. Note that the output unit 230 may include a display, an LED (Light Emitting Diode), etc. instead of or in addition to the speaker 231. The operation unit 240 includes operation buttons, a volume knob, etc. The operation unit 240 is an interface for receiving user operations such as turning on / off the power supply, adjusting the volume of the output sound, etc. The battery 250 stores the power used by the robot 200. When the robot 200 returns to the charging station, the battery 250 is charged by the charging station.

[0023] Next, the functional configuration of the control unit 110 will be described. As shown in Fig. 3, the control unit 110 functionally includes an event determination unit 111 which is an example of an event determination means, an operation control unit 112 which is an example of an operation control means, a data collection unit 113 which is an example of a data collection means, a learning unit 114 which is an example of a learning means, an identification unit 115 which is an example of an identification means, and a condition change unit 116 which is an example of a condition change means. In the control unit 110, the CPU functions as each of these units by reading a program stored in the ROM into the RAM and executing and controlling the program.

[0024] The event determination unit 111 determines whether or not an event has occurred based on an external stimulus detected by the sensor unit 210. Here, the external stimulus is a stimulus acting on the robot 200 from outside the robot 200. Specifically, the external stimulus is a touch detected by the touch sensor 211, an acceleration detected by the acceleration sensor 212, a sound detected by the microphone 213, an angular velocity detected by the gyro sensor 214, or a combination thereof.

[0025] The event determination unit 111 determines whether any of a plurality of events defined in an event table 121 has occurred based on detection values ​​of the touch sensor 211, the acceleration sensor 212, the microphone 213, and the gyro sensor 214 in the sensor unit 210. The event table 121 is a table that defines a plurality of events that may occur in the robot 200 and the conditions under which each event occurs. As an example, as shown in FIG. 4, the event table 121 defines events such as "a loud noise was made," "someone spoke to," "someone stroked," "someone hit," and "someone turned upside down."

[0026] The event determination unit 111 refers to the event table 121 and determines whether the detection value of the external stimulus by the sensor unit 210 satisfies the occurrence condition of any of the events. For example, when the microphone 213 detects a sound having a peak value equal to or greater than a first threshold TH1, the event determination unit 111 determines that the event "a loud noise was heard" has occurred. When the microphone 213 detects a sound having a peak value less than the first threshold TH1 and equal to or greater than a second threshold TH2, the event determination unit 111 determines that the event "someone spoke to me" has occurred. When the touch sensor 211 of the head 204 or the torso 206 detects a contact of less than a predetermined strength, the event determination unit 111 determines that the event "someone stroked me" has occurred. When the touch sensor 211 of the head 204 or the torso 206 detects a contact of a predetermined strength or greater, the event determination unit 111 determines that the event "someone hit me" has occurred.

[0027] The occurrence condition is not limited to the detection value of a single sensor, and may be determined by combining detection values ​​of multiple sensors in the sensor unit 210. For example, "being stroked on the head while in a horizontal position" is determined by the detection values ​​of the touch sensor 211, the acceleration sensor 212, and the gyro sensor 214 of the head 204. In this way, the event determination unit 111 determines whether or not the occurrence condition of any of the events defined in the event table 121 is met, based on the external stimulus detected by the sensor unit 210, and determines that the event has occurred if the occurrence condition of any of the events is met.

[0028] The movement control unit 112 controls the movement of the robot 200. Here, the movement of the robot 200 is realized by one or both of the motion by the drive unit 220 and the output by the output unit 230. Specifically, the motion by the drive unit 220 corresponds to rotating the head 204 by driving the twist motor 221 or the up / down motor 222. Furthermore, the output by the output unit 230 corresponds to outputting a cry from the speaker 231 or illuminating an LED. The movement of the robot 200 may also be called the gesture, behavior, etc. of the robot 200.

[0029] When an external stimulus is detected by the sensor unit 210, the movement control unit 112 causes the robot 200 to move in response to the detected external stimulus. More specifically, when the event determination unit 111 determines that any event has occurred, the movement control unit 112 causes the robot 200 to perform a corresponding movement corresponding to the event that has occurred. For example, when "a loud noise is heard," the movement control unit 112 causes the robot 200 to perform a surprised movement. When "someone speaks to" the robot 200, the movement control unit 112 causes the robot 200 to perform a movement that responds to the speech. When "the robot is turned upside down," the movement control unit 112 causes the robot 200 to perform a movement that indicates an unpleasant reaction. When "the robot is petted," the movement control unit 112 causes the robot 200 to perform a happy movement. When "the robot is hit," the movement control unit 112 causes the robot 200 to perform a sad movement.

[0030] Although not shown, the correspondence between events and corresponding actions is stored in advance as an action table in the storage unit 120. The action table defines, for each event, the amount and direction of rotation by the twist motor 221, the amount and direction of rotation by the up-down motor 222, and the type and volume of the cry to be output from the speaker 231. The action control unit 112 refers to the action table and causes the robot 200 to perform a corresponding action corresponding to the event that has occurred.

[0031] Returning to FIG. 3, the data collection unit 113 collects training data 122. Here, the training data 122 is training data for machine learning to identify a speaker from a voice. As will be described in detail later, the training data 122 is data used by the learning unit 114 to perform machine learning and generate a trained model 123. As shown in FIG. 5, the training data 122 includes multiple sets of voice data. FIG. 5 shows, as an example, a case where the training data 122 includes voice data 1 to 100.

[0032] The data collection unit 113 determines whether or not the microphone 213 of the sensor unit 210 has detected any sound as an external stimulus, such as the user's voice, the sound of the user clapping, environmental sound, or a sudden sound. If any sound is detected by the microphone 213, the data collection unit 113 determines whether or not the characteristics of the detected sound satisfy predetermined collection conditions. Here, the collection conditions are conditions for collecting the training data 122, and are preset conditions that make it easier to collect sound data suitable as the training data 122 from sounds detected by the microphone 213.

[0033] Specifically, the collection condition is satisfied when the peak value of the sound detected by the microphone 213 is equal to or less than a first threshold value TH1 and equal to or greater than a second threshold value TH2, and further, at least one feature amount of the sound detected by the microphone 213 is equal to or greater than a third threshold value TH3. Here, the peak value of the sound means the maximum value of the volume.

[0034] The first threshold TH1 is a threshold for determining whether a detected sound is a "loud sound." If the peak value of the sound detected by the microphone 213 is equal to or greater than the first threshold TH1, the data collection unit 113 determines that the detected sound is a "loud sound." A "loud sound" corresponds to a sound that is louder than a human voice, such as the sound of clapping hands or other noise. The first threshold TH1 is set to a value greater than the typical volume of a human voice so that such a "loud sound" can be determined.

[0035] In contrast, the second threshold TH2 is a threshold for determining whether the detected sound is a human voice. When the peak value of the sound detected by the microphone 213 is less than the first threshold TH1 and equal to or greater than the second threshold TH2, the data collection unit 113 determines that the detected sound corresponds to a voice. The second threshold TH2 is set to a value that is smaller than the first threshold TH1 and smaller than the typical volume of a human speaking voice. On the other hand, when the peak value of the sound detected by the microphone 213 is less than the second threshold TH2, the data collection unit 113 determines that the detected sound does not correspond to a "loud sound" or a voice, but is some other sound such as an environmental sound.

[0036] The third threshold TH3 is a threshold for determining whether the detected sound is suitable for use as training data 122. Speech suitable for use as training data 122 is a sound from which the speaker of the sound can be identified with high accuracy and in which characteristics are clearly apparent. When the peak value of the sound detected by microphone 213 is less than the first threshold TH1 and greater than or equal to the second threshold TH2, data collection unit 113 further determines whether at least one feature amount of the detected sound is greater than or equal to the third threshold TH3.

[0037] Specifically, the data collection unit 113 acquires the peak value and deviation of the detected sound as feature quantities of the detected sound. The deviation of the sound is the difference from a reference value such as the average value, median value, or mode. After acquiring the peak value and deviation, the data collection unit 113 determines whether the peak value is equal to or greater than a peak threshold TH3_1 and whether the deviation is equal to or greater than a deviation threshold TH3_2. The peak threshold TH3_1 is a value between the first threshold TH1 and the second threshold TH2. The peak threshold TH3_1 and the deviation threshold TH3_2 are specific examples of the third threshold TH3. Note that the data collection unit 113 may acquire not only the peak value or deviation as feature quantities of the sound detected by the microphone 213, but also other feature quantities such as changes in frequency components over time.

[0038] If it is determined that the peak value is equal to or greater than the peak threshold TH3_1 and the deviation is equal to or greater than the deviation threshold TH3_2, the data collection unit 113 determines that the detected sound satisfies the collection condition. In this case, the data collection unit 113 determines that the sound detected by the microphone 213 is suitable as training data 122. The data collection unit 113 then stores speech data representing the detected sound in the storage unit 120 as training data 122. For example, if the training data 122 includes multiple sets of speech data 1 to 100 as shown in FIG. 5, the data collection unit 113 adds speech data representing a newly detected sound to the training data 122 as new speech data 101.

[0039] On the other hand, if it is determined that the peak value is less than the peak threshold TH3_1 or the deviation is less than the deviation threshold TH3_2, the detected sound does not satisfy the collection condition. In this case, the data collection unit 113 determines that the sound detected by the microphone 213 is not suitable as learning data 122. The data collection unit 113 does not store audio data indicating the detected sound as learning data 122 in the storage unit 120. In this way, the data collection unit 113 stores audio data indicating audio that satisfies the predetermined collection condition, from among the audio detected by the microphone 213, in the storage unit 120 as learning data 122.

[0040] Returning to Fig. 3, the learning unit 114 performs machine learning based on the learning data 122 collected by the data collection unit 113 to generate a trained model 123. The trained model 123 is a model for identifying the speaker of a voice from the voice. More specifically, the trained model 123 is a model for accepting input of voice data and identifying, from the characteristics of the voice data, whether the speaker of the voice is a specific user corresponding to the owner or another user.

[0041] The learning unit 114 performs machine learning using a clustering technique, which is one type of unsupervised learning. Here, clustering is a technique for classifying a data set into multiple groups based on specific rules. The learning unit 114 classifies multiple sets of speech data included in the training data 122 into multiple groups using a known clustering technique, such as the k-means method or Ward's method.

[0042] More specifically, the learning unit 114 extracts a plurality of parameters indicating speech characteristics from each of the plurality of sets of speech data included in the training data 122. Then, the learning unit 114 maps the plurality of sets of speech data using the plurality of parameters, as shown in Fig. 5. In Fig. 5, one data point corresponds to one piece of speech data. The learning unit 114 classifies the plurality of sets of speech data into a plurality of groups (groups A to D in the example of Fig. 5) by classifying speech data with high similarity into the same group based on the similarity of the plurality of parameters.

[0043] The parameters extracted from the speech data may be parameters used in known speech recognition techniques. As an example, the learning unit 114 performs a fast Fourier transform on each of the sets of speech data included in the training data 122. The learning unit 114 may use the Fourier coefficients of the multiple frequency components obtained in this process as parameters. Note that, in FIG. 5, two parameters, parameter 1 and parameter 2, are used for ease of understanding, but this is just an example, and it is preferable to use more parameters to improve the recognition accuracy.

[0044] The learning unit 114 performs machine learning using such a clustering technique to classify multiple sets of speech data included in the training data 122 into multiple groups. As a result, the learning unit 114 generates a trained model 123 that outputs, in response to a speech input, information about the group to which the speech belongs. Note that the learning unit 114 performs machine learning using such a clustering technique when the number of pieces of data in the training data 122 collected by the data collecting unit 113 is equal to or greater than a reference value, thereby generating the trained model 123. Here, the number of pieces of data in the training data 122 is the number of pieces of speech data included in the training data 122. If the number of pieces of data in the training data 122 is small, effective machine learning cannot be performed, making it difficult to generate a trained model 123 that can accurately identify the speaker. Therefore, the learning unit 114 does not perform machine learning until the number of pieces of data in the training data 122 is equal to or greater than a predetermined reference value, and performs machine learning after a sufficient number of pieces of speech data has been collected by the data collecting unit 113.

[0045] Returning to FIG. 3 , when sound is detected as an external stimulus by the sensor unit 210, the identification unit 115 identifies the speaker of the detected sound using the trained model 123 generated by the learning unit 114. Specifically, when sound is detected by the microphone 213 and the detected sound satisfies the collection conditions, the identification unit 115 inputs sound data indicating the detected sound to the trained model 123. As an output for the input sound data, the trained model 123 outputs information indicating the group to which the sound of the input sound data is most likely to belong, out of multiple groups classified by clustering. The identification unit 115 identifies the speaker corresponding to the group indicated in the information output from the trained model 123 as the speaker of the detected sound.

[0046] More specifically, the identification unit 115 identifies whether the speaker of the detected voice is a specific user corresponding to the pseudo-owner of the robot 200. The pseudo-owner of the robot 200 (hereinafter simply referred to as "owner") is a user corresponding to the owner of a pet, such as the owner or manager of the robot 200. Since the owner is often present near the robot 200, it is assumed that the owner has more opportunities to talk to the robot 200 than other users. Therefore, the identification unit 115 determines that the group with the largest number of data items among the multiple groups classified by clustering is the group of voice data of the owner. Specifically, in the example of FIG. 5, the identification unit 115 determines that group A among groups A to D is the group of voice data of the pseudo-owner.

[0047] In this way, when the group indicated in the information output from trained model 123 is the group with the largest number of data, identification unit 115 identifies whether the speaker of the detected voice is a specific user corresponding to the owner. On the other hand, when the group indicated in the information output from trained model 123 is a group other than the group with the largest number of data, identification unit 115 identifies whether the speaker of the detected voice is a user other than the owner.

[0048] When the identification unit 115 identifies the speaker as a specific user, the operation control unit 112 causes the robot 200 to perform an operation different from that performed when the identification unit 115 identifies the speaker as a user other than the specific user. In other words, the operation control unit 112 causes the robot 200 to perform different operations as a response operation corresponding to "spoken to" in the event table 121 shown in Fig. 4 when the speaker is a specific user corresponding to the owner and when the speaker is not a specific user.

[0049] Specifically, when the identification unit 115 identifies the speaker as not a specific user, this corresponds to the robot 200 being spoken to by a user other than the owner. In this case, the movement control unit 112 causes the robot 200 to perform a movement that responds to the speech. On the other hand, when the identification unit 115 identifies the speaker as a specific user, this corresponds to the robot 200 being spoken to by the owner. In this case, the movement control unit 112 causes the robot 200 to perform a movement that responds with greater joy than when spoken to by a user other than the owner. For example, the movement control unit 112 rotates the twist motor 221 or the up / down motor 222 more vigorously or outputs a louder cry from the speaker 231 than when spoken to by a user other than the owner. In this way, by changing the movement in response to being spoken to by the owner and by others, it is possible to highly mimic the behavior of an actual pet and enhance the lifelikeness of the pet.

[0050] 3, when the sensor unit 210 detects an external stimulus that satisfies a specific condition as an external stimulus, the condition change unit 116 changes the collection conditions so that a newly detected voice by the microphone 213 is more likely to satisfy the collection conditions until a predetermined time has elapsed since the external stimulus that satisfies the specific condition is detected. Here, the specific condition is a condition that is set in advance to be satisfied when there is a high possibility that the owner is present around the robot 200, for the purpose of enabling the data collection unit 113 to efficiently collect voice data of the owner.

[0051] Specifically, the specific condition is satisfied when an event that a pet is likely to enjoy, such as "being spoken to" or "being petted," occurs among the multiple events defined in the event table 121. On the other hand, the specific condition is not satisfied when an event that a pet is likely to find unpleasant, such as "a loud noise is made" or "being hit," occurs. Therefore, the specific condition is satisfied when the microphone 213 detects a sound having a peak value less than the first threshold TH1 and equal to or greater than the second threshold TH2, which is the occurrence condition for "being spoken to." In other words, a sound having a peak value within a predetermined range (the range from the first threshold TH1 to the second threshold TH2) satisfies the specific condition. Furthermore, the specific condition is satisfied when the touch sensor 211 of the head 204 or the body 206 detects a contact of less than a predetermined strength, which is the occurrence condition for "being petted." In other words, a contact of less than a predetermined strength with respect to the robot 200 satisfies the specific condition. On the other hand, the specific condition is not satisfied when the occurrence conditions for "a loud noise is made" or "being hit" are met.

[0052] In addition to the occurrence conditions of "being spoken to" or "being stroked" as described above, the specific condition may be that the detection values ​​detected by the acceleration sensor 212 and the gyro sensor 214 are equal to or less than a predetermined value. This makes it possible to exclude cases where the device is spoken to or stroked while being turned upside down or strongly shaken from the specific condition.

[0053] When the sensor unit 210 detects an external stimulus that satisfies such a specific condition, it is highly likely that a user corresponding to the owner of the robot 200 is present near the robot 200. Therefore, the condition change unit 116 relaxes the collection conditions so that the characteristics of the voice detected by the microphone 213 are more likely to satisfy the collection conditions until a predetermined time has elapsed since the external stimulus that satisfies the specific condition was detected. The predetermined time is a predetermined length of time, such as 3 minutes, 5 minutes, etc. Relaxing the collection conditions during such a time makes it easier to collect voice data of the user corresponding to the owner as the learning data 122.

[0054] Specifically, the condition change unit 116 lowers the third threshold TH3 from the detection of an external stimulus satisfying a specific condition until a predetermined time has elapsed since the detection of the external stimulus satisfying the specific condition, compared to when an external stimulus satisfying the specific condition is not detected (hereinafter referred to as a "normal case"). In other words, the condition change unit 116 changes each of the peak threshold TH3_1 and the deviation threshold TH3_2, which are the third threshold TH3, to values ​​smaller than those in the normal case. As a result, even voice data with a smaller peak value and a smaller deviation than those in the normal case will satisfy the collection condition and will be collected by the data collection unit 113 as the learning data 122. In other words, voice data that does not satisfy the collection condition in the normal case will satisfy the collection condition until a predetermined time has elapsed since the detection of an external stimulus satisfying the specific condition. This makes it easier to collect voice data of the user corresponding to the owner as the learning data 122.

[0055] In this way, when an external stimulus that pleases the robot 200, such as "being spoken to" or "being petted," is detected, the condition change unit 116 relaxes the collection condition, which makes it easier for the identification unit 115 to identify the user who loves the robot 200 as the owner. In other words, the robot 200 can mimic the behavior of an actual pet, which is to recognize a person who loves it as the owner.

[0056] Next, the flow of the robot control process according to this embodiment will be described with reference to Fig. 7. The robot control process shown in Fig. 7 is executed by the control unit 110 of the control device 100 when the robot 200 is powered on. The robot control process shown in Fig. 7 is an example of a robot control method. When the robot control process starts, the control unit 110 executes an initialization process (step S1). In the initialization process, the control unit 110 sets various parameters used for controlling the robot 200 to initial values. After executing the initialization process, the control unit 110 determines whether any external stimulus has been detected by any of the touch sensor 211, acceleration sensor 212, microphone 213, and gyro sensor 214 in the sensor unit 210 (step S2).

[0057] If an external stimulus is detected (step S2; YES), the control unit 110 determines whether or not sound is detected by the microphone 213 as the external stimulus (step S3). If sound is detected (step S3; YES), the control unit 110 executes a data collection process (step S4). Details of the data collection process in step S4 will be described with reference to FIG. 8.

[0058] 8 starts, the control unit 110 determines whether the peak value of the sound detected in step S2 is equal to or greater than the first threshold value TH1 (step S41). If the peak value is equal to or greater than the first threshold value TH1 (step S41; YES), the control unit 110 determines that the detected sound corresponds to a "loud sound" (step S42), and ends the data collection process shown in FIG.

[0059] On the other hand, if the peak value is less than the first threshold value TH1 (step S41; NO), the control unit 110 next determines whether the peak value of the sound detected in step S2 is equal to or greater than the second threshold value TH2 (step S43). If the peak value is less than the second threshold value TH2 (step S43; NO), the control unit 110 determines that the detected sound is noise such as environmental sound, and ends the data collection process shown in FIG.

[0060] On the other hand, if the peak value is equal to or greater than the second threshold TH2 (step S43; YES), the control unit 110 determines that the detected sound corresponds to speech. In this case, the control unit 110 determines whether the peak value and deviation of the detected sound are equal to or greater than the third threshold TH3 (step S44). Specifically, the control unit 110 determines whether the peak value of the detected sound is equal to or greater than the peak threshold TH3_1 and whether the deviation of the detected sound is equal to or greater than the deviation threshold TH3_2.

[0061] If the peak value and the deviation are equal to or greater than the third threshold value TH3 (step S44; YES), the control unit 110 determines that the detected sound is sound suitable for the learning data 122. In this case, the control unit 110 functions as the data collection unit 113, and saves the sound data indicating the detected sound as the learning data 122 (step S45). In other words, the control unit 110 adds the sound data indicating the detected sound to the learning data 122 and stores it as new sound data.

[0062] When the voice data is saved as the training data 122, the control unit 110 determines whether the number of data items in the training data 122 is equal to or greater than a reference value (step S46). If the number of data items in the training data 122 is equal to or greater than the reference value (step S46; YES), the control unit 110 functions as the learning unit 114, performs machine learning using the training data 122, and generates a trained model 123 (step S47). At this time, if the trained model 123 already exists in the storage unit 120, the control unit 110 performs machine learning using the training data 122 to which the new voice data has been added, and updates the trained model 123. Note that if the trained model 123 already exists in the storage unit 120, the control unit 110 does not need to perform the process of generating the trained model 123 in step S47 every time new voice data is saved as the training data 122. For example, the control unit 110 may perform machine learning using the latest training data 122 when the number of data items in the training data 122 has increased to a certain extent, and update the trained model 123.

[0063] After performing machine learning, the control unit 110 functions as the identification unit 115 and uses the trained model 123 to identify the speaker of the voice detected in step S2 and determine whether the speaker is a specific user corresponding to the owner (step S48). If the speaker corresponds to the owner (step S48; YES), the control unit 110 determines that the sound detected in step S2 corresponds to "a conversation from the owner" (step S49). On the other hand, if the speaker does not correspond to the owner (step S48; NO), the control unit 110 determines that the sound detected in step S2 corresponds to simple "a conversation," that is, "a conversation" from a user other than the owner (step S50).

[0064] If the number of data items in the learning data 122 is less than the reference value in step S46 (step S46; NO), the control unit 110 determines that effective machine learning cannot be performed. In this case, the control unit 110 proceeds to step S50 without performing machine learning, and determines that the sound detected in step S2 corresponds to "talk" from a user other than the owner. If the peak value and deviation are less than the third threshold value TH3 in step S44 (step S44; NO), the control unit 110 determines that the sound detected in step S2 is not a sound suitable for the learning data 122. In this case, the control unit 110 proceeds to step S50 without updating the learning data 122, and determines that the sound detected in step S2 corresponds to "talk" from a user other than the owner. This completes the data collection process shown in FIG. 8.

[0065] 7, if sound is not detected as the external stimulus (step S3; NO), the control unit 110 skips step S4. Next, the control unit 110 determines whether the external stimulus detected in step S2 satisfies a specific condition (step S5). If the detected external stimulus satisfies the specific condition (step S5; YES), the control unit 110 functions as the condition change unit 116 and relaxes the collection condition for a predetermined time (step S6). Specifically, the control unit 110 lowers the third threshold TH3 compared to when the external stimulus does not satisfy the specific condition. This makes it more likely that the peak value and deviation of a sound newly detected by the microphone 213 are determined to be equal to or greater than the third threshold TH3 in the determination process in step S44 until the predetermined time has elapsed from that point. On the other hand, if the detected external stimulus does not satisfy the specific condition (step S5; NO), the control unit 110 skips step S6.

[0066] Next, the control unit 110 functions as the event determination unit 111 and determines whether or not an event based on the external stimulus has occurred (step S7). Specifically, the control unit 110 determines whether or not the external stimulus detected in step S2 has established a condition for the occurrence of any event defined in the event table 121. If an event has occurred (step S7; YES), the control unit 110 functions as the action control unit 112 and causes the robot 200 to perform an action corresponding to the event that has occurred (step S8).

[0067] More specifically, when the control unit 110 detects an external stimulus other than sound in step S2, it causes the robot 200 to perform a corresponding action corresponding to an event based on the external stimulus other than sound. For example, when the robot is "petted," the control unit 110 causes the robot 200 to perform a happy action. Alternatively, when the robot is "hit," the control unit 110 causes the robot 200 to perform a sad action. On the other hand, when the control unit 110 detects sound in step S2, it causes the robot 200 to perform an action corresponding to the determination result in the data acquisition process in step S4. For example, when it is determined in step S42 that the detected sound corresponds to a "loud sound," the control unit 110 causes the robot 200 to perform a startled action. Alternatively, when it is determined in step S50 that the detected sound corresponds to "talking" from someone other than the owner, the control unit 110 causes the robot 200 to perform an action responding to the talk. Furthermore, if it is determined in step S49 that the detected sound corresponds to "a call from the owner", the control unit 110 causes the robot 200 to perform an action of reacting with great joy to the call.

[0068] Furthermore, in step S2, if an external stimulus is not detected (step S2; NO) or if an event has not occurred (step S7; NO), the control unit 110 functions as the movement control unit 112 and causes the robot 200 to perform a spontaneous movement (step S9). A spontaneous movement is a movement that the robot 200 performs spontaneously without relying on an external stimulus, such as a breathing movement that simulates breathing or a movement that moves the body randomly. If no external stimulus is detected by any of the multiple sensors in the sensor unit 210, the control unit 110 causes the robot 200 to perform a spontaneous movement, for example, once every few seconds.

[0069] Thereafter, the control unit 110 returns the process to step S2. Then, the control unit 110 repeats the processes of steps S2 to S9 as long as the robot 200 is powered on and can operate normally. As a result, when a voice satisfying the collection condition is detected, the control unit 110 repeats the process of collecting voice data indicating the detected voice as the learning data 122.

[0070] As described above, when a voice satisfying a predetermined collection condition is detected, the robot 200 according to the first embodiment stores voice data indicating the detected voice as the learning data 122, and changes the collection condition so that the voice detected by the sensor unit 210 is more likely to satisfy the collection condition until a predetermined time has elapsed since the detection of an external stimulus satisfying the specific condition. In this way, the robot 200 according to the first embodiment can more easily collect voice data when there is a high possibility that a user corresponding to the owner is present near the robot 200, and can therefore efficiently collect voice data of a specific user corresponding to the owner as learning data.

[0071] Next, a second embodiment will be described. Descriptions of configurations and functions similar to those of the first embodiment will be omitted where appropriate. In the first embodiment, the robot 200 has the functions of the data collection unit 113, the learning unit 114, and the condition change unit 116. In contrast, in the second embodiment, a learning data collection device 300, which is an external device to the robot 200 and independent of the robot 200, has these functions. The learning data collection device 300 is an information processing device such as a general-purpose personal computer, a tablet terminal, a smartphone, or the like.

[0072] 9 , the training data collection device 300 according to the second embodiment includes a sensor unit 210, a communication unit 260, and a control device 100. The sensor unit 210 in the training data collection device 300 includes a touch sensor 211, an acceleration sensor 212, a microphone 213, and a gyro sensor 214, similar to the sensor unit 210 in the robot 200. The communication unit 260 includes a communication interface for communicating with devices external to the training data collection device 300. For example, the communication unit 260 communicates with external devices including the robot 200 in accordance with well-known communication standards such as wireless local area network (LAN), Bluetooth Low Energy (BLE), and near field communication (NFC).

[0073] The training data collection device 300 includes a control unit 110 of the control device 100, and functionally includes a data collection unit 113, a learning unit 114, and a condition change unit 116. These units are similar to the units included in the robot 200 in the first embodiment. Specifically, the data collection unit 113 stores, in the storage unit 120, speech data indicating speech that satisfies predetermined collection conditions among speech detected by the sensor unit 210 as training data 122. The condition change unit 116 changes the collection conditions so that speech newly detected by the microphone 213 is more likely to satisfy the collection conditions until a predetermined time has elapsed since an external stimulus that satisfies a specific condition was detected. The learning unit 114 performs machine learning based on the training data 122 collected by the data collection unit 113 to generate a trained model 123. After generating the trained model 123, the learning unit 114 communicates with the robot 200 via the communication unit 260 and transmits the generated trained model 123 to the robot 200.

[0074] Although not shown, the robot 200 according to the second embodiment functionally includes an event determination unit 111, an operation control unit 112, and an identification unit 115 in the control unit 110 of the control device 100, but does not include the data collection unit 113, the learning unit 114, or the condition change unit 116. The event determination unit 111 and the operation control unit 112 are the same as those in the first embodiment. The robot 200 communicates with the training data collection device 300 via a communication unit not shown, and acquires a trained model 123 from the training data collection device 300. The identification unit 115 then uses the trained model 123 acquired from the training data collection device 300 to identify the speaker of the voice detected by the sensor unit 210.

[0075] As described above, in the second embodiment, the robot 200 does not include the data collection unit 113, the learning unit 114, and the condition change unit 116, and the training data collection device 300, which is a device different from the robot 200, collects the training data 122 and generates the trained model 123 based on the collected training data 122. This simplifies the configuration of the robot 200 compared to the first embodiment. Furthermore, by transmitting the trained model 123 generated by the training data collection device 300 to multiple robots 200, the multiple robots 200 can accurately identify the speaker from the voice.

[0076] Although the embodiments of the present invention have been described above, the above embodiments are merely examples, and the scope of application of the present invention is not limited thereto. In other words, the embodiments of the present invention are applicable to various applications, and all embodiments are included in the scope of the present invention. For example, in the above embodiment, the learning unit 114 performs machine learning using a clustering technique to generate the trained model 123. However, the learning unit 114 may perform machine learning using other techniques, such as principal component analysis, instead of clustering.

[0077] In the above embodiment, the collection conditions for the data collection unit 113 to collect the learning data 122 are determined by three thresholds TH1 to TH3. However, the collection conditions are not limited to this and may be determined by any conditions. For example, the third threshold TH3 is not limited to the peak threshold TH3_1 and the deviation threshold TH3_2, but may be determined by thresholds for other features of the audio. Furthermore, if at least one feature is equal to or greater than the third threshold TH3, it may also be the case that at least one feature is a value within a predetermined range. In this case, the condition change unit 116 may change the collection conditions so that the features of the audio are more likely to satisfy the collection conditions by widening this predetermined range until a predetermined time has elapsed since an external stimulus satisfying a specific condition was detected.

[0078] In the above embodiment, the exterior 201 is formed in a cylindrical shape from the head 204 to the torso 206, and the robot 200 is in a prone position. However, the robot 200 is not limited to being modeled after a prone position creature. For example, the robot 200 may be modeled after a creature with arms and legs, and may be modeled after a creature that walks on four legs or two legs.

[0079] In the above embodiment, the control device 100 is built into the robot 200. However, the control device 100 may be a separate device (e.g., a server) rather than built into the robot 200. When the control device 100 is located outside the robot 200, the robot 200 communicates with the control device 100 via a communication unit to transmit and receive data to and from the control device 100. The control device 100 controls the robot 200 through such communication with the robot 200. Furthermore, in the second embodiment, the learning data collection device 300 includes the data collection unit 113, the learning unit 114, and the condition change unit 116. However, the learning unit 114 may be provided in a device separate from the learning data collection device 300. In other words, the data collection unit 113 and the learning unit 114 do not necessarily have to be provided in the same device, but may be provided in different devices.

[0080] In the above embodiment, the CPU in the control unit 110 executes a program stored in the ROM to function as the event determination unit 111, operation control unit 112, data collection unit 113, learning unit 114, identification unit 115, and condition change unit 116. However, in the present invention, the control unit 110 may include dedicated hardware, such as an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or various control circuits, instead of a CPU, and the dedicated hardware may function as each of these units. In this case, the functions of each unit may be realized by individual hardware, or the functions of each unit may be realized together by a single piece of hardware. Alternatively, some of the functions of each unit may be realized by dedicated hardware, and other parts may be realized by software or firmware.

[0081] It should be noted that the robot or learning data collection device can be provided with a configuration for realizing the functions according to the present invention, and by applying a program, an existing information processing device or the like can be made to function as the robot or learning data collection device according to the present invention. That is, by applying a program for realizing each functional configuration of the robot 200 or learning data collection device 300 exemplified in the above embodiment to a CPU or the like that controls the existing information processing device or the like so that it can be executed, the existing information processing device or the like can be made to function as the robot or learning data collection device according to the present invention.

[0082] Furthermore, the application method of such a program is arbitrary. The program can be applied by storing it on a computer-readable storage medium such as a flexible disk, a CD (Compact Disc)-ROM, a DVD (Digital Versatile Disc)-ROM, or a memory card. Furthermore, the program can be superimposed on a carrier wave and applied via a communication medium such as the Internet. For example, the program can be distributed by posting it on a bulletin board system (BBS) on a communication network. Then, the program can be started and executed under the control of an operating system (OS) in the same way as other application programs, thereby enabling the above-mentioned processing to be performed.

[0083] The above describes preferred embodiments of the present invention, but the present invention is not limited to the above-described embodiments, and various modifications and substitutions can be made to the above-described embodiments without departing from the scope of the claims. [Explanation of symbols]

[0084] 113...data collection unit, 116...condition change unit, 120...storage unit, 122...learning data, 200...robot, 210...sensor unit

Claims

1. A robot capable of collecting learning data for identifying a speaker from a voice, an external stimulus detection means for detecting an external stimulus; a data collection means for storing, in a storage means, voice data representing the detected voice as the learning data when the external stimulus detection means detects a voice satisfying a predetermined collection condition as the external stimulus; a condition changing means for changing the collection conditions when an external stimulus satisfying a specific condition is detected by the external stimulus detecting means so that the sound detected by the external stimulus detecting means is more likely to satisfy the collection conditions until a predetermined time has elapsed since the external stimulus satisfying the specific condition was detected; A robot comprising:

2. the collection condition includes that at least one feature amount of the sound detected by the external stimulus detection means is equal to or greater than a threshold value; the condition changing means lowers the threshold value during a period from when the external stimulus detecting means detects an external stimulus that satisfies the specific condition until the predetermined time has elapsed.

2. The robot according to claim 1 .

3. A sound whose peak value is within a predetermined range satisfies the specific condition.

3. The robot according to claim 1 or 2.

4. A contact with the robot having an intensity less than a predetermined intensity satisfies the specific condition.

3. The robot according to claim 1 or 2.

5. and an identification means for identifying the speaker from the detected voice using a trained model generated by machine learning when the external stimulus detection means detects a voice that satisfies the predetermined collection condition as the external stimulus.

3. The robot according to claim 1 or 2.

6. Further, the robot is provided with a motion control means for controlling the motion of the robot; When the identification means identifies the speaker as a specific user, the action control means causes the robot to perform an action different from that when the identification means identifies the speaker as a user other than the specific user.

6. The robot according to claim 5.

7. and a learning means for generating the trained model by performing the machine learning based on the training data when the number of data items in the training data collected by the data collection means becomes equal to or greater than a reference value.

6. The robot according to claim 5.

8. the learning means performs the machine learning using a clustering technique.

8. The robot according to claim 7.

9. A training data collection device that collects training data for identifying a speaker from a voice, comprising: an external stimulus detection means for detecting an external stimulus; a data collection means for storing, in a storage means, voice data representing the detected voice as the learning data when the external stimulus detection means detects a voice satisfying a predetermined collection condition as the external stimulus; a condition changing means for changing the collection conditions when an external stimulus satisfying a specific condition is detected by the external stimulus detecting means so that the sound detected by the external stimulus detecting means is more likely to satisfy the collection conditions until a predetermined time has elapsed since the external stimulus satisfying the specific condition was detected; A learning data collection device comprising:

10. A training data collection method for collecting training data for identifying a speaker from a voice, comprising: Detects external stimuli, When a sound satisfying a predetermined collection condition is detected as the external stimulus, sound data indicating the detected sound is stored in a storage means as the learning data; when an external stimulus satisfying a specific condition is detected, the collection condition is changed so that the detected sound is more likely to satisfy the collection condition until a predetermined time has elapsed since the external stimulus satisfying the specific condition was detected; A learning data collection method characterized by:

11. A computer capable of collecting learning data for identifying a speaker from a voice, an external stimulus detection means for detecting an external stimulus; a data collection means for storing, in a storage means, voice data representing the detected voice as the learning data, when the external stimulus detection means detects a voice satisfying a predetermined collection condition as the external stimulus; a condition changing means for changing the collection conditions when an external stimulus satisfying a specific condition is detected by the external stimulus detection means so that the sound detected by the external stimulus detection means is more likely to satisfy the collection conditions until a predetermined time has elapsed since the external stimulus satisfying the specific condition was detected; A program to function as a

Citation Information

Patent Citations

  • Learning system and learning method, and robot apparatus

    JP2003255989A

  • Identification device, robot, identification method, and program

    JP2020057300A

  • Device control device, device, device control method and program

    JP2022113701A

  • Robot, robot control method, and program

    JP2023049116A

  • Information processing device, information processing method, and robot device

    WO2019181144A1