Instruction device, robot system, and robot

The instruction device and robot system address the challenge of accommodating multiple user preferences by integrating feature quantities and instructing a robot to prompt photo-taking based on similarity, enabling effective shooting proposals that align with the preferences of multiple users.

WO2025109926A1PCT designated stage expired Publication Date: 2025-05-30PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/037325
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-10-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing imaging technologies cannot acquire photographs according to the preferences of multiple users simultaneously, limiting their ability to accommodate the preferences of two or more users.

Method used

An instruction device and robot system that integrate feature quantities from image data of multiple users, determine the similarity of these integrated features with external environment data, and instruct a robot to prompt photo-taking based on this similarity.

Benefits of technology

Enables the robot system to make shooting proposals that align with the preferences of multiple users, effectively prompting photo-taking operations that reflect the integrated preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024037325_30052025_PF_FP_ABST
    Figure JP2024037325_30052025_PF_FP_ABST
Patent Text Reader

Abstract

An instruction device comprising: a feature value integration unit that, on the basis of a first feature value of image data held by a first user and a second feature value of image data held by a second user, generates a third feature value in which the first feature value and the second feature value are integrated; a determination unit that determines whether or not to execute an output to the second user on the basis of the similarity between the third feature value and a fourth feature value extracted from the image data of an external environment; and an instruction unit that instructs execution of the output when it is determined that the output to the second user is to be executed.
Need to check novelty before this filing date? Find Prior Art

Description

Pointing device, robot system and robot

[0001] The present disclosure relates to a pointing device, a robot system, and a robot.

[0002] In recent years, techniques for imaging devices have been proposed that enable users to capture images of their choice without requiring special operations by the user. Patent Literature 1 proposes an imaging device that, based on data related to the captured images, assigns a higher weight to data related to captured images that are instructed by the user than to data related to automatically processed captured images.

[0003] JP 2019-106694 A

[0004] The technology described in Patent Document 1 only acquires photos that meet the preferences of a single user, and is unable to acquire photos that meet the preferences of more than one user when there are two or more users.

[0005] The present disclosure aims to provide an instruction device, a robot system, and a robot that can make photography suggestions according to the preferences of two or more users.

[0006] One aspect of the instruction device according to the present disclosure includes a feature integration unit that generates a third feature by integrating a first feature of image data held by a first user and a second feature of image data held by a second user, based on the first feature and the second feature; a determination unit that determines whether to execute output to the second user based on a similarity between the third feature and a fourth feature extracted from image data of an external environment; and an instruction unit that issues an instruction to execute the output when it is determined that output to the second user should be executed.

[0007] Furthermore, one aspect of a robot system according to the present disclosure includes the above-described instruction device and a robot that operates in response to instructions from the instruction device.

[0008] Furthermore, one aspect of a robot according to the present disclosure includes the above-described instruction device.

[0009] According to the present disclosure, it is possible to make photography suggestions according to the preferences of two or more users.

[0010] 1 is a diagram illustrating an overview of a robot system according to the present embodiment; FIG. 2 is a block diagram illustrating a functional configuration of a robot system according to the present embodiment; FIG. 3 is a flowchart illustrating an operation of the robot system; FIG. 4 is a diagram illustrating an example of creating a color histogram; FIG. 5 is a diagram illustrating an example of integrating feature amounts; FIG. 6 is a diagram illustrating calculation of similarity of feature amounts; and FIG. 7 is a diagram illustrating an example of operation according to the magnitude of similarity.

[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that each of the embodiments described below represents a specific example of the present disclosure. Therefore, the numerical values, shapes, materials, components, the arrangement and connection of each component, and each step and the order of each step shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components not recited in independent claims will be described as optional components.

[0012] Furthermore, each drawing is a schematic diagram and is not necessarily a precise illustration. In each drawing, substantially the same components are denoted by the same reference numerals, and redundant explanations will be omitted or simplified.

[0013] [Robot Overview] The configuration of a robot system according to this embodiment will be described below. Fig. 1 is a diagram showing an overview of a robot system according to this embodiment. Fig. 2 is a block diagram showing the functional configuration of a robot system 10 according to this embodiment.

[0014] The robot system 10 has a configuration in which a server device 20 and a robot 30 are connected via a network N. Note that the configuration of the robot system 10 is not limited to the robot 30, and may be a terminal or the like. The terminal may be any device that has an imaging unit and an output device, and may be, for example, a smartphone or a tablet.

[0015] As shown in FIG. 1, a server device 20 extracts user preferences from photo data 3 and 4 owned by two or more users 1 and 2, and determines whether photo data 5 obtained from a robot 30 matches the user preferences.

[0016] In addition, the photo data 3 and 4 owned by users 1 and 2 and the photo data 5 obtained from robot 30 may be image (video) data, and are not limited to still image data, but may also be video data obtained by filming a moving image.

[0017] If the photo data 5 acquired from the robot 30 matches the user's preferences, the robot 30 encourages the user 2 who owns the robot 30 to take the corresponding photo. By using the robot system 10 in this embodiment, it becomes possible to make photography suggestions that match the preferences of two or more users.

[0018] [Configuration of Robot System] The following describes the components of the robot system 10. The server device 20 is an instruction device that instructs the robot 30 to perform an action to capture an image of the external environment. The server device 20 includes a storage unit 21, a feature extraction unit 22, a feature integration unit 23, a similarity calculation unit 24, a determination unit 25, an instruction unit 26, and an image receiving unit 27.

[0019] The storage unit 21 stores photo data 3 and 4 of two or more users 1 and 2. The photo data 3 and 4 are managed in association with user IDs so that the owners can be identified. The storage unit 21 is realized by, for example, a semiconductor memory.

[0020] The feature extraction unit 22 extracts a first feature from the photo data 3 and a second feature from the photo data 4. The extracted feature may be, for example, the appearance frequency, color, texture, shape, or composition of the subject. The extracted feature is managed in association with the user ID. The extracted feature may also be stored in the storage unit 21.

[0021] The feature amount integration unit 23 generates a third feature amount by weighting and integrating the first feature amount and the second feature amount extracted from the photo data 3, 4 of two or more users 1, 2 extracted by the feature amount extraction unit 22. Then, the feature amount integration unit 23 stores the third feature amount in the storage unit 21.

[0022] In addition, if the memory unit 21 stores features extracted from the photo data 3 and 4 of two or more users 1 and 2, the feature integration unit 23 may read out the features extracted from the photo data of the two or more users from the memory unit 21 and integrate them.

[0023] The similarity calculation unit 24 calculates the similarity between the third feature amount and the feature amount of the photograph taken by the robot 30. The photograph taken by the robot 30 is sent from the robot 30 to the server 20 via the network N.

[0024] The determination unit 25 determines whether or not to execute an action that prompts the robot 30 to photograph the external environment, based on the similarity calculated by the similarity calculation unit 24. Specifically, the determination unit 25 determines whether or not the similarity calculated by the similarity calculation unit 24 satisfies the conditions for the robot 30 to execute an action.

[0025] When the determination unit 25 determines that an action to prompt the robot 30 to photograph the external environment should be performed, the instruction unit 26 instructs the robot 30 via the network N to perform an action to prompt the robot 30 to photograph the external environment.

[0026] The image receiving unit 27 receives the photo data 5 acquired by the camera 31 of the robot 30 from the image transmitting unit 34 of the robot 30 via the network N.

[0027] On the other hand, the robot 30 includes a camera 31, an instruction receiving unit 32, an operating unit 33, and an image transmitting unit 34. The camera 31 is an imaging device that takes pictures of the external environment of the robot 30. Note that the camera 31 is not limited to taking still images, but may also take moving images.

[0028] The instruction receiving unit 32 receives operation instructions for the robot 30 from the instruction unit 26 of the server device 20 via the network N. The operation instructions are information instructing the robot 30 to perform an operation such as changing the eye color of the robot 30, changing a facial expression, moving its body, emitting a sound, or sending a notification to a mobile device owned by the user 2.

[0029] When the instruction receiving unit 32 receives an instruction to operate the robot 30 from the server device 20, the operation unit 33 performs an operation to prompt the user 2 to take an image of the external environment of the robot 30.

[0030] The image transmitting unit 34 transmits the photo data 5 acquired by the camera 31 of the robot 30 to the instruction image receiving unit 27 of the server device 20 via the network N.

[0031] [Operation of Robot System] Next, the operation of the robot system 10 will be described. Fig. 3 is a flowchart showing the operation of the robot system 10. Note that the order of processing shown in the flowchart of Fig. 3 is an example. The order of processing may be changed, or multiple processing may be performed in parallel.

[0032] First, the storage unit 21 of the server device 20 receives the input of the photo data 3 and 4 uploaded to the server device 20 by the users 1 and 2, and stores the photo data 3 and 4 (step S1).

[0033] The photo data 3 and 4 may be, for example, photo data taken by the user with a smartphone camera or digital camera, or may be photo data captured from the screen of a PC, smartphone, or tablet. The photo data 3 and 4 may also be photo data downloaded from a social networking site or website, or photo data received from another person.

[0034] Next, the feature extraction unit 22 acquires the photo data 3 and 4 stored in the storage unit 21, extracts a first feature from the photo data 3, and extracts a second feature from the photo data 4 (step S2). Here, the extraction of features will be described in terms of a case where the feature extraction unit 22 extracts color information from the photo data of one user and creates a color histogram. Figure 4 shows an example of creating a color histogram.

[0035] The color histogram is created by examining the color information of each pixel, counting the number of each color, and expressing it as a histogram. When creating a color histogram using multiple (N) images 41, the feature extraction unit 22 adds up the color histograms created from each image 41 to create one color histogram.

[0036] Furthermore, in order to remove the influence of image size and the number of images 41 from the created color histogram, feature extraction unit 22 normalizes the color histogram so that its area is 1. Note that when extracting the appearance frequency, texture, shape, and composition of a subject from photograph data 3 and 4, feature extraction unit 22 creates a histogram in the same manner as when creating a color histogram.

[0037] Next, the feature integration unit 23 weights the feature extracted in step S2 for each user (step S3).Then, the feature integration unit 23 integrates the first feature and the second feature extracted from the photo data 3 and 4 of the two users, User 1 and User 2, based on the weight set for the feature of each user (step S4).

[0038] 5 is a diagram showing an example of integrating feature amounts. The feature amount integrating unit 23 assigns weights to the feature amounts of user 1 as p 1 In this case, the weight for the feature of user 2 is p 2 = 1 - p 1 That is, the sum of the weights is p 1 +p 2 = 1. The weight value is 0.0≦p 1 The user can arbitrarily set a weight within the range of ≦1.0, and when the weight of one user is set, the weight of the other user is automatically determined.

[0039] For example, p 1 = 0.5, p 2 = 0.5. This corresponds to the feature obtained by adding and averaging the feature amounts of user 1 and user 2. For example, if the owner of the robot 30 is user 2 and the preferences of user 1 and user 2 are different, the feature integration unit 23 calculates p 1 = 0.8, p2 It is preferable to set the weight of the feature amount of user 1 to be greater by setting the weight of the feature amount of user 1 to be greater.

[0040] This is because, although user 2 can take photos at his / her own discretion, user 1 has to rely on the robot 30's suggestions for taking photos that suit his / her preferences.

[0041] When weighting the features of three users, user 1 to user 3, the feature integration unit 23 assigns weights to the features of user 1 and user 2 as p 1 , p 2 and the weight p 3 o p 3 = 1 - (p 1 +p 2 ) That is, the sum of the weights is p 1 +p 2 +p 3 =1.

[0042] The weight value is 0.0≦p 1 ≦1.0 and 0.0≦p 1 +p 2 The user can arbitrarily set the weight within the range of ≦1.0, and once the weights of two of the three users have been set, the weight of the remaining user is automatically determined.

[0043] Meanwhile, the camera 31 mounted on the robot 30 takes pictures and acquires photo data 5 of the external environment (step S5). The camera 31 may acquire the photo data 5 at any timing, for example, one photo per minute or one photo per five minutes.

[0044] Next, the image transmitting unit 34 transmits the photo data 5 acquired by the camera 31 to the image receiving unit 27 of the server device 20 via the network N (step S6).

[0045] Then, the feature extraction unit 22 of the server device 20 extracts a fourth feature from the photo data 5 received from the image transmission unit 34 of the robot 30 (step S7). The feature extraction method is the same as the extraction method in step S2. When the image transmission unit 34 transmits the image to the server 20, the image may be compressed and transmitted in consideration of the load on the network N.

[0046] Then, the similarity calculation unit 24 calculates the similarity between the third feature generated from the photo data 3 and 4 of users 1 and 2 and the fourth feature extracted from the photo data 5 acquired by the robot 30 (step S8).

[0047] The similarity calculation unit 24 uses, for example, cosine similarity to calculate the similarity, i.e., calculates the cosine similarity between a feature vector formed from the third feature amount and a feature vector formed from the fourth feature amount.

[0048] 6 is a diagram showing calculation of the similarity of feature quantities. Here, in step S2 described above, the similarity calculation unit 24 vectorizes each of a plurality of histograms, such as a color histogram, a shape histogram, and a composition histogram, and calculates a feature quantity for each of them. Note that the feature quantity shown in FIG. 6 is the third feature quantity.

[0049] The similarity calculation unit 24 then calculates the similarity for each feature amount. Each feature amount is associated in advance with an action that prompts the robot 30 to take a picture of the external environment.

[0050] Next, the determination unit 25 performs a threshold determination to determine whether each similarity calculated in step S8 is greater than a threshold (step S9). If any similarity is greater than the threshold, the instruction unit 26 instructs the robot 30 via the network N to perform a predetermined action corresponding to the feature with the highest similarity (step S10).

[0051] The operation unit 33 of the robot 30 that has received the instruction executes an operation according to the magnitude of the similarity (step S11). Note that, if the similarity is less than the threshold in step S10, the instruction unit 26 may instruct the robot 30 not to operate, or may not send an instruction to the robot 30.

[0052] 7 is a diagram showing an example of an action depending on the magnitude of the similarity. For example, when the similarity is 0.7 or more, the operation unit 33 causes the robot 30 to perform an action of encouraging the user 2 to take a photo, when the similarity is 0.4 or more but less than 0.7, the operation unit 33 causes the robot 30 to perform an action of being happy, and when the similarity is less than 0.4, the operation unit 33 causes the robot 30 to not perform an action.

[0053] For example, if the similarity is 0.7 or higher, the operating unit 33 sends an image taken by the robot 30 and used to calculate the similarity to the smartphone of the user 2 who owns the robot 30, and prompts the user 2 to take a photo by having the robot 30 output audio such as "I want you to take a photo!" and "Why don't you take a photo?"

[0054] Furthermore, the operating unit 33 causes the robot 30 to perform the action of raising and lowering both hands, and also extracts the most frequent color from the image taken by the robot 30 and used to calculate the similarity, and changes the eye color of the robot 30 to that color.

[0055] If the similarity is equal to or greater than 0.4 and less than 0.7, the operation unit 33 causes the robot 30 to perform a happy action. The operation unit 33 also causes the robot 30 to output sounds such as "Yay!", "Oh!", and "Good!", and also causes the robot 30 to perform an action of raising and lowering one hand.

[0056] The actions that the operation unit 33 causes the robot 30 to perform are not limited to those shown in Fig. 7 and may be other actions. Furthermore, the thresholds are not limited to 0.4 and 0.7 and may be other values, or there may be multiple thresholds and the operation unit 33 may cause the robot 30 to perform various actions depending on the magnitude of the similarity.

[0057] [Effects, etc.] When the robot 30 acquires a photo according to the preferences of a single user, the robot 30 can only acquire a photo that matches the preferences of that user, and when there are two or more users, the robot 30 cannot acquire a photo that matches the preferences of two or more users.

[0058] In contrast, the robot system 10 includes a server device 20 including: a feature integration unit 23 that generates a third feature by integrating a first feature of image data held by a first user and a second feature of image data held by a second user, a determination unit 25 that determines whether to execute output to the second user based on the similarity between the third feature and a fourth feature extracted from image data of the external environment, and an instruction unit 26 that issues an instruction to output when it is determined that output to the second user should be executed; and a robot 30 that operates in response to instructions from the server device 20. As described above, the configuration of the robot system 10 is not limited to the robot 30 and may be a terminal or the like.

[0059] The robot system 10 can prompt users to take photos based on feature amounts extracted from the photo data of two or more users. Compared to using feature amounts extracted from the photo data of a single user, the robot system 10 can execute an operation to prompt users to take photos according to feature amounts extracted from the photo data of two or more users, making it possible to suggest photoshoots that meet the preferences of two or more users.

[0060] Furthermore, for example, the feature integration unit 23 generates the weighted third feature based on weights set for the first feature and the second feature, respectively. The robot system 10 can suggest photography that strongly reflects the preferences of the first user to the second user who is the owner of the robot or the terminal.

[0061] Furthermore, for example, the instructing unit 26 instructs the second user to perform an action that prompts the second user to take a photograph of the external environment, so that the robot system 10 can prompt the second user to take a photograph.

[0062] Furthermore, for example, the instruction unit 26 instructs the second user to perform different actions as output when the similarity is greater than the first threshold and when the similarity is less than the first threshold and greater than the second threshold, so that the robot system 10 can perform different actions when the similarity exceeds the first threshold and when it exceeds the second threshold.

[0063] Furthermore, for example, the first feature and the second feature have multiple types, and each type is associated with a different action that prompts the user to capture a picture of the external environment. The feature integration unit 23 generates a third feature corresponding to each type, and the instruction unit 26 instructs the robot system 10 to execute output to the second user corresponding to the type with the highest similarity. Therefore, the robot system 10 can store two or more parameters of the feature and execute an action according to the parameter of the feature that is determined to have a high similarity.

[0064] Although the embodiments have been described above, the present invention is not limited to the above-described embodiments. For example, in the above-described embodiments, the order of multiple processes may be changed, or multiple processes may be executed in parallel.

[0065] In the above embodiment, the configuration in which the server 20 and the robot 30 are connected via the network N has been described, but some or all of the components of the server 20 shown in Figure 2 may be included as components of the robot 30.

[0066] That is, the robot 30 may include a feature integration unit 23 that generates a feature by integrating the features of the image data 3 and 4 held by the users 1 and 2 based on the features of those image data 3 and 4, a determination unit 25 that determines whether or not to execute output to the second user based on the similarity between the integrated feature and the feature extracted from the image data of the external environment, and an instruction unit 26 that instructs the robot 30 to execute output to the second user when it is determined that output to the second user should be executed.

[0067] In the above-described embodiments, each component may be realized by executing a software program suitable for that component, or by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0068] Furthermore, the general or specific aspects of the present invention may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0069] In addition, the present invention also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art would think of, or forms realized by arbitrarily combining the components and functions of each embodiment within the scope of the present invention.

[0070] The present disclosure can be used in an instruction device, a robot system, and a robot that can make suggestions according to a user's preferences.

[0071] REFERENCE SIGNS LIST 1 User 3 Photo data 10 Robot system 20 Server device 21 Storage unit 22 Feature extraction unit 23 Feature integration unit 24 Similarity calculation unit 25 Determination unit 26 Instruction unit 27 Image reception unit 30 Robot 31 Camera 32 Instruction reception unit 33 Operation unit 34 Image transmission unit 41 Image

Claims

1. An instruction device comprising: a feature integration unit that generates a third feature by integrating a first feature of image data held by a first user and a second feature of image data held by a second user, based on the first feature and the second feature; a determination unit that determines whether or not to execute output to the second user based on a similarity between the third feature and a fourth feature extracted from image data of an external environment; and an instruction unit that issues an instruction to output when it is determined that output to the second user is to be executed.

2. The pointing device according to claim 1, wherein the feature integration unit generates the weighted third feature based on weights set for the first feature and the second feature.

3. The instruction device according to claim 1, wherein the output to the second user is an action that prompts the second user to take a photograph of the external environment.

4. The instruction device according to claim 1, wherein the instruction unit instructs the terminal to execute an output for the second user, and the first user is not the owner of the terminal, and the second user is the owner of the terminal.

5. The pointing device according to claim 2, wherein the weight of the first characteristic amount is set to be greater than the weight of the second characteristic amount.

6. The instruction device of claim 1, wherein the instruction unit instructs the second user to execute a first action as output to the second user when the similarity is greater than a first threshold, and instructs the second user to execute a second action different from the first action as output to the second user when the similarity is smaller than the first threshold and greater than a second threshold.

7. The instruction device of claim 1, wherein the first feature and the second feature each have a plurality of types, each of the plurality of types is associated with a different output for the second user, the feature integration unit generates the third feature corresponding to each of the plurality of types, and the instruction unit instructs to execute an output for the second user corresponding to the type having the highest similarity.

8. A robot system comprising: an instruction device according to claim 1; and a robot that operates according to instructions from the instruction device.

9. A robot equipped with the instruction device according to claim 1.

Citation Information

Patent Citations

  • Image management device and image management method, and control program

    JP2008146174A

  • Method and system for automatically extracting photography information

    US20090202157A1