System, program, and information processing method

The system enhances service efficiency and quality by using learning models to analyze customer attributes and autonomously operate an avatar robot, addressing inefficiencies in existing service provision methods.

WO2025115886A1PCT designated stage expired Publication Date: 2025-06-05AVATARIN INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041929
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-27
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing techniques for providing customer services are inefficient, lacking in quality and consistency, and do not effectively utilize customer attribute information to enhance service delivery.

Method used

A system that utilizes an image-attribute learning model and an attribute-action learning model to analyze customer attributes and determine appropriate actions for an avatar robot to perform, enhancing service provision efficiency and quality by replicating user behavior.

Benefits of technology

The system enables more efficient and high-quality service delivery by autonomously operating an avatar robot based on learned customer attributes, reducing labor costs and improving service consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024041929_05062025_PF_FP_ABST
    Figure JP2024041929_05062025_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To improve the efficiency with which a service is provided to customers. [Solution] A system 1 according to one embodiment of the present disclosure comprises: an acquisition unit 100 that acquires attribute information pertaining to a customer; a storage unit 14 that stores a learning model which has learned the relationships between attribute information pertaining to each of one or more other customers and an operation that a user has caused an avatar robot 4 to execute by means of an operation terminal device 3 in order to provide a service to each of the one or more other customers; and an execution unit 104 that, by inputting the attribute information pertaining to the customer to the learning model, causes the avatar robot 4 to execute an operation corresponding to the attribute information.
Need to check novelty before this filing date? Find Prior Art

Description

System, program and information processing method

[0001] The present invention relates to a system, a program, and an information processing method.

[0002] Conventionally, there are known techniques for acquiring customer attribute information. For example, Patent Literature 1 describes a technique for acquiring customer information related to services that the customer can use, regardless of whether the customer is a registered member or not.

[0003] Japanese Patent Application Laid-Open No. 2022-000794

[0004] However, the method disclosed in Patent Document 1 does not allow for sufficient efficiency in providing services to customers.

[0005] In view of the above, an object of the present invention is to improve the efficiency of providing services to customers.

[0006] A system according to one aspect of the present disclosure includes a memory unit that stores a learning model that has learned the relationship between attribute information of each of one or more other customers and actions that a user causes an avatar robot to perform using an operation terminal device in order to provide services to each of the one or more other customers, and an execution unit that inputs customer attribute information into the learning model and causes the avatar robot to perform actions corresponding to the attribute information.

[0007] A system according to another aspect of the present disclosure acquires information regarding the relationship between attribute information of each of one or more customers and actions that a user causes an avatar robot to perform using an operation terminal device in order to provide a service to each of the one or more customers.

[0008] A program according to another aspect of the present disclosure causes an avatar robot's computer to function as an acquisition means for acquiring customer attribute information, and an execution means for executing an action corresponding to the attribute information by inputting the customer attribute information into a learning model that has learned the relationship between the attribute information of each of one or more other customers and the actions that a user has the avatar robot perform using an operating terminal device to provide a service to each of the one or more other customers.

[0009] An information processing method according to another aspect of the present disclosure includes a step of causing a system including an avatar robot to acquire customer attribute information, and a step of causing the avatar robot to perform an action corresponding to the attribute information by inputting the customer attribute information into a learning model that has learned the relationship between the attribute information of each of one or more other customers and actions that a user has caused the avatar robot to perform using an operation terminal device in order to provide a service to each of the one or more other customers.

[0010] According to the present disclosure, it is possible to provide services to customers more efficiently.

[0011] FIG. 1 is a functional block diagram of the system 1. FIG. 2 is a diagram for explaining an example of the operation of the system 1 according to the first embodiment. FIG. 3 is a diagram for explaining an example of the display screen of the input / output device 5 according to the first embodiment. FIG. 4 is a diagram for explaining an example of the operation of the system 1 according to the second embodiment. FIG. 5 is a hardware configuration diagram of various devices included in the system 1. FIG. 6 is a diagram for explaining another example of the operation of the system 1. FIG. 7 is a diagram for explaining another example of the operation of the system 1.

[0012] Preferred embodiments of the present disclosure will be described with reference to the accompanying drawings. In each drawing, components with the same reference numerals have the same or similar configurations. The embodiments described below are merely examples of the present disclosure and are not intended to limit the scope of the present disclosure.

[0013] In this disclosure, the terms "unit," "means," "device," and "system" do not simply mean physical means, but also include cases where the functions of the "unit," "means," "device," and "system" are realized by software. Furthermore, the functions of one "unit," "means," "device," or "system" may be realized by two or more physical means, devices, or software, or the functions of two or more "units," "means," "device," or "system" may be realized by one physical means, device, or software.

[0014] 1. Overview of System 1 An overview of a system 1 according to the present disclosure (hereinafter simply referred to as "system 1") will be described. In one embodiment, the system 1 generates the following two types of learning models.

[0015] (1) Image-attribute learning model m1: A learning model that learns the relationship between image information generated by an imaging device capturing an image of a customer (hereinafter referred to as "customer image information") and the attribute information of that customer. According to the image-attribute learning model m1, by inputting customer image information, customer attribute information can be obtained as output.

[0016] (2) Attribute-action learning model m2: A learning model that learns the relationship between a customer's attribute information and the actions that a user causes the avatar robot 4 to perform using the operation terminal device 3 in order to provide a service to the customer. According to the attribute-action learning model m2, by inputting the customer's attribute information, it is possible to obtain as output information for causing the avatar robot 4 to perform an action to provide a service to the customer.

[0017] After generating the learning model, the system 1 acquires customer attribute information by inputting customer image information generated by the imaging device of the avatar robot 4 into an image-attribute learning model m1. Next, the system 1 acquires information about the behavior of the avatar robot 4 by inputting the customer attribute information into an attribute-behavior learning model m2. Next, the system 1 causes the avatar robot 4 to perform an action to provide a service to the customer based on the information about the behavior. The system 1 allows the avatar robot 4 to reproduce the user's method of providing a service.

[0018] In this disclosure, a service includes activities or benefits provided in response to an explicit or implicit request from a customer. A service does not have to be an independent commercial transaction. Providing a service to a customer includes, for example, serving customers, guiding customers visiting a facility to the facility, and providing products to customers.

[0019] In addition, in the present disclosure, improving the efficiency of providing services to customers includes providing higher quality services and reducing costs associated with providing services.

[0020] In the present disclosure, the generation of image information by an imaging device includes a process of converting a subject in real space into image data (e.g., a JPEG file, a GIF file, etc.) via an imaging element, etc. In the present disclosure, the generation of audio information by an audio input device includes a process of converting audio, such as a human voice, into audio data (e.g., an MP3 file, a WAV file, etc.) via a microphone, etc.

[0021] 2. Functional configuration of system 1 The system 1 includes an information processing device 2, an operation terminal device 3, an avatar robot 4, an input / output device 5, and a communication network 6. The information processing device 2, the operation terminal device 3, the avatar robot 4, and the input / output device 5 are configured to be able to communicate with each other via the communication network 6.

[0022] -Operation terminal device 3- The operation terminal device 3 is a terminal device used by the user to operate the avatar robot 4. The operation terminal device 3 may be a personal computer, a smartphone, a tablet terminal, a controller equipped with a joystick and buttons, or the like.

[0023] In one embodiment, the operation terminal device 3 includes an imaging device that captures images and a voice input device that inputs voice. Operating the operation terminal device 3 by a user includes having the imaging device of the operation terminal device 3 capture the user's facial expression. Operating the operation terminal device 3 by a user also includes having the voice input device of the operation terminal device 3 acquire the user's voice. Hereinafter, image information generated by the imaging device of the operation terminal device 3 capturing an image of the user will be referred to as "user image information." Voice information generated by the voice input device of the operation terminal device 3 acquiring the user's voice will be referred to as "user voice information."

[0024] -Avatar Robot 4- The avatar robot 4 is a device that provides services to customers in remote locations as an avatar of the user. The user can be said to be the operator of the avatar robot 4. In one embodiment, the avatar robot 4 includes a moving device for moving, an imaging device for capturing images, a display device for displaying an avatar, an audio output device for outputting audio, and an audio input device for inputting audio.

[0025] Examples of the moving device include wheels, legs, and propellers. The avatar displayed by the display device may be an image of the user himself / herself, as in a video call, or a 3D model generated by photographing the user. The "3D model" referred to here may be a 3D model of the user himself / herself, or a 3D model of a fictional person or character. The avatar robot 4 may also be called a telepresence robot or a telexistence robot. The avatar robot 4 is not limited to a humanoid robot, and its shape is not particularly limited. Hereinafter, the voice information generated by the voice input device of the avatar robot 4 acquiring the voice of the customer is referred to as "customer voice information."

[0026] In one embodiment, the avatar robot 4 receives information about an operation input to the operation terminal device 3 and operates based on the received information about the operation. For example, when an operation for specifying a movement direction is input to the operation terminal device 3, the avatar robot 4 moves in the movement direction.

[0027] Furthermore, the avatar robot 4 changes the facial expression of the avatar displayed on its display device based on the user image information, for example, to correspond to the facial expression of the user. In other words, the avatar robot 4 synchronizes the facial expression of the user with the facial expression of the avatar.

[0028] Furthermore, the avatar robot 4 outputs a sound corresponding to the user's voice from its own sound output device based on, for example, the user's voice information.

[0029] In one embodiment, information regarding the movement of the avatar robot 4 is transmitted to the information processing device 2. The information regarding the movement of the avatar robot 4 includes information indicating when and what movement the avatar robot 4 performed. In response to this, the information processing device 2 stores the received information regarding the movement in order to generate an attribute-movement learning model m2.

[0030] In one embodiment, the avatar robot 4 may be equipped with various instruments, and may transmit the measurement results of the various instruments to the information processing device 2 as part of information related to the operation of the avatar robot 4. The avatar robot 4 may be equipped with, for example, a position information sensor (e.g., GPS), an acceleration sensor, an angular acceleration sensor, a distance sensor, a temperature sensor, and a pressure sensor.

[0031] The input / output device 5 includes an output unit that outputs various information generated or acquired by the avatar robot 4 to the user, and an input unit that receives customer attribute information from the user. The input / output device 5 may be a personal computer, a smartphone, a tablet terminal, or the like.

[0032] In one embodiment, the output unit displays an image of the customer to the user based on the customer image information. Also, in one embodiment, the output unit outputs a sound corresponding to the customer's voice to the user based on the customer voice information. The output unit may output past records (e.g., video and audio recordings) of various information generated by the avatar robot 4, or may output various information in real time.

[0033] In one embodiment, the input unit receives input of attribute information of a customer from a user who has viewed an image of the customer.

[0034] - Information Processing Device 2 - The information processing device 2 is a device that generates an image-attribute learning model m1 and an attribute-action learning model m2, and also controls the actions of the avatar robot 4.

[0035] The information processing device 2 includes a control unit 10, a storage unit 14, a network interface unit 18, and a bus 16. The control unit 10, the storage unit 14, and the network interface unit 18 are electrically connected to one another via the bus 16.

[0036] Storage Unit 14 The storage unit 14 stores an image-attribute learning model m1 and an attribute-action learning model m2, as well as various programs executed by the control unit 10 (described later). The image-attribute learning model m1 and the attribute-action learning model m2 may be learning models including neural networks, and may be large-scale deep learning models such as Transformers. Data input to these learning models (learning data / determination data) may be data that has been subjected to appropriate dimensional compression and feature extraction.

[0037] (Image-attribute learning model m1) The image-attribute learning model m1 is a learning model generated by learning the relationship between customer image information and customer attribute information. In one embodiment, the image-attribute learning model m1 is generated by learning a plurality of training data in which the feature quantities of the customer image information are associated with the attribute information of the customer related to the customer image information. The customer attribute information can also be referred to as a label assigned to the customer.

[0038] (Attribute-action learning model m2) The attribute-action learning model m2 is a learning model generated by learning the relationship between customer attribute information and actions that users have made the avatar robot 4 perform using the operation terminal device 3 in order to provide services to the customers. In one embodiment, the attribute-action learning model m2 is generated by learning a plurality of training data in which information on actions that users have made the avatar robot 4 perform using the operation terminal device 3 is associated with customer attribute information. In other words, the attribute-action learning model m2 may learn, as training data, information indicating what actions users have made the avatar robot 4 perform for what type of customers.

[0039] In one embodiment, the action that the user causes the avatar robot 4 to perform using the operation terminal device 3 (i.e., the action that the attribute-action learning model m2 learns) includes moving the avatar robot 4 toward a customer. That is, the attribute-action learning model m2 may learn, as training data, information indicating to what kind of customer and how the user moved the avatar robot 4.

[0040] In one embodiment, the action that the user causes the avatar robot 4 to perform using the operation terminal device 3 includes changing the facial expression of the avatar displayed on the display device of the avatar robot 4. That is, the attribute-action learning model m2 may learn using, as training data, information indicating what facial expressions the user causes the avatar to make for what type of customer. Note that, when the avatar robot 4 changes the facial expression of the avatar based on user image information, the user changing the facial expression of the avatar includes the user having the imaging device of the operation terminal device 3 capture an image of the user's own facial expression.

[0041] In one embodiment, the action that the user causes the avatar robot 4 to perform using the operation terminal device 3 includes outputting a sound from the sound output device of the avatar robot 4. That is, the attribute-action learning model m2 may learn, as training data, information indicating what kind of sound the user caused the avatar robot 4 to output for what kind of customer. Note that, when the sound output device of the avatar robot 4 outputs a sound based on user voice information, the user causing the sound output device of the avatar robot 4 to output a sound includes the user causing the sound input device of the operation terminal device 3 to acquire the user's own voice.

[0042] Control Unit 10 The control unit 10 functions as an acquisition unit 100 , a reception unit 102 , and an execution unit 104 by executing various programs stored in the storage unit 14 .

[0043] (Acquisition Unit 100) The acquisition unit 100 acquires various information necessary for operation of the information processing device 2. In one embodiment, the acquisition unit 100 includes an image information acquisition unit 100a, an attribute information acquisition unit 100b, and an audio information acquisition unit 100c.

[0044] The image information acquisition unit 100a acquires customer image information. The customer image information is information about an image that shows, for example, the appearance, facial expression, belongings, location, and movement of a customer. The customer image information may also be information about a video.

[0045] In one embodiment, the image information acquisition unit 100a acquires customer image information generated by an imaging device of the avatar robot 4 capturing an image of the customer.

[0046] The attribute information acquisition unit 100b acquires customer attribute information. The customer attribute information includes information that characterizes the customer. In one embodiment, the customer attribute information includes information for identifying the customer's requests, questions, complaints, etc. In one embodiment, the customer attribute information includes information that is fixed and does not change about the customer for at least a predetermined period (e.g., while the customer is receiving a service) and information that may change from moment to moment while the customer is receiving a service. The information that is fixed about the customer for at least a predetermined period includes, for example, information about at least some of the customer's age, gender, accompanying persons, nationality, and belongings. The information that may change while the customer is receiving a service includes, for example, the customer's line of sight, behavior, satisfaction level, and location. The customer's location may change, for example, when the avatar robot 4 guides the customer to a predetermined location as part of a service.

[0047] In the present disclosure, customer satisfaction includes an index indicating the degree of satisfaction a customer has with a service, facility, etc., or the degree of excitement that exceeds expectations. Customer satisfaction includes, for example, an evaluation of the quality of the service and an evaluation of the ease of use of the facility. Customer satisfaction may include at least one of the customer's satisfaction before the avatar robot 4 provides the service to the customer and the customer's satisfaction after the avatar robot 4 provides the service to the customer. Customer satisfaction can be inferred, for example, from the customer's facial expression, tone of voice, gestures, etc.

[0048] In one embodiment, the attribute information acquiring unit 100b acquires customer attribute information by inputting the customer image information acquired by the image information acquiring unit 100a into the image-attribute learning model m1.

[0049] In the present disclosure, acquiring information includes making the information processable by the control unit 10. Acquiring information includes, for example, receiving the information from another device, reading the information from the storage unit 14, and generating the information by predetermined processing.

[0050] The voice information acquisition unit 100c acquires voice information related to voice input to the voice input device of the avatar robot 4. The voice information acquisition unit 100c also updates the customer's attribute information based on the voice information by voice recognition, natural language processing, etc. For example, if the voice information is related to a customer's voice saying, "I'm having trouble with XX, but...", the voice information acquisition unit 100c can add the information "I'm having trouble with XX" to the customer's attribute information.

[0051] (Receiving Unit 102) The receiving unit 102 receives information relating to an operation input to the operation terminal device 3 by the user.

[0052] (Execution unit 104) The execution unit 104 causes the avatar robot 4 to perform a predetermined action. In one embodiment, the execution unit 104 inputs the customer's attribute information acquired by the attribute information acquisition unit 100b into the attribute-action learning model m2, and causes the avatar robot 4 to perform an action corresponding to the attribute information.

[0053] In one embodiment, the execution unit 104 refers to the attribute-action learning model m2 and moves the avatar robot 4 toward a customer according to the customer's attribute information. For example, if the customer appears to be looking around in a hurry and searching for something, the execution unit 104 may move the avatar robot 4 straight toward the customer. Alternatively, for example, if the customer is walking while looking at a map on a smartphone, the execution unit 104 may move the avatar robot 4 in a parallel motion while slowly approaching the customer.

[0054] In one embodiment, the execution unit 104 refers to the attribute-action learning model m2 and changes the facial expression of the avatar displayed on the display device of the avatar robot 4 to a facial expression corresponding to the attribute information of the customer. For example, if the customer is smiling or has a relieved expression, the execution unit 104 may change the facial expression of the avatar to a smiling expression. For example, if the customer is angry, the execution unit 104 may change the facial expression of the avatar to an apologetic expression.

[0055] In one embodiment, the execution unit 104 refers to the attribute-action learning model m2 and causes the audio output device of the avatar robot 4 to output a voice corresponding to the attribute information of the customer. For example, if a customer is running around confusedly while carrying luggage, the execution unit 104 may cause the avatar robot 4 to output a voice saying, "Are you looking for a counter where I can leave my luggage?". For example, if a customer looks confused in front of a device for performing a predetermined procedure (e.g., an airplane ticket vending machine, a parking fee payment machine at a shopping center, or a ticket machine at a restaurant), the execution unit 104 may cause the avatar robot 4 to output a voice saying, "Are you having trouble using this device?"

[0056] In one embodiment, when a predetermined condition is satisfied, the execution unit 104 causes the avatar robot 4 to perform an action corresponding to the operation received by the receiving unit 102, instead of or in addition to the action corresponding to the customer's attribute information. That is, the execution unit 104 may basically cause the avatar robot 4 to operate autonomously based on the attribute-action learning model m2, and may also cause the avatar robot 4 to operate based on a user's operation when a predetermined condition is satisfied.

[0057] In one embodiment, a case where the predetermined condition is satisfied means a case where it is difficult for the avatar robot 4 to autonomously provide a service to a customer. For example, if a customer's request (an example of customer attribute information) is unusual and the attribute-behavior learning model m2 has not learned such a request as training data, the system 1 determines that the predetermined condition is satisfied. Also, for example, if the customer's emotions (an example of customer attribute information) identified from the customer's facial expression and tone of voice are significantly negative, the system 1 determines that the predetermined condition is satisfied.

[0058] In one embodiment, notification to the customer may be limited as to whether the execution unit 104 has the avatar robot 4 perform an action corresponding to the customer's attribute information or an action corresponding to the operation received by the receiving unit 102. In other words, the system 1 may be configured so that the customer cannot distinguish whether the avatar robot 4 is operating autonomously based on the attribute-action learning model m2 or based on the user's operation.

[0059] In the present disclosure, the autonomous operation of the avatar robot 4 includes the operation of the avatar robot 4 without being influenced by a user's operation (more specifically, an operation by the user on the operation terminal device 3). In other words, the autonomous operation of the avatar robot 4 includes the operation of the avatar robot 4 as a result of cooperation between the avatar robot 4 and the information processing device 2.

[0060] -Communication Network 6- The communication network 6 realizes communication between the information processing device 2, the operation terminal device 3, the avatar robot 4, and the input / output device 5. In one embodiment, the communication network 6 realizes communication between the devices using the TCP / IP protocol. The communication network 6 may be a closed network isolated from the Internet, or may be a private cloud network.

[0061] 3. Operation of System 1 - First Embodiment - An example of the operation of System 1 will be described with reference to FIGS. 2-4. In the following description, the avatar robot 4 is placed in airport AR10 and provides services to customers at airport AR10 while moving around within the airport AR10. The user is also assumed to be in an operation room AR20 separate from airport AR10. In this example, the user causes the avatar robot 4 to perform a predetermined operation (see S100-102, described below). Thereafter, the user is assumed to input customer attribute information while visually viewing a recorded video of the customer captured by the imaging device of the avatar robot 4 (see S104, described below).

[0062] Method for Generating Learning Models FIG. 2C is a diagram for explaining an example of a method by which the system 1 generates an image-attribute learning model m1 and an attribute-action learning model m2.

[0063] First, the avatar robot 4 transmits customer image information generated by capturing an image of a customer with its own imaging device to the input / output device 5 (S100). The input / output device 5 displays an image corresponding to the received customer image information to the user in real time. The input / output device 5 also stores the customer image information for the user to input customer attribute information (described later).

[0064] Next, the avatar robot 4 transmits the customer image information to the information processing device 2 (S101). The information processing device 2 at least temporarily stores the customer image information in order to allow the image-attribute learning model m1 to learn the relationship between the customer image information and customer attribute information received in S104, which will be described later.

[0065] Next, the user operates the operation terminal device 3 to cause the avatar robot 4 to perform a predetermined action (S102) while visually checking the customer image information displayed in real time on the input / output device 5. The predetermined action includes an action for providing a service to the customer.

[0066] In this example, the user operates the operation terminal device 3 to move the avatar robot 4 close to the customer, change the expression of the avatar displayed on the display device of the avatar robot 4 to a smiling face, and cause the voice output device of the avatar robot 4 to output a voice message saying, "Do you need any help checking in your baggage?"

[0067] Next, the avatar robot 4 transmits information about the action performed by the user to the information processing device 2 (S103). The information about the action of the avatar robot 4 may include at least a part of the action log of the avatar robot 4. The information processing device 2 at least temporarily stores the information about the action in order to allow the attribute-action learning model m2 to learn the relationship between the information about the action of the avatar robot 4 and the customer attribute information received in S105, which will be described later.

[0068] After the provision of the service to the customer is completed, the user inputs attribute information of the customer while viewing the recorded video of the customer on the input / output device 5 (S104-105).

[0069] Fig. 3 is a diagram illustrating an example of a display screen of the input / output device 5. The example display screen of Fig. 3 displays an image d100, a customer attribute input area d101, a time interval designation tool d108, a playback location d109, and a record button d110.

[0070] Image d100 is displayed based on customer image information generated by capturing an image of a customer with an imaging device of avatar robot 4. In this example, image d100 shows a customer carrying luggage operating a touch panel of an automated baggage check machine. Note that in this example, image d100 is assumed to be part of a recorded video.

[0071] The customer attribute input area d101 is an area where the user inputs customer attribute information. In the example display screen of Fig. 3, the customer attribute input area d101 displays a scene selection area d102, a customer category selection area d104, and a customer satisfaction level selection area d106.

[0072] The scene selection area d102 displays check boxes that allow the user to select a scene corresponding to the image d100. In the example display screen of FIG. 3, the scenes that the user can select include "facility information," "boarding pass change," "refund-related," "points-related," "baggage-related," "irregularity-related," and "complaint-related." In the example display screen of FIG. 3, two of these scenes, "facility information" and "baggage-related," have been selected. Note that the scenes correspond to the customer's location and / or the procedure the customer is about to perform, and are therefore included in the customer's attribute information.

[0073] The customer category selection area d104 displays pull-down buttons that allow the user to select a category of the customer associated with the image d100. In the example display screen of FIG. 3 , pull-down buttons for the items "gender," "age," "nationality," and "other characteristics" are displayed. The user can select a setting value for one of the pull-down buttons by pressing the pull-down button. For example, if the user presses the "age" pull-down button, the input / output device 5 displays "teens," "twenties," "thirties," and the like as setting value candidates for "age." In response, the user can set the selected setting value in the pull-down button by selecting one of the setting values. In this example, the setting values ​​for the items "gender," "age," and "nationality" are set to "male," "twenties," and "Japanese," respectively.

[0074] Furthermore, the user can input customer attribute information other than gender, age, and nationality by pressing the pull-down menu for the item "Other characteristics." For example, the user can associate the following setting values ​​for customer attribute information with accompanying person, line of sight, behavior, luggage, and location: "none (alone)," "looking at the touch panel of the automated baggage checker," "using the touch panel of the automated baggage checker," "one piece of luggage on the shoulder," and "in front of the automated baggage checker."

[0075] The customer satisfaction level selection area d106 displays radio buttons that allow the user to select the satisfaction level of the customer associated with image d100. From the grim expression of the customer associated with image d100, the user can infer that the customer does not know how to use the touch panel of the automated baggage check machine and is dissatisfied with the ease of use of the facility. Therefore, in the example display screen of FIG. 3, the level of "dissatisfied" is selected from three levels: "satisfied," "average," and "dissatisfied."

[0076] The time interval designation tool d108 is an element for designating a time interval from a predetermined time interval of the recorded video (in this example, from "now" to "6 hours ago") to which the content entered in the customer attribute input area d101 is to be associated. That is, by operating the time interval designation tool d108, the user can designate the time interval of the customer image information to be used for learning. In the example display screen of FIG. 3, the time interval from "2 hours ago" to "2.5 hours ago" is designated as the time interval of the customer image information to be used for learning.

[0077] The playback point d109 is an element that indicates the time point corresponding to the image d100 within a predetermined time period of the recorded video. In the example display screen of Fig. 3, the playback point d109 indicates the midpoint between "2 hours ago" and "2.5 hours ago." The user can change the image d100 to one that corresponds to a different time period by pressing and holding the playback point d109 and moving it along the time period designation tool d108.

[0078] The record button d110 is a button that, when pressed by the user, allows the content input in the customer attribute input area d101 (i.e., customer attribute information) to be sent to the information processing device 2.

[0079] Based on the user pressing the record button d110, the input / output device 5 transmits to the information processing device 2 information indicating that the image d100 for the time interval selected in the time interval designation tool d108 should be associated with the input customer attribute information and trained in the image-attribute learning model m1 (S104).

[0080] In the example display screen of FIG. 3, when the user presses the record button d110, the image-attribute learning model m1 learns by associating the following (1) and (2): (1) Customer image information relating to the images d100 from "2 hours ago" to "2.5 hours ago" (2) Customer attribute information indicating that the "gender," "age," and "nationality" are "male," "20s," and "Japanese," respectively, and that the customer is "dissatisfied" with the "facility guide" and / or "baggage-related" scenes

[0081] In addition, based on the user pressing the record button d110, the input / output device 5 transmits to the information processing device 2 information regarding the behavior of the avatar robot 4 during the time interval selected in the time interval designation tool d108, associating it with the input customer attribute information, and having the attribute-behavior learning model m2 learn it (S105).

[0082] In the example display screen of FIG. 3, when the user presses the record button d110, the attribute-action learning model m2 learns by associating the following (3) and (4): (3) Attribute information of a customer whose "gender," "age," and "nationality" are "male," "20s," and "Japanese," respectively, and who is "dissatisfied" with the "facility guide" and / or "baggage-related" scenes; (4) Information about the actions of the avatar robot 4 from "2 hours ago" to "2.5 hours ago"

[0083] The system 1 generates an image-attribute learning model m1 and an attribute-action learning model m2 by repeating the same operations as in S100-105 for a plurality of customers.

[0084] To summarize the operation for generating a learning model by system 1, information processing device 2 generates image-attribute learning model m1 by learning the relationship between customer image information (see S101) generated by the imaging device of avatar robot 4 photographing a customer and customer attribute information (see S104) related to the customer image information input by the user to input / output device 5. In addition, information processing device 2 generates attribute-action learning model m2 by learning the relationship between customer attribute information (see S105) related to the customer image information input by the user to input / output device 5 and information related to the action the user caused avatar robot 4 to perform to provide a service to the customer (see S103).

[0085] 4 is a diagram illustrating an example of a method in which the system 1 uses the image-attribute learning model m1 and the attribute-action learning model m2 to cause the avatar robot 4 to operate without user operation. First, the avatar robot 4 transmits customer image information generated by its own imaging device by capturing an image of a customer to the information processing device 2 (S200).

[0086] Next, the information processing device 2 inputs the received customer image information into the image-attribute learning model m1 to identify the attribute information of the customer related to the customer image information, and inputs the identified customer attribute information into the attribute-action learning model m2 (S202).

[0087] Next, the information processing device 2 inputs the customer's attribute information into the attribute-action learning model m2, and causes the avatar robot 4 to perform an action corresponding to the customer's attribute information identified by the input (S204).

[0088] Second Embodiment Another example of the operation of the system 1 will be described with reference to Fig. 5. In the first embodiment, an example in which the avatar robot 4 operates based on a user's operation was described with reference to Fig. 2, and an example in which the avatar robot 4 operates based on an image-attribute learning model m1 and an attribute-action learning model m2 without a user's operation was described with reference to Fig. 4. In contrast, in the second embodiment, an example of the system 1 will be described in which the avatar robot 4 is controlled to operate based on a user's operation or independently of a user's operation, based on whether a predetermined condition is satisfied.

[0089] In the following example, in the initial state, the image-attribute learning model m1 and the attribute-action learning model m2 are already generated, and the avatar robot 4 operates without user operation as described with reference to Figure 4.

[0090] First, the information processing device 2 acquires customer attribute information (S300). The customer attribute information may be identified based on, for example, customer image information and the image-attribute learning model m1.

[0091] Next, the information processing device 2 causes the avatar robot 4 to perform an action corresponding to the customer's attribute information (S302). The action corresponding to the customer's attribute information may be identified, for example, based on the customer's attribute information identified in S300 and the attribute-action learning model m2. In this example, the information processing device 2 causes the avatar robot 4 to output a voice message saying, "Is there something I can help you with?"

[0092] Next, the information processing device 2 acquires voice information relating to the customer's response from the avatar robot 4 (S304), and updates the customer's attribute information based on the voice information (S306).

[0093] Next, the information processing device 2 determines whether a predetermined condition is satisfied (S308). In this example, the information processing device 2 determines whether the attribute-action learning model m2 can be used for the updated attribute information of the customer. Specifically, the information processing device 2 determines whether the updated attribute information of the customer corresponds to an "outlier" identified by the attribute information of one or more other customers. More specifically, the information processing device 2 compares the distribution of the features of the attribute information of each of the one or more other customers with the features of the updated attribute information of the customer, and determines whether the occurrence probability of the updated attribute information of the customer is equal to or less than a predetermined threshold.

[0094] If the information processing device 2 determines that the predetermined condition is satisfied (YES in S308), the information processing device 2 sends a notification to the user (S312). Next, the information processing device 2 receives information regarding the user's operation of the operation terminal device 3 (S314) and causes the avatar robot 4 to perform an action corresponding to the operation (S316). Note that the customer does not need to be notified that the avatar robot 4 has switched from a state in which it is operating independently of the user's operation to a state in which it is operating based on the user's operation. In other words, the system 1 may be configured so that the customer cannot tell whether the avatar robot 4 is being operated by the user or autonomously.

[0095] On the other hand, if the information processing device 2 determines that the predetermined condition is not satisfied (S308 NO), the information processing device 2 causes the avatar robot 4 to perform an action corresponding to the updated attribute information of the customer based on the attribute-action learning model m2 (S310).

[0096] 4. Effects of System 1 System 1 according to one aspect of the present disclosure includes an acquisition unit 100 that acquires attribute information of a customer, a storage unit 14 that stores an attribute-action learning model m2 that has learned the relationship between the attribute information of each of one or more other customers and actions that a user has caused an avatar robot 4 to perform using an operation terminal device 3 in order to provide a service to each of the one or more other customers, and an execution unit 104 that inputs the attribute information of the customer into the attribute-action learning model m2 and causes the avatar robot 4 to perform an action corresponding to the attribute information. In one embodiment, the customer attribute information includes information regarding at least a part of the customer's age, gender, line of sight, action, accompanying person, nationality, belongings, satisfaction level, and location.

[0097] System 1 can improve the efficiency of providing services to customers. System 1 autonomously operates avatar robot 4 based on the user's past service provision records (e.g., what kind of services were provided to what kind of customers). In other words, system 1 can make avatar robot 4 reproduce the user's method of providing services. This reduces the need for the user to operate avatar robot 4 using operation terminal device 3, reduces labor costs, and enables services to be provided to multiple customers simultaneously.

[0098] In one embodiment, the attribute information of the customer is identified based on image information generated by an imaging device of the avatar robot 4 photographing the customer.

[0099] Generally, when a user operates the avatar robot 4 using the operation terminal device 3, the user determines the customer's attribute information while visually viewing an image captured by the imaging device of the avatar robot 4 using the input / output device 5. According to the above configuration, the avatar robot 4 can reproduce the method of determining the customer's attribute information by the user.

[0100] In one embodiment, the avatar robot 4 is a device that can move within a predetermined area.

[0101] In order to provide high-quality service to customers, it is preferable to have the avatar robot 4 move toward the customer rather than having the customer move toward the avatar robot 4. According to the above configuration, the avatar robot 4 can be moved toward the customer, thereby providing higher quality service to the customer.

[0102] In one embodiment, the action that the user causes the avatar robot 4 to perform using the operation terminal device 3 includes moving the avatar robot 4 toward each of one or more other customers, and the action that the execution unit 104 causes the avatar robot 4 to perform includes moving the avatar robot 4 toward the customer in accordance with the customer's attribute information.

[0103] In order to provide high-quality service to customers, the way in which the customer is approached is important. For example, if a customer appears to be in a hurry, approaching them slowly may not satisfy the customer's needs. With the above configuration, the system 1 can cause the avatar robot 4 to reproduce the way in which the customer is approached based on the user's operation.

[0104] In one embodiment, the avatar robot 4 may include a display device that displays an avatar, and the action that the user causes the avatar robot 4 to perform using the operation terminal device 3 includes changing the facial expression of the avatar, and the action that the execution unit 104 causes the avatar robot 4 to perform includes changing the facial expression of the avatar to an expression corresponding to the attribute information of the customer. Also, in one embodiment, the user changes the facial expression of the avatar based on image information generated by an imaging device of the operation terminal device 3 capturing an image of the user.

[0105] In order to provide high-quality service to customers, facial expressions used by the user are important. For example, providing service with a smile to an angry customer may reduce customer satisfaction. With the above configuration, the system 1 can make the avatar reproduce the way the user uses facial expressions.

[0106] In one embodiment, the action that the user causes the avatar robot 4 to perform using the operation terminal device 3 includes outputting a sound from the sound output device of the avatar robot 4, and the action that the execution unit 104 causes the avatar robot 4 to perform includes outputting a sound corresponding to the customer's attribute information from the sound output device of the avatar robot 4. Also, in one embodiment, the user causes the sound output device of the avatar robot 4 to output a sound using the operation terminal device 3 based on sound information generated by the sound input device of the operation terminal device 3 acquiring the user's voice.

[0107] In order to provide high-quality service to customers, it is important to speak to them. With the above configuration, the system 1 can make the avatar robot 4 reproduce the speech of the user.

[0108] In one embodiment, the system 1 further includes a receiving unit 102 that receives information regarding an operation input by a user to the operation terminal device 3, and the execution unit 104 causes the avatar robot 4 to perform an action corresponding to the operation received by the receiving unit 102, instead of or in addition to an action corresponding to the customer's attribute information, in response to a predetermined condition being met.

[0109] For example, when a customer's attribute information is unique or when an emergency response is required, it may be difficult for the avatar robot 4 to provide autonomous services. According to the above configuration, even if the avatar robot 4 basically acts autonomously, the avatar robot 4 will operate based on the user's operation when certain conditions are met, and therefore, it can flexibly respond to unforeseen circumstances.

[0110] In one embodiment, notification to the customer is limited as to whether the execution unit 104 is causing the avatar robot 4 to perform an action corresponding to the customer's attribute information or an action corresponding to a user's operation.

[0111] In order to give customers the impression that the service is consistent, it may be preferable not to indicate to customers whether the avatar robot 4 is operating independently of the user's operation or based on the user's operation. With the above configuration, it is possible to give customers the impression that the service is consistent.

[0112] The effects of the system 1 described above are merely examples and do not limit the scope of application of the present disclosure.

[0113] 6, an example of a hardware configuration in which the various devices described above are realized by a computer 70 will be described. Note that the functions of each device can also be realized by dividing them into multiple devices.

[0114] As shown in FIG. 6, the computer 70 includes a processor 700 , a storage device 702 , an input I / F 704 , a data I / F 706 , a communication I / F 708 , and a display device 710 .

[0115] The processor 700 controls various processes in the computer 70 by executing programs stored in the storage device 702. For example, the various functional units and the like provided in the control units of various devices can be realized by the processor 700 executing the programs stored in the storage device 702.

[0116] The storage device 702 is a storage medium such as a RAM (Random Access Memory), etc. The RAM temporarily stores the program code of the program executed by the processor 700 and data required when the program is executed.

[0117] The storage device 702 may also be a non-volatile storage medium such as a hard disk drive (HDD) or flash memory. The storage device 702 stores an operating system and various programs for implementing the above-described configurations. The storage medium storing the various programs may be a non-transitory computer-readable medium. In addition, the storage device 702 may also store tables that register various types of information and a database that manages the tables. Such programs and data are loaded into the storage device 702 as needed and referenced by the processor 700.

[0118] The input I / F 704 is a device for receiving input from a user. Specific examples of the input I / F 704 include a camera, a button, a microphone, a keyboard, a mouse, a touch panel, various sensors, and a wearable device. The input I / F 704 may be connected to the computer 70 via an interface such as a USB (Universal Serial Bus).

[0119] The data I / F 706 is a device for inputting data from outside the computer 70. A specific example of the data I / F 706 is a drive device for reading data stored in various storage media. The data I / F 706 may be provided outside the computer 70. In this case, the data I / F 706 is connected to the computer 70 via an interface such as a USB.

[0120] The communication I / F 708 is a device for performing data communication via the communication network 6, either wired or wirelessly, with devices external to the computer 70. The communication I / F 708 may be provided external to the computer 70. In this case, the communication I / F 708 is connected to the computer 70 via an interface such as a USB.

[0121] The display device 710 is a device for displaying various types of information. Specific examples of the display device 710 include a liquid crystal display, an organic EL (Electro-Luminescence) display, and a display of a wearable device. The display device 710 may be provided outside the computer 70. In this case, the display device 710 is connected to the computer 70 via, for example, a display cable. Furthermore, when a touch panel is used as the input I / F 704, the display device 710 can be configured as an integral part of the input I / F 704.

[0122] Furthermore, the components of the various devices described in the above embodiments are assumed to realize predetermined processing in cooperation with other hardware by the processor 700 executing a program stored in the storage device 702. In other words, these components are assumed to be software or firmware, as well as corresponding hardware, and in both of these concepts, they are also referred to as "functions," "means," "parts," "processing circuits," "units," or "modules," and can be interpreted as such.

[0123] 6. Modifications The system 1 described in the above embodiment is merely an example, and can be modified as appropriate within the scope of not causing any contradiction.

[0124] [Device storing image-attribute learning model m1 and attribute-action learning model m2] In the above embodiment, the image-attribute learning model m1 and attribute-action learning model m2 are described as being stored in the storage unit 14 of the information processing device 2, but this is not limiting. The image-attribute learning model m1 and attribute-action learning model m2 may be installed in the avatar robot 4, or may be stored in a cloud server device other than the information processing device 2.

[0125] [Path for Acquiring Customer Image Information] In the above embodiment, the image information acquisition unit 100a has been described as acquiring customer image information generated by an imaging device of the avatar robot 4 capturing an image of the customer, but this is not limited to this. The image information acquisition unit 100a may acquire customer image information generated by an imaging device (e.g., a fixed camera) installed in a predetermined area where the avatar robot 4 may be active (e.g., an airport, a train station, a shopping mall, a restaurant, a convenience store, etc.) capturing an image of the customer. Alternatively, the image information acquisition unit 100a may acquire customer image information generated by an imaging device of a smartphone owned by the customer capturing an image of the customer themselves.

[0126] [Path for Acquiring Customer Attribute Information] In the above embodiment, the attribute information acquisition unit 100b acquires customer attribute information by inputting customer image information acquired by the image information acquisition unit 100a into the image-attribute learning model m1. However, this is not limited to this. The attribute information acquisition unit 100b may acquire customer attribute information based on input from the operation terminal device 3. The attribute information acquisition unit 100b may also acquire customer attribute information generated by the avatar robot 4. The attribute information acquisition unit 100b may also acquire customer attribute information stored in a smartphone owned by the customer. The attribute information acquisition unit 100b may also acquire customer attribute information from another system that manages customer information. The attribute information acquisition unit 100b may also acquire customer attribute information based on customer identification information (e.g., biometric information of the customer, a ticket owned by the customer, the customer's smartphone, etc.).

[0127] [Information on Actions Executed by Avatar Robot 4] In the above embodiment, the attribute-action learning model m2 is described as learning information on actions executed by the user in the avatar robot 4 (for example, the action log of the avatar robot 4) by associating it with customer attribute information, but this is not limited to this. Instead of information on the actions of the avatar robot 4, the attribute-action learning model m2 may learn information on operations input by the user to the operation terminal device 3 (for example, the operation log of the operation terminal device 3) by associating it with customer attribute information.

[0128] [Method of Generating Learning Data] In the above embodiment, the user inputs the time period of the customer image information to be used for learning and the customer attribute information into the input / output device 5 (see FIG. 3 and S104-105) while visually viewing the recorded video after completing the provision of service to the customer (see S100-103), but this is not limited to this. The user may input the time period of the customer image information to be used for learning and the customer attribute information into the input / output device 5 in real time while providing the service to the customer (i.e., while operating the operation terminal device 3). In this case, the input / output device 5 and the operation terminal device 3 may be the same device.

[0129] The operation of the system 1 and an example of the display screen of the input / output device 5 when a user inputs the time period of customer image information used for learning and the attribute information of the customer into the input / output device 5 in real time while providing a service to the customer will be described with reference to Figures 7 and 8.

[0130] FIG. 7 is a diagram illustrating the operation and information flow of the system 1 in this case. In the above embodiment, as shown in S101 and S103 of FIG. 2 , the avatar robot 4 transmits customer image information and information about the avatar robot 4's movements to the information processing device 2, and the information processing device 2 stores this information for learning in the image-attribute learning model m1 and the attribute-movement learning model m2. Meanwhile, in the example of FIG. 7 , the avatar robot 4 transmits customer image information and information about the avatar robot 4's movements to the input / output device 5 (S200). In response, the input / output device 5, while displaying the customer image information, accepts from the user a designation of a time period for the customer image information to be used for learning and input of customer attribute information (described below with reference to FIG. 8 ). The input / output device 5 then transmits the input customer attribute information, the customer image information for the designated time period, and information about the avatar robot 4's movements for that time period to the information processing device 2 (S204 and S206). That is, the input / output device 5 may acquire various information from the avatar robot 4, extract information to be used for learning, and then transmit the extracted information to the information processing device 2. This makes it possible to avoid a burden on the communication capacity within the system 1 and a burden on the storage capacity of the information processing device 2.

[0131] FIG. 8 is a diagram illustrating an example of a display screen of the input / output device 5 in the example of FIG. 7 . In the above embodiment, the user operates the time interval designation tool d108 to designate a past time interval as the time interval of the customer image information to be used for learning. Meanwhile, in the example of the display screen of FIG. 8 , a recording start / stop button d200 is displayed instead of the time interval designation tool d108. The user can designate the start and end of the time interval of the customer image information to be used for learning by pressing the recording start / stop button d200. For example, if the user presses the recording start / stop button d200 when the customer image information for "12:00" is displayed in the image d100, and then presses the recording start / stop button d200 again when the customer image information for "12:20" is displayed in the image d100, customer image information corresponding to the 20 minutes from "12:00" to "12:20" is transmitted to the information processing device 2. At the same time, information about the movement of the avatar robot 4 within that time interval and the customer attribute information input in the customer attribute input area d101 are transmitted to the information processing device 2. The transmitted information is used in the information processing device 2 to learn the image-attribute learning model m1 and the attribute-movement learning model m2.

[0132] [Learning User Gaze] In the above embodiment, the image-attribute learning model m1 may further learn the relationship between customer image information and the user's gaze toward the image related to the customer image information. In this case, the information processing device 2 may, for example, determine weights corresponding to one or more coordinates or regions of the image related to the customer image information based on the user's attention to each of the coordinates or regions (e.g., the amount of time the user gazed at each coordinate or region), and the image-attribute learning model m1 may further learn the relationship between the customer image information and the weights corresponding to each of the one or more coordinates or regions of the image related to the customer image information. This relationship corresponds to information indicating which part of the image related to the customer image information the user focuses on when determining the customer's attribute information. In other words, the image-attribute learning model m1 may learn "where the user focuses on an image related to a given customer image information, and what customer attribute information is determined based on that information." This configuration enables the information processing device 2 to identify customer attribute information based on appropriate parts of the image related to the customer image information, thereby enabling the identification of customer attribute information with greater accuracy. In addition, information regarding the user's gaze may be acquired by the information processing device 2 via an imaging device provided in the input / output device 5, or may be acquired by the information processing device 2 via another device (e.g., a dedicated gaze observation device).

[0133] [Configuration of Learning Model] The image-attribute learning model m1 and the attribute-action learning model m2 may each include one or more learning models.

[0134] [Configuration of Avatar Robot 4] In the above embodiment, the avatar robot 4 has been described as including a moving device, but this is not limited thereto. The avatar robot 4 in the present disclosure may be replaced with a general-purpose or dedicated terminal device such as a tablet terminal. In this case, the terminal device may be configured to be movable by being connected to an external moving device.

[0135] [User Skill Level] In the above embodiment, the user may have a certain level of skill in providing services. If the user is an expert, higher quality training data is learned in the image-attribute learning model m1 and the attribute-action learning model m2, allowing the avatar robot 4 to provide higher quality services to customers even when operating autonomously.

[0136] [Learning in Virtual Space (Digital Twin)] In the above embodiment, the user operates the avatar robot 4 in a real space, such as airport AR10. However, this is not limiting. Specifically, the user may operate an avatar object corresponding to the avatar robot 4 in a virtual space (particularly, the virtual space of a digital twin) to virtually provide a service to a customer object corresponding to a customer in the virtual space. The system 1 may generate a learning model by treating information in the virtual space equivalent to information in the real space. More specifically, the system 1 may execute the following steps: acquiring attribute information of a customer object in the virtual space; and inputting attribute information of a customer in the real space into a learning model that has learned the relationship between the attribute information of each of one or more other customer objects in the virtual space and actions that the user has caused the avatar object in the virtual space to perform using the operation terminal device 3 to virtually provide a service to each of the one or more other customer objects.

[0137] [Operation Method of Avatar Robot 4] In the above embodiment, the avatar robot 4 is described as performing either an action based on a user's operation or an action not based on a user's operation (i.e., an action based on a learning model), but this is not limiting. Specifically, the avatar robot 4 may operate autonomously and also operate based on a user's operation. For example, the avatar robot 4 may autonomously perform an action related to movement based on a learning model, and perform an action related to voice output based on a user's operation.

[0138] [Utilization of Information Acquired by Information Processing Device 2] The information acquired by the information processing device 2 can also be used, for example, for customer service training. Specifically, the image-attribute learning model m1 learns the relationship between customer image information and customer attribute information. Based on this relationship, it is possible to create teaching materials that show, for example, "what kind of trouble a customer with a specific appearance is having." Furthermore, the attribute-motion learning model m2 learns the relationship between customer attribute information and information related to the movements of the avatar robot 4. Based on this relationship, it is possible to create teaching materials that show, for example, "how to serve a customer with a specific trouble." In summary, based on the information acquired by the information processing device 2, it is also possible to create teaching materials that show, for example, "how to serve a customer with a specific appearance." In this way, the information processing device 2 makes it possible to formalize customer service know-how, which traditionally tends to be tacit knowledge and personalized, as data.

[0139] [Estimation of Utterance Content and Feedback Thereon] In one embodiment, the input / output device 5 displays estimated data of the content of a user's utterance and accepts input of feedback from the user regarding the estimated data. In one example, the estimated data displayed on the input / output device 5 is the result of transcription of the user's utterance, and the feedback thereon may include corrections for discrepancies between the transcription result and the user's actual utterance. In another example, the estimated data displayed on the input / output device 5 includes estimated data regarding the outline, purpose, etc. of the utterance, and the feedback thereon may include corrections for discrepancies between the estimated data and the actual outline, purpose, etc. of the utterance. Note that the process of generating estimated data of the content of the utterance from audio data (e.g., waveform data) of the user's utterance (e.g., a transcription process and a process for estimating the outline, purpose, etc. of the utterance) may be executed by, for example, the information processing device 2, the input / output device 5, or another device. Furthermore, the process may be executed based on an appropriate speech recognition algorithm and a natural language processing algorithm.

[0140] In the above embodiment, the attribute-action learning model m2 is described as being capable of learning, as training data, data that associates information indicating what kind of voice the user has made the avatar robot 4 output with customer attribute information. In one embodiment, the information indicating what kind of voice the user has made the avatar robot 4 output may include estimated data on the content of the user's utterance (e.g., transcription results, estimated data regarding the outline and purpose of the utterance, etc.), and may also include information obtained by applying feedback to the estimated data (e.g., transcription results corrected by the user, the outline and purpose of the utterance corrected by the user, etc.). The information processing device 2 can configure the attribute-action learning model m2 using high-quality training data by having the attribute-action learning model m2 learn information obtained by applying feedback to the estimated data.

[0141] The information processing device 2 may store and output information regarding feedback on the estimated data in association with each of a plurality of users. In one embodiment, the information regarding feedback on the estimated data may include information regarding the number of times feedback has been provided. In one embodiment, the information regarding feedback on the estimated data may include information regarding the number of times or amount of time information obtained by applying feedback has been used as learning data for the attribute-action learning model m2. That is, the information processing device 2 may store and output information indicating how many useful feedback each of a plurality of users has provided to the estimated data. As described above, the information obtained by applying feedback to the estimated data may be high-quality training data. With this configuration, it is possible to identify users who are actively generating such high-quality training data from a plurality of users, and incorporate their achievements into, for example, personnel evaluations.

[0142] [Learning Subjective Information] The image-attribute learning model m1 and the attribute-action learning model m2 may further learn information based on the user's subjectivity in providing services via the avatar robot 4. In one embodiment, the information based on the user's subjectivity may include information about the customer's emotions based on the user's subjectivity. The customer's emotions based on the user's subjectivity may be, for example, emotions such as "I made them happy," "I scared them," or "I wasn't convinced." In one embodiment, the information based on the user's subjectivity may include information about the user's own emotions. The user's own emotions may be, for example, emotions such as "I was happy," "I felt scared," or "I was anxious."

[0143] In the above embodiment, the image-attribute learning model m1 has been described as learning training data in which customer image information and customer attribute information are associated. Also, in the above embodiment, the attribute-action learning model m2 has been described as learning training data in which customer attribute information and information on the actions of the avatar robot 4 for providing a service to the customer are associated. In one embodiment, at least one of these training data may be associated with information based on the user's subjective opinion.

[0144] The information based on the user's subjective opinion can be used as weights for the training data. The weights for the training data can be the magnitude (e.g., the step size when updating the weights between nodes) or the direction (e.g., whether to decrease or increase the weights between nodes) that influences the input / output process of the image-attribute learning model m1 or the attribute-action learning model m2 when the training data is used for the model.

[0145] In one example, if information based on a user's subjectivity associated with certain training data indicates a positive emotion (e.g., "happy," "they were convinced," etc.), the training data can be assigned a positive weight (or a weight with a relatively large value) and learned by the image-attribute learning model m1 or the attribute-action learning model m2. In another example, if information based on a user's subjectivity associated with certain training data indicates a negative emotion (e.g., "sad," "they seemed angry," etc.), the training data can be assigned a negative weight (or a weight with a relatively small value) and learned by the image-attribute learning model m1 or the attribute-action learning model m2.

[0146] Training data that is not associated with information based on the user's subjectivity can be assigned a neutral weight and learned by the image-attribute learning model m1 or the attribute-action learning model m2. Similarly, training data that is associated with information based on the user's subjectivity and that indicates a neutral emotion can be assigned a neutral weight and learned by the image-attribute learning model m1 or the attribute-action learning model m2.

[0147] According to this configuration, the training data regarding the provision of services when a user feels positive emotions, or the training data regarding the provision of services when a customer feels positive emotions, is positively reflected in the image-attribute learning model m1 and the attribute-action learning model m2, which strengthens the tendency of the image-attribute learning model m1 and the attribute-action learning model m2 to execute output in accordance with the training data.

[0148] In contrast, training data relating to the provision of services when a user has negative emotions, or training data relating to the provision of services when a customer has negative emotions, will be negatively reflected in the image-attribute learning model m1 and the attribute-action learning model m2. As a result, the image-attribute learning model m1 and the attribute-action learning model m2 will be more likely to output inconsistent with the training data or to output in contradiction to the training data.

[0149] 7. Other Embodiments The following aspects are also included within the scope of the present disclosure. [Supplementary Note 1] A system 1 according to one aspect of the present disclosure includes an acquisition unit 100 that acquires attribute information of a customer, a storage unit 14 that stores a learning model (attribute-action learning model m2) that has learned the relationship between the attribute information of each of one or more other customers and actions that a user causes an avatar robot 4 to perform using an operation terminal device 3 in order to provide a service to each of the one or more other customers, and an execution unit 104 that inputs the attribute information of customers into the learning model and causes the avatar robot 4 to perform an action corresponding to the attribute information.

[0150] [Supplementary Note 2] In the system 1 described in Supplementary Note 1, the customer attribute information may include information regarding at least a part of the customer's age, gender, line of sight, behavior, companions, nationality, belongings, satisfaction level, and location.

[0151] [Supplementary Note 3] In the system 1 described in Supplementary Note 1 or Supplementary Note 2, the attribute information of a customer may be identified based on image information generated by an imaging device of the avatar robot 4 photographing the customer.

[0152] [Supplementary Note 4] The system described in Supplementary Note 3 includes a first device that displays an image related to the image information, and a second device that measures a user's gaze toward the image related to the image information, and the customer attribute information may be identified further based on another learning model that has learned the relationship between the image information and the user's gaze toward the image related to the image information.

[0153] [Supplementary Note 5] In the system 1 according to any one of Supplementary Note 1 to Supplementary Note 4, the avatar robot 4 may be a device that can move within a predetermined area.

[0154] [Appendix 6] In the system 1 described in Appendix 5, the action that the user causes the avatar robot 4 to perform using the operation terminal device 3 may include moving the avatar robot 4 toward each of one or more other customers, and the action that the execution unit 104 causes the avatar robot 4 to perform may include moving the avatar robot 4 toward the customer in accordance with the customer's attribute information.

[0155] [Appendix 7] In the system 1 described in any one of Appendices 1 to 6, the avatar robot 4 may be equipped with a display device that displays an avatar, the action that the user causes the avatar robot 4 to perform using the operation terminal device 3 may include changing the facial expression of the avatar, and the action that the execution unit 104 causes the avatar robot 4 to perform may include changing the facial expression of the avatar to an expression that corresponds to the customer's attribute information.

[0156] [Supplementary Note 8] In the system 1 described in Supplementary Note 7, the user may change the facial expression of the avatar based on image information generated by an imaging device of the operation terminal device 3 capturing an image of the user.

[0157] [Appendix 9] In the system 1 described in any one of Appendices 1 to 8, the action that the user causes the avatar robot 4 to perform using the operation terminal device 3 may include outputting sound from the sound output device of the avatar robot 4, and the action that the execution unit 104 causes the avatar robot 4 to perform may include outputting sound corresponding to the customer's attribute information from the sound output device of the avatar robot 4.

[0158] [Appendix 10] In the system 1 described in Appendix 9, a user may use the operation terminal device 3 to output sound from the sound output device of the avatar robot 4 based on sound information generated by the sound input device of the operation terminal device 3 acquiring the user's voice.

[0159] [Appendix 11] The system 1 described in any one of Appendices 1 to 10 may further include a receiving unit 102 that receives information regarding an operation input by a user to the operation terminal device 3, and the execution unit 104 may cause the avatar robot 4 to perform an action corresponding to the operation received by the receiving unit 102 instead of or in addition to an action corresponding to the customer's attribute information, in response to a predetermined condition being satisfied.

[0160] [Appendix 12] In the system 1 described in Appendix 11, notification to the customer may be limited as to whether the execution unit 104 is causing the avatar robot 4 to perform an action corresponding to the customer's attribute information or an action corresponding to the user's operation.

[0161] [Supplementary Note 13] In the system 1 according to any one of Supplementary Note 1 to Supplementary Note 12, the avatar robot 4 may be a device arranged in a facility related to an airport.

[0162] [Appendix 14] The system 1 according to another aspect of the present disclosure acquires information regarding the relationship between the attribute information of each of one or more customers and the actions that the user causes the avatar robot 4 to perform using the operation terminal device 3 in order to provide a service to each of the one or more customers.

[0163] [Appendix 15] A program according to another aspect of the present disclosure causes the computer of the avatar robot 4 to function as an acquisition means for acquiring customer attribute information, and an execution means for executing an action corresponding to the attribute information by inputting the customer attribute information into a learning model that has learned the relationship between the attribute information of each of one or more other customers and the action that a user has caused the avatar robot 4 to perform using the operation terminal device 3 in order to provide a service to each of the one or more other customers.

[0164] [Appendix 16] An information processing method according to another aspect of the present disclosure includes a step of causing a system 1 including an avatar robot 4 to acquire customer attribute information, and a step of causing the avatar robot 4 to perform an action corresponding to the customer attribute information by inputting the customer attribute information into a learning model that has learned the relationship between the attribute information of each of one or more other customers and the actions that a user has caused the avatar robot 4 to perform using an operation terminal device 3 to provide a service to each of the one or more other customers.

[0165] REFERENCE SIGNS LIST 1 System 2 Information processing device 3 Operation terminal device 4 Avatar robot 10 Control unit 14 Storage unit 70 Computer 100 Acquisition unit 100a Image information acquisition unit 100b Attribute information acquisition unit 100c Audio information acquisition unit 102 Reception unit 104 Execution unit

Claims

1. A system comprising: an acquisition unit that acquires customer attribute information; a memory unit that stores a learning model that has learned the relationship between the attribute information of each of one or more other customers and actions that a user has caused an avatar robot to perform using an operation terminal device in order to provide a service to each of the one or more other customers; and an execution unit that inputs the customer's attribute information into the learning model, thereby causing the avatar robot to perform an action corresponding to the attribute information.

2. The system according to claim 1, wherein the customer attribute information includes information regarding at least a portion of the customer's age, gender, gaze, behavior, companions, nationality, belongings, satisfaction level, and location.

3. The system according to claim 1, wherein the customer attribute information is identified based on image information generated by an imaging device of the avatar robot photographing the customer.

4. The system described in claim 3, comprising: a first device that displays an image related to the image information; and a second device that measures the user's gaze with respect to the image related to the image information, wherein the customer attribute information is identified further based on another learning model that has learned the relationship between the image information and the user's gaze with respect to the image related to the image information.

5. The system according to claim 1, wherein the avatar robot is a device capable of moving within a predetermined area.

6. The system described in claim 5, wherein the action that the user causes the avatar robot to perform using the operation terminal device includes moving the avatar robot toward each of the one or more other customers, and the action that the execution unit causes the avatar robot to perform includes moving the avatar robot toward the customer in accordance with attribute information of the customer.

7. The system described in claim 1, wherein the avatar robot is provided with a display device that displays an avatar, the actions that the user causes the avatar robot to perform using the operation terminal device include changing the facial expression of the avatar, and the actions that the execution unit causes the avatar robot to perform include changing the facial expression of the avatar to an expression that corresponds to the attribute information of the customer.

8. The system according to claim 7, wherein the user changes the facial expression of the avatar based on image information generated by an imaging device of the operation terminal device photographing the user.

9. The system described in claim 1, wherein the action that the user causes the avatar robot to perform using the operation terminal device includes outputting sound from an audio output device of the avatar robot, and the action that the execution unit causes the avatar robot to perform includes outputting sound corresponding to the customer attribute information from the audio output device of the avatar robot.

10. The system described in claim 9, wherein the user uses the operation terminal device to output voice from the voice output device of the avatar robot based on voice information generated by an audio input device of the operation terminal device acquiring the voice of the user.

11. The system described in claim 1, further comprising a receiving unit that receives information regarding an operation input by the user to the operation terminal device, wherein the execution unit causes the avatar robot to perform an action corresponding to the operation instead of or in addition to an action corresponding to the customer's attribute information when a predetermined condition is satisfied.

12. The system described in claim 11, wherein notification to the customer is limited as to whether the execution unit is causing the avatar robot to perform an action corresponding to the customer's attribute information or an action corresponding to the operation.

13. The system according to any one of claims 1 to 12, wherein the avatar robot is a device located in an airport-related facility.

14. A system that acquires information regarding the relationship between attribute information of each of one or more customers and actions that a user causes an avatar robot to perform using an operation terminal device in order to provide services to each of the one or more customers.

15. A program that causes an avatar robot's computer to function as: an acquisition means for acquiring customer attribute information; and an execution means for executing an action corresponding to the attribute information by inputting the customer's attribute information into a learning model that has learned the relationship between the attribute information of each of one or more other customers and the action that a user has caused the avatar robot to perform using an operation terminal device in order to provide a service to each of the one or more other customers.

16. An information processing method in which a system including an avatar robot executes the following steps: acquiring customer attribute information; and causing the avatar robot to perform an action corresponding to the attribute information by inputting the customer attribute information into a learning model that has learned the relationship between the attribute information of each of one or more other customers and an action that a user has caused the avatar robot to perform using an operation terminal device in order to provide a service to each of the one or more other customers.

Citation Information

Patent Citations

  • Information processing device and control method

    JP2022000794A

  • Robot reacting on basis of user behavior and control method therefor

    EP3725470A1

  • Action control system and program

    JP2016012340A

  • Communication robot and control program for communication robot

    JP2020067799A

  • Robot interaction method and device

    JP2020534567A