Server device, server device control method, and program
The server device addresses the limitation of existing technologies by storing past encounters with associated ages and generating virtual persons for conversation, enhancing user satisfaction through personalized reminiscing experiences.
Patent Information
- Application Number
- JP2024004978
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-30
AI Technical Summary
Existing technologies, such as those described in Patent Document 1, do not allow users to select the specific age or time when they wish to converse with a person met in the past, failing to meet the needs of elderly users who often want to reminisce with friends or coworkers from particular periods in their lives.
A server device that stores facial images of past encounters in association with the age at which they were acquired, enabling the generation of a virtual person matching the desired age for conversation, thereby allowing users to converse with a virtual representation of individuals from their past.
Enhances user satisfaction by facilitating conversations with virtual persons that reflect specific periods in their lives, improving the experience of reminiscing with friends or coworkers from the past.
Smart Images

Figure 2025110929000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a server device, a method for controlling a server device, and a program. [Background technology]
[0002] There is technology that allows you to virtually converse with a deceased loved one.
[0003] For example, Patent Document 1 states that it provides a method and device for virtual dialogue with a deceased person, which allows the dialogue to proceed while the dialogue between the dialogue participant and the deceased person mutually influences each other. Patent Document 1 discloses a method for virtual dialogue with a deceased person, in which the dialogue participant operates a computer to engage in dialogue with the deceased. In this virtual dialogue method, the computer synthesizes a virtual image and a virtual voice of the deceased based on at least data related to the deceased, and the dialogue proceeds through the virtual image and / or the virtual voice. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2002-024371 Summary of the Invention [Problem to be solved by the invention]
[0005] Some users wish to converse with people they met in the past. This need is particularly prevalent among elderly users. Users have periods (ages) in their lives when they were particularly happy and fulfilled. Elderly users who wish to converse with people they met in the past often wish to converse with friends, coworkers, and others they met during these periods.
[0006] Patent Document 1 does not disclose the ability to select the time (age) when the user met the person with whom they wish to converse, and therefore cannot meet the needs of the elderly and the like.
[0007] A primary object of the present invention is to provide a server device, a server device control method, and a program that contribute to improving the satisfaction of users who wish to converse with people they have met in the past. [Means for solving the problem]
[0008] According to a first aspect of the present invention, there is provided a server device comprising: a storage means for storing facial images of a plurality of people that a user has met in the past, in association with the age of the user when each facial image was acquired; and a conversation control means for, when the user wishes to converse with a virtual person, acquiring the age of the user when the user met the person with whom the user wishes to converse, identifying a person corresponding to the acquired age from the plurality of stored people, generating the virtual person using the facial image of the identified person, and conversing with the user via the generated virtual person, when the user wishes to converse with a virtual person.
[0009] According to a second aspect of the present invention, there is provided a control method for a server device, comprising: a storage step of storing facial images of a plurality of people that a user has met in the past, in association with the age of the user at the time each facial image was acquired; and a conversation control step of, when the user wishes to converse with a virtual person, acquiring the age of the user when the user met the person with whom the user wishes to converse, identifying a person corresponding to the acquired age from the plurality of stored people, generating the virtual person using the facial image of the identified person, and conversing with the user via the generated virtual person.
[0010] According to a third aspect of the present invention, a computer mounted on a server device stores, in association with each other, face images of a plurality of persons that a user has encountered in the past and the age of the user when each face image was acquired. When the user wishes to converse with a virtual person, the age of the user when the user encountered the person with whom the user wishes to converse is acquired, a person corresponding to the acquired age is specified from among the plurality of stored persons, a virtual person is generated using the face image of the specified person, and the user converses with the virtual person. A program for causing the computer to execute a storage process and a conversation control process is provided.
Advantages of the Invention
[0011] According to each aspect of the present invention, a server device, a control method for the server device, and a program are provided that contribute to improving the satisfaction of a user who wishes to converse with a person encountered in the past. Note that the effects of the present invention are not limited to the above. Instead of or together with the above effects, other effects may be achieved by the present invention.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
MODE FOR CARRYING OUT THE INVENTION
[0013] First, an overview of an embodiment will be described. Note that the reference numerals in the drawings appended to this overview are for convenience of each element as an example to assist understanding, and the description of this overview is not intended to be limiting in any way. Also, unless otherwise specified, the blocks described in each drawing represent a configuration of functional units, not a configuration of hardware units. The connection lines between the blocks in each figure include both bidirectional and unidirectional ones. The one-way arrow schematically shows the flow of the main signal (data) and does not exclude bidirectionality. In this specification and the drawings, for elements that can be similarly described, duplicate descriptions may be omitted by assigning the same reference numerals.
[0014] The server device 100 according to an embodiment includes a storage unit 101 and a conversation control unit 102 (see FIG. 1). The storage unit 101 stores, in association with each other, the face images of a plurality of persons that the user has encountered in the past and the age of the user when each face image was acquired (step S1 in FIG. 2). When the user wishes to have a conversation with a virtual person, the conversation control unit 102 acquires the age of the user when the user met the person with whom the user wishes to have a conversation (step S2). The conversation control unit 102 identifies the person corresponding to the age acquired from among the plurality of stored persons, and generates a virtual person using the face image of the identified person (step S3). The conversation control unit 102 has a conversation with the user via the generated virtual person (step S4).
[0015] The server device 100 acquires the age (the age of the user; for example, the teens, twenties) in which the user wants to have a conversation going back in time. The server device 100 generates a virtual person from the face images of the persons who met the user in the past at the age of the user. The server device 100 has a conversation with the user via the virtual person. The user can have a conversation with a virtual person imitating friends or company colleagues at that time by inputting the age of a happy or fulfilling period in the past into the server device 10. As a result, the satisfaction of the user who wishes to have a conversation with the persons met in the past is improved.
[0016] Hereinafter, specific embodiments will be described in more detail with reference to the drawings.
[0017] [First Embodiment] The first embodiment will be described in more detail with reference to the drawings.
[0018] [Configuration of the System] As shown in FIG. 3, the information processing system according to the first embodiment includes a server device 10.
[0019] The server device 10 is managed by an operator who provides a "conversation service" to users. The conversation service is a service mainly targeted at the elderly and the like. The conversation service is a service in which a user (elderly person) converses with virtual characters imitating friends, acquaintances, etc. that the user has met in the past.
[0020] The server device 10 is a server that realizes the main functions of the conversation service. The server device 10 may be installed inside the building of the service provider that provides the conversation service, or may be installed on the network (in the cloud).
[0021] The user has a terminal 20. The user operates the terminal 20 to input various information into the server device 10 and the like, or acquire various information from the server device 10 and the like.
[0022] Also, when receiving the conversation service, the user wears a device for VR (Virtual Reality) or AR (Augmented Reality). For example, the user wears an HMD (Head Mounted Display) 21. The HMD 21 is equipped with a camera, a microphone, a speaker, etc. The HMD 21 supports gaze input and voice input, and the user can operate the HMD 21 using gaze and voice.
[0023] The devices shown in Fig. 3 are interconnected. Specifically, the server device 10 and the devices used by the user (terminal 20, HMD 21) are connected by wired or wireless communication means and configured to be able to communicate with each other.
[0024] 3 is an example and is not intended to limit the configuration. For example, the information processing system may include multiple server devices 10. The multiple server devices 10 may achieve load balancing and redundancy.
[0025] [General operation] Next, the general operation of the information processing system according to the first embodiment will be described.
[0026] <Create an account> To receive the conversation service, the user creates an account with the service provider. Specifically, the user operates the terminal 20 to access the server device 10. The user enters login information (ID, password), name, sex, date of birth, address, biometric information, etc. into the user registration site provided by the server device 10.
[0027] Examples of biometric information include data (features) calculated from physical characteristics unique to an individual, such as a face, fingerprint, voiceprint, veins, retina, and iris pattern. Alternatively, the biometric information may be image data such as a face image or fingerprint image. The biometric information may be any information that includes the user's physical characteristics. In the embodiment disclosed herein, a case where biometric information related to a person's "face" (a face image or features generated from a face image) is used will be described.
[0028] After acquiring the user's name, etc., the server device 10 generates a user ID for identifying the user. The server device 10 associates and stores the user ID, login information, name, gender, date of birth, biometric information, etc. The server device 10 stores the user ID, login information, name, gender, biometric information, etc. in a user management database. The user management database will be described in detail later.
[0029] By generating an account with the service provider, the user can enjoy the conversation service from the service provider.
[0030] <Registration of Conversation Partner Candidates> When an account is generated, the user registers with the server device 10 the persons who are candidates for conversation partners. Specifically, the user operates the terminal 20 to register (upload) image data such as photos and videos stored in the terminal 20 to the server device 10.
[0031] The server device 10 analyzes the uploaded image data and extracts the persons shown in the image data. For example, if 10 persons are shown in 100 photos (still image data), the server device 10 extracts the 10 persons. The server device 10 assigns a conversation partner ID for identifying the extracted persons to the persons.
[0032] Here, when the server device 10 extracts the persons shown in the image data, it memorizes the relationship between the extracted persons and the image data. For example, the server device 10 memorizes the fact that person A is shown in image 1, image 3, and image 5.
[0033] The server device 10 stores the information of conversation partner candidates in the conversation partner management database. Specifically, the server device 10 stores in association the conversation partner IDs of the persons extracted from the image data and the face images etc. of the persons. The details of the conversation partner management database will be described later.
[0034] Furthermore, the server device 10 identifies the attribute information (name, gender, the position of the person from the viewpoint of the user, etc.) of the persons (the persons extracted from the uploaded image data) stored in the conversation partner management database.
[0035] First, the server device 10 requests the user to input the name, sex, and position (e.g., friend, coworker, teacher, parent, child, etc.) of each person extracted from the image data from the user. In the above example, the server device 10 requests the user to input the name, sex, and position (relationship to the user; e.g., friend, coworker) of 10 people extracted from the image data.
[0036] For example, the server device 10 presents an image 1 in which a person A appears to the user, and acquires the name, sex, position, etc. of the person A from the user.
[0037] Furthermore, the server device 10 calculates the generation (age) of the user when each uploaded image data was acquired. For example, the server device 10 calculates the generation (teens, twenties, etc.) of the user when each image data was taken based on the photographic date and time (photographed date and time embedded in the image data) and the user's date of birth.
[0038] The server device 10 reflects the user's age calculated (specified) from the image data in the conversation partner management database. For example, if the image data in which person A appears was taken when the user was in his / her teens and twenties, the server device 10 sets the user's age to "teens" and "twenties" in the entry for person A.
[0039] By tagging in this way by the server device 10, the conversation partner management database stores, for each person, the position of the conversation partner candidate from the user's perspective and the time when the user met the conversation partner candidate (the user's generation).
[0040] <Conversation service> After completing the registration, the user wears the HMD 21 to receive the conversation service. The user logs in to their account using the HMD 21.
[0041] The server device 10 provides a conversation service in a virtual space to a user wearing an HMD 21. When starting the conversation service, the server device 10 acquires information about a person (a virtual person; an avatar) with whom the user wishes to have a conversation. Hereinafter, information about a person with whom the user wishes to have a conversation will be referred to as "conversation partner information."
[0042] First, the server device 10 acquires the age group (the user's age group) of the user who would like to go back in time and talk with acquaintances, friends, etc. from that time. Next, the server device 10 acquires the status of the conversation partner (for example, friend, colleague, etc.) that the user desires.
[0043] For example, the user may select "teens; friends" or "thirties; colleagues." Information such as "teens; friends" corresponds to conversation partner information.
[0044] The server device 10 identifies a person who matches the acquired conversation partner information from among the conversation partner candidates stored in the conversation partner management database. Furthermore, the server device 10 identifies image data (face images) that match the user's desired age group from among image data that shows the identified person.
[0045] For example, when the server device 10 acquires conversation partner information of "Teenager; Friend," the server device 10 identifies friends of the user when they were teenagers from among the conversation partner candidates stored in the conversation partner management database. Furthermore, the server device 10 identifies image data in which the friends from their teenage years are photographed.
[0046] The server device 10 generates a virtual conversation partner using the specified image data (face image). The server device 10 generates a virtual person (a so-called avatar) who will converse with the user.
[0047] For example, the server device 10 generates a virtual person from the specified image data using a generation AI (Artificial Intelligence). In the above example, the server device 10 generates a virtual person with the appearance of a teenager.
[0048] The server device 10 displays the generated virtual character on the HMD 21. The user has a conversation with the virtual character displayed via the HMD 21 (see FIG. 4). The server device 10 realizes the conversation between the user and the virtual character using a learning model (AI model) such as a language model.
[0049] Subsequently, details of each device included in the information processing system according to the first embodiment will be described.
[0050] [Server device] FIG. 5 is a diagram showing an example of the processing configuration (processing modules) of the server device 10 according to the embodiment disclosed in the present application. Referring to FIG. 5, the server device 10 includes a communication control unit 201, an account control unit 202, a registration control unit 203, a conversation control unit 204, and a storage unit 205.
[0051] The communication control unit 201 is a means for controlling communication with other devices. For example, the communication control unit 201 receives data (packets) from the terminal 20. Also, the communication control unit 201 transmits data to the terminal 20. The communication control unit 201 delivers the data received from other devices to other processing modules. The communication control unit 201 transmits the data acquired from other processing modules to other devices. In this way, other processing modules perform data transmission and reception with other devices via the communication control unit 201. The communication control unit 201 has a function as a receiving unit that receives data from other devices and a function as a transmitting unit that transmits data to other devices.
[0052] The account control unit 202 is a means for controlling the user's account. The account control unit 202 acquires the user's login information (ID, password), name, gender, date of birth, address, phone number, email address, biometric information (e.g., face image), etc. at the user registration site.
[0053] When acquiring the name, etc., the account control unit 202 generates a user ID for identifying the user. The user ID may be any information as long as it can uniquely identify the user. For example, the account control unit 202 may assign a unique value each time an account is generated and use it as the user ID.
[0054] Also, when acquiring a face image as biometric information, the account control unit 202 generates feature amounts from the acquired face image.
[0055] Regarding the generation process of the feature amounts, existing technologies can be used, so detailed descriptions thereof are omitted. For example, the account control unit 202 extracts eyes, nose, mouth, etc. as feature points from the face image. Thereafter, the account control unit 202 calculates the positions of each feature point and the distances between each feature point as feature amounts (generates a feature vector composed of a plurality of feature amounts).
[0056] The account control unit 202 stores the user ID, name, gender, date of birth, address, biometric information (face image, feature amounts), etc. in the user management database (see FIG. 6). Note that the user management database shown in FIG. 6 is an example and is not intended to limit the items to be stored.
[0057] The account control unit 202 authenticates a user who logs in to the account using the login information. Alternatively, the account control unit 202 may authenticate a user who logs in to the account by face authentication (biometric authentication).
[0058] The registration control unit 203 is a means for performing control related to the registration of candidates for a person who converses with the user (conversation partner candidates).
[0059] FIG. 7 is a flowchart showing an example of the operation of the registration control unit 203 according to the embodiment disclosed in the present application. Referring to FIG. 7, the operation of the registration control unit 203 will be described.
[0060] The registration control unit 203 acquires image data stored in the terminal 20 held by the user (step S101). For example, the registration control unit 203 displays a GUI (Graphical User Interface) or the like on the terminal 20 for acquiring image data stored inside the terminal 20. The registration control unit 203 acquires all or part of the image data stored inside the terminal 20.
[0061] The registration control unit 203 may acquire image data selected by the user from the image data stored in the terminal 20, or may acquire all image data stored in the terminal 20.
[0062] The registration control unit 203 analyzes the acquired image data (uploaded image data) and registers people appearing in the image data as conversation partner candidates in the conversation partner management database (registration of conversation partner candidates; step S102).
[0063] Specifically, the registration control unit 203 attempts to extract a face image from each image data.
[0064] Note that the facial image extraction process by the registration control unit 203 can use existing technology, and therefore detailed description thereof will be omitted. For example, the registration control unit 203 may extract a facial image (face region) from image data using a learning model trained by a CNN (Convolutional Neural Network). Alternatively, the registration control unit 203 may extract a facial image using a method such as template matching.
[0065] If the image data does not contain a facial image (if a facial image cannot be extracted), the registration control unit 203 discards the image data that does not contain the facial image. If the image data contains a facial image, the registration control unit 203 assigns an image ID to the image data that contains the facial image. The registration control unit 203 also assigns a facial image ID to the extracted facial image.
[0066] The registration control unit 203 stores the image ID and the face image ID in association with each other. For example, the registration control unit 203 stores the correspondence relationship between the image ID and the face image ID using the table information as shown in FIG. 8.
[0067] In addition, when a plurality of face images are included in one piece of image data, the registration control unit 203 may store the position of each face image in the image data (position in the image coordinate system) in the table information in addition to the image ID and the like. For example, consider the case where face image E and face image F are included in image 5 (see FIG. 9). In this case, the registration control unit 203 may store information such as face image E = (X1, Y1) and face image F = (X2, Y2) in the entry of image 5. Note that the center coordinates of the face image can be used as the position of the face image.
[0068] When the extraction process of the face image from each image data is completed, the registration control unit 203 generates feature amounts from the extracted face images.
[0069] Thereafter, the registration control unit 203 registers the person corresponding to the generated feature amount (biometric information) in the conversation partner management database (see FIG. 10). As shown in FIG. 10, the conversation partner management database stores the name, gender, position, biometric information (feature amount), age of the user at the time of acquiring the image data, combination of the image data and the face image in which the person appears, and the like of the person extracted from the image data. Note that the conversation partner management database shown in FIG. 10 is an example and is not intended to limit the items to be stored.
[0070] The registration control unit 203 determines whether or not the person corresponding to the generated feature amount is registered in the conversation partner management database. Specifically, the registration control unit 203 sets the generated feature amount as the collation side and the feature amount stored in the conversation partner management database as the registration side, and executes one-to-N collation (N is an integer; the same applies hereinafter).
[0071] The registration control unit 203 calculates the similarity between the feature amount to be collated and each of the feature amounts on the registration side. For the similarity, the chi-square distance, the Euclidean distance, or the like can be used. Note that the greater the distance, the lower the similarity, and the closer the distance, the higher the similarity.
[0072] If there is no feature amount whose similarity with the feature amount to be collated among the feature amounts registered in the conversation partner management database is equal to or greater than a predetermined value, the registration control unit 203 determines that the collation process has failed.
[0073] If there is a feature amount whose similarity with the feature amount to be collated among the feature amounts registered in the conversation partner management database is equal to or greater than a predetermined value, the registration control unit 203 determines that the collation process has succeeded. When the collation process succeeds, the user of the entry with the highest similarity is identified as the person corresponding to the extracted face image.
[0074] When the collation process fails, the registration control unit 203 determines that a new conversation partner has been detected and generates a conversation partner ID. Further, the registration control unit 203 adds a new entry to the conversation partner management database.
[0075] The registration control unit 203 stores the generated conversation partner ID and the feature amount used as the collation side during the collation process in the added entry. In addition, the registration control unit 203 stores the combination of the face image ID of the face image corresponding to the stored feature amount and the image ID of the image data including the face image in the added entry (see the lowermost row of FIG. 10). Note that at the stage when a new entry is added to the conversation partner management database, the name and the like of the added person are not described in the database.
[0076] When the collation process succeeds, the registration control unit 203 further adds the combination of the face image ID of the face image corresponding to the feature amount used in the collation process and the image ID of the image data including the face image to the entry specified by the collation process.
[0077] The registration control unit 203 executes the above registration process (face image allocation process) for each face image extracted from the image data uploaded to the server device 10.
[0078] When registering a conversation partner candidate in the database, the registration control unit 203 registers the attribute information of the person registered in the conversation partner management database (step S103 in FIG. 7). The registration control unit 203 tags the persons registered in the conversation partner management database.
[0079] First, the registration control unit 203 requests the user to input information such as the name, gender, and the position of the person from the user's perspective (e.g., friend, colleague, etc.) for each person registered in the conversation partner management database. For example, the registration control unit 203 displays a GUI as shown in FIG. 11 on the terminal 20 to obtain the name, gender, position, etc.
[0080] The registration control unit 203 stores the obtained name, gender, position, etc. in the conversation partner management database. For information that cannot be obtained, the registration control unit 203 may set the corresponding entry and field in the conversation partner management database to blank.
[0081] Furthermore, the registration control unit 203 calculates the age of the user at the time when the user had communication with each person registered in the conversation partner management database. Specifically, the registration control unit 203 uses the date and time when the face images (image data including face images) of each of the plurality of persons stored in the conversation partner management database were obtained and the user's date of birth to calculate the age of the user when the face images of each of the plurality of persons were obtained.
[0082] The registration control unit 203 stores the calculated age of the user in the corresponding person entry (entry in the conversation partner management database). More specifically, the registration control unit 203 stores the calculated age of the user in the user age field corresponding to the image data.
[0083] The registration control unit 203 registers, in the conversation partner management database, for each person extracted from the image data, the position of the conversation partner candidate as seen by the user and the period during which the user met the conversation partner candidate (the age of the user).
[0084] In this way, the registration control unit 203 acquires a plurality of image data from the terminal 20 possessed by the user. The registration control unit 203 extracts a plurality of persons that the user has met in the past based on the face images included in each of the acquired plurality of image data. The registration control unit 203 acquires the attribute information of each of the extracted plurality of persons from the user. More specifically, the registration control unit 203 acquires at least the name of each of the extracted plurality of persons and the position of each of the extracted plurality of persons as seen by the user.
[0085] The conversation control unit 204 is a means for controlling the conversation service provided to the user. The conversation control unit 204 provides the conversation service to the user who has logged in to its own account.
[0086] When the user wishes to talk to a virtual person, the conversation control unit 204 acquires the age of the user who met the person the user wishes to talk to. The conversation control unit 204 identifies the person corresponding to the acquired age from among the plurality of persons stored in the conversation partner management database. The conversation control unit 204 generates a virtual person who talks to the user in the virtual space using the face image of the identified person. The conversation control unit 204 talks to the user via the generated virtual person.
[0087] FIG. 12 is a flowchart showing an example of the operation of the conversation control unit 204 according to the embodiment disclosed in the present application. Referring to FIG. 12, the operation of the conversation control unit 204 will be described.
[0088] When providing the conversation service, the conversation control unit 204 acquires information (conversation partner information) about the person (virtual conversation partner) that the user wishes to talk to (step S201).
[0089] For example, the conversation control unit 204 acquires the age of the user who wishes to talk with acquaintances, friends, relatives, etc. from that time using a GUI such as that shown in Fig. 13. For example, a user who wishes to talk with a friend from elementary school selects "teens."
[0090] Thereafter, the conversation control unit 204 acquires the user's desired conversation partner's status (e.g., friend, colleague, etc.) using a GUI such as that shown in Fig. 14. For example, a user who wishes to talk with a friend from elementary school selects "friend." Alternatively, a user who wishes to talk with a teacher from elementary school selects "teacher."
[0091] Upon acquiring conversation partner information (e.g., the user's age and the conversation partner's position), the conversation control unit 204 identifies a conversation partner that matches the conversation partner information from among the conversation partner candidates registered in the conversation partner management database (conversation partner identification; step S202).
[0092] The conversation control unit 204 refers to the conversation partner management database and extracts conversation partner candidates for which the age of the user included in the conversation partner information is set. Furthermore, the conversation control unit 204 extracts conversation partner candidates for which the position included in the conversation partner information is set from among the extracted people.
[0093] When there are multiple candidates that match the conversation partner information, the conversation control unit 204 extracts one person from the multiple candidates. For example, the conversation control unit 204 may extract one conversation partner candidate based on a predetermined rule, or may extract conversation partner candidates randomly. For example, the conversation control unit 204 may select as a conversation partner a person with a large amount of image data (face images) registered. The fact that a large amount of image data is registered indicates that the person has a close relationship with the user.
[0094] If a conversation partner candidate corresponding to the conversation partner information is not registered in the conversation partner management database, the conversation control unit 204 notifies the user to that effect and acquires new conversation partner information. For example, the conversation control unit 204 acquires new conversation partner information using a GUI similar to FIG. 13 or the like.
[0095] When the conversation partner is identified, the conversation control unit 204 generates a virtual character corresponding to the identified conversation partner (step S203).
[0096] For example, the conversation control unit 204 generates a virtual character of the conversation partner using the face image of the identified conversation partner and an image generation AI. The conversation control unit 204 generates a virtual character having the face of the identified conversation partner.
[0097] For example, the conversation control unit 204 inputs the face image and information for defining (designating) the virtual character generated using the face image to the image generation AI. For example, if the conversation partner information is "teenager; friend" and the conversation partner is male, the conversation control unit 204 inputs an instruction of "male in his teens" to the image generation AI together with the face image. The image generation AI generates a virtual character having the acquired face and the appearance of a male in his teens.
[0098] When the virtual character is generated, the conversation control unit 204 displays the generated virtual character on the HMD 21.
[0099] When the virtual character is displayed on the HMD 21, the conversation control unit 204 conducts a conversation with the user through the virtual character (step S204).
[0100] The conversation control unit 204 acquires the user's speech (voice data) via the HMD 21. The conversation control unit 204 converts the voice data into text data. The conversation control unit 204 generates a query to be input to the language model from the text data. The conversation control unit 204 inputs the generated query to a pre-prepared language model and acquires a reply (response) of the virtual character to the user's speech.
[0101] The language model may be, for example, what is called a large language model (LLM), but is not limited to this.
[0102] A language model is a machine learning model (also called a generative model) that takes language as input and outputs language. A language model learns the relationships between words in a sentence and generates related strings from a target string. By using a language model that has been trained on sentences and sentences from various contexts, it is possible to generate related strings with appropriate content that are related to the target string.
[0103] For example, a case where a language model is used in question answering will be described. The language model receives an input question such as "What kind of country is Japan?" as a target string. The language model generates a string such as "Japan is an island country in the northern hemisphere" as an answer to the question.
[0104] The learning method of the language model is not particularly limited, but as an example, the language model may be learned to output at least one sentence including an input string. As a specific example, the language model is a GPT (Generative Pre-Transformer) that outputs a sentence including an input string by predicting a string that is likely to follow the input string.
[0105] For example, when the conversation control unit 204 acquires a user's utterance, "Yamada-kun, how are you doing?", it generates a query to be input to a language model based on the utterance. For example, the conversation control unit 204 generates a query such as "Please answer the question, 'Yamada-kun, how are you doing?'" and inputs it to the language model.
[0106] Alternatively, the conversation control unit 204 may generate a query to be input into the language model using conversation partner information acquired from the user or user attributes (e.g., name). For example, in the above example, the conversation control unit 204 generates a query such as "In response to the question, 'Yamada, how are you?' please answer on the assumption that the asker is a teenager named Suzuki and the answerer is a male friend of the asker."
[0107] The conversation control unit 204 converts the text data (answer) acquired from the language model into voice data and transmits it to the HMD 21. For example, when the conversation control unit 204 receives the answer "I'm doing well," it transmits the voice data of the answer to the HMD 21.
[0108] The conversation control unit 204 acquires a response to the user's utterance acquired from the HMD 21 from the language model, and transmits the voice data of the acquired response to the HMD 21. By repeating this control, the conversation control unit 204 realizes a conversation between the user and a virtual person (an avatar corresponding to a conversation partner selected by the user).
[0109] In this way, the conversation control unit 204 acquires the age of the user who met the person with whom the user wishes to converse and the attribute information of the person with whom the user wishes to converse. The conversation control unit 204 identifies a person corresponding to the acquired age and attribute information from among multiple people stored in the conversation partner management database. The conversation control unit 204 generates a virtual character corresponding to the identified person. The conversation control unit 204 uses a language model to control the conversation between the user and the virtual character.
[0110] The storage unit 205 is a means for storing information necessary for the operation of the server device 10. The storage unit 205 stores face images of each of a plurality of people whom the user has met in the past, in association with the age of the user when each face image was acquired. Furthermore, the storage unit 205 stores face images of each of a plurality of people, in association with the age and attribute information of the user.
[0111] [Device] A detailed description of the terminal 20 will be omitted. Examples of the terminal 20 include mobile terminal devices such as smartphones, mobile phones, game machines, tablets, and computers (personal computers, notebook computers). The terminal 20 can be any device or apparatus as long as it can receive a user's operation and communicate with the server device 10.
[0112] [HMD] A detailed description of the HMD 21 will be omitted. An existing device can be used for the HMD 21. The HMD 21 displays a virtual character generated by the server device 10 on a display. The HMD 21 acquires the user's voice (voice data) using a microphone and transmits the voice data to the server device 10. The HMD 21 acquires the voice data of the virtual character from the server device 10 and outputs the acquired voice data from a speaker.
[0113] [Operation of the System] Subsequently, the operation of the information processing system according to the first embodiment will be described.
[0114] FIG. 15 is a sequence diagram showing an example of the operation of the information processing system according to the embodiment disclosed in the present application. With reference to FIG. 15, the operation of the information processing system according to the first embodiment will be described.
[0115] At the start of the conversation service, the server device 10 acquires conversation partner information from the HMD 21. The HMD 21 transmits the conversation partner information to the server device 10 in response to a user's operation (step S01).
[0116] The server device 10 uses the acquired conversation partner information to identify a conversation partner desired by the user from among a plurality of conversation partner candidates stored in the conversation partner management database (step S02).
[0117] The server device 10 generates a virtual character to be the user's conversation partner using the acquired conversation partner information and the like (step S(03)).
[0118] The server device 10 converses with the user using the generated virtual character (step S04).
[0119] Subsequently, a modification example according to the first embodiment will be described.
[0120] <Modification Example 1> In the above embodiment, the registration control unit 203 has been described for the case of acquiring the name, gender, position, etc. of the conversation partner candidate as attribute information. However, the registration control unit 203 may acquire other information instead of or in addition to the above attribute information. For example, the registration control unit 203 may acquire the age of the conversation partner candidate from the user.
[0121] Alternatively, the registration control unit 203 may acquire more detailed information of the acquired attribute information. For example, when the user selects "friend" as the position of the conversation partner candidate, the registration control unit 203 may acquire information such as "elementary school friend" and "junior high school friend" (information on the group to which the user and the conversation partner candidate belong). Alternatively, when the user selects "company colleague" as the position of the conversation partner candidate, the registration control unit 203 may acquire the specific name of the company, etc.
[0122] <Modification Example 2> The method for specifying the conversation partner described in the above embodiment is an example, and any method can be adopted.
[0123] For example, when a plurality of conversation partner candidates that match the conversation partner information acquired from the user are stored in the conversation partner management database, the conversation control unit 204 may display a list of the plurality of candidates and make them selectable by the user. For example, the conversation control unit 204 may display a GUI as shown in FIG. 16 on the HMD 21 and acquire the conversation partner that the user wishes to converse with.
[0124] Alternatively, the conversation control unit 204 may display on the HMD 21 the number of conversation partner candidates corresponding to the age selected by the user and the position of the conversation partner in order to assist the user in selecting a conversation partner. For example, the conversation control unit 204 may display a screen as shown in FIG. 17 on the HMD 21.
[0125] Alternatively, the conversation control unit 204 may provide an interface that allows the user to directly input information about a desired conversation partner. For example, the conversation control unit 204 may directly acquire the name, etc., of the person with whom the user wishes to talk by voice input.
[0126] In this way, when identifying a conversation partner for a user, the conversation control unit 204 may present to the user the number of conversation partner candidates who have attributes that match the conversation partner information entered by the user, as well as specific information about the candidates (e.g., names).
[0127] <Variation 3> In the above embodiment, the case where the user converses with one conversation partner (virtual person) has been described. However, the conversation control unit 204 may realize a conversation between the user and multiple conversation partners (virtual people).
[0128] In this case, the conversation control unit 204 identifies multiple conversation partners with whom the user wishes to have a conversation, using a GUI similar to that shown in Fig. 16, etc. The conversation control unit 204 generates virtual characters for each of the identified multiple conversation partners and displays them on the HMD 21.
[0129] The conversation control unit 204 realizes a conversation between the user and multiple virtual characters using the user's utterances and a language model. The conversation control unit 204 realizes a conversation between the user and multiple virtual characters by appropriately generating queries to be input into the language model.
[0130] For example, consider a case where a user is having a conversation with virtual person A and virtual person B. The name of the conversation partner corresponding to virtual person A is "Yamada," and the name of the conversation partner corresponding to virtual person B is "Tanaka."
[0131] The conversation control unit 204 generates a query to be input to the language model from the user's utterance. For example, when obtaining the user's utterance "Is everyone doing well?", the conversation control unit 204 generates a query such as "Please answer the question 'Is everyone doing well?' as a person named Yamada." and obtains an answer from the language model. Furthermore, the conversation control unit 204 generates a query such as "Please answer the question 'Is everyone doing well?' as a person named Tanaka." and obtains an answer from the language model.
[0132] The conversation control unit 204 converts the answers (text data) obtained from the two queries into voice data and transmits it to the HMD 21.
[0133] In this way, the conversation control unit 204 may realize a conversation between the user and a plurality of virtual characters.
[0134] <Modification Example 4> In the above embodiment, the case where a conversation between the user and a virtual character is realized via the HMD 21 has been described. However, the conversation may also be realized via the terminal 20.
[0135] The conversation control unit 204 may display a virtual character corresponding to the conversation partner selected by the user on the terminal 20 and output the voice of the virtual character from the terminal 20.
[0136] <Modification Example 5> In the above embodiment, the case where the user inputs the attributes (e.g., name, gender, position) of the conversation partner candidates to the server device 10 has been described. However, all or part of the attributes may be estimated by the server device 10 based on the image data.
[0137] Specifically, the registration control unit 203 may input the image data in which the conversation partner candidate appears to the learning model and obtain the gender, position, etc. of the person appearing in the image data from the learning model.
[0138] <Modification Example 6> When starting a conversation between a user and a virtual character, the conversation control unit 204 may determine the theme of the conversation. For example, the conversation control unit 204 may determine the theme of the conversation based on the image data including the face image used for generating the virtual character. For example, the conversation control unit 204 may obtain the theme using a learning model that outputs a theme when inputting the image data. For example, when the learning model obtains image data showing children playing at school, it outputs "play at school" as the theme.
[0139] Alternatively, the conversation control unit 204 may determine the theme of the conversation based on the age group selected by the user. For example, the conversation control unit 204 may set a symbolic event, incident, etc. that occurred in the age group selected by the user as the theme of the conversation. For example, when the user selects "teenagers", the conversation control unit 204 sets an event that occurred when the user was a teenager as the theme of the conversation.
[0140] <Modification Example 7> For realizing the conversation between the user and the virtual character, the server device 10 may use a language model prepared for each virtual character (each conversation partner candidate).
[0141] Here, the answer of the language model is determined based on the teacher data (learning data) used for generating the model. In other words, by appropriately selecting the teacher data used for generating the language model, the answer of the language model will have characteristics (personality). For example, a language model generated based on the speech of user A outputs an answer that strongly reflects the way of thinking of user A.
[0142] Utilizing such characteristics of the language model and teacher data, the server device 10 may generate a language model for each conversation partner candidate (person extracted from the image data uploaded to the server device 10) and utilize it in the conversation.
[0143] In this case, when the registration control unit 203 acquires the attribute information of the conversation partner candidate, it also acquires the account information of the SNS (Social Networking Service) used by the conversation partner candidate and information related to the blog (for example, URL (Uniform Resource Locator)).
[0144] The registration control unit 203 accesses the acquired SNS account and the URL of the blog, and collects the conversation partner candidate's online statements and the like. The registration control unit 203 generates a language model using the collected statements (text data) as teacher data.
[0145] The generated language model is used as a language model dedicated to the sender (conversation partner candidate) of the data on which the language model is based. When a conversation partner candidate for whom a dedicated language model is prepared is selected, the conversation control unit 204 acquires the response of the virtual character corresponding to the conversation partner candidate using the dedicated language model. As a result, the conversation control unit 204 can provide a response as if the conversation partner selected by the user actually speaks.
[0146] As described above, the server device 10 according to the first embodiment analyzes the image data uploaded by the terminal 20, and extracts the person shown in the image data stored in the user's smartphone or the like. The server device 10 stores the extracted person in the database as a conversation partner candidate for the user. Further, the server device 10 requests the user to tag each of the extracted persons. The server device 10 sets as a conversation partner a person who meets the age desired by the user and has the attributes specified by the user, and generates a virtual character using the face information of the conversation partner. The server device 10 displays the generated virtual character on the HMD 21 worn by the user, and converses with the user via the virtual character and the HMD 21.
[0147] By inputting to the server device 10 the era of a past enjoyable time and the attributes of the person with whom the user wishes to converse, the user can have a conversation with a virtual person modeled after friends, company colleagues, etc. at that time. As a result, the satisfaction of users who wish to converse with people they met in the past is improved. That is, the server device 10 contributes to alleviating the sense of loneliness of elderly people and the like.
[0148] Subsequently, the hardware of each device constituting the information processing system will be described. FIG. 18 is a diagram showing an example of the hardware configuration of the server device 10.
[0149] The server device 10 can be configured by an information processing device (so-called computer) and has the configuration illustrated in FIG. 18. For example, the server device 10 includes a processor 311, a memory 312, an input / output interface 313, a communication interface 314, and the like. The components such as the processor 311 are connected by an internal bus or the like and are configured to be able to communicate with each other.
[0150] However, the configuration shown in FIG. 18 is not intended to limit the hardware configuration of the server device 10. The server device 10 may include hardware not shown, or may not be provided with the input / output interface 313 as necessary. Also, the number of components such as the processor 311 included in the server device 10 is not intended to be limited to the example shown in FIG. 18. For example, a plurality of processors 311 may be included in the server device 10.
[0151] The processor 311 is, for example, a programmable device such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or a DSP (Digital Signal Processor). Alternatively, the processor 311 may be a device such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The processor 311 executes various programs including an operating system (OS; Operating System).
[0152] The memory 312 is, for example, a RAM (Random Access Memory), a ROM (Read Only Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), or the like. The memory 312 stores an OS program, application programs, and various data.
[0153] The input / output interface 313 is an interface for a display device and an input device (not shown). The display device is, for example, a liquid crystal display or the like. The input device is a device that receives user operations such as a keyboard and a mouse.
[0154] The communication interface 314 is a circuit, a module, or the like that communicates with other devices. For example, the communication interface 314 includes a NIC (Network Interface Card) or the like.
[0155] The functions of the server device 10 are realized by various processing modules. The processing modules are realized, for example, by the processor 311 executing a program stored in the memory 312. Further, the program can be recorded on a computer-readable storage medium. The storage medium can be a non-transitory one such as a semiconductor memory, a hard disk, a magnetic recording medium, or an optical recording medium. That is, the present invention can also be embodied as a computer program product. Further, the above program can be downloaded via a network or updated using a storage medium storing the program. Furthermore, the above processing module may be realized by a semiconductor chip.
[0156] Note that the terminal 20 and the HMD 21 can also be configured by an information processing device in the same manner as the server device 10, and the basic hardware configuration thereof is not different from that of the server device 10, so the description thereof is omitted.
[0157] The server device 10, which is an information processing device, is equipped with a computer, and the functions of the server device 10 can be realized by causing the computer to execute a program. Further, the server device 10 executes a control method of the server device 10 by the program. Similarly, the terminal 20 and the HMD 21 are equipped with a computer, and the functions of the terminal 20 and the HMD 21 can be realized by causing the computer to execute a program. Further, the terminal 20 and the HMD 21 execute a control method of the terminal 20 and the HMD 21 by the program.
[0158] [Modification Example] Note that the configuration, operation, etc. of the information processing system described in the above embodiment are examples and are not intended to limit the configuration of the system.
[0159] For example, account generation may be performed not by the user himself / herself but by the user's agent. For example, a child may generate an account on behalf of an elderly parent.
[0160] When the number of pieces of image data acquired from the user's terminal 20 is large, the registration control unit 203 of the server device 10 may perform thinning (sampling) of the image data for extracting candidate conversation partners. At that time, the registration control unit 203 may perform thinning so that there is no bias in the shooting date and time (acquisition time of the image data) of the image data. For example, the registration control unit 203 may perform thinning so that the number of pieces of image data is equal for each age group of the user.
[0161] When there are a plurality of face images of the same person, the registration control unit 203 may select a representative face image from the plurality of face images and register it in the conversation partner management database. For example, when there are a plurality of face images of the same age group of the user, the registration control unit 203 may select one face image and store it in the conversation partner management database. At that time, the registration control unit 203 may select the face image to be stored in the conversation partner management database based on the quality of each of the plurality of face images. The registration control unit 203 may select the face image with the highest quality as the face image for registration.
[0162] When acquiring information about conversation partner candidates (when requesting the user to tag conversation partner candidates), the registration control unit 203 may avoid presenting image data in which multiple people appear. Alternatively, when only image data in which multiple people appear is present, the registration control unit 203 may clearly indicate the people to be tagged and display them on the terminal 20. For example, the registration control unit 203 may display on the terminal 20 a screen in which a frame is set around the people to be tagged, as shown in FIG. 19. Note that the registration control unit 203 may use the information (combination of the facial image ID and the coordinates of the facial image) described with reference to FIG. 9 regarding the coordinates of the facial image of the person to be tagged.
[0163] When registering a conversation partner candidate in the conversation partner management database, if the image data contains a facial image of the user himself / herself, the registration control unit 203 may exclude the facial image of the user himself / herself from the registration targets. Specifically, the registration control unit 203 performs one-to-one authentication using features generated from the facial image and the features of the user himself / herself registered in the user management database. If the one-to-one authentication is successful, the registration control unit 203 determines that the facial image shown in the image data is the facial image of the user himself / herself, and excludes it from the registration targets in the conversation partner management database.
[0164] The conversation control unit 204 may give the virtual character facial expressions, movements, etc. depending on the content of the conversation. For example, if the confidence level regarding the answer of the language model is high, the conversation control unit 204 makes the virtual character's facial expression appear confident. On the other hand, if the confidence level regarding the answer of the language model is low, the conversation control unit 204 makes the virtual character's facial expression appear unconfident.
[0165] The conversation control unit 204 may use image data showing a conversation partner selected by the user to generate a background for the user to use in a virtual conversation. For example, if a school is shown in the image data showing the conversation partner, the conversation control unit 204 may generate a background with the school as its motif.
[0166] The conversation control unit 204 may generate the voice of the virtual character using voice data of a conversation partner candidate corresponding to the virtual character. When video data of a conversation partner candidate is acquired from the terminal 20, the registration control unit 203 generates a learning model using voice data obtained from the video data. The conversation control unit 204 inputs text data obtained from the language model into the learning model to acquire voice data that imitates the voice of the conversation partner. The conversation control unit 204 transmits the acquired voice data to the HMD 21.
[0167] In the above embodiment, the server device 10 extracts conversation partner candidates (analyzes image data), but the process of extracting conversation partner candidates may be performed by the terminal 20 of the user.
[0168] In the above embodiment, the server device 10 generates a virtual person with the appearance of a conversation partner of an age specified by the user. However, the server device 10 (conversation control unit 204) may generate a virtual person with an appearance corresponding to the current age rather than the age specified by the user, or a virtual person with an appearance corresponding to the age of the conversation partner desired by the user. The conversation control unit 204 may generate a virtual person with the current age or the like by age simulation using an image generation AI. Specifically, the conversation control unit 204 generates a virtual person with the appearance of the specified age (age) by inputting the age (age) of the conversation partner obtained by simulation along with a facial image of the conversation partner into the image generation AI.
[0169] In the above embodiment, a case has been described in which a user information database, etc., is configured inside the server device 10, but the database may also be configured on an external database server, etc. In other words, some of the functions of the server device 10 may be implemented on another server. More specifically, the above-described "registration control unit (registration control means)," "conversation control unit (conversation control means)," etc. may be implemented on any of the devices included in the system.
[0170] The form of data transmission and reception between each device (for example, the server device 10 and the terminal 20) is not particularly limited, but the data transmitted and received between these devices may be encrypted. Between these devices, image data and the like are transmitted and received, and in order to appropriately protect this information, it is desirable that encrypted data be transmitted and received.
[0171] In the flowcharts (flowcharts, sequence diagrams) used in the above description, a plurality of steps (processes) are described in order, but the execution order of the steps executed in the embodiment is not limited to the described order. In the embodiment, for example, each process can be executed in parallel, and the order of the illustrated steps can be changed within a range that does not interfere with the content.
[0172] The above embodiments have been described in detail for ease of understanding of the disclosure of the present application, and it is not intended that all the configurations described above are necessary. Also, when a plurality of embodiments are described, each embodiment may be used alone or in combination. For example, it is also possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace other configurations for a part of the configuration of the embodiment.
[0173] From the above description, the industrial applicability of the present invention is clear, but the present invention is suitably applicable to an information processing system that realizes a conversation between a user and a virtual person.
[0174] Some or all of the above embodiments may be described as follows in the appended claims, but are not limited thereto.
[0175] [Appended Claim 1] Storage means for associating and storing the face images of a plurality of persons that the user has met in the past with the age of the user when each face image was acquired; a conversation control means for, when the user desires to converse with a virtual person, acquiring the age of the user when the user met the person with whom the user desires to converse, identifying a person corresponding to the acquired age from among the plurality of stored persons, generating the virtual person using a face image of the identified person, and conversing with the user via the generated virtual person; A server device comprising: [Appendix 2] The server device described in Appendix 1 further comprises a registration control means for acquiring multiple image data from a terminal carried by the user and extracting multiple people that the user has met in the past based on facial images contained in each of the acquired multiple image data. [Appendix 3] 3. The server device according to claim 2, wherein the registration control means acquires attribute information of each of the extracted plurality of persons from the user. [Appendix 4] The server device according to claim 3, wherein the registration control means acquires at least the name of each of the extracted people and the position of each of the extracted people from the perspective of the user. [Appendix 5] the storage means stores the facial images of the plurality of people, the ages of the users, and the attribute information in association with each other; The conversation control means Acquire the age of the user with whom the user has met a person with whom the user wishes to have a conversation and the attribute information of the person with whom the user wishes to have a conversation; The server device according to claim 3, wherein a person corresponding to the acquired era and the acquired attribute information is identified from the plurality of stored people. [Appendix 6] The server device described in Appendix 2, wherein the registration control means calculates the age of the user when the facial image of each of the plurality of persons was acquired using the date and time when the facial image of each of the plurality of persons was acquired and the user's date of birth. [Appendix 7] The server device according to any one of Appendices 1 to 6, wherein the conversation control means controls a conversation between the user and the virtual person using a language model. [Appendix 8] The server device according to Appendix 7, wherein the conversation control means displays the virtual person on an HMD (Head Mounted Display) worn by the user. [Appendix 9] A storage step of associating and storing face images of a plurality of persons the user has met in the past with the age of the user when each face image was acquired; When the user wishes to have a conversation with a virtual person, the age of the user when the user met the person with whom the user wishes to have a conversation is acquired, a person corresponding to the acquired age is specified from among the plurality of stored persons, a virtual person is generated using the face image of the specified person, and a conversation is conducted between the user and the virtual person. A conversation control step; A control method for a server device, comprising: [Appendix 10] The control method for a server device according to Appendix 9, further comprising a registration control step of acquiring a plurality of image data from a terminal possessed by the user and extracting a plurality of persons the user has met in the past based on face images included in each of the acquired plurality of image data. [Appendix 11] The control method for a server device according to Appendix 10, wherein the registration control step acquires attribute information of each of the plurality of extracted persons from the user. [Appendix 12] The control method for a server device according to Appendix 11, wherein the registration control step acquires at least the name of each of the plurality of extracted persons and the position of each of the plurality of extracted persons from the viewpoint of the user. [Appendix 13] The storage step stores the face images of the plurality of persons, the age of the user, and the attribute information in association with each other. The conversation control step is as follows: Obtain the age of the user when the user meets a person with whom the user wishes to have a conversation and the attribute information of the person with whom the user wishes to have a conversation. A control method for a server device according to Supplementary Note 11, which specifies a person corresponding to the obtained age and the obtained attribute information from among the plurality of stored persons. [Supplementary Note 14] The registration control step calculates the age of the user at the time when the face images of each of the plurality of persons are obtained, using the date and time when the face images of each of the plurality of persons are obtained and the birth date of the user. A control method for a server device according to Supplementary Note 10. [Supplementary Note 15] The conversation control step controls the conversation between the user and the virtual person using a language model. A control method for a server device according to any one of Supplementary Notes 9 to 14. [Supplementary Note 16] The conversation control step displays the virtual person on an HMD (Head Mounted Display) worn by the user. A control method for a server device according to Supplementary Note 15. [Supplementary Note 17] On a computer mounted on a server device, A storage process of storing, in association with each other, the face images of a plurality of persons the user has met in the past and the age of the user at the time when each face image is obtained, and When the user wishes to have a conversation with a virtual person, obtain the age of the user when the user meets a person with whom the user wishes to have a conversation, specify a person corresponding to the obtained age from among the plurality of stored persons, generate the virtual person using the face image of the specified person, and have a conversation with the user via the generated virtual person. A conversation control process, A program for causing the execution. [Supplementary Note 18] A program according to Supplementary Note 17, which further causes the execution of a registration control process of obtaining a plurality of image data from a terminal possessed by the user and extracting a plurality of persons the user has met in the past based on the face images included in each of the obtained plurality of image data. [Supplementary Note 19] The registration control process is a program described in Supplementary Note 18 that acquires the attribute information of each of the plurality of extracted persons from the user. [Supplementary Note 20] The registration control process is a program described in Supplementary Note 19 that acquires at least the name of each of the plurality of extracted persons and the position of each of the plurality of extracted persons from the perspective of the user. [Supplementary Note 21] The storage process stores the face images of each of the plurality of persons, the age of the user, and the attribute information in association with each other. The conversation control process acquires the age of the user who has encountered a person with whom the user wishes to have a conversation and the attribute information of the person with whom the user wishes to have a conversation, and is a program described in Supplementary Note 19 that identifies a person corresponding to the acquired age and the acquired attribute information from among the plurality of stored persons. [Supplementary Note 22] The registration control process is a program described in Supplementary Note 18 that calculates the age of the user at the time when the face image of each of the plurality of persons was acquired, using the date and time when the face image of each of the plurality of persons was acquired and the date of birth of the user. [Supplementary Note 23] The conversation control process is a program described in any one of Supplementary Notes 17 to 22 that controls the conversation between the user and the virtual person using a language model. [Supplementary Note 24] The conversation control process is a program described in Supplementary Note 23 that displays the virtual person on an HMD (Head Mounted Display) worn by the user.
[0176] In addition, some or all of the configurations described in Appendices 2 to 8 that are subordinate to Appendix 1 described above may be subordinate to Appendices 9 and 17 in the same subordinate relationship as Appendices 2 to 8. Furthermore, not limited to Appendices 1, 9, and 17, within the scope not departing from the above-described embodiments, similarly, some or all of the configurations described as appendices can be made subordinate to various hardware, software, various recording means for recording software, or systems.
[0177] Note that the disclosures of each of the above-cited prior art documents are incorporated herein by reference. As described above, the embodiments of the present invention have been explained, but the present invention is not limited to these embodiments. It will be understood by those skilled in the art that these embodiments are merely examples, and various modifications can be made without departing from the scope and spirit of the present invention. That is, the present invention naturally includes all disclosures including the claims, and various modifications and corrections that can be made by those skilled in the art according to the technical idea.
Explanation of Reference Numerals
[0178] 10 Server device 20 Terminal 21 HMD 100 Server device 101 Storage means 102 Conversation control means 201 Communication control unit 202 Account control unit 203 Registration control unit 204 Conversation control unit 205 Storage unit 311 Processor 312 Memory 313 Input / output interface 314 Communication interface
Claims
1. A storage means for associating and storing the face images of a plurality of persons that the user has encountered in the past with the age of the user when each face image was acquired; When the user wishes to converse with a virtual person, the age of the user when the user met the person with whom the user wishes to converse is acquired, a person corresponding to the acquired age is identified from among the plurality of stored persons, a virtual person is generated using the face image of the identified person, and conversation is carried out between the user and the virtual person via the generated virtual person, a conversation control means; A server device comprising the above.
2. The server device according to claim 1, further comprising a registration control means for acquiring a plurality of image data from a terminal possessed by the user and extracting a plurality of persons that the user has encountered in the past based on the face images included in each of the acquired plurality of image data.
3. The server device according to claim 2, wherein the registration control means acquires attribute information of each of the plurality of extracted persons from the user.
4. The server device according to claim 3, wherein the registration control means acquires at least the name of each of the plurality of extracted persons and the position of each of the plurality of extracted persons from the perspective of the user.
5. The storage means stores the face images of each of the plurality of persons, the age of the user, and the attribute information in association with each other, The conversation control means, acquires the age of the user when the user met the person with whom the user wishes to converse and the attribute information of the person with whom the user wishes to converse, The server device according to claim 3, which identifies a person corresponding to the acquired age and the acquired attribute information from among the plurality of stored persons.
6. The server device according to claim 2, wherein the registration control means calculates the age of the user when each of the face images of the plurality of persons was acquired using the date and time when each of the face images of the plurality of persons was acquired and the date of birth of the user.
7. The server device according to any one of claims 1 to 6, wherein the conversation control means controls the conversation between the user and the virtual person using a language model.
8. The server device according to claim 7, wherein the conversation control means displays the virtual person on an HMD (Head Mounted Display) worn by the user.
9. A storage step of associating and storing the face images of a plurality of persons the user has encountered in the past with the age of the user when each face image was acquired; When the user wishes to converse with a virtual person, obtaining the age of the user when the user encountered the person with whom the user wishes to converse, identifying the person corresponding to the obtained age from among the plurality of stored persons, generating the virtual person using the face image of the identified person, and conversing with the user via the generated virtual person, a conversation control step; A control method for a server device comprising the above.
10. On a computer mounted on a server device, A storage process of associating and storing the face images of a plurality of persons the user has encountered in the past with the age of the user when each face image was acquired; When the user wishes to converse with a virtual person, obtaining the age of the user when the user encountered the person with whom the user wishes to converse, identifying the person corresponding to the obtained age from among the plurality of stored persons, generating the virtual person using the face image of the identified person, and conversing with the user via the generated virtual person, a conversation control process; A program for causing the above to be executed.
Citation Information
Patent Citations
Method and device for having virtual conversation with the deceased
JP2002024371A