Information processing apparatus and program
The information processing apparatus addresses the challenge of providing natural and consistent virtual customer service by quantifying and replicating human staff's content features, enhancing the realism and effectiveness of virtual interactions.
Patent Information
- Application Number
- JP2024032461
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2025-06-16
- Estimated Expiration
- 2043-12-04
AI Technical Summary
Existing virtual staff systems struggle to provide natural and consistent customer service interactions in online environments, as they often lack the ability to accurately quantify and replicate the content features of human staff's comments and speeches.
An information processing apparatus that includes a receiving module for user comments, a natural language analysis module, and a feature quantity calculation module. This apparatus assigns coordinates in an N-dimensional space based on parts of speech decomposition, allowing for the quantification and matching of content features in user interactions.
The solution enables virtual staff to generate speech that is content-consistent and naturally aligned with human staff's speech patterns, improving the realism and effectiveness of customer service interactions in virtual environments.
Smart Images

Figure 2025089981000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus and a program.
Background Art
[0002] Currently, EC sites (online sales) such as apparel have become widespread, and it is assumed that in the future, there will be an increasing number of opportunities to open virtual stores in the shopping malls of virtual spaces (metaverses) on the Internet. In such a situation, as in the case of physical stores, it is important from the perspective of sales promotion to provide a one-on-one customer service where staff introduce recommended products through interaction with users, dig out users' needs, materialize the products that users desire, and respond to various inquiries.
[0003] Generally, the number of users visiting EC sites is much larger than the number of customers visiting physical stores. Therefore, it is not realistic for actual staff to serve each user visiting the EC site individually. Thus, it is assumed that virtual staff (also called digital staff) is constructed on a computer, and the digital staff serves users one-on-one instead of actual staff.
[0004] Various characters are set for digital staff, but it is not easy to generate characters with customer service capabilities. Therefore, characters modeled after actual staff are sometimes set.
[0005] However, no matter how much the character of the virtual staff is modeled after that of the actual staff, there are often obvious differences from the actual staff in various aspects such as the content and diction of their speeches.
Summary of the Invention
Problems to be Solved by the Invention
[0006] An object is to realize quantifying the content features regarding comments and speeches.
Means for Solving the Problem
[0007] The information processing apparatus according to the present embodiment includes a receiving means for receiving data of a user's comment or utterance from an external information processing apparatus, a means for performing natural language analysis processing on the comment or the utterance, and a feature quantity calculating means for calculating a feature quantity that quantifies the content feature regarding the comment or the utterance by assigning coordinates in an N-dimensional space based on the contents of a plurality of parts of speech decomposed by the natural language analysis processing.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiment for Carrying Out the Invention
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In this embodiment, as a community-type service that promotes and supports connections between people, a social networking service (hereinafter referred to as "SNS"), a blog site where users post their daily matters, thoughts, etc. as articles and display them in chronological order, a word-of-mouth site where users write and view evaluations of products and services, and an EC site that sells products and services on the Internet. By referring to comments (referred to as texts such as posts, article texts, word-of-mouth texts, evaluation texts, messages, etc.) actually created and posted by real users themselves, the generation AI as a speech generation device generates a speech text in response to the speech of other users, so that the speech content matches the speech of other users, and moreover, the content of the speech of the virtual user is natural and without a sense of incongruity in the diction that a real user would use as if the real user represented by it would send it.
[0010] For example, representative examples of SNS include "Instagram (registered trademark)", "Facebook (registered trademark)", "Twitter (registered trademark) which is currently X", etc. Representative examples of word-of-mouth sites include "Tabelog (registered trademark)", "Price.com (registered trademark)", etc. Also, as for blog sites, there are many sites such as blogs on FC2 (registered trademark).
[0011] In the following description, taking the scenario of selling clothes etc. as an example, it will be explained via an EC (electronic commerce) site that develops a service to sell products and services on a web site on the Internet via a network such as the Internet. Also, in this scenario, a virtual user (hereinafter referred to as a virtual staff) that represents or acts on behalf of a real user (hereinafter referred to as a real staff) will be used to explain the situation where they interact in a chat room with a real user (other user) who visits the EC site and try to sell clothes etc. by serving the other user.
[0012] As shown in FIG. 1, for the information processing apparatus 1 according to the present embodiment, a speech generation apparatus 2 that functions as a generation AI, an SNS server 3-1 that provides an SNS, a blog site server 3-2 that provides a blog service, a word-of-mouth site server 3-3 that operates a word-of-mouth site, an EC site server 4 that operates an EC site, and a user terminal 5 used by an actual user are connected via a typical communication line network, i.e., the Internet line network 6.
[0013] As shown in FIG. 2, the information processing apparatus 1 as an interactive apparatus has a RAM 12, a ROM 13, a storage unit 14, an input device 15, a display 16, and a communication unit 17 connected to a processor 11 via a system bus 10. The processor 11 is composed of, for example, a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The processor 11 executes a program loaded from the storage unit 14 and the ROM 13 into the RAM 12, and executes an interactive process for realizing an interaction with an actual user. The RAM 12 functions as a main memory, a work area, etc. of the processor 11. The ROM 13 or the storage unit 14 stores a BIOS (Basic Input Output System), an operating system program (OS), an interactive process program according to the present embodiment, programs for realizing various other functions, and various data required for those processes.
[0014] The input device 15 consists of a keyboard (KB), a pointing device such as a mouse or a touch panel, etc. The display 19 is typically realized by an LCD (Liquid Crystal Display). In the storage unit 14, in addition to the interactive process program, data related to comments required for interactive processes, respective feature amounts, profiles of actual staff, interactive messages between actual users and virtual staff, and their order (interaction history), etc. are stored.
[0015] As shown in FIG. 3, the processor 11 functions as a control unit 20, a profile acquisition unit 21, a posted comment collection unit 22, a natural language analysis processing unit 23, a feature quantity calculation processing unit 24, a user message reception unit 25, a posted comment selection processing unit 26, a speech generation request unit 27, a speech reception unit 28, and a message transmission unit 29 by executing an interaction processing program.
[0016] The profile acquisition unit 21 transmits a profile transmission request together with each account of the actual staff to each of the SNS server 3-1, the blog site server 3-2, the word-of-mouth site server 3-3, and the EC site server 4, and receives profile data from each of the SNS server 3-1, the blog site server 3-2, the word-of-mouth site server 3-3, and the EC site server 4. These profiles are edited into a single profile and stored in the storage unit 14 in association with the identification number of the virtual staff acting as the proxy for the actual staff.
[0017] The posted comment collection unit 22 transmits a transmission request for comments (such as posts, articles, word-of-mouth texts, evaluation texts, messages, etc.) posted by the actual staff together with each account of the actual staff to each of the SNS server 3-1, the blog site server 3-2, the word-of-mouth site server 3-3, and the EC site server 4, and receives data of comments actually posted by the actual staff from each of the SNS server 3-1, the blog site server 3-2, and the word-of-mouth site server 3-3. These posted comments are stored in the storage unit 14 in association with the identification number (ID) of the virtual staff acting as the proxy for the actual staff.
[0018] The natural language analysis processing unit 23 performs natural language analysis processing on the posted comments and messages, decomposes them into parts of speech, and extracts noun strings. The feature quantity calculation processing unit 24 calculates a feature quantity that quantifies the content features for each posted comment based on the extracted nouns. The feature quantity is typically a coordinate in an N-dimensional space. A plurality of types of classifications corresponding to the theme categories are respectively associated with the N axes that make up the N-dimensional space. A plurality of nouns related to each classification are assigned to each axis of the N-dimensional space together with coordinate values. For example, for each category such as apparel, music, health food, books, food, and home appliances, a plurality of types of classifications are predefined. In the case of the apparel category, as illustrated in FIG. 8, classifications of generic upper concepts such as time period, inventory, size, design, item, color tone, price range, age, and event are associated with each axis. Further, for each classification, a number of nouns (character strings) representing specific lower concept matters included in each classification are set. The numerous nouns set on each of these axes are ordered according to semantic proximity, and numerical values (coordinate values) are assigned in advance. The noun string extracted from the comment (text) is queried against the nouns on the N axes of the N-dimensional space. If there is a corresponding noun string, the coordinate value corresponding to the noun string is given. Of course, when there is no corresponding noun string, a zero value is given to the coordinate value of that coordinate axis. Comments with similar feature quantities can be said to be close in content and have similar content features.
[0019] Note that as the feature quantity calculation processing, a learned model that has been pre-trained to output data of feature quantities that quantify the content features of comments and messages using the comments and messages as input data may be used.
[0020] The user message receiving unit 25 receives a message (utterance) of an actual user from the user terminal 5 directly or indirectly via the EC site server 4. Actually, the message of the actual user is indirectly received via at least one of the SNS server 3-1, the blog site server 3-2, the word-of-mouth site server 3-3, and the EC site server 4, which is a platform where the actual user interacts with the virtual staff. This message is subjected to natural language analysis by the natural language analysis processing unit 23, and its feature amount is calculated by the feature amount calculation processing unit 24.
[0021] The posted comment selection processing unit 26 selects one comment that is closest to the feature amount of the message of the actual user, or selects a predetermined number of upper comments close to the feature amount of the message of the actual user. The selected comment has content features that are approximate to the message (utterance) of the actual user.
[0022] The utterance generation request unit 27 transmits the data of the message (utterance) of the actual user to the utterance generation device 2 together with a request for generating an utterance of a virtual user according to the message of the actual user. In addition to the message, the utterance generation request unit 27 transmits the profile of the actual staff to the utterance generation device 2. Further, in addition to the message of the actual user and the profile of the actual staff, the utterance generation request unit 27 attaches at least one comment that the actual staff actually posted on SNS or the like and has content features approximate to the message of the actual user as reference information to the generation request and transmits it to the utterance generation device 2.
[0023] Here, in order for a normal interaction to be established between the actual user and the virtual staff, and for the interaction with the virtual staff to be recognized by the actual user without discomfort whether it is an interaction with the actual staff represented by the virtual staff, the following requirements are necessary.
[0024] 1) The utterance of the virtual staff is content-consistent so as to respond to the utterance (message) of the actual user.
[0025] 2) The content spoken by the virtual staff is close to the content that an actual staff would speak.
[0026] 3) The diction of the virtual staff reflects the diction of the actual staff.
[0027] In addition to the statements (messages) of actual users and the profiles of actual staff, these three requirements are realized by sending, as reference information, the posting comments actually posted by actual staff on SNS, etc. and having content approximately similar to the messages of actual users to the utterance generation device 2, and the utterance generation device 2 utilizes these reference information to generate an utterance text of the virtual staff in response to the message of the actual user.
[0028] The utterance text receiving unit 28 receives the data of the utterance text generated by the utterance text generation device 2 from the utterance text generation device 2. The message sending unit 29 sends the utterance text received from the utterance text generation device 2 as it is or after appropriate modification as a message of the virtual staff to the user terminal 5 or the EC site server 4 for the actual user.
[0029] FIG. 4 shows, together with a data flow, the dialogue processing procedure centered on the information processing device 1 as the dialogue device according to the present embodiment. First, as a preprocessing, data of the accounts of actual staff assuming proxy by virtual staff regarding each of SNS, blog service, and review site are provided from the terminal of the actual staff to the information processing device 1 together with a virtual staff dialogue processing request. In the information processing device 1, an identification number (ID) for identifying the virtual staff is issued (S11).
[0030] A profile transmission request is sent from the profile acquisition unit 21 of the information processing apparatus 1 to the SNS server 3-1, the blog site server 3-2, the word-of-mouth site server 3-3, and the EC site server 4, together with the respective account information. Profile data of actual staff registered by the actual staff when using their respective services is returned from the SNS server 3-1, the blog site server 3-2, and the word-of-mouth site server 3-3 to the information processing apparatus 1. As illustrated in FIG. 5, these profiles are edited into a single profile. In addition to basic information such as name, gender, and age, the profile includes, for example, current occupation, characteristics of the corporate brand, and characteristics such as the personality and hobbies of the actual staff. The edited profile is stored in the storage unit 14 with a virtual staff ID associated therewith (S12).
[0031] Next, a post comment transmission request is sent from the post comment collection unit 22 to each of the SNS server 3-1, the blog site server 3-2, the word-of-mouth site server 3-3, and the EC site server 4, together with the respective accounts of the actual staff. Data of comments (posts, etc.) posted by the actual staff is received from each of the SNS server 3-1, the blog site server 3-2, the word-of-mouth site server 3-3, and the EC site server 4. These post comments are stored in the storage unit 14 with the ID of the virtual staff acting on behalf of the actual staff associated therewith (S13).
[0032] As illustrated in FIG. 6, each comment is subjected to natural language analysis processing by the natural language analysis processing unit 23, decomposed into parts of speech, and nouns are extracted (S14). Then, based on the extracted nouns, a feature quantity that quantifies the content feature of the comment is calculated by the feature quantity calculation processing unit 24 (S15). Feature quantity calculation is repeated for all comments.
[0033] Coordinates in an N-dimensional space are determined as feature quantities. As illustrated in FIG. 8, a plurality of divisions are respectively associated with the N axes of the N-dimensional space. For example, in the case of apparel categories, a plurality of divisions such as time, inventory, size, design, item, color tone, price range, age, event, etc. are respectively associated with a plurality of axes. A plurality of specific nouns (character strings) included in each division are associated therewith. The plurality of specific nouns included in the same division are ordered according to the approximate nature of their semantic content, and numerical values (coordinate values) are respectively assigned thereto. The nouns extracted from the comment are queried against the specific nouns on each axis of the N-dimensional space, and if a corresponding character string exists, the axis number and coordinate value corresponding to that character string are given.
[0034] For example, for Comment No. 2, "I plan to go to work tomorrow wearing a blue jacket with a Glen check pattern pants.", the nouns "tomorrow", "Glen check", "pants", "blue", "jacket", "work", "plan" are extracted. For example, on the axis associated with the time division, a large number of specific nouns related to time such as yesterday, today, tomorrow, January, February, next month, the month after next, spring, autumn, etc. are associated, and coordinate values are respectively assigned thereto. The noun "tomorrow" extracted from the comment is given the coordinate value assigned to the specific noun "tomorrow" on the axis associated with the time division. Similarly, on the axis associated with the item division, specific nouns related to clothing such as pants, skirt, blouse, shirt, dress, suit, jacket, blazer, coat, etc. are associated, and coordinate values are respectively assigned thereto. The noun "pants" extracted from the comment is given the coordinate value of the corresponding specific noun on the item division axis. On the event division axis, nouns related to events such as Father's Day, Mother's Day, sports meet, birthday, date, etc. are associated, and coordinate values are respectively assigned thereto.
[0035] The calculated feature quantity is associated with the comment and stored in the storage unit 14 (S16).
[0036] New comments are collected at a predetermined cycle such as once a day, feature amounts are calculated for the new comments, and the calculated feature amounts are accumulated. With the above, the pre - preparation process is completed.
[0037] When starting the dialogue, for example, a request to open a chat room (chat bot) for one - on - one dialogue is sent from the user terminal 5 to the EC - site server 4 that operates the EC site. Of course, the dialogue form is not limited to the chat room. In the information processing apparatus 1, a virtual staff member who will dialogue with the user is selected, and a chat room between the user and the virtual staff member is opened (S17). For example, if a virtual staff member who has had a dialogue with the same real user in the past is selected, or if a real staff member has served a real user in a physical store, a virtual staff member who acts as an agent for that real staff member is selected. The method of selecting the virtual staff member is not limited to these methods.
[0038] For example, a message (utterance) of a real user input from the user terminal 5 shown in FIG. 7 is received by the user message receiving unit 25 from the user terminal 5 via the EC - site server 4. This message is subjected to natural language analysis by the natural language analysis processing unit 23 (S18), and its feature amount is calculated by the feature amount calculation processing unit 24 (S19). The natural language analysis processing and the feature amount calculation processing for the message are the same as those for the comment.
[0039] Next, one or a predetermined number of comments are selected from the comments of actual staff by the posted comment selection processing unit 26 (S20). Specifically, one comment having the feature amount closest to the feature amount of the message from the actual user, or a predetermined number of comments having feature amounts close to the feature amount of the message from the actual user are selected. More specifically, since the feature amount is a coordinate in the N-dimensional space, the distance between the feature amount (coordinate) of the message from the actual user and the feature amount (coordinate) of the comment of the actual staff is calculated, and one comment with the shortest distance is selected, or a predetermined number of comments are selected in ascending order of the distance. Since the feature amount is a quantification of the content feature, the selected posted comment is content-wise approximated to the message of the actual user. For example, as illustrated in FIG. 9, when the message of the actual user is about asking for a recommendation of a jacket for traveling, comments related to jackets and holidays are selected.
[0040] It is transmitted from the speech generation request unit 27 to the speech generation device 2 together with the speech generation request. The speech generation request is attached with the message of the actual user or the history of the dialogue between the actual user and the actual staff. Further, the speech generation request is attached with the profile of the actual staff represented by the virtual staff. Furthermore, the speech generation request is attached with, as reference information, a comment content-wise close to the message of the actual user, selected from the comments actually posted by the actual staff on SNS or the like.
[0041] In the utterance generation request unit 27, an utterance of the virtual staff in response to the message of the real user is generated (S21) by referring to the profile of the real staff and a comment that is content - close to the message of the real user selected from the comments actually posted by the real staff on SNS or the like. When generating the utterance of the virtual staff, since a comment that is content - close to the message of the real user is referred to, the utterance of the virtual staff is generated with content close to the content that the real staff would respond with. For example, when a real user asks for a jacket recommendation, an answer recommending "a blue tweed jacket" is generated from the comments posted by the real staff in the past. Also, the diction of the virtual staff applies not only the profile of the real staff but also the diction used in the comments posted by the real staff in the past.
[0042] As described above, in generating the utterance of the virtual staff in response to the message of the real user, in addition to the profile of the real staff, a comment that is content - close to the message of the real user selected from the comments actually posted by the real staff on SNS or the like can be referred to, so that the following three requirements can be satisfied with high precision.
[0043] 1) The utterance of the virtual staff is content - consistent with the utterance (message) of the real user.
[0044] 2) The content uttered by the virtual staff is the same as or close to the content that the real staff would utter.
[0045] 3) The diction of the virtual staff reflects the diction of the real staff.
[0046] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, as well as in the invention described in the claims and its equivalent scope.
Description of Reference Numerals
[0047] 1... Information processing device (dialogue device), 2... Utterance generation device (generative AI), 3-1... SNS server, 3-2... Blog site server, 3-3... Word-of-mouth site server, 4... EC site server, 5... User terminal.
Claims
1. A receiving means for receiving data of a user's comment or statement from an external information processing device; means for performing natural language analysis processing on the comment or statement; and a feature calculation means for calculating a feature that quantifies content features related to the comment or statement by assigning coordinates in an N-dimensional space based on the content of the multiple parts of speech broken down by the natural language analysis processing.
2. The information processing device according to claim 1 , wherein the receiving means receives the data of the user's comments or remarks directly from a user terminal or indirectly via at least one of a SNS server, a blog site server, a word-of-mouth site server, and an EC site server.
3. 2. The information processing device according to claim 1, wherein said N-dimensional space has N axes to which a plurality of types of divisions represented by parts of speech correspond, and each axis is assigned a plurality of coordinate values to which a plurality of specific contents included in each division are assigned.
4. An information processing device as described in claim 1, further comprising a means for calculating a distance in the N-dimensional space between the feature relating to the comment or statement and another feature relating to the other comment or statement calculated by the feature calculation means as a numerical value representing the content similarity between the comment or statement and another comment or statement.
5. The information processing device of claim 1, wherein the feature calculation means uses a trained model that has been trained in advance using data on multiple parts of speech decomposed by the natural language analysis process relating to data on learning comments or learning utterances and data on features representing content features of the learning comments or learning utterances as training data, and outputs the features that quantify content features related to the comments or utterances from the multiple parts of speech decomposed by the natural language analysis process for the comments or utterances.
6. Computer, A receiving means for receiving data of a user's comment or statement from an external information processing device; means for performing natural language analysis processing on the comment or statement; a program that functions as a means for calculating feature quantities that quantify content features related to the comment or statement by assigning coordinates in an N-dimensional space based on the contents of the multiple parts of speech decomposed by the natural language analysis processing;
Citation Information
Patent Citations
Device and method for information filtering
JP2002024274A
Text sorting device, text sorting method, text sort program, and recording medium with its program recorded thereon
JP2008282328A
Text summarization device, text summarization method, and program
JP2015088064A
Inter-set relationship calculation device and information processing device
JP2023000605A