Apparatus, method, and program for generating text
By using an attribute database to set characteristics for language models, the apparatus and method enhance personalized text generation, addressing the limitations of existing models and improving interaction relevance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-04-07
AI Technical Summary
Existing language models lack the ability to generate personalized and contextually relevant natural language outputs based on individual attributes and characteristics, limiting their effectiveness in applications such as customer support and text creation.
An apparatus and method that utilizes an attribute database to set characteristics for a language model based on individual attributes, enabling personalized text generation through a database connection, characteristic setting, and text generation units, allowing for natural language dialogue and personality diagnosis.
Enhances the language model's ability to generate contextually relevant and personalized text, improving interactions and responses in applications like customer support and text creation.
Smart Images

Figure 2026059579000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus, a method, and a program for generating text.
Background Art
[0002] In recent years, language models (also referred to as "large language models (LLMs)", "natural language processing models (NLP models)", "(text) generation AIs", etc.) that generate natural language output texts in response to natural language input texts, such as ChatGPT (registered trademark), Microsoft Copilot, Google BERT, etc., have begun to be commonly used. Such language models are utilized in various applications, such as outputting answers in natural language in response to questions input in natural language, summarizing and outputting natural language input texts or data, etc., and outputting programs that realize functions instructed in natural language. In addition, such language models are being widely used, for example, in automatic responses in customer support, and various application scenarios such as text creation support or information search support are being realized or considered.
Summary of the Invention
[0003] In a first aspect of the present invention, there is provided an apparatus including: a database connection unit connected to an attribute database for storing a plurality of attribute values corresponding to a plurality of attributes for each of a plurality of subjects; a characteristic setting unit that sets characteristics based on at least one attribute value of at least one subject stored in the attribute database for a language model that generates a natural language output text in response to a natural language input text; and a text generation unit that generates an output text corresponding to the input text using the language model with the characteristics set.
[0004] The above-described device includes a database search unit that searches the attribute database for at least one subject from among the multiple subjects that meets specified search conditions, and the characteristic setting unit may set the characteristics for the language model based on at least one attribute value for the at least one subject.
[0005] In any of the above-described devices, the characteristic setting unit may set the characteristic based on the at least one attribute value in the language model by inputting an explanatory text describing the characteristic based on the at least one attribute value for the at least one subject into the language model.
[0006] In any of the above-described devices, the characteristic setting unit sets candidate characteristics for the language model, and includes a subject selection unit that selects at least one subject from among the plurality of subjects based on the degree of similarity between at least one attribute value estimated to be possessed by the language model on which the characteristic candidate is set and the corresponding at least one attribute value stored in the attribute database, and the characteristic setting unit may use the characteristic candidate as the characteristic of the selected at least one subject when at least one subject is selected according to the characteristic candidate.
[0007] Any of the above devices may allow a natural language dialogue to take place between the language model, which has the characteristics set, and the selected at least one subject.
[0008] In any of the above-described devices, the characteristic setting unit may set or update the characteristics based on the content of the dialogue.
[0009] In any of the above-described devices, the characteristic setting unit may set or update the characteristics based on the at least one attribute value stored in the attribute database that indicates the purchase history of a product or service.
[0010] Any of the above devices may include a personality diagnosis unit that performs a personality diagnosis of the language model, which has its characteristics set, by engaging in dialogue with the language model, which has its characteristics set.
[0011] In any of the above-described devices, the characteristic setting unit may include a questionnaire processing unit that sets one or more characteristics for one or more of the multiple subjects for the language model, inputs a questionnaire text to the language model for which each of the one or more characteristics has been set, and collects output texts that the language model outputs in response to the questionnaire.
[0012] A second embodiment of the present invention provides a method comprising: an apparatus accessing an attribute database for storing multiple attribute values corresponding to multiple attributes for each of a plurality of subjects; the apparatus setting a characteristic for a language model that generates a natural language output sentence corresponding to a natural language input sentence, based on at least one attribute value for at least one subject stored in the attribute database; and the apparatus generating an output sentence corresponding to an input sentence using the language model on which the characteristic has been set.
[0013] In a third aspect of the present invention, a program is provided which is executed by a computer and causes the computer to function as a database connection unit connected to an attribute database for storing multiple attribute values corresponding to multiple attributes for each of multiple subjects; a characteristic setting unit that sets characteristics for a language model that generates natural language output text corresponding to natural language input text, based on at least one attribute value for at least one subject stored in the attribute database; and a text generation unit that uses the language model with the characteristics set to generate output text corresponding to input text.
[0014] It should be noted that the above summary of the invention does not enumerate all of its features. Furthermore, subcombinations of these features may also constitute an invention. [Brief explanation of the drawing]
[0015] [Figure 1] The configuration of system 10 according to this embodiment is shown. [Figure 2] An example of the data structure stored in the attribute DB20 according to this embodiment is shown. [Figure 3] This diagram shows the operation flow of the language processing device 100 according to this embodiment. [Figure 4] An example of questionnaire processing by the questionnaire processing unit 180 according to this embodiment is shown. [Figure 5] The configuration of the language processing device 400 according to the first modified example of this embodiment is shown. [Figure 6] The operation flow of the language processing device 400 according to the first modified example is shown. [Figure 7] The configuration of the language processing device 700 according to a second modified example of this embodiment is shown. [Figure 8] The operation flow of the language processing device 700 according to the second modified example is shown. [Figure 9] An example of the persona selection screen related to the second modified version is shown. [Figure 10] An example of a discussion screen related to the second modified example is shown. [Figure 11] Examples of a computer 2200 in which multiple aspects of the present invention may be embodied in whole or in part are shown. [Modes for carrying out the invention]
[0016] The present invention will be described below through embodiments of the invention, but these embodiments are not intended to limit the invention as defined in the claims. Furthermore, not all combinations of features described in the embodiments are necessarily essential to the solution of the invention.
[0017] FIG. 1 shows the configuration of the system 10 according to this embodiment. The system 10 registers and manages the attribute data (also referred to as "attribute information") of subscribers or members such as point services, or other target persons in an attribute database (attribute DB) 20. By setting characteristics based on the attribute values of the target persons stored in the attribute DB 20 in a language model, the system 10 generates and makes available a virtual personality of the target person having characteristics corresponding to the attribute values of the target person and capable of engaging in dialogue in natural language.
[0018] The system 10 includes an attribute DB 20, an attribute DB management device 30, and a language processing device 100. The system 10 may be a computer such as a PC (personal computer), a tablet computer, a smartphone, a workstation, a server computer, or a general-purpose computer, or may be a computer system to which a plurality of computers are connected. Such a computer system is also a computer in a broad sense. Further, the system 10 may be implemented by a virtual computer environment that can be executed one or more times within the computer. Instead of this, the system 10 may be a dedicated computer designed for the management and use of the attribute database, or may be dedicated hardware realized by a dedicated circuit. The system 10 according to this embodiment manages, as an example, the attribute information of target persons such as members in a point system provided by one operator, a common point system shared by a plurality of operators, a credit card, electronic money, and other arbitrary membership services.
[0019] The attribute DB 20 stores attribute information including a plurality of attribute values corresponding to a plurality of attributes for each of a plurality of target persons. The attribute DB 20 may be realized by at least a part of the storage area of an external storage device such as a hard disk drive connected to the computer that processes the system 10, or may be realized by an external storage device of the system 10 provided by, for example, a cloud storage service.
[0020] The attribute DB management device 30 is connected to the attribute DB 20. The attribute DB management device 30 may be implemented by the same computer as the system 10 or dedicated hardware, or may be implemented by a computer or dedicated hardware different from the language processing device 100 etc. within the system 10. The attribute DB management device 30 manages the attribute DB 20 and performs marketing processing by utilizing the attribute DB 20. The attribute DB management device 30 includes an attribute information acquisition unit 40 and a marketing processing unit 60.
[0021] The attribute information acquisition unit 40 acquires the attribute information of the target person and stores it in the attribute DB 20. The attribute information acquisition unit 40 may have a transmission / reception circuit and communicate with a collection device such as a POS terminal installed in a store etc. that collects member information, or other devices to receive the attribute information of the target person. The source of the attribute information of the target person is, as an example, information that the target person must or optionally fills in or inputs during new member registration, the response of the target person to a member questionnaire, settlement information corresponding to the purchase of goods etc. by the target person in a store etc., settlement information corresponding to the purchase of goods etc. by the target person on an e-commerce site, information on the Web site accessed by the target person, information on an Internet advertisement (Web ad) clicked by the target person on the Web site, or information on a TV program viewed by the target person etc., and is provided upon receiving the consent of the target person. The attribute DB 20 may actively collect unknown attribute values by accessing the attribute DB 20 at an arbitrary timing such as periodically, searching for a target person whose attribute value for at least one attribute is unknown, and conducting a member questionnaire via a Web site etc.
[0022] The marketing processing unit 60 performs marketing processing using the attribute DB 20. In this embodiment, the marketing processing unit 60 performs product or service recommendation processing using the attribute DB 20 as an example. The marketing processing unit 60 selects whether or not to recommend a product or service to a target based on at least one attribute value for each of a group of target individuals. Here, the marketing processing unit 60 may decide to recommend a product to a target individual if, for each of the group of target individuals, the attribute value of an attribute indicating preference for a particular product or service is above a threshold. Alternatively, the marketing processing unit 60 may decide to recommend a product or service based on the attribute values of one or more basic attributes, lifestyle attributes, and attributes indicating at least some of the inclinations (e.g., place of residence, whether or not they own a car, whether or not they are luxury-oriented or frugal).
[0023] The marketing processing unit 60 then recommends products, etc., to the target audience members who have been selected from among the multiple target audience members to receive recommendations. For example, the marketing processing unit 60 may provide the target audience with emails, direct mail, and internet advertisements that include advertisements for the products, etc., or it may provide television commercials that include advertisements for the products, etc., to viewers including the target audience, or it may provide services such as coupons, discounts, and point rewards that offer preferential treatment for purchasing the products, etc. Note that the attribute DB management device 30 does not need to have a marketing processing unit 60 if it does not need to perform marketing processing.
[0024] The language processing device 100 is connected to the attribute DB 20. The language processing device 100 may be implemented using the same computer or dedicated hardware as the system 10, or it may be implemented using a different computer or dedicated hardware than the attribute DB management device 30, etc., within the system 10. The language processing device 100 sets characteristics in the language model according to the attribute values of the target person stored in the attribute DB 20 and makes them available. The language processing device 100 includes a DB connection unit 110, a DB search unit 120, a characteristic setting unit 130, a language model storage unit 140, and a text generation unit 150.
[0025] The DB connection unit 110 is connected to the attribute DB 20 and processes access to the attribute DB 20 from various parts within the language processing unit 100. The DB search unit 120 is connected to the DB connection unit 110. The DB search unit 120 may have input / output circuits or transmit / receive circuits and may receive search conditions specified by the user of the language processing unit 100 via the input / output circuits or transmit / receive circuits. The DB search unit 120 searches the attribute DB 20 for at least one subject that satisfies the specified search conditions from among multiple subjects registered in the attribute DB 20. For example, the DB search unit 120 searches the attribute DB 20 for subjects that match the search conditions specified by the user of the language processing unit 100.
[0026] The characteristic setting unit 130 is connected to the DB search unit 120. The characteristic setting unit 130 sets characteristics for the language model that generates natural language output text corresponding to natural language input text, based on at least one attribute value for at least one subject stored in the attribute DB 20. The characteristic setting unit 130 may set characteristics for the language model based on one or more attribute values stored in the attribute DB 20 for subjects retrieved by the DB search unit 120.
[0027] The language model storage unit 140 is connected to the characteristic setting unit 130. The language model storage unit 140 stores a language model in which the characteristics of the subject have been set by the characteristic setting unit 130. If the language processing device 100 uses an external language model provided, for example, by a cloud service, the language model storage unit 140 may be located outside the language processing device 100. In this case, the language processing device 100 may use a language processing system (cloud service, etc.) that includes the language model storage unit 140 via a network.
[0028] The text generation unit 150 is connected to the language model storage unit 140. The text generation unit 150 generates output text corresponding to the input text using a language model whose characteristics have been set by the characteristic setting unit 130. In this specification, for the sake of clarity, inputting explanatory text (or input text, etc.) into the processing system (hardware circuit, computer, cloud server, or other device) of the text generation unit 150 that processes the language model may also be expressed as "inputting input text (or explanatory text) into the language model." Similarly, for the sake of clarity, processing by the processing system of the text generation unit 150 that processes the language model, or processing directed to the processing system of the text generation unit 150, may be described as processing of the language model or processing to the language model, omitting the processing system that actually performs the information processing.
[0029] The language processing device 100 may include at least one of the following: a dialogue processing unit 160, a personality diagnosis unit 170, and a questionnaire processing unit 180. The dialogue processing unit 160 is connected to the text generation unit 150 and the characteristic setting unit 130. The dialogue processing unit 160 causes a language model with set characteristics to engage in natural language dialogue with the user of the language processing device 100, at least one subject selected by the DB search unit 120 (i.e., a subject with attribute values that formed the basis of the characteristics set in the language model), or other subjects. The dialogue processing unit 160 may have input / output circuits or transmission / reception circuits, and may exchange text with the user of the language processing device 100 via input / output devices (such as keyboards and display devices) used by the user of the language processing device 100 or terminal devices used by the user of the language processing device 100. The dialogue processing unit 160 may supply the content of the dialogue between the subject selected by the DB search unit 120 and the language model to the characteristic setting unit 130. As a result, the characteristic setting unit 130 may set or update the characteristics of the subject based on the content of the dialogue with the subject.
[0030] The personality assessment unit 170 is connected to the text generation unit 150 and the DB connection unit 110. The personality assessment unit 170 performs a personality assessment of a language model whose characteristics are set based on the attribute values of the subject selected by the DB search unit 120 by engaging in dialogue with the language model whose characteristics are set based on those characteristics. The personality assessment unit 170 may store the results of the personality assessment in the attribute DB 20, associating them with new or existing attributes of the subject.
[0031] The survey processing unit 180 is connected to the text generation unit 150 and the DB connection unit 110. The survey processing unit 180 uses a language model to identify two or more characteristics of two or more subjects from among multiple subjects and performs a simulated survey on two or more subjects.
[0032] Figure 2 shows an example of a data structure stored in the attribute DB 20 according to this embodiment. The attribute DB 20 stores, for each of multiple subjects, personal identification information (personal ID) that identifies the individual, and attribute information (also referred to as "attribute values") for multiple attributes that the individual possesses. The attribute DB 20 may store attribute information for a hypothetical subject that integrates the attribute information of two or more individuals, or that represents two or more individuals. In this case, the attribute DB 20 may store identification information for a hypothetical subject instead of personal identification information, and may store attribute information that integrates the attribute information of each individual or attribute information that represents the attribute information of each individual instead of the attribute information of each individual. The attribute information for a hypothetical subject may be the mean, median, sum, or distribution of the attribute information of two or more individuals.
[0033] "Personal ID" is an identifier used in System 10 to identify individual subjects, such as a membership number or login ID for services provided by System 10. Alternatively, Attribute DB 20 may use the subject's name, email address, address, telephone number, identification information of the subject's mobile device, or information generated based on at least one combination of these as "Personal ID".
[0034] "Attribute information" refers to attribute values for various attributes possessed by the subject. The attributes stored in the attribute DB20 according to this embodiment are broadly classified into general attribute data, purchase potential data, and recommendation potential data.
[0035] "General attribute data" is a set of attributes that represent the characteristics of each subject, and in particular, it generally represents the characteristics of the subject themselves. "General attribute data" may include at least one of the following: one or more attributes classified as basic attributes, one or more attributes classified as lifestyle attributes, and one or more attributes classified as orientations.
[0036] "Basic attributes" are the basic information of each subject, and include at least one of the following attributes classified as basic attributes: name, date of birth, age or age group, gender, address, telephone number, etc. "Basic attributes" are mainly entered when a subject is newly registered or when registration details are changed, but at least some attributes may be optional and may be used as prediction targets using a language model.
[0037] "Lifestyle attributes" refer to information about the lifestyle of the subject, and may include at least one of the following attributes classified as "lifestyle attributes": marital status, housing type, household income, personal income, occupation, car ownership, and home ownership. Attribute values related to "lifestyle attributes" may be collected at the time of new registration, collected by various methods such as questionnaires, and may be used as prediction targets in language models.
[0038] "Orientation" refers to information indicating the preferences, tendencies, and / or tastes of the subject, and may include at least one of the following attributes classified as "orientation": for example, quality orientation / challenge orientation / practical orientation / brand orientation for clothing, luxury orientation / frugality orientation / discount orientation for food, convenience store orientation / urban orientation / region-focused orientation for housing, and other attributes such as health orientation, career orientation, and global orientation. Alternatively, orientation may include at least one attribute relating to the subject's tastes, such as whether or not they have a preference for various hobbies such as driving, gourmet food, travel, and sports, whether or not they have a preference for various products, etc., and whether or not they have a preference for various websites, etc. Attributes relating to "orientation" may be added in various surveys depending on the purpose of the survey. Attribute values relating to "orientation" may be collected when new registrations are made, may be collected by various methods such as questionnaires to the subject, and may be used as a target for prediction using language models.
[0039] "Purchase potential data" is a set of attributes that indicate the purchase potential of each target individual for each of several products, or groups of products or services. "Purchase potential data" may include attributes associated with each product or service corresponding to each genre, type, or classification of products, such as entertainment, food, and daily necessities. Each attribute of the "Purchase potential data" may be a preference attribute that indicates the target individual's preference for the product associated with that attribute.
[0040] For example, "purchase potential data" includes attributes for each of the numerous products, etc., that are subject to sales management by the membership service provided by System 10. "Purchase potential data" may also include attributes corresponding to each code value of, for example, a JAN code that identifies each product, etc. In this case, one or more attributes may be assigned to each product, etc. to which a JAN code is assigned. If the membership service provided by System 10 provides a common point system that spans multiple businesses, the purchase potential data may include hundreds of thousands to millions or even more attributes. Furthermore, "purchase potential data" may have attributes corresponding to groups of products or services. For example, "purchase potential data" may have attributes corresponding to groups of products such as beer, alcoholic beverages, and / or other beverages.
[0041] The "purchase potential data" may store known attribute values for each target individual, such as purchase history (whether or not a purchase was made, purchase quantity, purchase timing, purchase location, etc.) and / or preference scores that quantify the target individual's response to product questionnaires or advertisements. In addition, the "purchase potential data" may store predicted values predicted using a language model as at least part of the attribute values.
[0042] "Recommendation potential data" is a set of attributes that indicate the characteristics of each target audience regarding recommendations. "Recommendation potential data" may include, for example, one or more attributes related to media response, one or more attributes related to incentive response, and one or more attributes related to churn potential.
[0043] "Media response" includes one or more attributes indicating the effectiveness or evaluation value of recommendations using a particular medium for a target audience, for each media used for recommendations, such as direct mail, email, advertisements printed on receipts, internet advertisements, and TV advertisements. For example, if a target audience does not respond to direct mail, the marketing processing unit 60 may decrease the attribute value of the attribute related to the effectiveness or evaluation value of direct mail, or set a lower attribute value. Also, if a target audience clicks on an internet advertisement more frequently than a standard frequency, or purchases a product via an internet advertisement, the marketing processing unit 60 may increase the attribute value of the attribute related to the effectiveness or evaluation value of the internet advertisement, or set a higher attribute value.
[0044] "Incentive response" includes one or more attributes for each incentive, such as discounts, coupons, points, point increases, and promotional items, that indicate the effectiveness or evaluation value of recommendations using that incentive for the target person. For example, if a target person does not purchase a product even after a discount is offered, the marketing processing unit 60 may decrease the attribute value of the attribute related to the effectiveness or evaluation value of the discount, or set a lower attribute value. Also, if a target person purchases a product for which points were awarded, the marketing processing unit 60 may increase the attribute value of the attribute related to the effectiveness or evaluation value of the point award, or set a higher attribute value.
[0045] "Churrency" includes, for example, one or more attributes for each sales company and / or each product manufacturer or service provider that indicate the likelihood of the customer ceasing to use the products of that company. For example, if a target customer has not purchased products from a particular sales company for a certain period of time, the marketing processing unit 60 considers that the customer has churned from that sales company and stores an attribute value indicating this churning in the attribute DB 20. Alternatively, the marketing processing unit 60 may store a churning evaluation value calculated based on the time elapsed since the customer last used the sales company as an attribute value in the attribute DB 20.
[0046] Figure 3 shows the operation flow of the language processing device 100 according to this embodiment. In step 300 (S300), the DB search unit 120 acquires the search conditions for the target user. As an example, the DB search unit 120 may acquire the search conditions by receiving or inputting the search conditions entered by the user to a terminal device or input device used by the user of the language processing device 100, using a communication circuit or input / output circuit, etc.
[0047] The search criteria for target individuals are used to narrow down one or more target individuals from among multiple target individuals registered in the attribute DB20. The search criteria may specify attribute values or ranges of attribute values for each of one or more attributes. For example, the search criteria may be conditions related to basic attributes such as "men in their 30s," conditions related to lifestyle attributes such as "married people who own a car," conditions related to preferences such as "people who are urban-oriented and have a hobby of gourmet food," conditions related to purchase potential data such as "people who have purchased product A," conditions related to recommendation potential data such as "people who value incentives," conditions related to media response such as "people who have purchased a product via internet advertising," conditions related to incentive response such as "people who have purchased a product at a discount," conditions related to churn potential such as "people who have a product with a repeat purchase rate of X% or higher," or conditions that are a combination of at least two of these (combinations using AND, OR, NOT or other set operations) (for example, "men in their 30s who own a car, have a hobby of gourmet food, and value incentives"). Furthermore, the search criteria may include other conditions such as specifying or specifying a range of individual IDs, or conditions that limit the number of people to be included (for example, a total of 5 people who meet the attribute value conditions mentioned above).
[0048] The search criteria for target individuals may specify one or more target individuals from among the multiple target individuals registered in the attribute DB20. Alternatively, the language processing device 100 may perform the processing from S310 onward on all target individuals registered in the attribute DB20.
[0049] In S310, the DB search unit 120 searches the attribute DB 20 for one or more subjects that meet the specified search criteria from among the multiple subjects registered in the attribute DB 20. The DB search unit 120 may also reduce the number of subjects by thinning (sampling) the subjects that meet the search criteria, such as by random selection.
[0050] In S320, the characteristic setting unit 130 reads at least one attribute value for one or more retrieved subjects from the attribute DB 20 and sets a characteristic in the language model based on at least one attribute value. The characteristic setting unit 130 may set a characteristic based on attribute values for one or more attributes specified by the user, or for one or more predetermined attributes, or it may set a characteristic based on attribute values for all attributes. The characteristic setting unit 130 may also set a characteristic based on at least one attribute value stored in the attribute DB 20 that indicates the purchase history of a product or service.
[0051] The characteristic setting unit 130 may generate a language model in which characteristics are set for each of the two or more searched subjects based on the attribute values of that subject. Alternatively, the characteristic setting unit 130 may generate a language model in which characteristics are set based on attribute values that combine the attribute values of each of the two or more searched subjects (e.g., the average attribute values of "men in their 30s who own a car"). When combining the attribute values of each of the two or more subjects, the characteristic setting unit 130 may calculate the mean, median, maximum and / or minimum, distribution, or other statistical values of the attribute values of the two or more subjects and consider them as attribute values for the two or more subjects.
[0052] Here, the characteristic setting unit 130 can obtain a language model having characteristics corresponding to the attribute values by converting the attribute values into a format that can be set in the language model and providing it to the language model, or by further training the language model according to the attribute values, and as a result, characteristics can be set in the language model based on the attribute values. As an example, the characteristic setting unit 130 may set characteristics based on attribute values in the language model by one of the following methods.
[0053] (1) Setting characteristics using explanatory text The characteristic setting unit 130 may set a characteristic based on at least one attribute value in the language model by inputting an explanatory text into the language model that describes a characteristic based on at least one attribute value for at least one subject. That is, for example, the characteristic setting unit 130 may create an explanatory text that describes a set of attribute values of the subject for one or more attributes that should be reflected in the characteristic and input it into the language model. The characteristic setting unit 130 may create an explanatory text that represents the position, background, role, setting, or characteristics corresponding to the attribute value and input it into the language model, for example, "Please output a response to the following text from the following position: ·30s ·Male ·Married ·Owns a car ·Preference for product A is 6 out of 10..." For example, when using ChatGPT (registered trademark) as the language model, the characteristic setting unit 130 may input an explanatory text using prompts (#...) into the language model. When using other language models, if the language model supports a function equivalent to prompts, the characteristic setting unit 130 may use that function to input an explanatory text into the language model. As a result, the text generation unit 150, which uses a language model with defined characteristics, can generate text based on the defined characteristics as preconditions.
[0054] (2) Setting characteristics using files, etc. The characteristic setting unit 130 may set characteristics in the language model by generating a file (text, CSV, PDF, or other file) that describes characteristics based on at least one attribute value for at least one subject and loading it into the language model.
[0055] (3) Setting characteristics through learning The characteristic setting unit 130 may set characteristics in the language model by having the language model learn (additional learning) characteristics based on at least one attribute value for at least one subject. For example, the characteristic setting unit 130 may determine a set of document data to be used for additional learning regarding a certain attribute according to the attribute value of that attribute, and provide it to the language model to perform additional learning. If a certain attribute is a specific attribute value, the characteristic setting unit 130 may provide the language model with more document data related to that specific attribute value compared to the case where the attribute is not a specific attribute value to perform learning. For example, if the subject's address is Osaka, the characteristic setting unit 130 may provide the language model with more document data related to Osaka compared to the case where the subject's address is not Osaka to perform learning.
[0056] The characteristic setting unit 130 may, when the attribute value of an attribute representing preference for a certain matter is higher, provide the language model with more document data containing positive content about that matter compared to when the attribute value of that attribute is lower, and allow it to learn. Similarly, the characteristic setting unit 130 may, when the attribute value of an attribute representing preference for a certain matter is lower, provide the language model with more document data containing negative content about that matter compared to when the attribute value of that attribute is higher, and allow it to learn. For each of the one or more attributes in the attribute DB 20, the characteristic setting unit 130 may store or make available a predetermined set of document data to be used for learning for each attribute value. Furthermore, for each of the one or more attributes in the attribute DB 20, the characteristic setting unit 130 may store or make available a set of document data to be used more often when the attribute value is higher, and / or a set of document data to be used more often when the attribute value is lower. As a result, the text generation unit 150, which uses a language model with set characteristics, can generate text using a language model that has characteristics learned intensively according to the attribute values of the target person.
[0057] The characteristic setting unit 130 may set characteristics for the base language model itself. Here, the base language model may be one that has been pre-trained by the developer or provider of the language model, the administrator or user of the language processing device 100, or other person or company, etc., to the point where it can generate output text from input text at least to a certain level.
[0058] Alternatively, the characteristic setting unit 130 may set characteristics on a copy or instance of the base language model (such as an object that performs processing of the language model). This allows the characteristic setting unit 130 to later generate copies or instances of the language model with different characteristics set, or to generate multiple copies or instances of the language model, each having different characteristics, in parallel.
[0059] The characteristic setting unit 130 stores the language model with set characteristics in the language model storage unit 140. Here, the language model consists of a machine learning model body, such as a Transformer, and a set of parameter values included in the machine learning model, the values of which are determined or adjusted by machine learning. The language model storage unit 140 may store the machine learning model body (or information indicating the structure of the machine learning model body) and the set of parameter values. If the structure of the machine learning model body to be used is predetermined, the language model storage unit 140 may not store information indicating the structure of the machine learning model body, but may store the set of parameter values included in the machine learning model.
[0060] In S330, the text generation unit 150 generates output text corresponding to the input text using a language model in which characteristics based on the attribute values of the target person are set. The text generation unit 150 may be applied to various fields that utilize natural language input and output. As an example, the text generation unit 150 generates output text corresponding to the input text in the following uses or forms.
[0061] (1) Interaction between a language model with defined characteristics and the user of the language processing device 100 The dialogue processing unit 160 may perform a process to conduct a natural language dialogue between a language model with defined characteristics and the user of the language processing device 100 or other people. The dialogue processing unit 160 supplies the text input by the person as input text to the text generation unit 150. The text generation unit 150 generates output text corresponding to the input text using a language model with defined characteristics based on the attribute values of the target person. The dialogue processing unit 160 conveys the generated output text to the person by displaying or reading it aloud. The dialogue processing unit 160 supplies the text input by the person who received the output text as the next input text to the text generation unit 150. The text generation unit 150 generates output text corresponding to the next input text using the language model. By repeating this process, the dialogue processing unit 160 enables a natural language dialogue between the language model with defined characteristics and the user of the language processing device 100 or other people.
[0062] (2) A language model with characteristics set based on the attribute values of the target person, and a dialogue between that target person and the model. The dialogue processing unit 160 facilitates a natural language dialogue between a language model with the subject's characteristics set and the subject (i.e., a subject possessing the attribute values that formed the basis of the characteristics set in the language model). By having the dialogue processing unit 160 interact with a language model that is closer to the subject than the base language model, based on the subject's attribute values, the dialogue processing unit 160 can obtain text created by the real subject, i.e., text containing information, reactions, writing style / speaking style, etc., of the real subject. Furthermore, by making the language model with characteristics based on the subject's attribute values available to the subject, the dialogue processing unit 160 can provide the subject with a language model that responds to input text based on the subject's characteristics. As a result, the language processing unit 100 can provide responses to the dialogue processing unit 160 based on the subject's basic attributes such as address, preferences such as food tastes, or other various attributes, even if the real subject does not provide the language model with preconditions about themselves. Furthermore, the language processing device 100 can reduce the processing load of setting preconditions and other information in the language model while interacting with the user, thereby increasing the effective amount of dialogue that can be processed per unit time.
[0063] (3) Personality assessment of a language model with characteristics set based on the attribute values of the subject The personality assessment unit 170 performs a personality assessment of a language model with defined characteristics by engaging in dialogue with the language model. The personality assessment unit 170 may provide one or more predetermined personality assessment questions as input text to the text generation unit 150, and assess the output text output by the text generation unit 150 in response to the input text, thereby assessing the personality of the language model with defined characteristics. Here, the personality assessment unit 170 may use the same questions used for assessing a person's personality as the one or more predetermined personality assessment questions. The personality assessment unit 170 may assess the personality of the language model with defined characteristics by judging and classifying each output text output by the text generation unit 150 in response to one or more personality assessment questions, in the same way as in a person's personality assessment.
[0064] It should be noted that language models created using machine learning are realized through massive computation and data processing, and are generally understood not to possess human-specific characteristics such as personality, character, or emotions. In this regard, language models created using machine learning learn from a large amount of human-generated text in response to input texts, and the texts created by humans take on content and expression that reflect their personality, etc. Therefore, if a language model is trained to generate the same output texts as those created by a person with a certain personality, etc., for various input texts, it becomes impossible to distinguish between that person and the language model based solely on the input and output texts. A language model trained in this way will appear to have the same personality, etc., as the person in question, at least superficially (based on the input and output texts). In this specification, such superficial personality, etc., that appears in the input and output texts of a language model will also be expressed as the language model's "characteristics," "personality," "virtual personality," or "persona." Against this backdrop, the personality diagnosis unit 170 can diagnose the personality that the target language model appears to possess by performing a personality diagnosis on the language model with set characteristics.
[0065] If the answer to the question is a yes / no choice or a five-point scale, the personality assessment unit 170 may aggregate the selected answers and determine or classify the personality of the language model with defined characteristics. If the answer to the question is in sentence form, the personality assessment unit 170 may use a predetermined personality assessment algorithm to determine or classify the personality of the language model with defined characteristics based on specific words or phrases included in the answer. Alternatively, the personality assessment unit 170 may use a personality assessment model that has been trained to take a sentence-form answer as input and output a personality assessment result to determine or classify the personality of the language model with defined characteristics. The personality assessment unit 170 may provide the results of the personality assessment to the corresponding subject or user of the language processing device 100 (or the subject / user's terminal) by displaying or transmitting them. By performing a personality assessment on a language model with characteristics based on the subject's attribute values, the personality assessment unit 170 can obtain a personality assessment result that is closer to the subject's personality than using a base language model. The language processing device 100 can provide the subject with new insights by feeding back such personality assessment results to the subject. Furthermore, System 10 can utilize these personality assessment results for marketing and other purposes.
[0066] (4) Questionnaire processing using a language model with characteristics set based on the attribute values of the target individuals. When processing the questionnaire, the DB search unit 120 may search in S310 for one or more subjects that meet the search criteria from among the multiple subjects registered in the attribute DB 20. The characteristic setting unit 130 may set one or more characteristics (characteristics for each subject) for the language model for one or more of the searched subjects in S320. The characteristic setting unit 130 may include all subjects in the questionnaire processing, in which case it may set the characteristics and language model for each of the subjects in S320.
[0067] In S330, the survey processing unit 180 receives input text for a survey conducted on language models, each of which has one or more characteristics set, and collects output texts that the language models output in response to the survey. For example, the survey processing unit 180 may receive a question such as "Where would you like to travel?" from each of several language models, each of which has characteristics set based on the attribute values of several subjects in their "20s," and collect the texts output by the language models. Each language model reflects attribute values for attributes such as food preferences, whether or not a car is owned, address, hobbies, etc., depending on the type of attribute used to set the characteristics, and the language model outputs output text corresponding to the attribute values if the question content has any relation to such attributes. Therefore, by conducting a survey on language models with set characteristics, the survey processing unit 180 can obtain answers that are closer to what the subjects actually answer compared to conducting a survey on a base language model.
[0068] The survey processing unit 180 may provide the responses from each language model to the survey by displaying or transmitting them to the user of the language processing unit 100 (or the user's terminal). The survey processing unit 180 may also provide the responses from each language model to the survey by aggregating them and displaying or transmitting them to the user of the language processing unit 100 in the form of a table or graph.
[0069] In S340, at least one of the dialogue processing unit 160, the personality diagnosis unit 170, or the questionnaire processing unit 180 may reflect attributes corresponding to the output text of the text generation unit 150 in the attribute DB 20, or reflect characteristics corresponding to the output text of the text generation unit 150 in the language model. For example, in S330(1), if a user of the language processing device 100 interacts with a language model in which characteristics based on the attribute values of a target person for marketing purposes are set, and as a result the text generation unit 150 generates output text corresponding to the characteristics of the language model, the dialogue processing unit 160 may store the attribute values corresponding to the output text as attribute values of the corresponding attribute in the attribute DB 20. For example, if it is unknown whether a certain target person owns a car, the dialogue processing unit 160 may set attribute values (or predicted values of attributes) for the car ownership attribute of that target person in accordance with the answer output by the text generation unit 150 to a question such as "Do you own a car?".
[0070] For example, in S330(2), the characteristic setting unit 130 may set or update the characteristics set in the language model for a subject based on the content of the dialogue between the subject and the language model, which has characteristics set based on the subject's attribute values. In such a dialogue, the actual subject receives the output text from the language model and inputs a response or other text to the language processing device 100. Therefore, the characteristic setting unit 130 can make the language model closer to the actual subject by training the language model with pairs of output text from the language model and the subject's response or other text as training data. Thus, the characteristic setting unit 130 can overwrite or update the characteristics set in the language model by performing a language model training process using the content of the dialogue.
[0071] For example, the personality diagnosis unit 170 may store the personality of the language model, which has been diagnosed in S330(3) and whose characteristics are set based on the subject's attribute values, in the attribute DB 20 as the subject's attribute values, i.e., the attribute values of "estimated personality attributes". The questionnaire processing unit 180 may store the questionnaire responses obtained in S330(4) from the language model, which has been diagnosed and whose characteristics are set based on the subject's attribute values, in the attribute DB 20 as the subject's attribute values for the attribute corresponding to that questionnaire.
[0072] Furthermore, the characteristic setting unit 130 may store the language model in which characteristics based on the subject's attribute values are set, the parameters of the language model, or the characteristics set in the language model in the record associated with that subject in the attribute DB 20. This reduces the processing load and processing time that the language processing device 100 would have to set the characteristics for each subject in the language model each time.
[0073] According to the language processing device 100 described above, a language model with characteristics based on the attribute values of the subject can be generated and made available for use. As a result, the language processing device 100 can obtain a language model that has a higher probability of outputting output sentences that are closer to the actual responses of the subject to the input sentence, compared to the base language model.
[0074] The language processing device 100 can generate and make available a language model that takes into account the purchase history of a target person, by setting characteristics based on attribute values that indicate the purchase history of one or more products or services of the target person. This allows the language processing device 100 to obtain a language model that, compared to a base language model, has a higher probability of outputting output text that corresponds to the products or services preferred by the target person. The language processing device 100 can then use such a language model to provide marketing opportunities.
[0075] Furthermore, if a new product or service purchase history is stored in the attribute DB20 after the characteristic setting unit 130 has set characteristics for the target person, the characteristic setting unit 130 may update the characteristics of the language model based on at least one attribute value indicating the new product or service purchase history. The characteristic setting unit 130 may create additional explanatory text explaining the attribute value indicating the new product or service purchase history and input it into the language model, similar to (1) in S320. Alternatively, the characteristic setting unit 130 may load a file explaining the attribute value indicating the new product or service purchase history into the language model, similar to (2) in S320. The characteristic setting unit 130 may further train the language model according to the attribute value indicating the new product or service purchase history, similar to (3) in S320. Similarly, if attribute values for other attributes of the target person are stored in the attribute DB20, the characteristic setting unit 130 may update the characteristics of the language model for that target person based on those attribute values. This allows the characteristic setting unit 130 to reflect additions and changes to the attribute values of the target person registered in the attribute DB 20 in the language model and make them available.
[0076] Figure 4 shows an example of questionnaire processing by the questionnaire processing unit 180 according to this embodiment. In this example, the questionnaire processing unit 180 inputs predefined questionnaire items (questions, etc.) into a language model with characteristic A set, and collects output texts that the language model outputs in response to the questionnaire items. In this example, a question bot implemented by the questionnaire processing unit 180 asks a language model with characteristic A set for a certain subject a question about two or more attributes of that subject. The question bot inputs the question sentence "Do you tend to choose foods that are health-conscious?" (a question about the subject's attribute related to "health consciousness in food") into the language model. The text generation unit 150, using the language model with characteristic A set, responds with "Neither" in response to this input sentence. Based on this result, the questionnaire processing unit 180 stores the attribute value of this subject regarding "health consciousness in food" as 50% (moderate) in the attribute DB 20.
[0077] Similarly, the questionnaire processing unit 180 may ask two or more questions about two or more attributes (in the example shown in this figure, "preference for casual clothing," "preference for renting leisure activities," etc.) to a language model in which the characteristics of a certain subject have been set. The questionnaire processing unit 180 may receive the output text output by the text generation unit 150 in response to these questions and determine attribute values from the output text (e.g., "moderate preference for casual clothing," "high preference for renting leisure activities").
[0078] In the example shown in this figure, the questionnaire processing unit 180 can estimate the attribute value of a subject for at least one attribute by conducting a questionnaire on a language model in which the characteristics of the subject have been set, using the output text produced by this language model. The questionnaire processing unit 180 can store the attribute value estimated in this way as the attribute value of the subject in the attribute DB 20. Therefore, the questionnaire processing unit 180 can fill in the attribute values of attributes that have not been set in the attribute DB 20, and allow the attribute DB management device 30 to refer to such attribute values.
[0079] Figure 5 shows the configuration of a language processing device 400 according to the first modified example of this embodiment. Since the language processing device 400 is a modified example of the language processing device 100, components having the same function and configuration as those in Figure 1 are denoted by the same reference numerals, and descriptions are omitted below except for differences. The language processing device 400 according to this modified example differs from the language processing device 100 in Figure 1 in that it includes a characteristic setting unit 430 instead of the characteristic setting unit 130 in Figure 1, and also includes a target user selection unit 435.
[0080] The characteristic setting unit 430 sets candidate characteristics for the language model. The characteristic setting unit 430 may have the same functions as the characteristic setting unit 130 shown in Figures 1 to 4. The following will mainly describe the differences between the characteristic setting unit 430 and the characteristic setting unit 130. The characteristic setting unit 430 may use characteristics pre-set in the language processing unit 400 as candidate characteristics, or it may use randomly generated characteristics as candidate characteristics. Similar to the characteristic setting unit 130, the characteristic setting unit 430 may set candidate characteristics based on at least one attribute value for one or more subjects selected from a plurality of subjects registered in the attribute DB 20. The characteristic setting unit 430 stores the language model with the candidate characteristics set in the language model storage unit 140.
[0081] The subject selection unit 435 is connected to the language model storage unit 140. The subject selection unit 435 estimates the attribute values of the language models for which candidate characteristics have been set. For example, the subject selection unit 435 may estimate the attribute values of the language models by conducting a questionnaire on the language models for which candidate characteristics have been set, similar to the questionnaire processing unit 180.
[0082] The subject selection unit 435 selects at least one subject from among multiple subjects registered in the attribute DB 20 based on the degree of similarity between at least one attribute value estimated to be possessed by the language model for which the characteristic candidate has been set and the corresponding at least one attribute value stored in the attribute DB 20. The subject selection unit 435 may select a subject if the degree of similarity between the set of attribute values estimated to be possessed by the language model for which the characteristic candidate has been set and the set of attribute values stored in the attribute DB 20 for the subject satisfies the selection criteria. The subject selection unit 435 may select a subject having a set of attribute values that has the greatest degree of similarity to the set of attribute values estimated to be possessed by the language model, select a subject having a set of attribute values whose degree of similarity to the set of attribute values estimated to be possessed by the language model is equal to or greater than a predetermined threshold, or use other selection criteria to prioritize the selection of subjects with a high degree of similarity to attribute values. Here, the degree of approximation between pairs of attribute values may be the sum of the absolute values of the differences between each attribute value, the sum of the squares or mean squares of the absolute values of the differences between each attribute value, the difference between attribute vectors (or normalized attribute vectors) whose elements are each attribute value, or any other index value that indicates error, which increases when the index value indicating error is small and decreases when the index value indicating error is large. The characteristic setting unit 130 uses the characteristic candidate as the characteristic of the selected at least one subject when at least one subject is selected according to the characteristic candidate.
[0083] Figure 6 shows the operation flow of the language processing device 400 according to the first modified example. In S600, the characteristic setting unit 430 sets one or more characteristic candidates for the language model. The characteristic setting unit 430 stores the language model with the characteristic candidates set in the language model storage unit 140.
[0084] In S610, the subject selection unit 435 estimates one or more attribute values for each of the one or more candidate characteristics that the language model to which the characteristic candidate is set possesses. For each candidate characteristic, the subject selection unit 435 selects one or more subjects from among the multiple subjects based on the degree of similarity between each attribute value estimated to be possessed by the language model to which the characteristic candidate is set and the attribute value of the corresponding attribute stored in the attribute DB 20.
[0085] In S620, the characteristic setting unit 430 sets a language model using the characteristic candidate for each characteristic if one or more subjects are selected according to the characteristic candidate. The characteristic setting unit 430 may use the language model in which the characteristic candidate is set as a language model in which characteristics based on attribute values for that subject are set.
[0086] By performing the processing from S600 to S620, the language processing device 400 can prepare a language model with defined characteristics for at least one of the multiple subjects registered in the attribute DB20. The language processing device 400 may proceed to S300 in Figure 3 after S620. If a language model with defined characteristics for a subject retrieved by the DB search unit 120 in S310 has already been prepared, the characteristic setting unit 430 may provide that language model to the text generation unit 150 in S320.
[0087] According to the language processing device 400 described above, candidate characteristics are set in a language model, and if the set of attribute values of a language model with candidate characteristics set approximates a set of attribute values of any subject, that language model can be selected as the language model for that subject. In this way, a language model with a high degree of agreement with the subject's characteristics can be selected.
[0088] Figure 7 shows the configuration of a language processing device 700 according to a second modified example of this embodiment. Since the language processing device 700 according to this modified example is a modified version of the language processing device 100 shown in Figure 1 and the language processing device 400 shown in Figure 5, components having the same functions and configuration as those in Figures 1 and 5 are denoted by the same reference numerals, and descriptions are omitted below except for differences. The language processing device 700 according to this modified example simulates a discussion or group discussion by repeatedly exchanging sentences between language models, each of which has multiple characteristics set.
[0089] The language processing unit 700 is connected to the virtual personality candidate DB 710 and includes a characteristic setting unit 720, a language model storage unit 140, an agenda acquisition unit 752, a text generation unit 150, a repetition processing unit 754, a characteristic selection unit 756, and a summary processing unit 758. The virtual personality candidate DB 710 stores the characteristics of a plurality of predefined virtual personality candidates.
[0090] The characteristic setting unit 720 and the characteristic setting unit 430 may have the same functions as the characteristic setting unit 130 shown in Figures 1 to 4, or the characteristic setting unit 430 shown in Figures 5 to 6. Below, the differences between the characteristic setting unit 720 and the characteristic setting units 130 and 430 will be mainly explained. The characteristic setting unit 720 sets each of multiple characteristics for the language model. As a result, the characteristic setting unit 720 generates N language models (where N is an integer greater than or equal to 2), such as a language model with a first characteristic set, a language model with a second characteristic set, ... a language model with the Nth characteristic set. In this modified example, the characteristic setting unit 720 sets each of the multiple characteristics for the language model, which are characteristics associated with a virtual personality candidate specified by the user from among multiple virtual personality candidates that are predefined and stored in the virtual personality candidate DB 710.
[0091] The language model storage unit 140 is the same as the language model storage unit 140 in Figures 1 and 4. The language model storage unit 140 is connected to the characteristic setting unit 720 and stores N language models generated by the characteristic setting unit 720.
[0092] The agenda acquisition unit 752 acquires agenda items to be discussed using the language processing device 700 from the user of the language processing device 700. The agenda acquisition unit 752 may have an input / output circuit or a transmit / receive circuit, and may receive agenda items from the user of the language processing device 700 via an input / output device used by the user of the language processing device 700 or a terminal device used by the user of the language processing device 700.
[0093] The text generation unit 150 is the same as the text generation unit 150 in Figures 1 and 4. The text generation unit 150 is connected to the language model storage unit 140. The text generation unit 150 generates output texts for each of the multiple characteristics (N characteristics) of the input text using the N language models stored in the language model storage unit 140.
[0094] The repetition processing unit 754 is connected to the agenda acquisition unit 752 and the text generation unit 150. The repetition processing unit 754 uses multiple language models, each with multiple characteristics set, executed by the text generation unit 150, to conduct discussions on the agenda supplied by the agenda acquisition unit 752. The repetition processing unit 754 uses the output text output according to one of the multiple characteristics as input text to be given to the language model according to the other characteristics, and repeatedly generates the next output text according to the other characteristics. In this way, the repetition processing unit 754 causes the text generation unit 150 to generate output texts from the language model from the perspective of other characteristics (i.e., the perspective of other virtual personalities) for the output text output by the language model from the perspective of one characteristic (i.e., the perspective of one virtual personality), and collects statements according to each of the multiple perspectives.
[0095] The characteristic selection unit 756 is connected to the repetition processing unit 754. Based on the output sentences output by the repetition processing unit 754, the characteristic selection unit 756 selects the next characteristic to be used to generate the next output sentence from among multiple characteristics. When the sentence generation unit 150 outputs an output sentence according to one characteristic, the characteristic selection unit 756 selects a characteristic of the language model to be used to generate the next output sentence based on that output sentence or the most recent one or more output sentences and instructs the characteristic setting unit 720. The characteristic setting unit 720 sets the characteristic instructed by the characteristic selection unit 756 in the language model, or selects a language model in which the characteristic instructed by the characteristic selection unit 756 has already been set, and supplies it to the sentence generation unit 150 via the language model storage unit 140.
[0096] The summarization processing unit 758 is connected to the repetition processing unit 754. The summarization processing unit 758 receives output sentences generated according to each of the multiple characteristics by the repetition processing of the repetition processing unit 754 and summarizes the discussion. The summarization processing unit 758 may use a language model to generate an output sentence that summarizes the output sentences generated according to each characteristic and output it to a display device or terminal device used by the user of the language processing device 700.
[0097] Figure 8 shows the operation flow of the language processing device 700 according to the second modified example. In S800, the characteristic setting unit 720 determines a plurality of characteristics to be set for the language model. In this modified example, the characteristic setting unit 720 determines the characteristics to be set for the language model that are associated with a virtual personality candidate specified by the user of the language processing device 700 from among a plurality of virtual personality candidates stored in the virtual personality candidate DB 710 in advance.
[0098] In S810, the agenda acquisition unit 752 acquires the agenda entered by the user of the language processing device 700. In S820, the repetition processing unit 754 generates an input document based on the agenda acquired by the agenda acquisition unit 752. The repetition processing unit 754 may use the agenda entered by the user as the first input document, or it may generate an input document in which instruction documents that instruct the language model, for which each characteristic has been set, on the agenda document, or it may add instruction documents that instruct the language model, for which each characteristic has been set, on the agenda document. For example, the repetition processing unit 754 may add instruction documents to the agenda document such as "The output document shall be no more than 800 characters," "The time to respond shall be no more than 1 minute," or "The response may be made considering each document after this text." using instruction documents that are specified or selected by the user of the language processing device 700, or instruction documents that are pre-set in the language processing device 700. The repetition processing unit 754 may input the agenda to a language model that has been set with characteristics as a moderator, thereby generating a document containing the minutes of the meeting according to the agenda, and using this document as input to the language model with each set of characteristics.
[0099] In S830, the characteristic selection unit 756 selects the next characteristic (or the first characteristic) to be used to generate the next output document (or the first output document if no output documents corresponding to any position have been generated). The characteristic selection unit 756 may randomly or in a fixed order select the next characteristic from among several characteristics used in the discussion. The characteristic selection unit 756 may prioritize selecting the characteristic with the smallest total amount of output documents generated so far as the next characteristic, so that the total amount of output documents generated for each characteristic (e.g., number of characters, words, items, etc.) approaches equality. In this way, the characteristic selection unit 756 can function as a facilitator, moderator, or administrator of the discussion and move the discussion forward.
[0100] The characteristic selection unit 756 may select the next characteristic to be used to generate the next output sentence from among multiple characteristics based on the output sentence output by the repetitive processing by the repetitive processing unit 754. The characteristic selection unit 756 may also select the next characteristic to be used to generate the next output sentence based on the relationship between the output sentence from a language model with a certain characteristic set and each previous output sentence from each language model with other characteristics set.
[0101] For example, the characteristic selection unit 756 may select another characteristic as the next characteristic if the output text from a language model with a certain characteristic set contains content opposite to that of a previous output text from a language model with a different characteristic set, i.e., content that shows the opposite facts, content that shows the opposite viewpoint, or content that shows the opposite conclusion. The characteristic selection unit 756 may also select another characteristic as the next characteristic if the output text from a language model with a certain characteristic set contains content that is identical or has a higher degree of similarity to a previous output text from a language model with a different characteristic set, and may increase or decrease the priority of selecting such another characteristic as the next characteristic compared to cases where these output texts are different or have a lower degree of similarity. This allows the characteristic selection unit 756 to act as a facilitator, moderator, or administrator of a discussion, and to guide the discussion more appropriately according to its flow.
[0102] The characteristic setting unit 720 sets the selected characteristics in the language model. The characteristic setting unit 720 may set characteristics in the language model using various methods shown in relation to S320 in Figure 3. For example, the characteristic setting unit 720 may set characteristics in the language model by inputting an explanatory text describing the characteristics to be set into the language model (see (1) in S320 in Figure 3). Alternatively, the characteristic setting unit 720 may set characteristics in the language model by generating a file describing the characteristics and having the language model read it (see (2) in S320 in Figure 3). Alternatively, the characteristic setting unit 720 may set characteristics in the language model by having the language model learn the characteristics to be set (see (3) in S320 in Figure 3).
[0103] The characteristic setting unit 720 may use a language model shared for at least two characteristics. In this case, the characteristic setting unit 720 may switch the characteristics of the language model to the other characteristics when providing an output sentence output according to one characteristic as an input sentence to a language model corresponding to another characteristic. For example, each time the characteristic selection unit 756 selects the next characteristic, the characteristic setting unit 720 may reset the language model to the selected next characteristic.
[0104] The characteristic setting unit 720 may store multiple language models, each with multiple characteristics set, in the language model storage unit 140, and allow the text generation unit 150 to use a language model corresponding to the next characteristic from among the multiple language models. The characteristic setting unit 720 may set each of the multiple characteristics for each of the multiple instances of the language model, and allow the text generation unit 150 to use an instance corresponding to the next characteristic from among the multiple instances.
[0105] In S840, the repetition processing unit 754 provides an input sentence to the sentence generation unit 150, which processes a language model with the following characteristics set. The sentence generation unit 150 uses the language model with the following characteristics set to generate the next output sentence from the input sentence according to the following characteristics.
[0106] In S850, the iteration processing unit 754 determines whether or not to terminate the discussion among the language models for which each of the multiple characteristics has been set. The iteration processing unit 754 may determine to terminate the iteration when the processing from S830 to S840 has been repeated a predetermined number of times, i.e., for N or more times, an integer multiple of N (an integer multiple of 2 or more), or any other arbitrary number of times.
[0107] The repetition processing unit 754 may determine to terminate the repetition if the newly generated output text is the same as the content of the previously generated set of output texts. For example, the repetition processing unit 754 may determine to terminate the repetition if the number of words in the newly generated output text that are not included in the previously generated set of output texts is less than or equal to a predetermined threshold. The repetition processing unit 754 may determine to terminate the repetition if the number of words in the newly generated output text that are not included in the previously generated set of output texts is less than or equal to a predetermined threshold is repeated a predetermined number of times (for example, N or more times). This allows the repetition processing unit 754 to terminate the discussion when the discussion is converging and it is becoming difficult to make new statements (output texts). Therefore, the repetition processing unit 754 can reduce the processing load on the language processing unit 700 caused by continuing the discussion when it is difficult to make new statements, thereby reducing the processing load on the language processing unit 700. When terminating the repetition, the repetition processing unit 754 proceeds to S860. If the repetition continues, the repetition processing unit 754 proceeds to S830.
[0108] In S860, the summarization processing unit 758 summarizes the multiple output sentences output according to each of the multiple characteristics. The summarization processing unit 758 may use a language model to input the output sentences output according to each of the multiple characteristics and generate a summarized output sentence. Specifically, the summarization processing unit 758 may input the output sentences output according to each of the multiple characteristics in chronological order to the sentence generation unit 150 and generate a summarized output sentence (also referred to as the "summarization sentence") using a language model. The summarization processing unit 758 may generate the summarized output sentence using a language model other than the language model for which characteristics have been set for discussion, such as a base language model or a language model with settings for summarizing output sentences. The summarization processing unit 758 may output the generated summarized sentence to the user of the language processing device 700.
[0109] The summarization processing unit 758 may summarize the results of the discussion by taking a majority vote or the like for the language models to which each characteristic has been set. The summarization processing unit 758 may input questions such as "Based on the results of the discussion so far, do you agree or disagree with XXXX?" to the text generation unit 150 that processes the language models to which each characteristic has been set, collect answers according to each characteristic, and determine a conclusion by aggregating the answers. The summarization processing unit 758 may collect and aggregate answers in the same manner as the questionnaire processing unit 180 shown in Figures 1 to 4. The summarization processing unit 758 may output the generated summary result to the user of the language processing device 700.
[0110] According to the language processing device 700 described above, by setting each of multiple characteristics in a language model, it is possible to have discussions on a given topic take place between multiple language models, each with a pseudo-personality, virtual personality, or persona corresponding to each of the multiple characteristics. As a result, the language processing device 700 can provide the results of discussions between language models, each with its own unique personality, character, or characteristics, instead of gathering multiple people and having them discuss the topic. Such a language processing device 700 can provide the results of discussions on a given topic on demand at the time needed by the user of the language processing device 700, and can output more discussion results faster compared to having multiple people discuss the topic.
[0111] Figure 9 shows an example of a persona selection screen according to the second modified example. In this example, the characteristic setting unit 720 displays a selection screen to the user of the language processing device 700 for selecting a virtual personality (persona) to participate in the discussion from among a plurality of predefined virtual personality candidates (persona candidates). The characteristic setting unit 720 may display the persona selection screen on an input / output device (keyboard, display, etc.) used by the user and receive instructions from the user, or it may display the persona selection screen on a terminal device used by the user via a network and receive instructions from the user.
[0112] In the example shown in this figure, the characteristic setting unit 720 displays multiple virtual personality candidates (personas A to I) stored in the virtual personality candidate DB 710 on the screen in icon format. The virtual personality candidate DB 710 stores identification information (ID, etc.) and characteristics to be set in the language model for each virtual personality candidate. The virtual personality candidate DB 710 may also store at least one of the following for each virtual personality candidate: name, appearance image (photo or avatar, etc.), description, or other information. In the example shown in this figure, the characteristic setting unit 720 displays the identification information, name, and appearance image for each virtual personality candidate, such as persona A.
[0113] The characteristic setting unit 720 may display a description or other information stored in the virtual personality candidate DB 710 for each virtual personality candidate. The characteristic setting unit 720 may also display detailed information (description, characteristics, etc.) of a virtual personality candidate in response to a user instruction to display the detailed information of that virtual personality candidate. The characteristic setting unit 720 may also display the characteristics of a virtual personality candidate using a table or graph, etc., that shows attribute values for one or more attributes, as shown in the lower part of Figure 4, as an example.
[0114] In the example shown in this figure, the user instructs the characteristic setting unit 720 to select personas A, D, and E. The characteristic setting unit 720 sets characteristics for the language model that correspond to one or more virtual personality candidates specified by the user from among multiple virtual personality candidates. By providing predefined virtual personality candidates that can be selected by the user, the language processing unit 700 can reduce the processing time required to generate the characteristics of virtual personalities after receiving instructions from the user. Furthermore, by presenting the user with multiple virtual personality candidates, the language processing unit 700 can reduce the effort required for the user to customize the virtual personalities to participate in the discussion, as well as the processing load on the language processing unit 700.
[0115] The characteristic setting unit 720 may also assign one or more virtual personality candidates from among the multiple virtual personality candidates to specific roles such as moderator, facilitator, or other roles that facilitate or direct discussions or dialogues between language models with set characteristics. For example, the characteristic setting unit 720 may set characteristics for a moderator in the language model such as "summarize recent statements, present new perspectives that have not been discussed, etc."
[0116] The language processing device 700 may determine each characteristic to be set in the language model by other means. For example, the language processing device 700 may determine the characteristics to be set in the language model in the same manner as the language processing device 100 shown in Figures 1 to 4. That is, the language processing device 700 may include a DB connection unit 110 connected to the attribute DB 20, similar to the language processing device 100 in Figure 1. The characteristic setting unit 720 may function as a characteristic setting unit 130. The characteristic setting unit 720 is connected to the DB connection unit 110 and may set characteristics in the language model based on at least one attribute value for at least one subject stored in the attribute DB 20, similar to S300 to S320 in Figure 3. The language processing device 700 may then include the language model with the characteristics set in this manner in the discussion.
[0117] The language processing unit 700 may determine the characteristics to be set in the language model in the same manner as the language processing unit 400 shown in Figures 5 and 6. That is, the language processing unit 700 may include a DB connection unit 110, a DB search unit 120, and a target user selection unit 435, similar to the language processing unit 100 in Figure 5. The characteristic setting unit 720 may function as a characteristic setting unit 430. The characteristic setting unit 720 sets one or more characteristic candidates for the language model, similar to S600 in Figure 6, and stores the language model with the characteristic candidates set in the language model storage unit 140. The target user selection unit 435 estimates the attribute values that the language model with the characteristic candidates has for each of the one or more characteristic candidates, similar to S610 in Figure 6. The subject selection unit 435 selects one or more subjects from among multiple subjects based on the degree of similarity between the attribute values estimated to be present in the language model for which the characteristic candidate is set and the corresponding attribute values stored in the attribute DB 20.
[0118] The characteristic setting unit 430, similar to S620 in Figure 6, sets a language model for each characteristic candidate by using that characteristic candidate as the characteristic of the selected subject when one or more subjects are selected according to the characteristic candidate. The characteristic setting unit 430 may use the language model in which the characteristic candidate is set as a language model in which characteristics based on the attribute values of that subject are set, and have it participate in the discussion. In addition, the subject selection unit 435 may identify subjects from among multiple subjects registered in the attribute DB 20 that are close to the selected virtual personality candidate or virtual personality (e.g., those with the highest degree of similarity or those with a degree of similarity above a predetermined threshold) based on the degree of similarity to the virtual personality candidate stored in the virtual personality candidate DB 710 or the virtual personality selected by the user.
[0119] Figure 10 shows an example of a discussion screen according to the second modified example. In this example, the language processing device 700 performs the process of displaying a discussion screen to the user for inputting topics, displaying discussion content, and displaying a summary. The language processing device 700 may display the discussion screen on an input / output device (keyboard, display, etc.) used by the user and receive instructions from the user, or it may display the discussion screen on a terminal device used by the user via a network and receive instructions from the user.
[0120] The agenda acquisition unit 752 may acquire agenda items entered by the user in an agenda input field such as "(Agenda Input)" in the example shown in Figure 8, as shown in S810. The repetition processing unit 754 generates input text for presenting agenda items based on the agenda items acquired by the agenda acquisition unit 752, as shown in S820 in Figure 8. In the example shown in Figure 8, the repetition processing unit 754 inputs agenda items to a language model (referred to as "Moderator Bot" in the figure) with characteristics as a moderator, which is executed by the text generation unit 150, generates text for presenting agenda items, and supplies it to the language model with each characteristic set. The repetition processing unit 754 may perform a process to display the text for presenting agenda items in an agenda presentation field such as "(Agenda Presentation)" in the example shown in Figure 8.
[0121] As shown in S830 of Figure 8, the characteristic selection unit 756 selects the next virtual personality to be used to generate the next output text. In the example shown in this figure, the characteristic selection unit 756 selects the virtual personalities in the order of Persona A, Persona E, and Persona D. When the characteristic selection unit 756 selects a virtual personality, the characteristic setting unit 720 sets the characteristics of the selected virtual personality in the language model.
[0122] As shown in S840 of Figure 8, the repetition processing unit 754 inputs the output text of a language model with the characteristics of the previous virtual personality (or two or more previous virtual personalities) to the text generation unit 150, which processes a language model with the characteristics of the next virtual personality set. In this example, the repetition processing unit 754 supplies the output text of the previous persona E (indicated as "(Opinion 2)" in the figure), the output texts of the previous personas A and E (indicated as "(Opinion 1)" and "(Opinion 2)" in the figure), or the output texts of the previous personas A and E and the moderator Bot (indicated as "Presentation of Agenda," "(Opinion 1)," and "(Presentation of Agenda)" in the figure) to the language model with the characteristics of persona D. As a result, the text generation unit 150 can generate output text corresponding to the previous output texts using a language model with the characteristics of the next virtual personality set. The repetition processing unit 754 may perform the process of displaying the output sentences generated using the language model with each characteristic set in the opinion presentation fields, such as "(Opinion 1)", "(Opinion 2)", and "(Opinion 3)" in the example shown in the figure.
[0123] The repetition processing unit 754 determines whether or not to terminate the discussion between language models, each of which has multiple characteristics set, as shown in S850 of Figure 8. In the example shown in this figure, the repetition processing unit 754 determines to terminate the discussion after the language model with the characteristics of Persona D has output the output sentence ("(Opinion 3)").
[0124] The summarization processing unit 758 summarizes the multiple output sentences output according to each of the multiple virtual personalities, as shown in S860 of Figure 8. In the example shown in this figure, the summarization processing unit 758 may generate a summary sentence from the multiple output sentences output according to each of the multiple virtual personalities and display it in the summary bulletin board in the form of a summary by the moderator bot, as shown as "(summary)" in the figure.
[0125] The language processing device 700 described above allows for discussions on a given topic among multiple language models, each with characteristics corresponding to multiple virtual personalities. The language processing device 700 provides a user interface that allows users to easily select multiple virtual personalities to participate in discussions using a persona selection screen as shown in Figure 9 and a discussion screen as shown in Figure 10, and to easily review the opinions and summaries of the discussion from each virtual personality's perspective.
[0126] In the language processing devices 100, 400, and 700 described above, the text generation unit 150 outputs output text from one language model or one instance of a language model, each with characteristics corresponding to a virtual personality. Alternatively, the characteristic setting unit 130, 430, or 720 may set characteristics to be set (characteristics corresponding to a virtual personality) from among multiple characteristics to two or more language models or two or more instances of a language model. In this case, the text generation unit 150 may output a text that summarizes the output texts from two or more language models or two or more instances with the characteristics set, as the output text corresponding to those characteristics.
[0127] For example, the characteristic setting unit 720 in the language processing device 700 may set the characteristics of multiple subjects in their "20s" stored in the attribute DB 20 into the language model, set the characteristics of multiple subjects in their "30s" into the language model, and set the characteristics of multiple subjects in their "40s" into the language model. When the text generation unit 150 generates output text from the perspective of a "20s" person, it may summarize the output texts output by the language model in which the characteristics of multiple subjects in their "20s" have been set and output it as output text corresponding to the characteristics of a "20s" person. For example, the text generation unit 150 may extract content from each output text that is included in a predetermined proportion of the output texts and include it in the summary text, or it may simply concatenate each output text as is to create the summary text, or it may generate the summary text by other methods. In this way, the text generation unit 150 can generate and output output text related to common characteristics from output texts with multiple characteristics that share at least some of those characteristics.
[0128] Various embodiments of the present invention may be described with reference to flowcharts and block diagrams, where a block may represent (1) a stage in a process in which an operation is performed or (2) a section of a device having the role of performing the operation. Specific stages and sections may be implemented by dedicated circuits, programmable circuits supplied with computer-readable instructions stored on a computer-readable medium, and / or processors supplied with computer-readable instructions stored on a computer-readable medium. Dedicated circuits may include digital and / or analog hardware circuits, and may include integrated circuits (ICs) and / or discrete circuits. Programmable circuits may include reconfigurable hardware circuits, including logical AND, logical OR, logical XOR, logical NAND, logical NOR, and other logic operations, flip-flops, registers, memory elements such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), etc.
[0129] Computer-readable media may include any tangible device capable of storing instructions to be executed by a suitable device, and as a result, computer-readable media having instructions stored therein will comprise a product containing instructions that can be executed to create means for performing operations specified in a flowchart or block diagram. Examples of computer-readable media may include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, etc. More specific examples of computer-readable media may include floppy disks, diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disk read-only memory (CD-ROM), digital multipurpose disc (DVD), Blu-ray® disc, memory stick, integrated circuit card, etc.
[0130] Computer-readable instructions may include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as SMALLTALK®, JAVA®, C++, and traditional procedural programming languages such as the C programming language or similar programming languages.
[0131] Computer-readable instructions are provided locally or via a wide area network (WAN) such as a local area network (LAN) or the internet to the processor or programmable circuit of a programmable data processing device such as a computer, and may be executed to create means for performing operations specified in a flowchart or block diagram. Here, the computer may be a PC (personal computer), tablet computer, smartphone, workstation, server computer, general-purpose computer, or special-purpose computer, and may also be a computer system in which multiple computers are connected. Such a computer system in which multiple computers are connected is also called a distributed computing system and is a computer in a broad sense. In a distributed computing system, multiple computers execute a program collectively by each computer executing a part of the program and passing data during program execution between computers as needed.
[0132] Examples of processors include computer processors, central processing units (CPUs), processing units, microprocessors, digital signal processors, controllers, and microcontrollers. A computer may have one or more processors. In a multiprocessor system with multiple processors, each processor executes a portion of the program, and the processors collectively execute the program by passing program execution data between them as needed. For example, in the execution of multitasking, each of the multiple processors may execute a portion of each task in small chunks by switching tasks at each time slice. In this case, which part of a program each processor executes changes dynamically. Which part of a program each of the multiple processors executes may also be statically determined by multiprocessor-aware programming.
[0133] Figure 11 shows an example of a computer 2200 in which multiple aspects of the present invention may be embodied in whole or in part. A program installed on the computer 2200 can cause the computer 2200 to function as an operation or one or more sections of an apparatus according to an embodiment of the present invention, or to execute such operation or one or more sections, and / or to cause the computer 2200 to execute a process or a stage of such process according to an embodiment of the present invention. Such a program may be executed by the CPU 2212 to cause the computer 2200 to perform a particular operation associated with some or all of the blocks in the flowcharts and block diagrams described herein.
[0134] The computer 2200 according to this embodiment includes a CPU 2212, RAM 2214, a graphics controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes a communication interface 2222, storage devices such as a hard disk drive 2224, an input / output unit such as a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.
[0135] The CPU 2212 operates according to programs stored in the ROM 2230 and RAM 2214, thereby controlling each unit. The graphics controller 2216 retrieves image data generated by the CPU 2212 from a frame buffer provided in RAM 2214 or from itself, and displays the image data on the display device 2218.
[0136] The communication interface 2222 communicates with other electronic devices via the network. Storage devices such as the hard disk drive 2224 store programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads programs or data from the DVD-ROM 2201 and provides them to storage devices such as the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from the IC card and / or writes programs and data to the IC card.
[0137] The ROM 2230 stores boot programs and / or programs that depend on the computer 2200's hardware, which are executed by the computer 2200 when activated. The input / output chip 2240 may also connect various input / output units to the input / output controller 2220 via parallel ports, serial ports, keyboard ports, mouse ports, etc.
[0138] The program is provided on a computer-readable medium such as a DVD-ROM 2201 or an IC card. The program is read from the computer-readable medium and installed on a storage device such as a hard disk drive 2224 (an example of a computer-readable medium), RAM 2214, or ROM 2230, and executed by the CPU 2212. The information processing described within these programs is read by the computer 2200, resulting in coordination between the program and the various types of hardware resources described above. The apparatus or method may be configured to realize the manipulation or processing of information in accordance with the use of the computer 2200.
[0139] For example, when communication is performed between a computer 2200 and an external device, the CPU 2212 may execute a communication program loaded into RAM 2214 and, based on the processing described in the communication program, instruct the communication interface 2222 to perform communication processing. Under the control of the CPU 2212, the communication interface 2222 reads transmission data stored in a transmission buffer processing area provided in a storage device such as RAM 2214, a hard disk drive 2224, a DVD-ROM 2201, or a recording medium such as an IC card, transmits the read transmission data to the network, or writes received data received from the network to a receive buffer processing area provided on the recording medium.
[0140] Furthermore, the CPU 2212 may read all or necessary parts of files or databases stored on external storage media such as a hard disk drive 2224, a DVD-ROM drive 2226 (DVD-ROM 2201), or an IC card into the RAM 2214, and perform various types of processing on the data in the RAM 2214. The CPU 2212 then writes the processed data back to the external storage media.
[0141] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and subjected to information processing. The CPU 2212 may perform various types of processing on the data read from RAM 2214, including various types of operations, information processing, conditional judgments, conditional branching, unconditional branching, information retrieval / replacement, etc., as described throughout this disclosure and specified by the program instruction sequence, and write the results back to RAM 2214. The CPU 2212 may also retrieve information in files, databases, etc., within the recording medium. For example, if multiple entries are stored in the recording medium, each having an attribute value of a first attribute associated with an attribute value of a second attribute, the CPU 2212 may search among the multiple entries for an entry that matches the condition for which the attribute value of the first attribute is specified, read the attribute value of the second attribute stored in that entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.
[0142] The programs or software modules described above may be stored on or near computer 2200 on a computer-readable medium. Alternatively, recording media such as hard disks or RAM provided within a server system connected to a dedicated communication network or the Internet can be used as computer-readable media, thereby providing programs to computer 2200 via the network.
[0143] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications or improvements can be made to the above embodiments. It will be clear from the claims that such modified or improved forms may also be included in the technical scope of the present invention.
[0144] It should be noted that the execution order of operations, procedures, steps, and stages in the apparatus, systems, programs, and methods shown in the claims, specifications, and drawings is not explicitly stated as "before," "prior to," etc., and that these can be implemented in any order unless the output of a previous process is used in a later process. Even if the operation flow in the claims, specifications, and drawings is described using phrases such as "first," "next," etc. for convenience, it does not mean that it is essential to perform the operations in that order. [Explanation of Symbols]
[0145] 10 System, 20 Attribute DB, 30 Attribute DB Management Device, 40 Attribute Information Acquisition Unit, 60 Marketing Processing Unit, 100 Language Processing Unit, 110 DB Connection Unit, 120 DB Search Unit, 130 Characteristic Setting Unit, 140 Language Model Storage Unit, 150 Text Generation Unit, 160 Dialogue Processing Unit, 170 Personality Diagnosis Unit, 180 Questionnaire Processing Unit, 400 Language Processing Unit, 430 Characteristic Setting Unit, 435 Target Selection Unit, 700 Language Processing Unit, 710 Virtual Personality Candidate DB, 720 Characteristic Setting Unit, 752 Agenda Acquisition Unit, 754 Repetition Processing Unit, 756 Characteristic Selection Unit, 758 Overall Processing Unit, 2200 Computer, 2201 DVD-ROM, 2210 Host Controller, 2212 CPU, 2214 RAM, 2216 Graphics Controller, 2218 Display Device, 2220 Input / Output Controller, 2222 Communication Interface, 2224 Hard Disk Drive, 2226 DVD-ROM Drive, 2230 ROM, 2240 Input / Output Chip, 2242 Keyboard
Claims
1. A database connection unit that connects to an attribute database for storing multiple attribute values corresponding to multiple attributes for each of multiple subjects, A characteristic setting unit sets characteristics based on at least one attribute value for at least one subject stored in the attribute database for a language model that generates natural language output text corresponding to natural language input text. A text generation unit that generates output text corresponding to input text using the language model with the above characteristics set, A device equipped with the following features.
2. The system includes a database search unit that searches the attribute database for at least one of the aforementioned multiple subjects that meets the specified search criteria, The characteristic setting unit sets the characteristic for the language model based on the at least one attribute value for the at least one subject. The apparatus according to claim 1.
3. The apparatus according to claim 1, wherein the characteristic setting unit sets the characteristic based on the at least one attribute value in the language model by inputting an explanatory text describing the characteristic based on the at least one attribute value for the at least one subject into the language model.
4. The characteristic setting unit sets candidate characteristics for the language model, The subject selection unit selects at least one subject from among the plurality of subjects based on the degree of similarity between at least one attribute value estimated to be possessed by the language model in which the candidate characteristics of the aforementioned characteristics are set and the corresponding at least one attribute value stored in the attribute database. The apparatus according to claim 1, wherein the characteristic setting unit, when at least one subject is selected according to the characteristic candidates, uses the characteristic candidates as the characteristic of the selected at least one subject.
5. The apparatus according to claim 1, further comprising a dialogue processing unit that causes a natural language dialogue to take place between the language model, which has the aforementioned characteristics set, and the selected at least one subject.
6. The apparatus according to claim 5, wherein the characteristic setting unit sets or updates the characteristics based on the content of the dialogue.
7. The apparatus according to claim 1, wherein the characteristic setting unit sets or updates the characteristic based on the at least one attribute value indicating the purchase history of a product or service stored in the attribute database.
8. The apparatus according to claim 1, further comprising a personality diagnosis unit that performs a personality diagnosis of the language model in which the characteristics have been set by engaging in dialogue with the language model in which the characteristics have been set.
9. The characteristic setting unit sets one or more characteristics for one or more of the multiple subjects to the language model, The system includes a survey processing unit that takes input text for a survey as input to the language model, each of which has been configured with one or more of the aforementioned characteristics, and collects output text that the language model outputs in response to the survey. The apparatus according to claim 1.
10. The device accesses an attribute database to store multiple attribute values corresponding to multiple attributes for each of multiple subjects, The device sets characteristics for a language model that generates natural language output text corresponding to natural language input text, based on at least one attribute value for at least one subject stored in the attribute database. The device generates output text corresponding to input text using the language model whose characteristics have been set. A method that includes this.
11. It is executed by a computer, and the computer, A database connection unit that connects to an attribute database for storing multiple attribute values corresponding to multiple attributes for each of multiple subjects, A characteristic setting unit sets characteristics based on at least one attribute value for at least one subject stored in the attribute database for a language model that generates natural language output text corresponding to natural language input text. A text generation unit that generates output text corresponding to input text using the language model with the above characteristics set, A program that makes something work.