Information processing device and nonvolatile storage medium

The information processing device uses generation AI to generate virtual survey answers based on SNS and review site posts, addressing cost and time inefficiencies and bias in internet surveys, ensuring high reliability.

WO2025258590A1PCT designated stage Publication Date: 2025-12-18AIQ INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/020941
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-10
Filing Date
2025-06-10
Publication Date
2025-12-18

AI Technical Summary

Technical Problem

Internet surveys face challenges such as high costs, lengthy response times, and biased respondent attributes, leading to reduced reliability due to user discretion and anonymity.

Method used

An information processing device that utilizes a generation AI to generate 'virtual answers' based on posts from social networking services (SNS), blogs, and review sites, ensuring the answers are similar to those a real user would give, thereby reducing costs and time while enhancing reliability.

Benefits of technology

This approach reduces survey costs and time, eliminates response bias, and improves answer reliability by generating virtual answers that closely mirror real user responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025020941_18122025_PF_FP_ABST
    Figure JP2025020941_18122025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device 1 is connected to a service server 3 such as an SNS server, a questionnaire survey request terminal 4, and a generative AI 2 via a communication line. The information processing device 1 comprises: a means for receiving a questionnaire survey request from an information processing terminal; a means for collecting, from the server, data of a plurality of posted texts posted by a plurality of users to a service such as SNS; a feature quantity calculation means for calculating a content-based feature quantity for each of the questions and the plurality of posted texts; a means for comparing the feature quantities of the posted texts with the feature quantities of the questions, and extracting a posted text for each user; a means for generating a prompt for requesting that answers to the questions be generated for each user on the basis of the extracted posted text; a means for transmitting the prompt to a generation device; a means for receiving, from the generation device, data of the answers generated in accordance with the prompt by the generation device; and a means for transmitting the received answers and the result of aggregation to the information processing terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and nonvolatile storage medium

[0001] The present disclosure relates to an information processing device and a nonvolatile storage medium.

[0002] A questionnaire survey is a method of research in which specific consumers are asked to answer standardized questions in order to understand the opinions and preferences of the survey subjects, and the survey results are used for product development, for example. Questionnaire surveys can be conducted in the form of mail surveys, street surveys, door-to-door surveys, or internet surveys. Internet surveys have become particularly popular in recent years.

[0003] Although Internet surveys are advantageous in terms of cost reduction and time required compared to mail surveys, street surveys, and in-person surveys, they require a waiting period for responses from the required number of users.

[0004] Furthermore, because the decision to answer a survey is up to each user, the attributes of respondents may be biased. Furthermore, the high level of anonymity can reduce the reliability of the responses.

[0005] There is a desire to reduce costs, shorten survey times, and improve reliability of questionnaire surveys.

[0006] The information processing device of this embodiment is connected via a communication line to a server that provides at least one of an SNS service, a blog service, and a word-of-mouth service, an information processing terminal that requests a survey, and a generation device (generation AI), and is equipped with: a means for receiving a survey request for a survey including survey questions and survey subjects from the information processing terminal; a means for collecting data on multiple posts posted to at least one of the SNS service, the blog service, and the word-of-mouth service by multiple users that match the survey subjects from the server; a feature calculation means for calculating features that quantify the content features of the question and the multiple posts; an extraction means for comparing multiple features related to the multiple posts with the features of the question and extracting at least one post from the multiple posts for each user; a means for generating a prompt to request that an answer to the question be generated for each extracted user based on the extracted posts; a means for transmitting the prompt to the generation device; a means for receiving from the generation device data on answers generated by the generation device in accordance with the prompt; and a means for transmitting at least one of the received answers and the aggregation results to the information processing terminal.

[0007] FIG. 1 is a configuration diagram of an entire system including an information processing device according to this embodiment. FIG. 2 is a diagram showing the physical configuration of the information processing device of FIG. 1. FIG. 3 is a diagram showing the functional configuration of the information processing device of FIG. 1. FIG. 4 is a diagram showing an example of user information. FIG. 5 is a diagram showing an example of questionnaire questions and their feature quantities. FIG. 6 is a diagram showing an example of an N-dimensional space representing feature quantities in this embodiment. FIG. 7 is a diagram showing the data flow of the entire system of FIG. 1 and the processing of the information processing device. FIG. 8 is a diagram showing an example of a posted message by a certain user collected in step S16 of FIG. 7. FIG. 9 is a diagram showing the feature quantities of each posted message collected in step S18 of FIG. 7 in an N-dimensional space. FIG. 10 is a diagram showing the posted messages extracted in step S19 of FIG. 7 and their respective feature quantities.

[0008] An embodiment of the present invention will now be described with reference to the drawings. An information processing device according to this embodiment executes a survey and executes a process of reporting the results to a survey requester. In this embodiment, rather than collecting answers to survey questions from real users, a generation AI generates answers (referred to as virtual answers) that a given user would give to the survey questions, and the virtual answers are collected for multiple users. This embodiment is characterized in that, to ensure that the virtual answers are similar to the answers of real users, i.e., to increase the accuracy of the virtual answers, one or more posts whose content characteristics are similar to the survey questions are extracted from many posts posted by real users on social networking services (hereinafter referred to as "SNS"), and the generation AI generates answers (virtual answers) that a user whose profile was collected from an SNS, etc. would give to the survey questions based on the extracted posts.

[0009] As shown in FIG. 1, an information processing device 1 as a computer according to this embodiment is connected via a communication network, typically an Internet network 6, to the following: a virtual answer generation device 2 that functions as a generation AI such as ChatGPT (registered trademark); an SNS server 3-1 that provides a social networking service (hereinafter referred to as "SNS") as a community-based service that promotes and supports connections between people; a blog site server 3-2 that provides a blog service in which users post articles about everyday matters and thoughts and display them in chronological order; a review site server 3-3 that operates a review site where users can write and view evaluations of products and services; and a questionnaire survey request terminal (information processing device) 4 that requests a questionnaire survey from the information processing device 1.

[0010] For example, typical examples of SNS include "Instagram (registered trademark)," "Facebook (registered trademark)," and "Twitter (registered trademark)," which is now X." Typical examples of review sites include "Tabelog (registered trademark)" and "Kakaku.com (registered trademark)." There are also many blog sites, such as FC2 (registered trademark) blogs. Note that SNS, blog services, and review sites are examples of services that provide users with an opportunity to post their own writings, and are not limited to the SNS server 3-1, blog site server 3-2, and review site server 3-3, as long as they provide the services.

[0011] As shown in FIG. 2 , the information processing device 1 has a processor 11 connected to a RAM 12, a ROM 13, a storage unit 14, an input device 15, a display 16, and a communication unit 17 via a system bus 10. The processor 11 is composed of, for example, a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The processor 11 executes a program loaded into the RAM 12 from the storage unit 14, the ROM 13, a CD-ROM, or a USB memory device, which serve as a non-volatile storage medium, to perform a questionnaire survey process. The RAM 12 functions as the processor 11's main memory, work area, etc. The ROM 13 or the storage unit 14 stores a BIOS (Basic Input Output System), an operating system program (OS), a questionnaire survey processing program, and other programs for implementing various functions, as well as various data required for these processes.

[0012] The input device 15 includes a keyboard (KB), a mouse, a touch panel, or other pointing device. The display 19 is typically implemented as an LCD (Liquid Crystal Display). The storage unit 14 stores, in addition to the questionnaire response processing program, data related to posted messages required for questionnaire response processing, their respective features, user profiles, questionnaire questions, and their survey subjects.

[0013] As shown in FIG. 3, by executing the questionnaire response processing program, the processor 11 functions as a control unit 20, a user information collection unit 21, a questionnaire survey request receiving unit 22, a user extraction unit 23, a posted message collection unit 24, a natural language analysis processing unit 25, a feature calculation unit 26, a posted message extraction unit 27, a virtual answer generation command creation unit 28, a command sending unit 29, a virtual answer receiving unit 30, a virtual answer tallying unit 31, and a tally result sending unit 32.

[0014] As shown in Figures 4(a) and 4(b), the user information includes user account information for each SNS, blog site, review site, etc., and user profile information registered on each SNS, blog site, review site, etc. The user information collection unit 21 collects user account information for all users from the servers 3-1, 3-2, and 3-3, and also collects user profile information for each user based on the user account information. The collected user account information and user profile information data is stored in the storage unit 14.

[0015] The survey request receiving unit 22 receives a survey request including the survey subjects and questions exemplified in Fig. 5 from the survey request terminal 4. The received data on the survey questions and survey subjects is stored in the storage unit 14.

[0016] The user extraction unit 23 extracts users to be included in the survey target based on the user profile information from all users whose user account information and user account information has been collected and stored in the storage unit 14. For example, the survey target may be "women in their 30s living in Tokyo," and the question may be "Please tell us about the fashion items you recently purchased, their colors, and where you purchased them."

[0017] The posted message collection unit 24 uses the user account information of the extracted users to collect posted messages (referred to as posted messages, articles, word-of-mouth messages, evaluation messages, messages, and other messages) posted by each of the extracted users from the SNS server 3-1, blog site server 3-2, word-of-mouth site server 3-3, etc. The collected data on the posted messages of each user is stored in the storage unit 14.

[0018] The natural language analysis processing unit 25 performs natural language analysis on the questions and posted messages, breaking them down into parts of speech and extracting noun strings. The feature calculation unit 26 calculates feature quantities that quantify the content characteristics of each question and posted message based on the extracted nouns. The feature quantities are given as coordinates in an N-dimensional space. An N-dimensional space is set in advance for each genre, such as apparel, music, health foods, books, food, and home appliances. A genre is identified based on the questionnaire question, and an N-dimensional space corresponding to the identified genre is selected.

[0019] A plurality of types that classify nouns according to their meanings are associated with the N axes that make up the N-dimensional space. A plurality of nouns related to each type are assigned to the N axes of the N-dimensional space, each with its own coordinate value. In the case of the apparel genre, as shown in FIG. 6, types such as color, product type, size, material, price range, time of year, place of purchase, purpose, fashion style, source of fashion information, and budget are associated with each axis. On each axis, a large number of nouns (character strings) that represent specific subordinate concepts included in each type are ordered according to their semantic proximity, and each is assigned a numerical value (coordinate value) in advance.

[0020] Nouns extracted from a posted message (text) are compared with nouns on the N-axis of an N-dimensional space, and if a corresponding noun string exists, the coordinate value corresponding to that noun string is assigned. If a corresponding noun string does not exist, the coordinate value on that coordinate axis is assigned a value of zero. Posts with similar features are considered to be close in content, with similar content features. For example, the noun "blue" is assigned the coordinate value assigned to "blue" on the coordinate axis corresponding to color. For example, the noun "skirt" is assigned the coordinate value assigned to "skirt" on the coordinate axis corresponding to product type.

[0021] In addition, the feature calculation process may be performed using a pre-trained model that uses questions and posted messages as input data and outputs feature data that quantifies the content features of those messages.

[0022] Furthermore, as the feature calculation process, any method conventionally used in natural language analysis processing may be used, such as "Bag of Words (BoW)" which counts the number of times a word appears in a sentence, "TF-IDF (Term Frequency-Inverse Document Frequency)" which calculates the importance of a word by weighting it based on the frequency of appearance of the word in the sentence and the inverse document frequency, or "Word Embedding" which vectorizes words taking into account the semantic similarity of the words.

[0023] The posted message extraction unit 27 compares the feature quantities of the user's posted messages with the feature quantities of the questionnaire question, and extracts at least one posted message from the collected multiple posted messages. This extraction process is performed individually for each user. Specifically, one or a predetermined number of posted messages having feature quantities closest to the feature quantities of the question are extracted. Alternatively, one or more posted messages in which the difference between the feature quantities of the posted message and the feature quantities of the question is less than a predetermined value are extracted. As described above, feature quantities are quantified content features, so the extracted posted messages are similar in content to the questionnaire question.

[0024] The virtual answer generation command creation unit 28 uses the survey questions, the extracted posted text, and the user's profile to create a command (prompt) for each user to request the virtual answer generation device 2 to create a virtual answer to the survey question. For example, "You are a general consumer with the following profile. #Command sentence: Please generate your answer to the following survey question based on your statement.

[0025] #Profile AAA #Your comment BBB #Survey question bbb" Answers (virtual answers) to survey questions are generated based on comments that the user has actually posted on SNS etc., along with the profile of the actual user, and which are similar in content to the survey questions. This means that the virtual answers are close to the answers that the user would give to the survey questions, meaning that the accuracy of the virtual answers can be improved.

[0026] The command sending unit 29 sends the command created by the virtual answer generation command creating unit 28 to the virtual answer generation device 2. This transmission is repeated for each user. The virtual answer generation device 2 generates answers (virtual answers) that a general consumer with profile AAA would give to the questionnaire questions based on the usual utterances BBB. The generation of these virtual answers is repeated for each user.

[0027] The virtual answer receiving unit 30 receives data on the virtual answers of each extracted user, which are generated in response to prompts by the virtual answer generation device 2, from the virtual answer generation device 2. The virtual answer counting unit 31 counts the received virtual answers of the extracted users as "frequency," "ratio," etc. The counting result sending unit 32 sends the virtual answers of the extracted users and the counting results as a questionnaire survey report to the questionnaire survey request terminal (information processing device) 4.

[0028] 7 shows the processing steps of a questionnaire survey by the information processing device 1 according to this embodiment, along with data flows for the servers 3-1, 3-2, and 3-3, the questionnaire survey request terminal 4, and the virtual answer generation device 2. The user information collection unit 21 of the information processing device 1 periodically visits the servers 3-1, 3-2, and 3-3 and collects account and profile data as user information for each of a plurality of users registered with each service (S11). The collected user account and profile data is stored in the storage unit 14 (S12). Steps S11 and S12 are preparation processes for the questionnaire survey.

[0029] The questionnaire survey process is started when the questionnaire survey request receiving unit 22 receives a questionnaire survey request from the questionnaire survey request terminal 4. The questionnaire survey request includes a question and a survey subject.

[0030] The user extraction unit 23 extracts users who match the questionnaire survey target based on their profiles from all users stored in the storage unit 14 (S13). As shown in the example of Figure 5, if the survey target is "women in their 30s living in Tokyo," users whose profiles include the age of 30, living in Tokyo, and female are extracted.

[0031] Using the user account information of the extracted users, all of the messages posted by each of the extracted users are collected from the servers 3-1, 3-2, and 3-3 by the message collection unit 24 (S14). The collected data on the messages posted by each user is stored in the storage unit 14. Figure 8 shows an example of a message posted by one user.

[0032] The natural language analysis processing unit 25 performs natural language analysis on each of the survey questions and posted messages, breaking them down into parts of speech and extracting nouns (S15). As illustrated in FIG. 9, the feature calculation unit 26 calculates feature quantities that quantify content characteristics based on the extracted nouns for each of the survey questions and posted messages (S16). For example, for posted message No. 2, the nouns extracted from the posted message—“tomorrow,” “glen plaid,” “pants,” “blue,” “jacket,” and “plan”—are each assigned a coordinate value on the N-axis. For example, “tomorrow” is replaced with the season of autumn and assigned a coordinate value of “5,” which is assigned to autumn on the sixth axis representing time. For example, “pants” is assigned a coordinate value of “7,” which is assigned to pants on the second axis representing product type. Similarly, corresponding coordinate values ​​are assigned to other nouns. If a noun of a type assigned to a coordinate does not appear in the posted message, a zero value is assigned to that axis. Features are similarly calculated for the survey questions.

[0033] Preferably, in order to replace multiple nouns that have the same meaning but are written in different ways or phrased in different ways with a unified noun, a "synonym dictionary," a "word variation dictionary," a glossary created by an original developer, etc. are stored in advance in the memory unit 14, and the feature calculation unit 26 performs a preprocessing process of using these dictionaries to replace nouns extracted from posted texts with nouns that are written in a consistent, unified way.

[0034] Next, as illustrated in FIGS. 9 and 10 , the feature quantities of the multiple posted messages are compared with the feature quantities of the question, and one or a predetermined number of posted messages having feature quantities closest to the feature quantities of the question are extracted by the posted message extraction unit 27 (S17). In other words, one posted message having feature quantities that are the smallest difference from the feature quantities of the question, or a predetermined number of posted messages having feature quantities that are the smallest difference from the feature quantities of the question, are extracted. Note that the method of extracting posted messages is not limited to these methods. For example, posted messages whose feature quantities differ from the feature quantities of the question by less than a predetermined value may be extracted. Since feature quantities are quantified content features, posted messages whose feature quantities are close to those of the question are similar in content to the survey question. The posted message extraction process is performed for each user. Posted messages whose content is similar to the survey question are extracted for each user.

[0035] Using the survey questions, the user profiles, and the extracted posted text, the virtual answer generation command creating unit 28 creates a command (prompt) for each user to request the virtual answer generation device 2 to generate a virtual answer to the survey question (S18). An example of how the prompt is created is as described above. The created prompt is sent from the command sending unit 29 to the virtual answer generation device 2.

[0036] In the virtual answer generation device 2, a virtual answer that the user would give to the survey question is generated based on the user's profile and a post that the user actually posted on SNS etc. and that is similar in content to the survey question (S19). Since this virtual answer is generated based on the user's profile and a post that the user actually posted on SNS etc. and that is similar in content to the survey question, the virtual answer is close to the answer that the user would give to the survey question.

[0037] Data on a plurality of virtual answers corresponding to a plurality of users, which are generated by the virtual answer generation device 2, is transmitted from the virtual answer generation device 2 and received by the virtual answer receiving unit 30. The virtual answer counting unit 31 subjects the plurality of virtual answers to counting processing (S20), and the counting result transmitting unit 32 transmits the results to the questionnaire survey request terminal 4.

[0038] As described above, according to this embodiment, instead of collecting answers to survey questions from real users, the virtual answer generation device 2 generates virtual answers to survey questions using the survey questions, the profiles of real users, and the extracted posted text, so Internet surveys are less expensive and take less time than mail surveys, street surveys, and in-person surveys, and there is no bias in the attributes of respondents.Furthermore, there is no mixing of the arbitrary opinions of real users, so there is no reduction in the reliability of the answers.

[0039] In addition to the profiles of real users, posts that have similar content characteristics to the survey questions are extracted from posts that real users have actually posted to SNS etc., and the generation AI generates virtual answers to the survey questions based on the extracted posts.This makes it possible to improve the accuracy of the virtual answers and increase their reliability compared to when no posted posts are used, or even if posted posts are used, when the posted posts are not extracted based on the similarity of content characteristics to the survey questions.

[0040] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention described in the claims and their equivalents.

[0041] 1...information processing device, 2...virtual answer generation device (generation AI), 3-1...SNS server, 3-2...blog site server, 3-3...word-of-mouth site server, 4...questionnaire survey request terminal (information processing device).

Claims

1. An information processing device connected via communication lines to a server providing at least one of an SNS service, a blog service, and a word-of-mouth service, an information processing terminal as a requester of a survey, and a generation device (generation AI), comprising: means for receiving a survey request for the survey, including survey questions and survey subjects, from the information processing terminal; means for collecting data on a plurality of posts posted to at least one of the SNS service, the blog service, and the word-of-mouth service by a plurality of users who match the survey subjects, from the server; feature calculation means for calculating features that quantify content features of the question and the plurality of posts; extraction means for comparing a plurality of features related to the plurality of posts with the feature values ​​of the question, and extracting at least one post from the plurality of posts for each user; means for generating a prompt to request generation of an answer to the question for each user based on the extracted post; means for transmitting the prompt to the generation device; and means for receiving from the generation device data on the answer generated by the generation device in accordance with the prompt. an information processing device comprising: means for transmitting at least one of the received answers and the tabulation results to the information processing terminal.

2. The information processing device according to claim 1, wherein said extraction means extracts at least one posted message whose feature amount is closest to that of said question.

3. The information processing device according to claim 1, wherein the extraction means extracts at least one posted message in which the difference between the feature amount of the question and the feature amount of the posted message is less than a predetermined value.

4. The information processing device of claim 1, wherein the feature calculation means comprises: means for extracting nouns from the question and each of the posted messages; and means for assigning, based on the content of the extracted nouns, coordinates in an N-dimensional space in which multiple types of categories represented by the nouns are each associated with N axes and multiple nouns related to each category are assigned to each axis along with their coordinate values, as the feature.

5. An information processing device as described in claim 1, wherein the feature calculation means calculates the feature of the question and the feature of each of the plurality of posted messages using a trained model that has been trained in advance to input various messages and output features.

6. An information processing device as described in claim 1, further comprising: means for sending to the server a request for sending profile data registered by a user to at least one of the SNS service, the blog service, and the word-of-mouth service; means for storing the profile data sent from the server in response to the request; and means for extracting a plurality of users who match the survey subject based on the profile.

7. A computer processor connected via a communication line to a server providing at least one of an SNS service, a blog service, and a word-of-mouth service, an information processing terminal as a survey requester, and a generation device (generation AI), includes the steps of: receiving a survey request for the survey, including survey questions and survey subjects, from the information processing terminal; collecting data on a plurality of posts posted to at least one of the SNS service, the blog service, and the word-of-mouth service by a plurality of users who match the survey subjects, from the server; calculating features that quantify content characteristics of the question and the plurality of posts; comparing a plurality of features related to the plurality of posts with the features of the question, and extracting at least one post from the plurality of posts for each user; generating a prompt for requesting that an answer to the question be generated for each user based on the extracted post; transmitting the prompt to the generation device; receiving data on the answers generated by the generation device in accordance with the prompt from the generation device. a non-volatile storage medium storing a program for executing a step of transmitting at least one of the received answers and the tabulation results to the information processing terminal;

Citation Information

Patent Citations

  • Information processing system, information processing method, and program

    JP2022162850A

  • Computer system and method for market research using automation and virtualization

    US20210090097A1