Insight research system, insight research method, and insight research program
The insight research system enhances advertising evaluation by analyzing biometric data to generate questions and gather user responses, offering deeper insights into consumer behavior and motivations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- VIAGATE CO LTD
- Filing Date
- 2025-03-16
- Publication Date
- 2026-05-15
Smart Images

Figure 2026079663000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an insight investigation system, an insight investigation method, and an insight investigation program.
Background Art
[0002] Due to the impact of the COVID-19 pandemic, online transformation (in other words, digital transformation (DX)) has been accelerating in most aspects of social life. In particular, in the marketing field, digital marketing has become more mainstream than traditional marketing using paper media and the like. In the field of digital marketing, in addition to sales through e-commerce sites, methods such as displaying advertising content (images, videos, etc.) on various portal sites or uploading advertising content on video platforms such as YouTube (registered trademark) are becoming the current mainstream.
[0003] Conventionally, in an advertising evaluation system, the effect of the content has been evaluated according to the sales performance of whether the user actually purchased the product introduced in the advertising content. Therefore, in this advertising evaluation system, it was not possible to evaluate the advertising content at a stage where the actual sales performance of the product targeted for advertising was not yet available. On the other hand, on the corporate side that conducts digital marketing for products, there is a need to grasp detailed insights regarding the content before launching the advertising content. Here, "insight" means a deep understanding of customers' potential needs, emotions, motivations, values, behavior patterns, and trends, presented in a refined form through the analysis of customers' data and feedback.
[0004] Patent Documents 1 and 2 disclose systems that analyze consumer insights into content based on biometric information (such as gaze and facial expressions) when a person views the content, in order to meet the needs of the aforementioned companies. Specifically, Patent Document 1 targets video content such as commercials, and Patent Document 2 targets still image content such as web pages. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Patent No. 7398853 [Patent Document 2] Patent No. 7398854 [Overview of the project] [Problems that the invention aims to solve]
[0006] However, while the aforementioned patent documents provide insights into consumer behavior based on biometric information when people view content, it is difficult to obtain more in-depth information. For example, if it is found that consumers show a high level of attention or interest in a particular scene or object within the content, from a marketing perspective, it would be desirable to understand the reasons for this to improve the content, but this has not been achieved. In order to provide more useful value through insight research, it is important to not only understand what consumers focused on and were interested in, but also to deeply understand the background, including the reasons and the consumer's living environment.
[0007] This disclosure aims to reveal information that cannot be obtained from biometric data alone in content insight research. [Means for solving the problem]
[0008] An insight research system according to one aspect of the present disclosure is an insight research system for researching user insights into content, which includes displaying the content, acquiring biometric information when the user views the content, identifying points of interest in the content that the user focused on based on the biometric information, generating questions regarding the points of interest, displaying the questions, receiving answers from the user to the questions, and outputting the results of the insight research based on the answers. [Effects of the Invention]
[0009] According to this disclosure, it is possible to uncover information that cannot be obtained from biometric data alone in content insight research. [Brief explanation of the drawing]
[0010] [Figure 1] Overview of the Insight Research System [Figure 2] Viewer terminal hardware configuration [Figure 3] Server hardware configuration [Figure 4] Hardware configuration of enterprise terminals [Figure 5] Overview of the processes performed by the Insight Research System [Figure 6] Interview settings screen 1: Home [Figure 7] Interview settings screen 2: Brand information registration [Figure 8] Interview settings screen 3: Viewing brand information [Figure 9] Interview settings screen 4: Company information registration [Figure 10] Interview settings screen 5A: Interview creation A [Figure 11] Interview settings screen 5B: Interview creation B [Figure 12] Interview settings screen 5C: Create interview C [Figure 13] Interview Summary [Figure 14] List of Question Models [Figure 15] Interview Example 1: Main Question and Follow-up Question [Figure 16] Interview Example 2: Displaying Part of the Content at the Beginning [Figure 17] Interview Result Screen: Home (Top) [Figure 18] Interview Result Screen: Clustering Result [Figure 19] Interview Result Screen: Characteristics of Clusters [Figure 20] Interview Result Screen: Download of Table Data [Figure 21] Interview Result Screen: Attribute Distribution of Clusters [Figure 22] Interview Result Screen: Motivation - Scene - Action (Tree Diagram) [Figure 23] Interview Result Screen: Motivation - Scene - Action (Sunburst Diagram) [Figure 24] Interview Result Screen: Motivation - Scene - Action (Analysis Results by Attribute) [Figure 25] Interview Result Screen: Motivation - Scene - Action (Number of Occurrences of Combinations / Detailed Information by Label) [Figure 26] Interview Result Screen: Motivation - Scene - Action (User Information Confirmation) [Figure 27] Interview Result Screen: User List [Figure 28] Interview Result Screen: Detailed Information of Each User [ [Figure 29] Interview Result Screen: Content of Interviews of Each User
Mode for Carrying Out the Invention
[0011] 〔Definition〕 The definitions of terms according to this embodiment are as follows: "Insight" means a deep understanding of customers' latent needs, emotions, motivations, values, behavioral patterns, and tendencies, presented in a refined form by analyzing customer data and feedback. "Content" means images, videos, etc., created for people to view, and includes, for example, web pages on e-commerce sites, advertisements on various portal sites, and advertisements on video platforms such as YouTube®. "Viewer" and "User" mean people who view the content, meaning viewers if the content is a video, and viewers if the content is a web page, and represent representatives of general consumers who are subjects of the insight survey. "Operator" means a person in charge at a company that orders the insight survey, or a person in charge at a company that receives and conducts the insight survey. "Biometric information" means eye-tracking information, facial expression information, etc., obtained by measuring a person's gaze, facial expressions, etc., when they are viewing the content.
[0012] [System Overview] The insight survey system 1 according to this embodiment will be described below with reference to the drawings. Figure 1 is a diagram showing an example of the configuration of the insight survey system 1 according to this embodiment. As shown in Figure 1, the insight survey system 1 comprises viewer terminals 2a and 2b, a server 3, and a corporate terminal 4. These are connected to a communication network 8. Each of the viewer terminals 2a and 2b is connected to the server 3 via the communication network 8 in a communicative manner. The corporate terminal 4 is connected to the server 3 via the communication network 8 in a communicative manner. The communication network 8 consists of at least one of the following: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and a wireless core network.
[0013] Viewer terminal 2a is a terminal associated with viewer Va and is operated by viewer Va. Viewer terminal 2b is a terminal associated with viewer Vb and is operated by viewer Vb. In this embodiment, for the sake of explanation, viewer terminals 2a and 2b may be collectively referred to as viewer terminal 2. Similarly, viewers Va and Vb may be collectively referred to as viewer V. In this embodiment, a large number of viewer terminals 2 associated with a large number of viewers are provided in the insight survey system 1, but for the sake of explanation, only two of the large number of viewer terminals, viewer terminals 2a and 2b, are shown in Figure 1. The type of viewer terminal 2 is not particularly limited, and viewer terminal 2 may be, for example, a smartphone, a personal computer, a tablet, or a wearable device (e.g., a head-mounted display or an AR display).
[0014] [Viewer terminal configuration] Next, the hardware configuration of the viewer terminal 2 will be described below with reference to Figure 2. Figure 2 is a diagram showing an example of the hardware configuration of the viewer terminal 2. As shown in Figure 2, the viewer terminal 2 comprises a control unit 20, a storage device 21, an imaging unit 22, a communication unit 23, an input operation unit 24, a display unit 25, a speaker 26, and an RTC (Real Time Clock) 28. These elements constituting the viewer terminal 2 are connected to a communication bus 29.
[0015] The control unit 20 includes memory and a processor. The memory is configured to store computer-readable instructions (programs). For example, the memory consists of a ROM (Read Only Memory) in which various programs are stored, and a RAM (Random Access Memory) having multiple work areas in which various programs executed by the processor are stored. The processor consists of at least one of a CPU (Central Processing Unit), an MPU (Micro Processing Unit), and a GPU (Graphics Processing Unit). The CPU may consist of multiple CPU cores. The GPU may consist of multiple GPU cores. The processor may be configured to load a specified program from the various programs embedded in the storage device 21 or ROM onto the RAM and execute various processes in cooperation with the RAM.
[0016] The storage device 21 is, for example, a storage device such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory, and is configured to store programs and various data. The imaging unit 22 is configured to acquire video data showing the surrounding environment of the viewer terminal 2 through shooting. In particular, the imaging unit 22 is a camera configured to generate image data or video data showing the surrounding environment of the viewer terminal 2 through shooting, and comprises an image sensor (for example, a CCD sensor or a CMOS sensor) and an image sensor driving processing circuit. In this embodiment, the control unit 20 functions as a gaze tracking unit that detects changes in the viewer V's gaze based on the video data acquired by the imaging unit 22. Furthermore, the control unit 20 functions as a facial expression tracking unit that detects changes in the viewer V's facial expression based on the said video data.
[0017] The communication unit 23 includes a wireless communication module and / or a wired communication module for communicating with external devices connected to the communication network 8. The wireless communication module is configured to communicate wirelessly with external devices such as base stations and wireless LAN routers, and includes a transmitting and receiving antenna and a wireless transmitting and receiving circuit. The wireless communication module may be a wireless communication module compatible with short-range wireless communication standards such as Wi-Fi or Bluetooth, or it may be a wireless communication module compatible with the Xth generation mobile communication system (e.g., the 4th generation mobile communication system such as LTE) using a SIM (Subscriber Identity Module).
[0018] The input operation unit 24 is, for example, a touch panel, mouse, and / or keyboard superimposed on the video display of the display unit 25, and is configured to receive input operations from the viewer V and generate operation signals corresponding to those input operations. The display unit 25 is composed of, for example, a video display and a video display circuit that drives and controls the video display. The display unit 25 has a display screen 27 on which video is displayed.
[0019] Speaker 26 is configured to output the audio from the video to an external source based on the audio data contained in the video. RTC 28 is configured to acquire information indicating the current time.
[0020] Returning to Figure 1, Server 3 is connected to the viewer terminals 2 and the corporate terminal 4 via the communication network 8. Server 3 transmits video data to each of the multiple viewer terminals 2 via the communication network 8, and also transmits a composite video with a gaze point heatmap superimposed on the video to the corporate terminal 4. Server 3 may be composed of multiple servers. Server 3 functions as a web server configured to provide an insight survey application as a web application. In this respect, Server 3 is configured to transmit data (e.g., HTML files, CSS files, image and video files, program files, etc.) for displaying the insight survey screen in the web browser of the corporate terminal 4. In this way, Server 3 functions as a server for providing SaaS (System as a Service). Server 3 may be built on-premises or it may be a cloud server. Furthermore, Server 3 functions as a data management server that manages multiple video data, gaze point data for each viewer V, and facial expression data for each viewer V.
[0021] [Server Configuration] Referring to Figure 3, the hardware configuration of Server 3 will be described below. Figure 3 is a diagram showing an example of the hardware configuration of Server 3. As shown in Figure 3, Server 3 comprises a control unit 30, a storage device 31, an input / output interface 32, a communication unit 33, an input operation unit 34, and a display unit 35. These elements constituting Server 3 are connected to a communication bus 36.
[0022] The control unit 30 includes memory and a processor. The memory is configured to store computer-readable instructions. In particular, the memory may store an insight investigation program that causes the processor to execute a series of processes (insight investigation methods) performed by the server 3. The memory consists of ROM and RAM. The processor consists of at least one of a CPU, MPU, and GPU.
[0023] The storage device 31 is, for example, a storage device such as an HDD, SSD, or flash memory, and is configured to store programs and various data. The storage device 31 stores multiple video data, gaze point data for each viewer V, and facial expression data for each viewer V. The storage device 31 also stores a viewer information table related to the information of each viewer V and a user information table related to each user U who uses the insight research application. The viewer information table includes attribute information for each viewer V. For example, the viewer information table may include at least one of the following for each viewer V: identification information, gender information, age information, household size information, address information, and occupation information. The user information table may include identification information, attribute information, login information, etc., for each user U.
[0024] The input / output interface 32 is an interface that enables connection between an external device and the server 3, and includes an interface that conforms to a predetermined communication standard such as the USB standard or the HDMI® standard. The communication unit 33 may include various wired communication modules for communicating with external terminals on the communication network 8. The input operation unit 34 is, for example, a touch panel, mouse, and / or keyboard, and is configured to receive input operations from the operator and generate operation signals corresponding to the operator's input operations. The display unit 35 is, for example, composed of a video display display and a video display circuit.
[0025] The enterprise terminal 4 is a terminal operated by user U, who uses the insight research application provided by server 3. In this embodiment, multiple enterprise terminals 4 are provided in the insight research system 1 (in other words, multiple users U use the insight research application in this embodiment), but for the sake of explanation, only one enterprise terminal 4 is shown in Figure 1.
[0026] [Configuration of the enterprise terminal] Referring to Figure 4, the hardware configuration of the enterprise terminal 4 will be described below. Figure 4 is a diagram showing an example of the hardware configuration of the enterprise terminal 4. As shown in Figure 4, the enterprise terminal 4 may be, for example, a personal computer, a smartphone, a tablet, or a wearable device attached to user U. The enterprise terminal 4 has a web browser. The insight research application will run on the web browser of the enterprise terminal 4. The enterprise terminal 4 comprises a control unit 40, a storage device 41, an input / output interface 42, a communication unit 43, an input operation unit 44, and a display unit 45. These elements are connected to a communication bus 46.
[0027] The control unit 40 includes memory and a processor. The memory is configured to store computer-readable instructions (programs). For example, the memory consists of ROM and RAM. The processor consists of at least one of the following: CPU, MPU, and GPU.
[0028] The storage device 41 is, for example, a storage device such as an HDD, SSD, or flash memory, and is configured to store programs and various data. The input / output interface 42 is an interface (for example, USB or HDMI®) that enables connection between an external device and the enterprise terminal 4. The communication unit 43 is configured to connect the enterprise terminal 4 to the communication network 8 and includes a wireless communication module and / or a wired communication module. The input operation unit 44 is, for example, a touch panel, mouse, and / or keyboard, and is configured to receive input operations from user U and generate operation signals corresponding to user U's input operations. The display unit 45 is, for example, composed of a video display and a video display circuit.
[0029] [Overview of the process] Referring to Figure 5, the overview of the processes performed by the Insight Survey System 1 is described below. Figure 5 shows a sequence diagram of the processes performed between the company terminal 4, the server 3, and the viewer terminal 2. This process is broadly composed of acquiring viewer biometric information (S101-S105), setting up the interview (S111-S112), conducting the interview (S121-S128), and outputting the interview results (S131-S134). Solid lines represent the flow of the process, and dotted lines represent the references to information, etc. In this insight survey, after conducting the surveys in S101-S105 and S121-S128 with multiple viewers (subjects), the survey results are output in S131-S134 based on the data obtained.
[0030] #Acquiring viewer biometric information# In S101, the server 3 transmits content to the viewer terminal 2 using the communication unit 33. In S102, the viewer terminal 2 receives the content from the server 3 using the communication unit 23 and displays the content using the display unit 25. The viewer then views the displayed content. In S103, the viewer terminal 2 acquires the viewer's biometric information (such as gaze information and facial expression information) using the imaging unit 22 and transmits this biometric information to the server 3 using the communication unit 23. Here, the acquisition of biometric information can be done using the technology described in, for example, Japanese Patent No. 7398853. For example, the imaging unit 22 (camera) of the viewer terminal 2 is used to photograph the viewer's face while they are viewing the content, and the video is recorded (images are associated and stored for each time period). When acquiring gaze information, the video is used to analyze which part of the content the viewer was focusing on at each time period. When acquiring facial expression information, the video is used to analyze what kind of facial expression the viewer had at each time period (types of expressions such as joy, anger, sadness, etc. / the degree of those expressions). In S104, the server 3 receives biometric information from the viewer terminal 2 using the communication unit 33. In S105, the server 3 analyzes the biometric information using the control unit 30 and stores the analysis results (for example, information indicating the changes in the viewer's attention and interest in the content, and scenes or objects that were particularly attention-grabbing or of interest) using the storage device 31. Here, the analysis of biometric information can be performed using the technology described in, for example, Japanese Patent No. 7398853. For example, when analyzing gaze information, based on gaze information for each time period, if the viewer's gaze remained on a certain part for a predetermined time or longer, the corresponding part (scene and area if the content is a video / area if the content is a still image) is recorded as having a high level of attention, and the information is stored. When analyzing facial expression information, based on facial expression information for each time period, if the degree of the viewer's facial expression (for a certain type of expression / combining multiple types of expressions) exceeds a predetermined value, the corresponding part (scene and area if the content is a video / area if the content is a still image) is recorded as having a high level of attention, and the information is stored. The analysis results from S105 are used in S122, S132, etc.Regarding the details of S101 to S105, since the techniques described in, for example, Patent Document 1 and Patent Document 2 mentioned above can be used, the explanation will be omitted.
[0031] #Interview Settings# In S111, the company terminal 4 uses the input operation unit 44 to receive interview settings from the operator and uses the communication unit 43 to send those settings to the server 3. Details of S111 will be described later with reference to Figures 6 to 9. In S112, the server 3 uses the communication unit 33 to receive interview settings from the company terminal 4 and uses the storage device 31 to save those settings in association with the company ID, etc. The settings from S112 are used in S122, etc.
[0032] #Conducting an interview# In S121, the company terminal 4 uses the input operation unit 44 to receive instructions from the operator to execute the interview and uses the communication unit 43 to send those instructions to the server 3. Details of S121 will be described later with reference to Figures 10 to 12. In S122, the server 3 uses the communication unit 33 to receive instructions from the company terminal 4 to execute the interview and uses the control unit 30 to create interview questions. For example, it analyzes the video for scenes / areas that attracted the viewer's attention, identifies the subject (person / animal / their facial expressions / plant / object / weather / background, etc.), and then generates questions to confirm whether or not the viewer paid attention to that subject and why. Details of S122 will be described later with reference to Figures 13 to 14. In S123, the server 3 uses the communication unit 33 to send the interview questions to the viewer terminal 2. In S124, the viewer terminal 2 uses the communication unit 23 to receive the interview questions from the server 3 and uses the display unit 25 to display the questions. In response, the viewer answers the displayed questions. In S125, the viewer terminal 2 receives the interview answers using the input operation unit 24 and sends the answers to the server 3 using the communication unit 23. Details of S124 to S125 will be described later with reference to Figures 15 to 16. In S126, the server 3 receives the interview answers from the viewer terminal 2 using the communication unit 33. In S126, the server 3 uses the control unit 30 to decide whether or not to end the interview. Specifically, for example, if it is determined that all pre-set questions have been answered, the interview is ended. If the answer in S126 is YES, the process returns to S122. If the answer in S126 is NO, the process proceeds to S128. In S128, the server 3 uses the control unit 30 to analyze the interview answers and uses the storage device 31 to save the analysis results (answer result data) in association with the viewer ID, etc. To perform this analysis, for example, when using a generative AI, the prompt beforehand should state: "Perform text mining on the response, organize the obtained information according to predetermined criteria and format, and then output it."A prompt should be prepared with the gist of the following, and by using this prompt to input the answer to the AI, the desired output can be obtained. The analysis results (response result data) from S128 will be used in S132, etc. In actual insight research, the above processing (S101~S128) will be performed on multiple viewers (subjects / users), and the analysis results (response result data) corresponding to those multiple viewers (subjects / users) will be collected.
[0033] #Output of interview results# In S131, the company terminal 4 uses the input operation unit 44 to receive instructions from the operator to output the interview results and uses the communication unit 43 to send those instructions to the server 3. In S132, the server 3 uses the communication unit 33 to receive instructions from the company terminal 4 to output the interview results and uses the control unit 30 to create the interview results. In S133, the server 3 uses the communication unit 33 to send the interview results to the company terminal 4. In S134, the company terminal 4 uses the communication unit 43 to receive the interview results from the server 3 and uses the display unit 45 to display the interview results. Details of S134 will be described later with reference to Figures 17 to 20.
[0034] [Interview setting] Referring to Figures 6 to 9, the details of S111 (Interview Setup) in Figure 5 are explained below.
[0035] Figure 6 shows the home screen (top screen) of the interview settings screen. U100 (the left side of the screen) displays information and menus corresponding to the user ID. By default, "Home" is selected in the U100 menu, and U111 is displayed on the right side of the screen. U111 displays a list of surveys, each with information such as its name and status. To the right of this is a "Check Survey Results" button, and clicking it displays (outputs) the survey results. This click corresponds to S131 (instruction to output the interview) in Figure 5.
[0036] Figure 7 corresponds to the screen used when registering brand information within the interview settings screen. This screen is displayed when "Register Brand Information" is selected from the U100 menu. U121 displays input fields for setting information such as organization name, brand name, brand message, brand concept, brand target, and brand category for each brand. Once information is entered (written / selected) in each input field and the "Register" button is clicked, the entered information is stored (registered / set) in the storage device 31 of server 3 as brand information, associated with the user ID. In principle, multiple brand information entries are expected per user ID, but this number can be either a predetermined number or an arbitrary number.
[0037] Figure 8 corresponds to the screen used to view brand information within the interview settings screen. This screen is displayed when "View Brand Information" is selected in the U100 menu. U131 displays a list of registered brand information, with each displaying information such as the brand name and company name. To the right of this is a "Details" button, which, when clicked, displays the details of that brand information (all information entered in each input field).
[0038] Figure 9 corresponds to the screen used to register company information within the interview settings screen. This screen is displayed when "Company Information Registration" is selected from the U100 menu. U141 displays input fields for setting information such as company name, homepage URL, mission, vision, and values for each company. Once information is entered (written / selected) in each input field and the "Set" button is clicked, the entered information is stored (registered / set) in the storage device 31 of server 3 as company information, associated with the user ID. In principle, it is assumed that one piece of company information can be set per user ID.
[0039] [Instructions for conducting the interview] Referring to Figures 10 to 12, the details of S121 (instructions for conducting the interview) in Figure 5 are explained below.
[0040] Figure 10 shows the screen used to create an interview, which is part of the interview settings screen. This screen is displayed when "Create Interview" is selected from the U100 menu. U151 in Figure 10 displays input fields for setting information such as the purpose of the interview (including the brand to be investigated, the target audience to be investigated, the type of investigation, and additional interview information: details of the purpose), and interview settings (including the title, the number of boxes in the introduction, and the opening sentence). Scrolling this screen displays U152 in Figure 11.
[0041] U152 in Figure 11 displays input fields for setting interview settings (including language, organization, brand, end date, target sample size, and collaboration options) and sequence settings (including the default number of questions, multiple questions can be set, and the message on the completion page). Clicking on each "question" displays U153 in Figure 12. Once information is entered (written / selected) in each input field from U151 to U153 and the "Create Interview" button is clicked, the entered information is stored (registered / set) in the storage device 31 of Server 3 as interview information, associated with the user ID. This click corresponds to S121 (instruction to run the interview) in Figure 5.
[0042] In Figure 12, U153 allows you to input (write / select) the content of the question (including sequence, question text, answer format, and brand). Clicking "Select Sequence" allows you to select the type of question (main question / follow-up question) and the question model to be used when creating (generating) the question. If a main question is selected, you can enter the main question text, which will be described later, in "Question Text". In addition to simple predefined phrases, questions can also include information such as thumbnails showing scenes / areas that attracted the viewer's attention / interest in the content, based on, for example, S105 in Figure 5 (analysis results of insight survey based on biometric information), and then set up questions to delve deeper into that. In that case, the ON / OFF of this function can be set, for example, by selecting the type of question model mentioned above or by using a separate checkbox. The position in which the thumbnails are displayed can be set, for example, by selecting the position. Furthermore, if the "Question Text" contains text that is a transcription of the content of the thumbnails based on image analysis, its position can be indicated in the text with a designated symbol or similar. If a main question is selected, you can choose either a multiple-choice or open-ended response format in the "Answer Format" section.
[0043] [Drafting the question] Referring to Figures 13 and 14, the details of S122 (creating the question) in Figure 5 are explained below.
[0044] Figure 13 is a conceptual diagram illustrating the outline of the interview. In this interview, a set of questions is formed by combining a "main question" with several subsequent "follow-up questions," and an in-depth interview is conducted by repeating multiple sets of these question sets. The main question is created, for example, using the question model selected in the "sequence" U153 in Figure 12, based on the content entered in the "question text." The follow-up questions are generated by AI based on the answer to the main question, for example, using the question model selected in the "sequence" U153 in Figure 12. The number of follow-up questions can be, for example, a predetermined number of times, or the number of times until the waiting state for an answer times out. An example of the actual question structure is shown in the figure.
[0045] Figure 14 is a table showing a list of question models. There are multiple different question models for both main questions and follow-up questions. The outline of each question model is as shown in the figure. Note that the question model for main questions refers to formatting such as whether or not to display the aforementioned thumbnail when asking a question, while the question model for follow-up questions refers to the learning model used when the AI creates the question.
[0046] [Conducting the interview] Referring to Figures 15 and 16, the details of conducting the interview, specifically steps S124 (displaying the question) to S125 (receiving the response) in Figure 5, are explained below.
[0047] Figure 15 illustrates the relationship between the main question and follow-up questions as an example of an interview. As shown in the figure, after the main question, follow-up questions are repeated based on the answer to it.
[0048] Figure 16 shows an example of an interview, specifically when a portion of the content is displayed at the beginning. As illustrated, at the beginning of the interview, a thumbnail is displayed indicating a scene or area in the content that garnered high attention / interest, followed by a main question related to that thumbnail. If the content is a video such as a commercial, a thumbnail should be created for the scene that garnered high attention / interest. If a particularly high-attention / interest area can be identified, that part may be highlighted or enlarged. If the content is an image (still image) such as a website, a thumbnail should be created that highlights or enlarges the area in the content that garnered high attention / interest. If the content is a video, clicking the thumbnail may allow playback of the corresponding video (the whole or a portion before or after the thumbnail). In the main question, as shown in the underlined section, the content corresponding to the thumbnail may be analyzed using AI or similar technology to generate text that describes the characteristics of the thumbnail, and this text may be included as the subject of the question. The underlined portion may be replaced with a standard phrase such as "Scene shown in the thumbnail above: Scene that attracted your attention through AI analysis."
[0049] [Display of interview results] Referring to Figures 17 to 29, the details of S124 (display of interview results) in Figure 5 are explained below.
[0050] #Summary of Results# Figure 17 shows the home (top) screen of the interview results screen. This screen (U201) is displayed when the "Check Results" button is clicked in U111 of Figure 6. By switching the tabs at the top, the various survey results described later will be displayed. Note that the tab content and number of tabs may differ depending on the interview question model. Clicking the "View Clustering Results" button will perform user clustering based on the survey results data and display the results, etc. In the single-choice interview results display, the results of single-choice questions asked in the AI interview are displayed in graph format. Clicking the "Interview List Button" at the bottom will return to U111 (interview selection screen) in Figure 6. Note that some buttons and graph displays may not be displayed if there is no data or if the amount of data is large.
[0051] #Clustering# Figure 18 shows the clustering results screen within the interview results. This screen (U211) is displayed when the "View Clustering Results" button is clicked in U201 of Figure 17. A cluster is a grouping of users with similar characteristics within the data. User clustering is performed based on a summary of the interview conversation history. The number of clusters may be automatically determined to achieve optimal classification. Here, each user (viewer / subject) is classified into a cluster based on the response data and displayed in various formats. A cluster is, for example, a word or sentence (which may include not only identical but also similar ones) that appears a certain number of times or more in the response data, and corresponds to a criterion for classifying users with similar characteristics. To perform this clustering, for example, when using a generation AI, a prompt such as "Perform clustering on the response data, organize it according to the predetermined criteria and format, and output it" should be prepared in advance, and the response data should be input to the AI using this prompt to obtain the desired output. U211A shows the clustering results as a scatter plot. This scatter plot is created by vectorizing the text summary of the conversation history and then reducing its dimension to two. The X and Y axes are set to correspond to either of these two dimensions. Each user is plotted as a point, color-coded for each cluster. Clicking on each point may display the corresponding user information. U211B indicates the status and progress of the processing when outputting the clustering results. U211C shows the number of users in each cluster in the clustering results and their proportion to the total.
[0052] Figure 19 shows the interview results screen corresponding to the cluster characteristics. This screen (U212) is displayed together with U211 in Figure 18. U212A shows a list of names that concisely describe the characteristics of each cluster. Clicking on any of these names displays U212B. U212B shows the detailed information of the selected cluster (number of users belonging to that cluster, analysis results, characteristic information, etc.).
[0053] Figure 20 shows the screen corresponding to the download of table data from the interview results screen. This screen (U213) is displayed together with U211 in Figure 18. U213 shows the screen for downloading the response results data as table data in CSV format or other formats. On this screen, it is possible to filter (narrow down) the table data to be downloaded by specifying conditions such as cluster, gender, and age group. Below that, a list of filtered table data is displayed. When the user clicks the download button (not shown), the response results data filtered according to the specified conditions is downloaded.
[0054] Figure 21 shows the attribute distribution of clusters in the interview results screen. This screen (U214) is displayed together with U211 in Figure 18. U214A shows the proportion of each cluster by gender. U214B shows the proportion of each cluster by age group.
[0055] #Motivation, Scenes, Action# Figure 22 shows the Motivation, Scene, and Action (tree diagram) section of the interview results screen. This screen (U221) is displayed when "Motivation, Scene, and Action" is clicked in the top tab of U201 in Figure 17. Here, based on the response data, keywords included in the user's responses are classified from the perspective of motivation, scene, and action and displayed in various formats. Here, "motivation," "scene," and "action" refer to the attributes of keywords included in the response when the interview delves deeper into a situation where, for example, the user paid attention to (was interested in / had a high level of interest in) a part of the content. For example, "motivation" means "why they paid attention," "scene" means "what scene they associated it with," and "action" means "what action it led to." The menu at the top of the screen allows you to narrow down the population of the response data to be displayed from the perspective of generation / age, scene, day of the week, etc. In the tree diagram at the bottom of the screen, each keyword is organized and arranged according to which attribute it corresponds to: motivation, scene, or action (motivation: red bubble in the top row / scene: blue bubble in the middle row / action: green bubble in the bottom row *the size of each bubble is proportional to the frequency of occurrence of the keyword), and keywords with high correlation to each other are connected by lines. In the menu above the tree diagram, it is possible to narrow down the bubbles (keywords) to be displayed from the perspective of the age / gender of the user corresponding to the keyword, the frequency of occurrence of the keyword, etc.
[0056] Figure 23 corresponds to the Motivation, Scene, and Action (Sunburst Diagram) screen in the interview results. This screen (U222) is displayed together with U221 in Figure 22. The diagram at the top of the screen represents the same data as the tree diagram in Figure 22, but in a sunburst diagram. Here, corresponding keywords are classified in concentric circles in the order of Motivation, Scene, and Action, starting from the center of the circle. This order can be arbitrarily set according to instructions via the UI. When any keyword is clicked in the diagram at the top of the screen, the diagram at the bottom of the screen is displayed. This diagram places the specified keyword at the center of the circle and redraws a detailed sunburst diagram for the lower layers of that keyword (those located outside the concentric circles).
[0057] Figure 24 shows the interview results screen, specifically the section corresponding to Motivation, Scene, and Action (analysis results by attribute). This screen (U223) is displayed together with U221 in Figure 22. This screen displays various analysis results using keywords corresponding to the Motivation attribute as the population. Alternatively, similar analysis may be performed for keywords corresponding to each Scene and Action attribute, and the results displayed. Clicking the "View Clustering Analysis Results" button at the top of the screen extracts keywords corresponding to the Motivation attribute, performs clustering, and displays the results in a format similar to Figures 18-21. The "Word Cloud" visually displays keywords corresponding to the Motivation attribute according to their frequency of occurrence (for example, larger font size for frequently occurring keywords, similar meanings in similar colors, etc.). "Frequently Occurring Words" displays the most frequently occurring words in the Word Cloud in a graph format, ranked by frequency of occurrence. "Motivation Label Selection" allows you to select the desired motivation labels (keywords corresponding to that attribute, or classifications of those keywords from certain perspectives). The "Word Count Results" show the most frequent words in the word cloud, ranked in order of frequency (most frequent first), after narrowing the population to keywords corresponding to the selected motivation label.
[0058] Figure 25 shows the interview results screen corresponding to Motivation, Scene, and Action (frequency of occurrence of combinations / detailed information by label). This screen (U224) is displayed together with U221 in Figure 22. "Frequency of occurrence of each motivation, scene, and action combination" is calculated for each pattern of motivation, scene, and action combinations and displayed in a count table format in ranking order (most frequent occurrence first). "Detailed information by label" is displayed for each motivation, scene, and action, showing the frequency of occurrence for each label in a count table or graph format in ranking order (most frequent occurrence first).
[0059] Figure 26 shows the interview results screen, specifically the section corresponding to Motivation, Scene, and Action (User Information Confirmation). This screen (U225) is displayed together with U221 in Figure 22. Here, information for each user (interview respondent) is displayed in a list format, filtered by various labels. The lower part of the screen (V2) has more information and allows for more flexible filtering than the upper part (V1).
[0060] #User List# Figure 27 shows the user list screen within the interview results. This screen (U231) is displayed when "User List" is clicked in the top tab of U201 in Figure 17. Clicking the "Create Downloadable File" button in the upper left corner of the screen allows you to download all users' attribute information and conversation history in CSV format, etc. In the center of the screen, various scores are displayed for each user. These scores are calculated by scoring the response data according to the following criteria: The total score is the sum of all scores, and the higher the score, the higher the quality of the monitor / the more valuable the monitor is to listen to. The interview time score is higher the longer the time spent participating in the interview. The response accuracy score is higher the more accurately the interviewer answers the questions. The response length score is higher the longer the respondent's response. The response variation score is higher the more variation there is in the responses, and lower if the same thing is said repeatedly. The interview purpose suitability score is higher the more the interview is suited to the purpose. Clicking the "View User Details" button below the scores displays the screen shown in Figure 28. If the same user submits multiple responses, the various scores for each response may be displayed in a collapsible list format, as shown at the bottom of the screen, or in a non-illustrated diary format. Additionally, motivation labels may be provided, allowing users to select from motivation clusters to filter and display results.
[0061] Figure 28 corresponds to the detailed information for each user in the interview results screen. This screen (U232) is displayed when the "View User Details" button below each user's score in U231 of Figure 27 is clicked. The facial photograph is a persona generated using AI based on the user's response data. The information to the right of the facial photograph, such as age and gender, is attribute information that the user pre-registered as a monitor. To the right of that, "Estimated Inner Information from Emomiru's Past Response History" is an estimate of the user's lifestyle, values, personality, etc., based on the user's response data, using AI. To the right of that, "Estimated Person Profile / Summary from This Response History" is an estimate of the user's person profile / summary of the interaction content, based on the user's response data, using AI. When the "View Conversation Content" button below is clicked, Figure 29 is displayed.
[0062] Figure 29 shows the interview results screen corresponding to the content of each user's interview. This screen (U233) is displayed when the "View Conversation Content" button is clicked in U232 of Figure 28. This screen displays the actual AI interview exchange (the entire content of the answers) from S124 to S125 in Figure 5.
[0063] Although embodiments of the present invention have been described above, the technical scope of the present invention should not be interpreted as being limited by the description of these embodiments. These embodiments are examples, and it will be understood by those skilled in the art that various modifications to the embodiments are possible within the scope of the invention described in the claims. The technical scope of the present invention should be determined based on the scope of the invention described in the claims and the scope of its equivalents. [Explanation of Symbols]
[0064] 1: Insight survey system, 2,2a,2b: Viewer terminals, 3: Server, 4: Corporate terminal, 8: Communication network, 20: Control unit, 21: Storage device, 22: Imaging unit, 23: Communication unit, 24: Input operation unit, 25: Display unit, 26: Speaker, RTC: 28, 30: Control unit, 31: Storage device, 32: Input / output interface, 33: Communication unit, 34: Input operation unit, 35: Display unit, 40: Control unit, 41: Storage device, 42: Input / output interface, 43: Communication unit, 44: Input operation unit, 45: Display unit, U: User, V,Va,Vb: Viewer
Claims
1. An insight research system for investigating viewer insights regarding videos, Display the aforementioned video, By capturing the viewer's face when they watch the video, the gaze information or facial expression information of the viewer when they watch the video is obtained. By analyzing the aforementioned gaze information or facial expression information, the scenes in the video that attracted the viewer's attention are identified. By analyzing the aforementioned scenes of interest, questions related to those scenes of interest are generated. Display the aforementioned question, Answers to the aforementioned questions will be accepted from the aforementioned viewers. Insight research system.
2. The aforementioned questions and answers will be conducted in chat format. The insight research system according to claim 1.
3. Along with the aforementioned question, display an image showing the aforementioned scene of interest. The insight research system according to claim 1.
4. By analyzing the aforementioned scenes of interest, words that describe the characteristics of those scenes of interest are generated. The above question includes words that describe the above characteristics, The insight research system according to claim 1.
5. The aforementioned questions consist of a main question and follow-up questions. The insight research system according to claim 1.
6. The user selects the question model to be used when generating the aforementioned questions. The question is generated by inputting the video of the scene of interest to the aforementioned question model. The insight research system according to claim 1.
7. An insight research method in an insight research system for investigating user insights regarding videos, Display the aforementioned video, By capturing the viewer's face when they watch the video, the gaze information or facial expression information of the viewer when they watch the video is obtained. By analyzing the aforementioned gaze information or facial expression information, the scenes in the video that attracted the viewer's attention are identified. By analyzing the aforementioned scenes of interest, questions related to those scenes of interest are generated. Display the aforementioned question, Answers to the aforementioned questions will be accepted from the aforementioned viewers. Insight research methods.
8. An insight research program that causes an insight research system to execute the insight research method described in claim 7.