Insight investigation system, insight investigation method, and insight investigation program
The insight investigation system enhances advertising evaluation by analyzing biometric data to generate questions and accept user answers, offering deeper consumer insights into content engagement and motivations.
Patent Information
- Application Number
- JP2025041958
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-10-30
- Filing Date
- 2025-03-16
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2045-03-16
AI Technical Summary
Existing advertising evaluation systems fail to provide in-depth insights into consumer behavior and motivations beyond biometric data, lacking understanding of the reasons behind consumer attention to specific content elements.
An insight investigation system that acquires biometric information, identifies points of interest, generates questions, accepts user answers, and outputs detailed insights based on those answers to understand consumer motivations and behaviors.
Reveals information not obtainable from biometric data alone, providing deeper consumer insights into content engagement and motivations.
Smart Images

Figure 0007750587000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an insight investigation system, an insight investigation method, and an insight investigation program. [Background technology]
[0002] Due to the impact of the COVID-19 pandemic, most aspects of social life are now moving online (in other words, digital transformation (DX)) at an accelerating pace. In particular, in the field of marketing, digital marketing has become more mainstream than traditional marketing using paper media. In the field of digital marketing, in addition to sales through e-commerce sites, methods such as displaying advertising content (images, videos, etc.) on various portal sites and uploading advertising content to video platforms such as YouTube (registered trademark) are now becoming mainstream.
[0003] Conventionally, advertising evaluation systems have evaluated the effectiveness of advertising content based on sales results, i.e., whether users actually purchased the product introduced in the advertising content. Therefore, such advertising evaluation systems could not evaluate advertising content before there were any actual sales results for the advertised product. Meanwhile, companies engaged in digital marketing for products need to gain detailed insights into advertising content before launching it. Here, "insights" refers to a deep understanding of customers' latent needs, emotions, motivations, values, behavioral patterns, and tendencies, which are presented in a sophisticated form through the analysis of customer data and feedback.
[0004] In order to meet the needs of the companies, Patent Documents 1 and 2 disclose systems that analyze consumer insights into content based on biometric information (gaze, facial expression, etc.) when a person views the content. Specifically, Patent Document 1 targets video content such as commercials, while Patent Document 2 targets still image content such as web pages. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent No. 7398853 [Patent Document 2] Patent No. 7398854 Summary of the Invention [Problem to be solved by the invention]
[0006] However, while the aforementioned patent documents provide insights into consumers based on biometric information when people view content, it is difficult to obtain more in-depth information. For example, if it is determined that a certain scene or object in the content attracts a high level of consumer attention or interest, from a marketing perspective, it would be desirable to know the reason for this and use it to improve the content, but this has not yet been achieved. To provide more useful value through insight research, it is important to not only understand what consumers pay attention to and are interested in, but also to deeply understand the reasons for this and the background of the consumers, including their living environment, etc.
[0007] The present disclosure aims to clarify information that cannot be obtained from biometric information alone in an insight investigation into content. [Means for solving the problem]
[0008] An insight investigation system according to one aspect of the present disclosure is an insight investigation system that investigates a user's insights regarding content, displays the content, acquires biometric information when the user views the content, identifies points of interest in the content that the user focused on based on the biometric information, generates questions related to the points of interest, displays the questions, accepts answers to the questions from the user, and outputs the results of the insight investigation based on the answers. [Effects of the Invention]
[0009] According to the present disclosure, in an insight investigation into content, it is possible to reveal information that cannot be obtained from biometric information alone. [Brief explanation of the drawings]
[0010] [Figure 1] Overview of the Insight Survey System [Figure 2] Hardware configuration of viewer terminal [Figure 3] Server hardware configuration [Figure 4] Hardware configuration of the company's terminal [Figure 5] Overview of what the Insights investigation system does [Figure 6] Interview settings screen 1: Home [Figure 7] Interview setting screen 2: Brand information registration [Figure 8] Interview setting screen 3: View brand information [Figure 9] Interview setting screen 4: Company information registration [Figure 10] Interview Settings Screen 5A: Interview Creation A [Figure 11] Interview settings screen 5B: Interview creation B [Figure 12] Interview settings screen 5C: Interview creation C [Figure 13] Interview Summary [Figure 14] List of question models [Figure 15] Interview Example 1: Main Questions and Follow-Up Questions [Figure 16] Interview Example 2: Displaying a portion of the content at the beginning [Figure 17] Interview results screen: Home (top) [Figure 18] Interview results screen: Clustering results [Figure 19] Interview results screen: Cluster characteristics [Figure 20] Interview results screen: Download table data [Figure 21] Interview results screen: Cluster attribute distribution [Figure 22] Interview result screen: Motivation, Scene, Action (tree diagram) [Figure 23] Interview results screen: Motivation, Scene, Action (Sunburst diagram) [Figure 24] Interview results screen: Motivation, Scene, Action (analysis results for each attribute) [Figure 25] Interview result screen: Motivation, Scene, Action (number of occurrences of combinations / detailed information by label) [Figure 26] Interview result screen: Motivation, Scene, Action (User information confirmation) [Figure 27] Interview results screen: User list [Figure 28] Interview results screen: Detailed information for each user [Figure 29] Interview results screen: Interview details for each user DETAILED DESCRIPTION OF THE INVENTION
[0011] [Definition] The definitions of terms used in this embodiment are as follows: "Insights" refers to a deep understanding of a customer's latent needs, emotions, motivations, values, behavioral patterns, and tendencies, presented in a refined form through the analysis of customer data and feedback. "Content" refers to images, videos, and the like created for human viewing, including, for example, web pages on e-commerce sites, advertisements on various portal sites, and advertisements on video platforms such as YouTube (registered trademark). "Viewer" and "user" refer to a person who views content, and if the content is a video, refers to a viewer. If the content is a web page, refers to a viewer. They are representative of the general consumers who will be the subjects of the insight survey. "Operator" refers to a person in charge of a company that orders an insight survey, or a person in charge of a company that accepts and carries out the insight survey. "Biometric information" refers to gaze information, facial expression information, and the like obtained by measuring a person's gaze, facial expression, and the like while viewing content.
[0012] [System Overview] An insight investigation system 1 according to this embodiment will be described below with reference to the drawings. FIG. 1 is a diagram showing an example of the configuration of the insight investigation system 1 according to this embodiment. As shown in FIG. 1, the insight investigation system 1 comprises viewer terminals 2a and 2b, a server 3, and a company terminal 4. These are connected to a communication network 8. Each of the viewer terminals 2a and 2b is communicatively connected to the server 3 via the communication network 8. The company terminal 4 is communicatively connected to the server 3 via the communication network 8. The communication network 8 is composed of at least one of a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a wireless core network.
[0013] Viewer terminal 2a is a terminal associated with viewer Va and is operated by viewer Va. Viewer terminal 2b is a terminal associated with viewer Vb and is operated by viewer Vb. In this embodiment, for convenience of explanation, viewer terminals 2a and 2b may be collectively referred to as viewer terminal 2. Similarly, viewers Va and Vb may be collectively referred to as viewer V. In this embodiment, a large number of viewer terminals 2 associated with a large number of viewers are provided in the insight survey system 1, but for convenience of explanation, only two of the large number of viewer terminals, viewer terminals 2a and 2b, are shown in FIG. 1. The type of viewer terminal 2 is not particularly limited, and the viewer terminal 2 may be, for example, a smartphone, a personal computer, a tablet, or a wearable device (e.g., a head-mounted display or an AR display).
[0014] [Configuration of viewer terminal] Next, the hardware configuration of the viewer terminal 2 will be described below with reference to Fig. 2. Fig. 2 is a diagram showing an example of the hardware configuration of the viewer terminal 2. As shown in Fig. 2, the viewer terminal 2 includes a control unit 20, a storage device 21, an imaging unit 22, a communication unit 23, an input operation unit 24, a display unit 25, a speaker 26, and an RTC (Real Time Clock) 28. These elements constituting the viewer terminal 2 are connected to a communication bus 29.
[0015] The control unit 20 includes a memory and a processor. The memory is configured to store computer-readable instructions (programs). For example, the memory may include a read-only memory (ROM) storing various programs and a random access memory (RAM) having multiple work areas for storing various programs executed by the processor. The processor may include at least one of a central processing unit (CPU), a micro processing unit (MPU), and a graphics processing unit (GPU). The CPU may include multiple CPU cores. The GPU may include multiple GPU cores. The processor may be configured to load a specified program from various programs stored in the storage device 21 or the ROM onto the RAM and execute various processes in cooperation with the RAM.
[0016] The storage device 21 is a storage device (storage) such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory, and is configured to store programs and various data. The imaging unit 22 is configured to acquire video data showing the surrounding environment of the viewer terminal 2 through imaging. In particular, the imaging unit 22 is a camera configured to generate image data or video data showing the surrounding environment of the viewer terminal 2 through imaging, and includes an image sensor (e.g., a CCD sensor or a CMOS sensor) and an image sensor drive processing circuit. In this embodiment, the control unit 20 functions as an eye gaze tracking unit that detects changes in the line of sight of the viewer V based on the video data acquired by the imaging unit 22. Furthermore, the control unit 20 functions as an expression tracking unit that detects changes in the expression of the viewer V based on the video data.
[0017] The communication unit 23 includes a wireless communication module and / or a wired communication module for communicating with external devices connected to the communication network 8. The wireless communication module is configured to communicate wirelessly with external devices such as base stations and wireless LAN routers, and includes a transmitting / receiving antenna and a wireless transmitting / receiving circuit. The wireless communication module may be a wireless communication module compatible with short-range wireless communication standards such as Wi-Fi and Bluetooth, or may be a wireless communication module compatible with an X-generation mobile communication system (for example, a fourth-generation mobile communication system such as LTE) using a SIM (Subscriber Identity Module).
[0018] The input operation unit 24 is, for example, a touch panel, a mouse, and / or a keyboard arranged over the video display of the display unit 25, and is configured to accept input operations from the viewer V and generate operation signals in response to the input operations. The display unit 25 is, for example, configured with a video display and a video display circuit that drives and controls the video display. The display unit 25 has a display screen 27 on which moving images are displayed.
[0019] The speaker 26 is configured to output the audio of the video to the outside based on the audio data included in the video. The RTC 28 is configured to obtain information indicating the current time.
[0020] Returning to FIG. 1 , the server 3 is communicatively connected to the viewer terminal 2 and the company terminal 4 via a communication network 8. The server 3 transmits video data to each of the multiple viewer terminals 2 via the communication network 8, and also transmits a composite video in which a gaze point heat map is superimposed on the video to the company terminal 4. The server 3 may be composed of multiple servers. The server 3 functions as a web server configured to provide an insight survey application as a web application. In this regard, the server 3 is configured to transmit data (e.g., HTML files, CSS files, image and video files, program files, etc.) for displaying an insight survey screen on the web browser of the company terminal 4. In this way, the server 3 functions as a server for providing SaaS (System as a Service). The server 3 may be built on-premises or may be a cloud server. The server 3 also functions as a data management server that manages multiple video data, gaze point data of each viewer V, and facial expression data of each viewer V.
[0021] [Server configuration] The hardware configuration of the server 3 will be described below with reference to Fig. 3. Fig. 3 is a diagram showing an example of the hardware configuration of the server 3. As shown in Fig. 3, the server 3 includes a control unit 30, a storage device 31, an input / output interface 32, a communication unit 33, an input operation unit 34, and a display unit 35. These elements constituting the server 3 are connected to a communication bus 36.
[0022] The control unit 30 includes a memory and a processor. The memory is configured to store computer-readable instructions. In particular, the memory may store an insight investigation program that causes the processor to execute a series of processes (insight investigation method) executed by the server 3. The memory is configured with a ROM and a RAM. The processor is configured with at least one of a CPU, an MPU, and a GPU.
[0023] The storage device 31 is, for example, a storage device (storage) such as an HDD, SSD, or flash memory, and is configured to store programs and various data. The storage device 31 stores multiple video data, gaze point data for each viewer V, and facial expression data for each viewer V. The storage device 31 also stores a viewer information table related to information about each viewer V and a user information table related to each user U who uses the insight survey application. The viewer information table includes attribute information for each viewer V. For example, the viewer information table may include at least one of identification information, gender information, age information, household size information, address information, and occupation information for each viewer V. The user information table may include identification information, attribute information, login information, etc. for each user U.
[0024] The input / output interface 32 is an interface that enables connection between an external device and the server 3, and includes an interface conforming to a predetermined communication standard such as the USB standard or the HDMI (registered trademark) standard. The communication unit 33 may include various wired communication modules for communicating with external terminals on the communication network 8. The input operation unit 34 is, for example, a touch panel, a mouse, and / or a keyboard, and is configured to accept input operations by an operator and to generate operation signals in response to the input operations by the operator. The display unit 35 is, for example, configured by a video display and a video display circuit.
[0025] The company terminal 4 is a terminal operated by a user U who uses the insight investigation application provided by the server 3. In this embodiment, multiple company terminals 4 are provided in the insight investigation system 1 (in other words, in this embodiment, multiple users U use the insight investigation application), but for convenience of explanation, only one company terminal 4 is shown in FIG.
[0026] [Configuration of company terminal] The hardware configuration of the company terminal 4 will be described below with reference to FIG. 4. FIG. 4 is a diagram showing an example of the hardware configuration of the company terminal 4. As shown in FIG. 4, the company terminal 4 may be, for example, a personal computer, a smartphone, a tablet, or a wearable device attached to a user U. The company terminal 4 has a web browser. The insight survey application runs on the web browser of the company terminal 4. The company terminal 4 includes a control unit 40, a storage device 41, an input / output interface 42, a communication unit 43, an input operation unit 44, and a display unit 45. These elements are connected to a communication bus 46.
[0027] The control unit 40 includes a memory and a processor. The memory is configured to store computer-readable instructions (programs). For example, the memory is configured with a ROM and a RAM. The processor is configured with at least one of a CPU, an MPU, and a GPU.
[0028] The storage device 41 is, for example, a storage device such as an HDD, SSD, or flash memory, and is configured to store programs and various data. The input / output interface 42 is an interface (for example, USB, HDMI (registered trademark), etc.) that enables connection between an external device and the company terminal 4. The communication unit 43 is configured to connect the company terminal 4 to the communication network 8, and includes a wireless communication module and / or a wired communication module. The input operation unit 44 is, for example, a touch panel, a mouse, and / or a keyboard, and is configured to accept input operations by the user U and generate operation signals in response to the input operations by the user U. The display unit 45 is, for example, configured by a video display and a video display circuit.
[0029] [Processing Overview] An overview of the processing executed by the insight survey system 1 will be described below with reference to Figure 5. Figure 5 shows a sequence diagram executed between the company terminal 4, server 3, and viewer terminal 2. This processing can be broadly divided into the following steps: acquiring viewer biometric information (S101-S105), setting up an interview (S111-S112), conducting the interview (S121-S128), and outputting the interview results (S131-S134). Solid lines indicate the process flow, and dotted lines indicate references to information, etc. In this insight survey, the surveys of S101-S105 and S121-S128 are conducted on multiple viewers (subjects), and then the survey results are output in S131-S134 based on the data obtained.
[0030] #Getting viewer biometric information# In S101, the server 3 transmits content to the viewer terminal 2 using the communication unit 33. In S102, the viewer terminal 2 receives the content from the server 3 using the communication unit 23 and displays the content using the display unit 25. In response, the viewer watches the displayed content. In S103, the viewer terminal 2 acquires biometric information (such as gaze information and facial expression information) of the viewer using the imaging unit 22 and transmits the biometric information to the server 3 using the communication unit 23. Here, the biometric information can be acquired using, for example, technology described in Japanese Patent No. 7398853. For example, the imaging unit 22 (camera) of the viewer terminal 2 captures an image of the viewer's face while they are watching the content, and records the video (stores the image in association with each time). When acquiring gaze information, the video is used to analyze which part of the content the viewer was gazing at at each time. When acquiring facial expression information, the video is used to analyze what kind of facial expression the viewer was making at each time (type of facial expression such as joy, anger, sadness, or happiness / degree of each facial expression). In S104, the server 3 receives the biometric information from the viewer terminal 2 using the communication unit 33. In S105, the server 3 analyzes the biometric information using the control unit 30 and stores the analysis results (e.g., information indicating the transition of the viewer's attention and interest in the content, and scenes and objects that particularly attracted attention and interest, etc.) using the storage device 31. Here, the analysis of the biometric information may be performed using, for example, a technique described in Japanese Patent No. 7398853. For example, when analyzing gaze information, if the viewer's gaze stays on a certain part for a predetermined period of time based on the gaze information for each time, the corresponding part (a scene and area if the content is a video, or an area if the content is a still image) is determined to have received high attention, and information is stored. When analyzing facial expression information, if the level of the viewer's facial expression (for a certain type of facial expression / a combination of multiple types of facial expressions) reaches a predetermined value or more based on the facial expression information for each time, the corresponding part (a scene and area if the content is a video, or an area if the content is a still image) is determined to have received high attention, and information is stored. The analysis results of S105 are used in S122, S132, etc.Note that the details of S101 to S105 can be achieved by using the techniques described in Patent Document 1 and Patent Document 2, for example, and therefore a detailed description thereof will be omitted.
[0031] #Setting up an interview# In S111, the company terminal 4 uses the input operation unit 44 to accept settings related to the interview from the operator, and transmits the settings to the server 3 using the communication unit 43. Details of S111 will be described later with reference to FIGS. 6 to 9. In S112, the server 3 uses the communication unit 33 to receive the settings related to the interview from the company terminal 4, and uses the storage device 31 to store the settings in association with the company ID, etc. The settings in S112 are used in S122, etc.
[0032] #Conducting an Interview# In S121, the company terminal 4 receives an instruction to conduct an interview from the operator using the input operation unit 44, and transmits the instruction to the server 3 using the communication unit 43. Details of S121 will be described later using FIGS. 10 to 12. In S122, the server 3 receives an instruction to conduct an interview from the company terminal 4 using the communication unit 33, and creates interview questions using the control unit 30. For example, it analyzes video of a scene / area that has attracted the viewer's attention, identifies the subject (people / animals / their facial expressions / plants / objects / weather / background, etc.), and then generates questions to confirm whether the subject was of interest and the reason for the subject. Details of S122 will be described later using FIGS. 13 and 14. In S123, the server 3 transmits the interview questions to the viewer terminal 2 using the communication unit 33. In S124, the viewer terminal 2 receives the interview questions from the server 3 using the communication unit 23 and displays the questions using the display unit 25. In response, the viewer answers the displayed question. In S125, the viewer terminal 2 accepts the interview answers using the input operation unit 24 and transmits the answers to the server 3 using the communication unit 23. Details of S124 to S125 will be described later with reference to FIGS. 15 and 16. In S126, the server 3 receives the interview answers from the viewer terminal 2 using the communication unit 33. In S126, the server 3 determines, using the control unit 30, whether to end the interview. Specifically, for example, if it determines that answers have been received for all pre-set questions, the server 3 ends the interview. If the answer is YES in S126, the process returns to S122. If the answer is NO in S126, the process proceeds to S128. In S128, the server 3 analyzes the interview answers using the control unit 30 and stores the analysis result (answer result data) in association with the viewer ID, etc., using the storage device 31. To perform this analysis, for example, if generative AI is used, a prompt such as "Perform text mining on the answers, organize the obtained information according to a predetermined standard and format, and then output it."" or similar prompts, and by inputting answers into the AI using the prompts, the desired output can be obtained. The analysis results (answer result data) of S128 are used in S132, etc. In an actual insight survey, the above processes (S101 to S128) are executed for multiple viewers (subjects / users), and the analysis results (answer result data) corresponding to these multiple viewers (subjects / users) are collected.
[0033] #Outputting interview results# In S131, the company terminal 4 receives an instruction from the operator using the input operation unit 44 to output the interview results, and transmits the instruction to the server 3 using the communication unit 43. In S132, the server 3 receives an instruction from the company terminal 4 using the communication unit 33 to output the interview results, and creates the interview results using the control unit 30. In S133, the server 3 transmits the interview results to the company terminal 4 using the communication unit 33. In S134, the company terminal 4 receives the interview results from the server 3 using the communication unit 43, and displays the interview results using the display unit 45. Details of S134 will be described later using Figures 17 to 20.
[0034] [Interview setup] Details of S111 (interview setup) in FIG. 5 will be described below with reference to FIGS.
[0035] Figure 6 corresponds to the home screen (top screen) of the interview setting screen. U100 (left side of the screen) displays information, menus, etc. corresponding to the user ID. By default, "Home" is selected in the U100 menu, and U111 is displayed on the right side of the screen. U111 displays a list of surveys, each displaying information such as the name and status. To the right of that is a "Check survey results" button, which, when clicked, displays (outputs) the survey results. This click corresponds to S131 (interview output instruction) in Figure 5.
[0036] FIG. 7 corresponds to the screen displayed when registering brand information among the interview setting screens. This screen is displayed when "Register Brand Information" is selected from the U100 menu. U121 displays input fields for setting information such as the organization name, brand name, brand message, brand concept, brand target, and brand category for each brand. When information is entered (written / selected) in each input field and the "Register" button is clicked, the entered information is associated with the user ID as brand information and stored (registered / set) in the storage device 31 of the server 3. In principle, multiple pieces of brand information can be set per user ID, but this can be either a predetermined number or an arbitrary number.
[0037] Figure 8 corresponds to the screen for viewing brand information from the interview setting screen. This screen is displayed when "View Brand Information" is selected from the U100 menu. U131 displays a list of registered brand information, with information such as the brand name and company name for each. To the right of this is a "Details" button, and clicking this displays details of the brand information (all information entered in each input field).
[0038] Figure 9 corresponds to the interview setting screen when registering company information. This screen is displayed when "Register Company Information" is selected from the U100 menu. U141 displays input fields for setting information such as the company name, homepage URL, mission, vision, and values for each company. When information is entered (written / selected) in each input field and the "Set" button is clicked, the entered information is associated with the user ID as company information and stored (registered / set) in the storage device 31 of the server 3. In principle, the number of company information that can be set is assumed to be one per user ID.
[0039] [Instructions for conducting the interview] Details of S121 (instruction to execute interview) in FIG. 5 will be described below with reference to FIGS.
[0040] Figure 10 is the interview setting screen that corresponds to the screen used when creating an interview. This screen is displayed when "Create Interview" is selected from the U100 menu. U151 in Figure 10 displays input fields for setting information such as the purpose of the interview (including the brand you want to research, the user demographic you want to research, the type of research, and additional information about the interview: details of the purpose), interview settings (including the title, number of introduction boxes, and opening text). Scrolling this screen displays U152 in Figure 11.
[0041] U152 in Figure 11 displays input fields for setting information such as interview settings (including language, organization, brand, end date, target sample size, and collaboration options), sequence settings (including the default number of questions, questions (multiple questions can be set), and completion page messages). Clicking on each "Question" displays U153 in Figure 12. When information is entered (written / selected) in each input field of U151 to U153 and the "Create Interview" button is clicked, the entered information is associated with the user ID as interview information and stored (registered / set) in the storage device 31 of the server 3. This click corresponds to S121 in Figure 5 (instruction to execute the interview).
[0042] In U153 of FIG. 12, the content of the question (including the sequence, question text, answer format, and brand) can be input (written / selected). Clicking "Select Sequence" allows the user to select the type of question (main question / follow-up question) and the question model to be used when creating (generating) the question. If the main question is selected, the user can input the main question text described below in "Question Text." The question may not be a simple standard phrase, but may also include, for example, a question that digs deeper into a scene or area that has attracted the viewer's attention or interest in the content, based on S105 of FIG. 5 (the analysis results of the insight survey based on biometric information). In this case, the ON / OFF of this function may be set, for example, using the type of question model described above or a separate checkbox. The position where the thumbnail or other information is displayed may be set, for example, by selection. If the "Question Text" contains text that represents the content of the thumbnail or other information through image analysis, the position may be indicated by a predetermined symbol or other symbol in the text. When a main question is selected, the user can select either multiple choice or free description as the answer format for the question in the "Answer format" field.
[0043] [Creating questions] Details of S122 (creating a question) in FIG. 5 will be described below with reference to FIGS.
[0044] Figure 13 is a conceptual diagram showing an overview of the interview. In this interview, a "main question" is combined with multiple "follow-up questions" to form a set of questions, and this set of questions is repeated multiple times to conduct an in-depth interview. The main question is created based on the content entered in the "question statement," for example, using a question model selected in "sequence" of U153 in Figure 12. The follow-up questions are generated by AI in response to the answer to the main question, for example, using a question model selected in "sequence" of U153 in Figure 12. The number of follow-up questions may be, for example, a preset number, or the number of times until the waiting state for an answer times out. An example of the actual question configuration is as shown in the figure.
[0045] Figure 14 is a table showing a list of question models. There are multiple different question models for both main questions and follow-up questions. An overview of each question model is as shown in the figure. Note that the question model for main questions refers to the format, such as whether or not to display the aforementioned thumbnails when asking a question, while the question model for follow-up questions refers to the learning model used when creating questions using AI.
[0046] [Interview implementation] With reference to FIGS. 15 and 16, the details of conducting the interview: S124 (displaying questions) to S125 (accepting answers) in FIG. 5 will be described below.
[0047] Figure 15 shows the relationship between main questions and follow-up questions in Interview Example 1. As shown in the figure, after the main question is asked, follow-up questions are asked repeatedly in response to the answer.
[0048] Figure 16 illustrates a second example of an interview, in which a portion of the content is displayed at the beginning. As shown in the figure, thumbnails depicting scenes or areas of the content that attracted high attention or interest are displayed at the beginning of the interview, followed by a main question related to the thumbnail. If the content is a video (moving image) such as a commercial, thumbnails may be created for scenes that attracted high attention or interest. If a particular area of the content attracts high attention or interest, that area may be highlighted or enlarged. If the content is an image (still image) such as a website, thumbnails may be created highlighting or enlarging an area of the content that attracted high attention or interest. If the content is a video, clicking on the thumbnail may play the video (all or a portion of the video, such as the beginning or end) corresponding to the thumbnail. In the main question, as shown in the underlined portion, the content corresponding to the thumbnail may be analyzed using AI or other means to generate a sentence that reflects the characteristics of the thumbnail, and that sentence may be used as the subject of the question. The underlined portion may simply be replaced with a standard phrase such as "The scene shown in the thumbnail above: The scene that attracted your attention according to the AI analysis."
[0049] [Display interview results] Details of S124 (display of interview results) in FIG. 5 will be described below with reference to FIGS.
[0050] #Results Summary# Figure 17 corresponds to the home (top) screen of the interview results screen. This screen (U201) is displayed when the "Check Results" button is clicked in U111 of Figure 6. By switching between the tabs at the top, the survey results described below are displayed. Note that the tab content and number of tabs may vary depending on the interview question model. Clicking the "View Clustering Results" button will perform clustering of users based on the survey result data and display the results, etc. The interview single selection result display displays the results of single-choice questions asked in the AI interview in graph format. Clicking the "Interview List Button" button at the bottom returns to U111 (Interview Selection Screen) in Figure 6. Note that among the various buttons and graph displays, those with no data or large amounts of data may not be displayed.
[0051] #Clustering# Figure 18 shows the clustering results screen from the interview results screen. This screen (U211) is displayed when the "View Clustering Results" button is clicked in U201 of Figure 17. Clusters are groups of users with similar characteristics in the data. Users are clustered based on a summary of the interview conversation history. The number of clusters may be determined automatically to optimize classification. Here, each user (viewer / subject) is classified into a cluster based on the answer result data and displayed in various formats. A cluster is, for example, a word or phrase (which may include not only identical but similar words) that appears a certain number of times in the answer result data, and serves as a criterion for classifying users with similar characteristics. To perform this clustering, for example, when using a generative AI, a prompt such as "Perform clustering on the answer result data, organize it according to a specified standard, and output it." The desired output can be obtained by inputting the answer result data into the AI using this prompt. U211A shows the clustering results in a scatter plot. The scatter plot is created by dimensionally compressing the vectorized text of the conversation history summary into two dimensions. The X and Y axes are set to correspond to one of the two dimensions. Each user is plotted as a point, color-coded by cluster. Clicking on a point may display the corresponding user information. U211B shows the status and progress of the processing when outputting the clustering results. U211C shows the number of users in each cluster in the clustering results and their percentage of the total.
[0052] Figure 19 shows the interview results screen corresponding to the cluster characteristics. This screen (U212) is displayed together with U211 in Figure 18. U212A shows a list in which names that express the characteristics of each cluster in one word are set. Clicking on one of the names here will display U212B. U212B shows detailed information about the selected cluster (the number of users who fall into that cluster, analysis results, characteristic information, etc.).
[0053] Figure 20 shows the interview result screen that corresponds to downloading table data. This screen (U213) is displayed together with U211 in Figure 18. U213 shows a screen for downloading response result data as table data in CSV format or the like. On this screen, it is possible to filter (narrow down) the table data to be downloaded by specifying conditions such as cluster / gender / age, for example. A list of the filtered table data is displayed below. When the user clicks the download button (not shown), the response result data that has been filtered according to the specified conditions will be downloaded.
[0054] Figure 21 shows the cluster attribute distribution from the interview results screen. This screen (U214) is displayed together with U211 in Figure 18. U214A shows the proportion of each cluster by gender. U214B shows the proportion of each cluster by age group.
[0055] #Motivation Scene Action# Figure 22 shows the Motivation, Scene, and Action (tree diagram) section of the interview results screen. This screen (U221) is displayed when the "Motivation, Scene, and Action" tab at the top of U201 in Figure 17 is clicked. Based on the response data, keywords contained in the user's responses are categorized by motive, scene, and action, and displayed in various formats. Here, "motive," "scene," and "action" refer to the attributes of the keywords contained in the response when, for example, a user pays attention to a piece of content (is interested / highly interested), and the interview delves deeper into that. For example, "motivation" refers to "why the user paid attention," "scene" refers to "what scene the content reminds them of," and "action" refers to "what behavior the content leads to." The menu at the top of the screen allows you to narrow down the population of response data to be displayed by generation, age, scene, day of the week, etc. In the tree diagram at the bottom of the screen, each keyword is organized according to which attribute it corresponds to: motivation, scene, or action (motivation: red bubble on the top row; scene: blue bubble on the middle row; action: green bubble on the bottom row *the size of each bubble is proportional to the frequency of the keyword), with lines connecting keywords that are highly correlated.The menu above the tree diagram makes it possible to narrow down the bubbles (keywords) to be displayed based on factors such as the age / gender of the user corresponding to the keyword, and the frequency of the keyword's appearance.
[0056] Figure 23 corresponds to the Motivation, Scene, and Action (Sunburst Diagram) section of the interview results screen. This screen (U222) is displayed in conjunction with U221 in Figure 22. The upper diagram on the screen is a sunburst diagram representation of the same data as the tree diagram in Figure 22. Here, the corresponding keywords are classified concentrically in the order of Motivation, Scene, and Action from the center of the circle. This order may be set arbitrarily based on instructions via the UI, etc. When a keyword is clicked in the upper diagram, the lower diagram on the screen is displayed. This diagram places the specified keyword in the center of the circle and redraws a detailed sunburst diagram for the keyword's lower level (those located on the outer concentric circles).
[0057] Figure 24 shows the interview results screen, corresponding to the Motivation, Scene, and Action (analysis results by attribute). This screen (U223) is displayed in conjunction with U221 in Figure 22. This screen displays various analysis results using keywords corresponding to the Motivation attribute as the population. Note that a similar analysis can be performed on keywords corresponding to the Scene and Action attributes, and the results can be displayed. Clicking the "View Clustering Analysis Results" button at the top of the screen extracts keywords corresponding to the Motivation attribute, performs clustering, and displays the results in a format similar to Figures 18 to 21. The "Word Cloud" visually displays keywords corresponding to the Motivation attribute according to their frequency of occurrence (e.g., keywords with higher frequency of occurrence are displayed in larger font size, keywords with similar meanings are displayed in similar colors, etc.). The "Frequent Words" graphically displays the frequently occurring words in the word cloud in ranking order (by frequency of occurrence). The "Select Motivation Label" allows users to select the Motivation Label (keywords corresponding to that attribute or keywords classified according to some perspective) they wish to narrow down. The "word count results" narrow down the population to keywords corresponding to the selected motivation label, and then display the frequently occurring words contained in the word cloud in a count table format in ranking order (in order of frequency of occurrence).
[0058] Figure 25 is the interview results screen that corresponds to the motive, scene, and action (number of times each combination appears / detailed information by label). This screen (U224) is displayed together with U221 in Figure 22. "Number of times each combination of motive, scene, and action appears" calculates the number of times each combination of motive, scene, and action appears, and displays it in ranking order (most frequently appearing) in count table format. "Detailed information by label" displays the number of times each label appears for each motive, scene, and action in count table format or graph format, in ranking order (most frequently appearing).
[0059] Figure 26 shows the interview results screen, which corresponds to the Motivation, Scene, and Action (user information confirmation) screen. This screen (U225) is displayed together with U221 in Figure 22. Here, information about each user (interview respondent) is displayed in list format, filtered using various labels. The lower part of the screen (V2) contains more information than the upper part (V1), allowing for more flexible filtering.
[0060] #User List# Figure 27 shows the user list section of the interview results screen. This screen (U231) is displayed when the "User List" tab at the top of U201 in Figure 17 is clicked. Clicking the "Create Download File" button in the upper left corner of the screen allows you to download all users' attribute information and conversation history in CSV format. Various scores are displayed for each user in the center of the screen. These scores are calculated by scoring the response result data based on the following criteria. The total score is the sum of all individual scores, and the higher the total score, the higher the monitor's quality / worth listening to. The longer the interview time score, the higher the interview duration. The response accuracy score increases with the accuracy of the interviewer's questions. The response length score increases with the length of the respondent's answers. The response variation score increases with a wide variety of responses and decreases with repeated repetition. The interview purpose fit score increases with the interview's purpose. Clicking the "View User Details" button below the score displays the screen shown in Figure 28. If the same user provides multiple answers, the scores for each answer may be collapsed and displayed in a list format, as shown at the bottom of the screen, or displayed in a diary format (not shown). Alternatively, a motivation label may be provided, and users may be selected from motivation clusters to narrow down the results.
[0061] Figure 28 shows the interview results screen, which displays each user's detailed information. This screen (U232) is displayed when the "View user details" button below each user's score in U231 in Figure 27 is clicked. The facial photo is a persona generated using AI based on the user's response data. The information to the right of the facial photo, such as age and gender, is the attribute information the user pre-registered as a monitor. The "Estimated internal information from Emomil's past response history" to the right of that is an AI-generated estimate of the user's lifestyle, values, personality, etc. based on the user's response data. The "Personality profile / summary estimated from this response history" to the right of that is an AI-generated estimate of the user's personality and a summary of the interaction based on the user's response data. Clicking the "View conversation content" button below displays Figure 29.
[0062] Figure 29 shows the interview result screen, which corresponds to the content of each user's interview. This screen (U233) is displayed when the "View conversation content" button is clicked in U232 of Figure 28. This screen displays the actual exchanges (full response content) from the AI interview in S124 to S125 of Figure 5.
[0063] Although the embodiments of the present invention have been described above, the technical scope of the present invention should not be construed as being limited by the description of the present embodiments. The present embodiments are merely examples, and it will be understood by those skilled in the art that various modifications of the embodiments are possible within the scope of the invention described in the claims. The technical scope of the present invention should be determined based on the scope of the invention described in the claims and its equivalents. [Explanation of symbols]
[0064] 1: Insight survey system, 2, 2a, 2b: Viewer terminal, 3: Server, 4: Company terminal, 8: Communication network, 20: Control unit, 21: Storage device, 22: Imaging device, 23: Communication unit, 24: Input operation unit, 25: Display unit, 26: Speaker, RTC: 28, 30: Control unit, 31: Storage device, 32: Input / output interface, 33: Communication unit, 34: Input operation unit, 35: Display unit, 40: Control unit, 41: Storage device, 42: Input / output interface, 43: Communication unit, 44: Input operation unit, 45: Display unit, U: User, V, Va, Vb: Viewer
Claims
1. An insight survey system for surveying viewer insights about videos, Display the video, By photographing the face of the viewer when watching the video, gaze information or facial expression information of the viewer when watching the video is acquired; By analyzing the gaze information or facial expression information, a scene of interest corresponding to a time when the viewer was paying high attention to the video is identified; By inputting the scene of interest into AI, a question related to the scene of interest is generated; Displaying the question, accepting an answer to the question from the viewer; Insight Survey System.
2. The questions and answers are given in a chat format. The insight research system of claim 1 .
3. displaying an image showing the scene of interest together with the question; The insight research system of claim 1 .
4. generating words that indicate the characteristics of the scene of interest; the question includes a word that indicates the characteristic; The insight research system of claim 1 .
5. The questions consist of a main question and a follow-up question. The insight research system of claim 1 .
6. receiving from the viewer a selection of a question model to be used in generating the question; generating the question by inputting a video corresponding to the scene of interest into the question model; The insight research system of claim 1 .
7. An insight research method executed by an insight research system that researches viewer insights about a video, Display the video, By photographing the face of the viewer when watching the video, gaze information or facial expression information of the viewer when watching the video is acquired; By analyzing the gaze information or facial expression information, a scene of interest corresponding to a time when the viewer was paying high attention to the video is identified; By inputting the scene of interest into AI, a question related to the scene of interest is generated; Displaying the question, accepting an answer to the question from the viewer; Insight research methodology.
8. An insight investigation program that causes an insight investigation system to execute the insight investigation method according to claim 7.
Citation Information
Patent Citations
Image display system, digital photo-frame, information processing system, program, and information storage medium
JP2010224715A
Video viewing analysis system, video viewing analysis method, and video viewing analysis program
JP7398853B1
Web page browsing analysis system, web page browsing analysis method, and web page browsing analysis program
JP7398854B1
Information processing method, information processing device and information processing program
WO2023112745A1
Cited By
Information processing system, information processing method and program
JP7824617B1
Information processing system, information processing method and program
JP7824618B1