Information processing method, program, and information processor
The information processing method addresses the challenge of interpreting human behavior by calculating an interpretation matrix from action data, which supports the understanding of human actions and aids in decision-making and marketing strategies.
Patent Information
- Application Number
- JP2024069255
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-04-22
- Publication Date
- 2025-06-12
AI Technical Summary
It is challenging for humans to objectively interpret the reasons behind human behavior, such as why a product becomes popular, as human actions can be influenced by various factors including mood, physical condition, and prior actions, making it difficult to determine the reproducibility of actions based on presented information.
An information processing method that acquires sets of action data combining explanatory database vectors and target database vectors, arranges these vectors into matrices, calculates an interpretation matrix through vector product and generalized inverse matrix operations, and outputs charts related to this matrix to support the interpretation of human behavior reasons.
This method enables the interpretation of human behavior reasons by providing an interpretation matrix that illustrates the influence of various factors on human actions, aiding in decision-making and marketing strategies by identifying effective features for increasing product popularity.
Smart Images

Figure 2025089232000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing method, a program, and an information processing apparatus.
Background Art
[0002] A machine learning model generated by machine learning is a black box, and it is difficult for a user to interpret the behavior. XAI (Explainable Artificial Intelligence) technology has been proposed to show a reasonable ground for the result output by the machine learning model for human acceptance (Patent Document 1, Non-Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] By the way, a phenomenon in which the reason for the occurrence of a result cannot be easily interpreted by humans occurs not only in the output result of a machine learning model. For example, it is difficult to objectively interpret the reason why a product is a hit, that is, the reason why many people purchase the product.
[0006] On one side, it aims to provide an information processing method or the like that supports the interpretation of human behavior reasons.
Means for Solving the Problem
[0007] The information processing method acquires a plurality of sets of action data in which an explanatory database vector composed of a plurality of feature amounts related to information presented to a person and a target database vector obtained by quantifying the actions of a plurality of people to whom the information is presented are combined, arranges a plurality of sets of the explanatory database vectors included in the action data into an explanatory matrix, and calculates an interpretation matrix that is a vector product of the target matrix arranged in the order corresponding to the explanatory database vector and the generalized inverse matrix of the target matrix, and the computer executes a process of outputting a chart related to the interpretation matrix.
Effect of the Invention
[0008] On one side, it is possible to provide an information processing method or the like that supports the interpretation of human behavior reasons.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Mode for Carrying Out the Invention
[0010] [Embodiment 1] A machine learning model that receives input of explanatory data and outputs target data is generated by using various machine learning algorithms. When using the generated machine learning model, it is difficult for humans to interpret the judgment process from the input of the explanatory data to the output of the target data.
[0011] However, when applying the machine learning model to decision-making in the real world, it is important for humans to be able to interpret the judgment process of the machine learning model. For example, when the output target data seems to deviate significantly from human common sense, if humans can appropriately interpret the judgment process of the machine learning model, humans can also appropriately judge how to handle the target data and the machine learning model.
[0012] Regarding the target data output from the machine learning model, the technology for explaining the reason for the output is called XAI (Explainable AI). Non-Patent Document 1 discloses AIME (Approximate Inverse Model Explanations), which is a type of XAI. AIME is an information processing method that supports users so that they can interpret the behavior of various machine learning models, including black box models whose generation algorithms and the like are unknown, from various viewpoints.
[0013] By the way, as described above, phenomena whose causes cannot be easily interpreted by humans occur not only in the output results of machine learning models. For example, it is difficult to objectively interpret the reasons for human behavior. This is because humans may act based on logical judgment, but they may also act intuitively without being particularly aware of clear reasons.
[0014] In the present embodiment, by regarding a person himself / herself as a black box that receives information and acts, the AIME disclosed in Citation 1 is applied to the interpretation of the reasons for human behavior. That is, in the present embodiment, an information processing method for assisting a user regarding the interpretation of the reasons for human behavior is provided.
[0015] FIG. 1 is an explanatory diagram for explaining the outline of the interpretation matrix A†. A person receives various types of information such as images, videos, voices, music, or scents and then acts. The image includes cases where characters are superimposed on a photograph or an illustration. The video includes cases where characters are superimposed on a moving image such as a landscape or an animation, and cases where voices such as BGM (Back Ground Music) are added. In the following description, the case where the information is an image will be described as an example.
[0016] The action is, for example, clicking on an image or not clicking on it. When the image is an advertisement, the action is buying the advertised product or not buying it. The action may be classified into three or more values depending on, for example, the number of products purchased or the store where the product was purchased.
[0017] An individual's action is affected by various factors including the mood and physical condition at the time of receiving the information, and the individual's actions immediately before receiving the information. Therefore, based on only the presented information, it is not possible to determine how an individual's action will be, nor can reproducibility be expected.
[0018] However, in the case of a group composed of a large number of individuals, there may be a tendency towards reproducibility in the actions after the information is presented. Specifically, by presenting information to each individual belonging to the group and then observing their subsequent actions or conducting a questionnaire, it is possible to quantify the action tendencies such as the number or proportion of people who took a certain action. The group may be formed, for example, by age, gender, or residential area, etc. By comparing the action tendencies between groups, the characteristics of the action tendencies for each group can be clarified.
[0019] By modeling the action tendencies of a group when receiving various information, a black box model 22 can be created. The black box model 22 is a model that receives an input of an explanatory database vector xn representing the information provided to a person with various feature quantities and outputs a target database vector yn indicating the tendency of the people belonging to the group. The black box model 22 is generated using a supervised machine learning algorithm for a known classification model such as, for example, random forest, CNN (Convolutional Neural Network), or transformer.
[0020] The black box model 22 may be realized by inputting, for example, into a large language model (LLM) such as GPT-4 or Gemini, a query that combines the action tendencies of a group when receiving various information, the newly presented information, and the question text.
[0021] The procedure of inputting the explanatory database vector xn into the black box model 22 and obtaining the target database vector yn is repeated multiple times. In the following explanations, the pair of the explanatory database vector xn and the corresponding target database vector yn may be referred to as action data.
[0022] The explanatory database vector xn may include the explanatory database vector xn corresponding to the actually presented information. The explanatory database vector xn may be generated randomly or based on a predetermined rule.
[0023] By arranging a plurality of explanatory data vectors xn in the row direction, an explanatory matrix X which is a two-dimensional matrix is generated. Similarly, by arranging a plurality of target data vectors yn in the row direction, a target matrix Y which is a two-dimensional matrix is generated. Here, the array order of the explanatory data vectors xn and the array order of the corresponding target data vectors yn are the same. That is, the target data vectors yn are arranged in the same order as the corresponding explanatory data vectors xn.
[0024] By using the black box model 22, an explanatory matrix X and a target matrix Y are generated which contain more elements than the information obtained by actually presenting to a large number of people and observing their behavioral tendencies. Note that in order to execute the subsequent processes, the target data vectors yn need to be linearly independent. That is, the vector product of the target matrix Y and the transposed matrix of the target matrix Y needs to be a regular matrix.
[0025] Based on the explanatory matrix X and the target matrix Y, an interpretation matrix A† which is a matrix whose vector product with the target matrix Y is equal to the explanatory matrix X as shown in equation (1) is calculated. The details of the method for calculating the interpretation matrix A† will be described later. X = A†Y ‥‥‥ (1)
[0026] The interpretation matrix A† can be used to calculate the explanatory data vector x corresponding to the information that causes the target data vector y representing the desired behavior for the user. Taking the case where the information is a product advertisement as an example. The user can calculate the explanatory data vector xp that causes the target data vector yp indicating the behavior of purchasing a specific product based on equation (2). xp = A†yp ‥‥‥ (2)
[0027] The user can adjust the created advertisement so that the explanatory data vector xn approaches the explanatory data vector xp calculated based on equation (1), thereby realizing an advertisement with a high advertising effect, that is, an advertisement that contributes to an increase in the sales of the product. That is, it is desirable that each element constituting the explanatory data vector xn is an element that can be adjusted by the user as appropriate. Specific examples of each element will be described later.
[0028] Figure 2 is an explanatory diagram for explaining a method of calculating the interpretation matrix A†. In the following description, N is a natural number indicating the number of times the process of inputting the explanatory data vector xn to the black box model 22 and obtaining the target data vector yn is repeated. n is a natural number indicating which time the vector is input to the black box model 22 or output from the black box model 22.
[0029] The explanatory data vector xn has L elements from Ex1n to ExLn. The target data vector yn has M elements from Ob1n to ObMn. Here, L and M are natural numbers. In Figure 2, the explanatory data vector xn and the target data vector yn where n = 2 are shown surrounded by a dashed line. As described above, the pair of the explanatory data vector xn and the target data vector yn is the behavioral data.
[0030] As described above, by arranging the N explanatory data vectors xn obtained by the N processes in the row direction, an explanatory matrix X which is a two-dimensional matrix is created. The explanatory matrix X is a two-dimensional matrix with L rows and N columns. Similarly, by arranging the N target data vectors yn in the row direction, a target matrix Y which is a two-dimensional matrix is generated. The target matrix Y† is a two-dimensional matrix with M rows and N columns.
[0031] Regarding the target matrix Y, the Moore-Penrose generalized inverse matrix Y† is calculated. In the following description, Y† may be described as the target inverse matrix Y†. The target inverse matrix Y† is calculated by equation (3). The target inverse matrix Y† is a matrix with N rows and M columns. Y† = Y T (YY T )-1 ‥‥‥ (3)
[0032] The interpretation matrix A† is the vector product of the explanatory matrix X and the target inverse matrix Y†. The equation for calculating the interpretation matrix A† is shown in Equation (4). A† = XY† ‥‥‥ (4)
[0033] The interpretation matrix A† is a two-dimensional matrix with L rows and M columns, that is, the same number of rows as the number of elements in the explanatory data vector xn and the same number of columns as the number of elements in the target data vector yn. The element in the a-th row and b-th column of the interpretation matrix A† indicates the influence of the a-th element of the explanatory data vector xn on the b-th element of the target data vector yn.
[0034] As described above, even without information such as the algorithm for generating the black box model 22 and the training data, the interpretation matrix A† can be generated if there is an environment in which the black box model 22 can be used. That is, even when using the black box model 22 generated by a third party, the interpretation matrix A† can be generated.
[0035] For reference, the outline of the formula transformation for deriving Equation (4) from Equations (1) and (3) is shown below. First, multiply both sides of Equation (1) by the transpose matrix of the target matrix Y from the right to obtain Equation (5). XY T = A†YY T ‥‥‥ (5)
[0036] As described above, since the vector product of the target matrix Y and the transpose matrix of the target matrix Y is a regular matrix, the inverse matrix can be calculated. Multiply both sides of Equation (5) by this inverse matrix from the right to obtain Equation (6). XY T (YY T ) -1 = A†(YY T )(YY T ) -1 = A† ‥‥‥ (6)
[0037] After swapping the left side and the right side of Equation (6) and substituting Equation (3) into the right side, Equation (7) is obtained. Equation (4) is derived from both ends of Equation (7). A† = XY T (YY T ) -1 = XY† ‥‥‥ (7)
[0038] FIG. 3 is an explanatory diagram for explaining the configuration of the information processing apparatus 10. The information processing apparatus 10 includes a control unit 11, a main storage device 12, an auxiliary storage device 13, a communication unit 14, a display unit 15, an input unit 16, a reading unit 19, and a bus.
[0039] The control unit 11 is an arithmetic control device that executes the program of the present embodiment. One or more CPUs (Central Processing Units), GPUs (Graphics Processing Units), TPUs (Tensor Processing Units), or multi-core CPUs, etc. are used for the control unit 11. The control unit 11 is connected to each hardware part constituting the information processing apparatus 10 via a bus.
[0040] The main storage device 12 is a storage device such as SRAM (Static Random Access Memory), DRAM (Dynamic Random Access Memory), or flash memory. In the main storage device 12, information necessary during the processing performed by the control unit 11 and the program being executed by the control unit 11 are temporarily stored.
[0041] The auxiliary storage device 13 is a storage device such as SRAM, flash memory, hard disk, or magnetic tape. In the auxiliary storage device 13, the black box model 22, the program to be executed by the control unit 11, and various data necessary for the execution of the program are stored. The black box model 22 may be stored in an external storage device connected via a network.
[0042] The communication unit 14 is an interface that conducts communication between the information processing apparatus 10 and the network. The display unit 15 is, for example, a liquid crystal display device or an organic EL (Electro Luminescence) display device. The input unit 16 is an input device such as, for example, a keyboard, a mouse, a trackball, or a microphone.
[0043] The portable recording medium 96 is, for example, a USB (Universal Serial Bus) memory, a CD-ROM (Compact Disc Read only memory), a magneto-optical disk medium, other optical disk media, or an SD memory card, etc. The portable recording medium 96 stores a program 97 for realizing AIME.
[0044] The reading unit 19 is an interface capable of connecting the portable recording medium 96, such as, for example, a USB connector, a CD-ROM drive, or an SD memory reader, etc. The semiconductor memory 98 stores the program 97 and is a memory that can be installed inside the information processing apparatus 10.
[0045] The information processing apparatus 10 is a general-purpose personal computer, a tablet, a mainframe computer, a virtual machine operating on a mainframe computer, or a quantum computer. The information processing apparatus 10 may be constituted by a plurality of personal computers that conduct distributed processing, or hardware such as a mainframe computer. The information processing apparatus 10 may be constituted by a cloud computing system. The information processing apparatus 10 may be constituted by a plurality of personal computers that operate in cooperation, or hardware such as a mainframe computer.
[0046] The program 97 is recorded on the portable recording medium 96. The control unit 11 reads the program 97 via the reading unit 19 and stores it in the auxiliary storage device 13. Further, the control unit 11 may read out the program 97 stored in the semiconductor memory 98. Furthermore, the control unit 11 may download the program 97 from another server computer (not shown) connected via the communication unit 14 and a network (not shown) and store it in the auxiliary storage device 13.
[0047] Program 97 is installed as a control program for the information processing apparatus 10, loaded into the main storage device 12, and executed. The program 97 of the present embodiment is an example of a program product.
[0048] FIG. 4 is a flowchart for explaining the processing flow of a program for calculating the interpretation matrix A†. The control unit 11 determines the explanatory data vector xn to be input to the black box model 22 at the n-th time (step S501). When information used in a questionnaire or the like at the time of data collection can be obtained, the control unit 11 may generate the explanatory data vector xn based on the information. When the training data used for machine learning of the black box model 22 can be obtained, the control unit 11 may extract the explanatory data vector xn from the training data.
[0049] The control unit 11 may generate the explanatory data vector xn randomly or based on a predetermined rule. When the control unit 11 generates the explanatory data vector xn, it is desirable to generate the explanatory data vector xn within the range where the use of the black box model 22 is assumed. If an explanatory data vector xn including an unexpected element is generated, the behavior of the black box model 22 cannot be correctly interpreted. Similarly, it is desirable to match the distribution of the plurality of explanatory data vectors xn with the range where the use of the black box model 22 is assumed.
[0050] The control unit 11 may obtain information automatically generated by a generative AI (Artificial Intelligence) and calculate the explanatory data vector xn.
[0051] The control unit 11 inputs the explanatory data vector xn obtained in step S501 to the black box model 22 to obtain the target data vector yn (step S502). The control unit 11 associates the explanatory data vector xn and the target data vector yn and records them in the main storage device 12 or the auxiliary storage device 13 (step S503).
[0052] The control unit 11 determines whether or not to finish generating the pair of the explanatory data vector xn and the target data vector yn (step S504). For example, when the control unit 11 finishes processing the explanatory data vector xn recorded in the training data, it determines to end in step S504. The control unit 11 may determine to end the process when a predetermined number of pairs are generated. If it determines not to end (NO in step S504), the control unit 11 returns to step S501.
[0053] If it determines to end (YES in step S504), the control unit 11 generates an explanatory matrix X based on the data recorded in step S503 (step S505). The control unit 11 generates a target matrix Y based on the data recorded in step S503 (step S506). The control unit 11 calculates a target inverse matrix Y† which is the Moore-Penrose generalized inverse matrix based on the target matrix Y (step S507). The control unit 11 calculates an interpretation matrix A† which is the vector product of the explanatory matrix X and the target inverse matrix Y† (step S508). The control unit 11 ends the process.
[0054] As described above, the element in the a-th row and b-th column of the interpretation matrix A† indicates the influence of the a-th element of the explanatory data vector xn on the b-th element of the target data vector yn. That is, the element in the a-th row and b-th column of the interpretation matrix A† indicates the influence of the a-th element of the explanatory data vector xn, that is, the degree to which it became the reason for the action, on the action taken by the people belonging to the group corresponding to the b-th element of the target data vector yn.
[0055] A specific example of the usage of the interpretation matrix A† will be described later together with specific examples of the explanatory data vector xn and the target data vector yn.
[0056] According to this embodiment, an information processing method for assisting in the interpretation of human reasons for actions can be provided. By using the black box model 22, the interpretation matrix A† can be calculated using more data than the data obtained by actually presenting information to a group and investigating actions. If an appropriate black box model 22 is generated, an appropriate interpretation matrix A† can be calculated with fewer research institutions and less effort.
[0057] [Embodiment 2] This embodiment relates to an information processing method for generating the interpretation matrix A† without creating the black box model 22. Descriptions of parts common to Embodiment 1 are omitted.
[0058] FIG. 5 is an explanatory diagram for explaining the outline of the interpretation matrix A† of Embodiment 2. In this embodiment, a large amount of information and data obtained by observing the action tendencies of the group that received the information are prepared. Based on each piece of information, an explanatory data vector xn is calculated. Based on each action tendency, a target data vector yn is calculated.
[0059] By arranging the explanatory data vectors xn in the row direction, an explanatory matrix X, which is a two-dimensional matrix, is created. By arranging the target data vectors yn in the row direction, a target matrix Y, which is a two-dimensional matrix, is generated. The interpretation matrix A† is calculated from the explanatory matrix X and the target matrix Y.
[0060] FIG. 6 is a flowchart for explaining the processing flow of the program for calculating the interpretation matrix A† of Embodiment 2. The flowchart in FIG. 6 is executed with the explanatory data vector xn and the target data vector yn stored in the auxiliary storage device 13.
[0061] The control unit 11 acquires a plurality of explanatory data vectors xn stored in the auxiliary storage device 13 (step S511). The control unit 11 generates an explanatory matrix X in which the plurality of explanatory data vectors xn are arranged in the row direction (step S512). The control unit 11 acquires a plurality of target data vectors yn stored in the auxiliary storage device 13 (step S513). The control unit 11 generates a target matrix Y in which the plurality of target data vectors yn are arranged in the row direction (step S514). Note that the array order of the target data vectors yn is the same as the array order of the corresponding explanatory data vectors xn when generating the explanatory matrix X.
[0062] The control unit 11 calculates a target inverse matrix Y†, which is the Moore-Penrose generalized inverse matrix, based on the target matrix Y (step S515). The control unit 11 calculates an interpretation matrix A†, which is the vector product of the explanatory matrix X and the target inverse matrix Y† (step S516). The control unit 11 ends the process.
[0063] According to the present embodiment, the interpretation matrix A† can be calculated without creating the black box model 22. When a large amount of data obtained by actually presenting information to a group and investigating behaviors can be collected, an appropriate interpretation matrix A† can be calculated with low computational cost.
[0064] [Embodiment 3] This embodiment relates to a specific example of the interpretation matrix A† described with reference to FIGS. 1 and 5. Descriptions of parts common to Embodiment 1 are omitted.
[0065] FIG. 7 is an explanatory diagram for explaining the outline of the process of Embodiment 3. A video site 45 is provided where a user can view various videos. In the following description, a user who views videos using the video site 45 is referred to as a viewing user. The viewing user accesses the video site 45 using a playback device 46 that can be connected to a network, such as a smartphone, tablet, personal computer, or smart glasses.
[0066] In the upper right part of FIG. 7, there is shown an example of a screen of a video site 45 displayed on a playback device 46 when the playback device 46 is a smartphone. An image 47 being played is arranged at the upper part of the screen. Immediately below the image 47 being played, there is arranged a title bar 471 including the title of the image 47 being played, the number of plays, a summary text, and the like.
[0067] Below the title bar 471, a plurality of thumbnail images 48 are displayed in a vertical column. To the right of each thumbnail image 48, there is displayed a title bar 481 including the title of the video corresponding to the thumbnail image 48, the number of plays, a summary text, and the like.
[0068] When a browsing user clicks on a thumbnail image 48, the image 47 being played switches to the video corresponding to the clicked thumbnail image 48. The browsing user can view the video. Data such as the number of plays and the playback time of each video is recorded. The data is provided to a third party based on a predetermined rule.
[0069] In the following description, a user who analyzes the reasons for the actions of a browsing user using the interpretation matrix A† will be described as an analysis user. The analysis user acquires data satisfying predetermined conditions from the video site 45 and records it in a video information DB (Database) 34 (see FIG. 8). The record layout of the video information DB 34 will be described later.
[0070] The video information DB 34 includes a plurality of thumbnail images 48 and the number of plays of the videos corresponding to the respective thumbnail images 48. Here, the thumbnail images 48 are information presented to the browsing user. The number of plays is an action of the browsing user.
[0071] An explanatory data vector xn is created based on each thumbnail image 48. Table 1 shows the items of feature amounts used in the present embodiment. In the present embodiment, for each thumbnail image 48, thirty feature amounts shown in Table 1 are extracted. Therefore, the explanatory data vector xn has thirty elements, and the value of L in FIG. 2 is 30.
[0072]
Table 1
[0073] 1-1 to 1-6 are feature amounts related to the faces of the people in the thumbnail image 48. Among them, 1-3 to 1-6 are values obtained by averaging the detected faces after scoring on a five-point scale where each face shows a weak emotion as one point and a strong emotion as five points. When no face is detected, it is 0 point. The feature amounts related to the face can be calculated by, for example, known face detection technology and face expression determination technology. The detection of the face and the determination of the expression may also be performed by a person.
[0074] 2-1 to 2-7 are feature amounts related to the text displayed in the thumbnail image 48. The feature amounts related to the text can be calculated by known character recognition technology. The determination of the feature amounts related to the text may also be performed by a person.
[0075] 3-1 to 3-11 are feature amounts related to the color of the thumbnail image 48. The pixels constituting the thumbnail image 48 are classified into five clusters based on color. For cluster classification, for example, the k-means method is used. The color corresponding to the centroid of each cluster is the representative color of the cluster. The representative colors of the clusters are classified into any of the eleven basic color terms. The feature amounts of 3-1 to 3-11 are the numbers of the representative colors classified into each color. The feature amounts related to the color are calculated using software for cluster classification.
[0076] For example, a large feature amount indicated by reference numeral 3-1 indicates that the thumbnail image 48 has a large area occupied by red-based colors. Similarly, a large feature amount indicated by reference numeral 3-11 indicates that the thumbnail image 48 has a large area occupied by whitish colors.
[0077] 4-1 to 4-6 are texture feature quantities using a GLCM (Gray Level cooccurrence Matrix). Since all of them are well-known texture feature quantities, the detailed explanation will be omitted. The texture feature quantities are calculated using software for texture analysis.
[0078] Based on the number of playback times, the target data vector yn is generated. In this embodiment, the element of the target data vector yn is one. That is, in FIG. 2, the value of M is 1, and the target data vector yn is a scalar quantity indicating the number of people who viewed the video. Since the value of L is 30 and the value of M is 1, the interpretation matrix A† is a vector of thirty rows and one column.
[0079] For each country where the video site 45 is provided, the interpretation matrix A† is calculated. By comparing each element of the interpretation matrix A†, as will be described later, the characteristics of the reasons for the behavior of the viewing users for each country will be revealed.
[0080] FIG. 8 is an explanatory diagram for explaining the record layout of the video information DB34. The video information DB34 is a database in which video information acquired from the video site 45 is recorded. The video information DB34 is stored in an auxiliary storage device 13 or a mass storage device connected to the information processing device 10.
[0081] The video information DB34 has a country field, a thumbnail image field, a playback count field, and a video ID field. In the country field, the country where the viewing user resides or stays is recorded. In the thumbnail image field, the thumbnail image 48 is recorded. In the playback count field, the number of times the video file was played by selecting the thumbnail image 48 is recorded. In the video ID field, the video ID uniquely assigned to the video file played by selecting the thumbnail image 48 is recorded.
[0082] The video information DB34 has one record for one thumbnail image 48. Note that a plurality of thumbnail images 48 may be created for one video file. That is, the same video ID may be recorded in the video ID fields of a plurality of records in the video information DB34.
[0083] In the present embodiment, for videos distributed by popular video distributors, the number of views for each thumbnail image 48 one week after the video was published was obtained from the video site 45. Table 2 shows the number of thumbnail images 48 for each country used in the present embodiment.
[0084]
Table 2
[0085] Generally, as time passes after a video is published, various factors such as information dissemination by so-called influencers contribute to the number of views. In the present embodiment, in order to examine the influence of the characteristics of the thumbnail image 48 on the number of views, the number of views one week after the video was published was used.
[0086] FIG. 9 is an explanatory diagram for explaining the record layout of the behavior data DB35. The behavior data DB35 is a database in which the feature amounts of the thumbnail image 48 are recorded. The behavior data DB35 is stored in a mass storage device connected to the auxiliary storage device 13 or the information processing device 10.
[0087] The behavior data DB35 has a country field, a thumbnail image field, a feature amount field, and a playback count field. The feature amount field has thirty sub-fields corresponding to the symbols of the respective feature amounts described using Table 1, such as a 1-1 field and a 1-2 field.
[0088] In the country field, the country where the browsing user resides or stays is recorded. In the thumbnail image field, the thumbnail image 48 is recorded. In each sub-field of the feature quantity field, each feature quantity obtained from the thumbnail image 48 recorded in the thumbnail image field is recorded. In the playback count field, the number of times the video file has been played based on the selection of the thumbnail image 48 is recorded.
[0089] The behavior data DB35 has one record for one thumbnail image 48. The feature quantities recorded in the feature quantity field may be calculated by the control unit 11 or by another information processing device (not shown). The feature quantities may be provided from the video site 45. The feature quantities may be input by the analyzing user.
[0090] In the behavior data DB35, the elements of the explanatory data vector xn are recorded in the thirty feature quantity sub-fields included in one record, and the elements of the target data vector yn are recorded in the playback count field. That is, one piece of behavior data is recorded in one record.
[0091] The control unit 11 executes the flowchart described with reference to FIG. 4 or the flowchart described with reference to FIG. 6 to calculate the interpretation matrix A† of thirty rows and one column for each country. The control unit 11 displays the interpretation matrix A† in a chart or the like. The user views the displayed chart or the like.
[0092] FIG. 10 is a graph showing the interpretation matrix A† regarding Japan. FIG. 11 is a graph showing the interpretation matrix A† regarding Germany. The horizontal axes of FIGS. 10 and 11 are the values of the respective elements of the interpretation matrix A†. The vertical axes of FIGS. 10 and 11 are symbols indicating the feature quantities corresponding to the respective elements of the interpretation matrix A†. The symbols on the vertical axis are the same as the symbols shown in the symbol column of Table 1. To the left of the vertical axis, the ranges corresponding to "face", "text", "color", and "texture" are illustrated. FIGS. 10 and 11 are examples of charts regarding the interpretation matrix A†.
[0093] The longer the graph corresponding to each element is in the positive direction, the more it means that the element contributes to the behavior of the browsing user clicking on the thumbnail image 48 to view the video. Conversely, the longer it is in the negative direction, the more it means that the element contributes to the behavior of the browsing user not clicking on the thumbnail image 48 and not viewing the video.
[0094] An example of the information obtained by comparing FIGS. 10 and 11 will be described. In both Japan and Germany, the fact that there are many faces of people included in the thumbnail image 48 contributes to an increase in the number of views. In Japan, the fact that the score (symbols 1-3) indicating joy on the face is large does not contribute much to an increase in the number of views. In Germany, the score (symbols 1-3) indicating joy on the face, the score (symbols 1-4) indicating sadness on the face, the score (symbols 1-5) indicating anger on the face, and the score (symbols 1-6) indicating surprise on the face all contribute to an increase in the number of views. The fact that there is a lot of yellow text (symbol 2-5) displayed in the thumbnail image 48 contributes to an increase in the number of views in Japan, but contributes to a decrease in the number of views in Germany.
[0095] FIG. 12 is a graph showing the interpretation matrix A† for seven countries. The horizontal axis of FIG. 12 is a symbol indicating the feature amount corresponding to each element of the interpretation matrix A†. The symbols on the horizontal axis are the same as the symbols shown in the symbol column of Table 1. Below the horizontal axis, the ranges corresponding to "face", "text", "color", and "texture" are illustrated. The vertical axis of FIG. 12 is the respective countries shown in Table 2. In FIG. 12, the seven interpretation matrices A† calculated for each of the seven countries are displayed together on one sheet. FIG. 12 is an example illustration of a chart that displays a plurality of interpretation matrices A† together on one sheet.
[0096] For each feature amount in each country, positive values are illustrated with downward-left hatching, and negative values are illustrated with downward-right hatching. The finer the hatching, the larger the absolute value of the feature amount. For the part where the absolute value of the feature amount is less than 0.2, no hatching is applied.
[0097] For example, as shown in FIG. 10, in Japan, since the feature amounts of symbols 1-1 and 1-2 are positive values, downward-left hatching is displayed in FIG. 12. Similarly, in Japan, since the feature amounts of symbols 4-3 to 4-6 are negative values, downward-right hatching is displayed in FIG. 12.
[0098] With the display method shown in FIG. 12, for each of the seven countries, the magnitudes of the thirty feature amounts can be aggregated and displayed in one figure. Note that by using the shade of color instead of hatching, the magnitudes of the feature amounts can be visualized in more detail.
[0099] Taking an example of the information that can be grasped by the analysis user based on FIG. 12. In any of the seven countries, a large number of faces (symbol 1-1) displayed in the thumbnail image 48 contributes to an increase in the number of views. In the five countries excluding Japan and Germany, the ratio of the face area (symbol 1-2) contributes more to the increase in the number of views than the number of faces (symbol 1-1).
[0100] A large score of the face expressing joy (symbol 1-3) greatly contributes to an increase in the number of views in Germany, but does not contribute in Japan, Australia, and Canada. The ratio of the area occupied by the text region (symbol 2-2) contributes to an increase in the number of views in Japan and Germany, but does not contribute in the other five countries.
[0101] A large area occupied by a reddish color (symbol 3-1) contributes to an increase in the number of views in the five countries excluding the UK and Canada. Among the five countries, in Australia, it greatly contributes to an increase in the number of views. A large area occupied by gray (symbol 3-9) contributes to a decrease in the number of views in the five countries excluding the US and Canada. Among the five countries, in Australia, it greatly contributes to a decrease in the number of views.
[0102] Based on this information, the analysis user can expect an increase in the number of views in each country by adjusting the thumbnail image 48 for each country.
[0103] Figure 13 is a diagram showing the average value of the elements of the interpretation matrix A† for seven countries. The horizontal axis of Figure 13 represents the value of each element of the interpretation matrix A†. The vertical axis of Figure 13 is a symbol indicating the feature quantity corresponding to each element of the interpretation matrix A†. The symbol on the vertical axis is the same as the symbol shown in the symbol column of Table 1. In Figure 13, the feature quantities corresponding to "color" and "texture" are shown. Figure 3 is an illustration of a chart regarding the interpretation matrix A†.
[0104] Based on Figure 12, examples of information that can be grasped by the analysis user are given. The fact that the contrast (symbol 4-1) of the thumbnail image 48 is large and the amount of local change (symbol 4-2) is large contributes more to the increase in the number of views than the color tendency (symbols 3-1 to 3-11) of the entire thumbnail image 48.
[0105] The average value of the elements of the interpretation matrix A† for seven countries may be an average value weighted based on the number of data in each country shown in Table 2. The average value of the elements of the interpretation matrix A† for seven countries may be an average value weighted based on, for example, the number of active users in each country or the population.
[0106] The behavior of the browsing users obtained from the video site 45 may be, for example, the ratio of the people who viewed the video among the users who viewed the thumbnail image 48. The behavior of the browsing users obtained from the video site 45 may be, for example, the number of people who viewed the video and the number of people who did not view the video among the users who viewed the thumbnail image 48. When the behavior is binary, the interpretation matrix A† becomes a two-column matrix.
[0107] The feature quantity of the thumbnail image 48 is not limited to the feature quantities listed in Table 1. For example, a feature quantity indicating the age, gender, hairstyle, or clothing of the people included in the thumbnail image 48 may be used. The selection of the feature quantity is arbitrary, but it is desirable that the feature quantity can be reflected in the thumbnail image 48.
[0108] In addition to the feature amount of the thumbnail image 48, the feature amount regarding the title of the video may be used for the explanatory data vector xn. The feature amount regarding the title of the video is, for example, the number of characters in the title, the number of Chinese characters, the number of exclamation marks, and the like.
[0109] The country is an example of a plurality of groups that divide the browsing users. The interpretation matrix A† may be created for each group that divides the browsing users by attributes such as age or gender, for example. By adjusting the thumbnail image 48 according to the attributes of the target layer that is desired to be browsed, an increase in the number of views of the target layer can be expected.
[0110] [Modification Example 3-1] The information presented to the user is a poster for product advertising, and the user's behavior may be the amount of change in the sales volume of the product. An example will be described where a new poster regarding an existing product is posted. The explanatory data vector xn includes feature amounts regarding people and text in the poster, etc. The explanatory data vector xn may include feature amounts regarding the posting location and posting density of the poster, etc. The target data vector yn includes the amount of change in the sales volume of the product one week after the poster is posted.
[0111] By using the interpretation matrix A† calculated based on a set of a large number of explanatory data vectors xn and target data vectors yn, the user can determine the features of the poster that promotes the purchase of the product.
[0112] The information presented to the user may be an advertising video displayed on a TV, browser, or digital signage on the street, etc. The information presented to the user may be the package of the product. The information presented to the user may be the display method of the product in a retail store.
[0113] In any case, by determining the feature amount that promotes the purchase of the product using the interpretation matrix A†, an effect of increasing the sales volume of the product can be expected.
[0114] The information presented to the user may be, for example, prizes provided to customers in a campaign for sales promotion, the winning probability of the prizes, the eligibility for participating in the campaign, or the implementation period of the campaign. The actions of the user may be, for example, the number of applications to the campaign or the number of times the campaign is mentioned on SNS (Social Network Service).
[0115] In any case, by determining the feature amount that promotes the reaction to the campaign using the interpretation matrix A†, an effect of increasing the effectiveness of the campaign can be expected. That is, the interpretation matrix A† can be used for marketing activities such as sales promotion activities. [Modification Example 3-2] The information presented to the user may be the product itself, and the actions of the user may be the change amount of the sales volume of the product. The explanatory data vector xn includes feature amounts related to the product itself. For example, when the product is food, features such as size, weight, taste, smell, texture, and color can be used. The target data vector yn includes the sales volume of the product or the number of discarded products due to leftovers.
[0116] By determining the feature amount that promotes the purchase of the product using the interpretation matrix A†, it is possible to contribute to the development of a product with good sales. That is, the interpretation matrix A† can be used for marketing activities during product development.
[0117] The program is an exemplification of a program product. The computer program can be deployed to be executed on a single computer, or on one site, or distributed over a plurality of sites and executed on a plurality of computers interconnected by a communication network.
[0118] The technical features (constituent elements) described in each embodiment can be combined with each other, and new technical features can be formed by the combination. The embodiments disclosed herein should be considered illustrative in all respects and not restrictive. The scope of the present invention is shown not by the above description but by the claims, and it is intended that all modifications within the meaning and scope equivalent to the claims be included.
[0119] The independent claims and dependent claims described in the claims can be combined with each other in any combination regardless of the citation form. Further, although the claims use a form (multi-claim form) of describing a claim that cites two or more other claims, it is not limited thereto. A form of describing a multi-claim (multi-multi-claim) that cites at least one multi-claim may be used.
Description of Reference Numerals
[0120] 10 Information processing apparatus 11 Control unit 12 Main storage device 13 Auxiliary storage device 14 Communication unit 15 Display unit 16 Input unit 19 Reading unit 34 Video information DB 35 Behavior data DB 45 Video site 46 Reproduction device 47 Image during reproduction 471 Title bar 48 Thumbnail image 481 Title bar 96 Portable recording medium 97 Program 98 Semiconductor memory
Claims
1. Obtaining a plurality of sets of behavioral data each of which is a combination of an explanatory data vector composed of a plurality of feature amounts related to information presented to a person and a purpose data vector which quantifies the behavior of a plurality of people to whom the information is presented; Calculating an interpretation matrix which is a vector product of an explanation matrix in which a plurality of sets of the explanation data vectors included in the behavior data are arranged, and a generalized inverse matrix of a target matrix in which the target data vectors included in the behavior data are arranged in an order corresponding to the explanation data vectors; Output a chart related to the interpretation matrix An information processing method in which processing is performed by a computer.
2. The generalized inverse of the objective matrix is the Moore-Penrose generalized inverse of the objective matrix The information processing method according to claim 1 .
3. The objective data vector is a scalar quantity indicating the number of people who performed a particular action; Each element of the interpretation matrix indicates the degree to which each item of the feature quantity encourages the specific behavior. The information processing method according to claim 1 .
4. the information is an image, The feature amount includes the number of human faces included in the image, the ratio of the human faces to the image, and the number of characters included in the image. The information processing method according to claim 1 .
5. The behavioral data includes obtaining a plurality of said explanatory data vectors; The explanatory data vectors are input to a model that receives the explanatory data vectors and outputs a target data vector, and the target data vector output from the model is combined with the explanatory data vector input to the model to obtain the explanatory data vector. The information processing method according to claim 1 .
6. Calculating the interpretation matrix for each of a plurality of groups; The calculated interpretation matrices are output on a single chart.
6. The information processing method according to claim 1.
7. Obtaining a plurality of sets of behavioral data each of which is a combination of an explanatory data vector composed of a plurality of feature amounts related to information presented to a person and a purpose data vector which quantifies the behavior of a plurality of people to whom the information is presented; Calculating an interpretation matrix which is a vector product of an explanation matrix in which a plurality of sets of the explanation data vectors included in the behavior data are arranged, and a generalized inverse matrix of a target matrix in which the target data vectors included in the behavior data are arranged in an order corresponding to the explanation data vectors; Output a chart related to the interpretation matrix A program that causes a computer to carry out processing.
8. An information processing device including a control unit, The control unit is Obtaining a plurality of sets of behavioral data each of which is a combination of an explanatory data vector composed of a plurality of feature amounts related to information presented to a person and a purpose data vector which quantifies the behavior of a plurality of people to whom the information is presented; Calculating an interpretation matrix which is a vector product of an explanation matrix in which a plurality of sets of the explanation data vectors included in the behavior data are arranged, and a generalized inverse matrix of a target matrix in which the target data vectors included in the behavior data are arranged in an order corresponding to the explanation data vectors; Output a chart related to the interpretation matrix Information processing device.
Citation Information
Patent Citations
Medical image processing apparatus, endoscope system, medical image processing system, method of operating medical image processing apparatus, program, and storage medium
JP2023083555A