Information processing method, program, and information processor

The information processing method addresses the challenge of interpreting human reactions by calculating an interpretation matrix from time-series data and its derivatives, enabling the explanation of reaction influences and improving content preference understanding.

JP2025089244APending Publication Date: 2025-06-12EIGENBEATS LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120324
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-01
Filing Date
2024-07-25
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

It is challenging for humans to interpret the reasons behind a person's reaction to information, such as emotions or feelings towards specific content, which is similar to the interpretability issues faced with machine learning models.

Method used

An information processing method that associates and records explanatory data vectors from time-series data, its first-order and second-order differentials, with target data vectors representing human reactions. This method calculates an interpretation matrix through a vector product of an explanatory matrix and a generalized inverse matrix of a target matrix, and outputs a chart for interpretation.

Benefits of technology

The method supports the interpretation of human reactions by providing a chart that explains the influence of different features on the reaction, aiding in understanding preferences and improving content creation based on user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025089244000001_ABST
    Figure 2025089244000001_ABST
Patent Text Reader

Abstract

To provide an information processing method and the like for supporting interpretation concerning a person's reaction.SOLUTION: In an information processing method, a computer executes processing of: recording a plurality of sets of an explanatory data vector xn including time series feature amount data acquired from original data which is time series data, first order differential feature amount data of the time series feature amount data, and second order differential feature amount data of the time series feature amount data, and an objective data vector yn representing a person's reaction to the original data in association with each other; calculating an interpretation matrix (A dagger) which is a vector product of an explanatory matrix X in which the plurality of sets of explanatory data vectors xn are arranged and a generalized inverse matrix (Y dagger) of an objective matrix Y in which the objective data vectors yn are arranged in an order corresponding to the explanatory data vectors xn; and outputting a chart concerning the interpretation matrix (A dagger).SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing method, a program, and an information processing apparatus.

Background Art

[0002] A machine learning model generated by machine learning is a black box, and it is difficult for a user to interpret the behavior. XAI (Explainable Artificial Intelligence) technology has been proposed to show a reasonable ground for the result output by the machine learning model for human understanding (Patent Document 1, Non-Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] By the way, a phenomenon in which the reason for the occurrence of a result cannot be easily interpreted by humans occurs not only in the output result of a machine learning model. For example, it may be difficult for others as well as the person who reacted to interpret the reason for the reaction of a person who has come into contact with some information, that is, what kind of emotion or feeling the person has towards the information.

[0006] On one side, it aims to provide an information processing method or the like that supports the interpretation of human reactions.

Means for Solving the Problem

[0007] The information processing method associates and records multiple sets of an explanatory data vector including time-series feature quantity data obtained from original data that is time-series data, first-order differential feature quantity data of the time-series feature quantity data, and second-order differential feature quantity data of the time-series feature quantity data, and a target data vector representing a human reaction to the original data. The computer executes a process of calculating an interpretation matrix that is a vector product of an explanatory matrix in which multiple sets of the explanatory data vectors are arranged and a generalized inverse matrix of a target matrix in which the target data vectors are arranged in the order corresponding to the explanatory data vectors, and outputting a chart regarding the interpretation matrix.

Effect of the Invention

[0008] On one side, it is possible to provide an information processing method or the like that supports the interpretation of human reactions.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

[0010] [Embodiment 1] Using various machine learning algorithms, a machine learning model is generated that accepts the input of explanatory data and outputs target data. When using the generated machine learning model, it is difficult for humans to interpret the decision-making process from the input of explanatory data to the output of target data.

[0011] However, when applying the machine learning model to decision-making in the real world, it is important for humans to be able to interpret the decision-making process of the machine learning model. For example, when the output target data seems to deviate significantly from human common sense, if humans can appropriately interpret the decision-making process of the machine learning model, they can also appropriately determine how to handle the target data and the machine learning model.

[0012] Regarding the target data output from the machine learning model, the technology for explaining the reason for the output is called XAI technology. Non-Patent Document 1 discloses AIME (Approximate Inverse Model Explanations), which is a type of XAI technology. AIME is an information processing method that assists users in interpreting the behavior of various machine learning models, including black box models with unknown generation algorithms, from various perspectives.

[0013] By the way, as described above, it is difficult to interpret the reason why a person who has come into contact with some information reacts. In the present embodiment, an information processing method for assisting the interpretation of the subjective reaction of a person who has come into contact with music is provided.

[0014] FIG. 1 is an explanatory diagram for explaining an outline of a procedure for analyzing a user's reaction to a piece of music. The user listens to various pieces of music and answers questions prepared in advance for each piece of music. The question items are, for example, a binary choice between "like" and "dislike". The question items may be a binary choice between "like" and "not like".

[0015] The question items may be, for example, "How many points out of 10 points do you like it?" When the user answers, for example, "6 points" to such an item, the user's answer can be interpreted such that the sum of the positive answer and the negative answer, such as "like" 6 points and "not like" 4 points, is 10 points, which is the full score. The question items may be, for example, "I want to let a close friend listen to it", "I want to listen to it live", "I want to play it myself", "I felt nostalgic", "I felt gloomy", "I don't want to hear it again", etc.

[0016] The user listens to music using an information device such as a smartphone or a personal computer and answers the question items displayed on the screen of the same smartphone. A plurality of question items may be displayed for one piece of music. The user may also fill in the answers to the questions on a paper answer sheet. In addition, the user's answers are collected by any method.

[0017] The user's answer to each question is an example of the reaction of the user who listened to the music. In addition, biometric data such as the amount of increase in the user's heart rate may be used for the user's reaction. Based on the user's reaction to one piece of music, one target data vector yn is generated.

[0018] The user may answer questions in real time while listening to music. For example, for one piece of music, the user may answer "like" for the first 8 seconds and "dislike" for the next 10 seconds. Based on such a time-series change in the user's answers, the target data vector yn may be generated.

[0019] For each piece of music, the temporal change of the feature quantity is calculated. The feature quantity at time t is expressed as a function using time t as a variable, such as the first feature quantity f(t) and the second feature quantity g(t). t is, for example, the elapsed time after the start of the music. In the case of a music piece with singing, t may be the elapsed time after the start of the singing. Details of the feature quantity will be described later. Note that the feature quantity is not limited to two types, namely the first feature quantity f(t) and the second feature quantity g(t). One type or three or more types of feature quantities may be used.

[0020] For each feature quantity, the first-order derivative and the second-order derivative of the variable t can be calculated. In the following description, the first-order derivatives of the first feature quantity f(t) and the second feature quantity g(t) are denoted as f'(t) and g'(t). Similarly, the second-order derivatives of the first feature quantity f(t) and the second feature quantity g(t) are denoted as f"(t) and g"(t). f'(t), g'(t), f"(t) and g"(t) are all functions using time t as a variable. Note that derivatives of the third order or higher may be used.

[0021] The feature quantity, its first-order derivative, and its second-order derivative can all be expressed as a sequence of values for each time t. In the following description, the sequence of values of the feature quantity for each time t may be described as time-series feature quantity data. The time-series feature quantity data is data indicating a sequence of numerical values related to the time-series change of the feature quantity.

[0022] Similarly, the sequence of values of the first-order derivative of the feature quantity for each time t may be described as first-order derivative feature quantity data, and the sequence of values of the second-order derivative of the feature quantity for each time t may be described as second-order derivative feature quantity data. The first-order derivative feature quantity data and the second-order derivative feature quantity data can be calculated, for example, by numerically differentiating the time-series feature quantity data.

[0023] These multiple numerical sequences are arranged in a single column to generate an explanatory data vector xn. One explanatory data vector xn is a vector indicating the feature amount of one piece of music.

[0024] In the present embodiment, it is regarded as a kind of black box model that receives an input of music expressed by the explanatory data vector xn from a user and outputs a reaction expressed by the target data vector yn. By regarding it in this way, XAI technology can be utilized for the reaction of a user who has listened to music.

[0025] The lower half of FIG. 1 shows an example in the case of using AIME disclosed in Non-Patent Document 1 in order to explain the reaction of a user. An explanatory matrix X is generated based on the explanatory data vectors xn regarding a plurality of pieces of music listened to by the user. A target matrix Y is generated based on the target data vectors yn representing the reaction of the user to each piece of music. An interpretation matrix A† shown in Equation (1) is calculated based on the explanatory matrix X and the target matrix Y. Details of the explanatory matrix X, the target matrix Y, and the interpretation matrix A† will be described later. X = A†Y ‥‥‥ (1)

[0026] By utilizing the interpretation matrix A†, the reaction of the user, which is a black box model, can be explained from various viewpoints.

[0027] FIG. 2 is an explanatory diagram for explaining the temporal change of feature amounts. FIG. 2A schematically shows the sound pressure data of music. The horizontal axis of FIG. 2A is time t. The vertical axis of FIG. 2A is the sound pressure of music. Note that since FIG. 2A is a schematic diagram, the numerical values and units of the scales are omitted. The sound pressure data of music is an example of the original data that is time-series data.

[0028] Figure 2B schematically shows the temporal change of the feature amount of the music. A window 51 with a short time width is set for the sound pressure data shown in Figure 2A. The feature amount is calculated for each window 51. By sequentially moving the window 51 along the horizontal axis, the temporal change of the feature amount can be calculated as shown in Figure 2B. The horizontal axis of Figure 2B is the same time t as the horizontal axis of Figure 2A. The vertical axis of Figure 2B is the feature amount.

[0029] The feature amount is, for example, the tempo of the music, that is, BPM (Beats Per Minute). The feature amount may be the dynamics of the music. The dynamics are defined, for example, by the maximum signal level or the average signal level in the window 51. The feature amount may be defined as the peak frequency in the window 51. In addition, any defined feature amount can be used. In this embodiment, a case where a plurality of feature amounts f(t) and g(t) are calculated at the same time interval will be described as an example.

[0030] Figure 3 is an explanatory diagram for explaining the configuration of the information processing apparatus. The information processing apparatus 10 includes a control unit 11, a main storage device 12, an auxiliary storage device 13, a communication unit 14, a display unit 15, an input unit 16, a speaker 17, a reading unit 19, and a bus.

[0031] The control unit 11 is an arithmetic control device that executes the program of this embodiment. One or more CPUs (Central Processing Unit), GPUs (Graphics Processing Unit), TPUs (Tensor Processing Unit), or multi-core CPUs, etc. are used for the control unit 11. The control unit 11 is connected to each hardware part constituting the information processing apparatus 10 via a bus.

[0032] The main memory device 12 is a storage device such as SRAM (Static Random Access Memory), DRAM (Dynamic Random Access Memory), or flash memory. In the main memory device 12, information necessary during the processing performed by the control unit 11 and the program being executed by the control unit 11 are temporarily stored.

[0033] The auxiliary storage device 13 is a storage device such as SRAM, flash memory, hard disk, or magnetic tape. In the auxiliary storage device 13, a music DB (Database) 36, a program to be executed by the control unit 11, and various data necessary for the execution of the program are stored. The music DB 36 may be stored in an external storage device connected via a network.

[0034] The communication unit 14 is an interface that performs communication between the information processing device 10 and the network. The display unit 15 is, for example, a liquid crystal display device or an organic EL (Electro Luminescence) display device. The input unit 16 is, for example, an input device such as a keyboard, mouse, trackball, or microphone. The speaker 17 is used for playing music. The speaker 17 may be headphones.

[0035] The portable recording medium 96 is, for example, a USB (Universal Serial Bus) memory, CD-ROM (Compact Disc Read only memory), magneto-optical disk medium, other optical disk medium, or SD memory card, etc. The portable recording medium 96 stores a program 97 for realizing the processing of the present embodiment.

[0036] The reading unit 19 is an interface such as a USB connector, CD-ROM drive, or SD memory reader that can connect the portable recording medium 96. The semiconductor memory 98 stores the program 97 and is a memory that can be installed inside the information processing device 10.

[0037] The information processing apparatus 10 is a general-purpose personal computer, tablet, mainframe computer, virtual machine operating on a mainframe computer, or quantum computer. The information processing apparatus 10 may be composed of a plurality of personal computers performing distributed processing, or hardware such as a mainframe computer. The information processing apparatus 10 may be composed of a cloud computing system. The information processing apparatus 10 may be composed of a plurality of personal computers operating in cooperation, or hardware such as a mainframe computer.

[0038] The program 97 is recorded on the portable recording medium 96. The control unit 11 reads the program 97 via the reading unit 19 and stores it in the auxiliary storage device 13. Further, the control unit 11 may read out the program 97 stored in the semiconductor memory 98. Furthermore, the control unit 11 may download the program 97 from another server computer (not shown) connected via the communication unit 14 and a network (not shown) and store it in the auxiliary storage device 13.

[0039] The program 97 is installed as a control program for the information processing apparatus 10, loaded into the main memory device 12, and executed. The program 97 of the present embodiment is an example of a program product.

[0040] FIG. 4 is an explanatory diagram for explaining the record layout of the music DB 36. The music DB 36 is a database that records music data, feature amounts, first-order differentials of the feature amounts, second-order differentials of the feature amounts, and the reactions of users who listened to the music in association with each other. The music DB 36 has a music data field, a music analysis field, and a reaction field.

[0041] The music analysis field has an arbitrary number of sub-fields such as a first analysis field and a second analysis field. In the following description, the first analysis field, the second analysis field, etc. may be referred to as analysis fields. Each analysis field has a feature amount field, a first-order differential field, and a second-order differential field.

[0042] The reaction field has sub-fields such as the Q1 field and the Q2 field, which correspond to the reactions of the user to be recorded. The music database 36 has one record for each piece of music that the user appreciates.

[0043] The music data field stores music data recorded in an arbitrary format such as WAV (Waveform Audio File Format). The feature quantity sub-field of the first analysis field stores a sequence of numbers indicating the first feature quantity f(t). The first-order differential field of the first analysis field stores a sequence of numbers obtained by differentiating the first feature quantity f(t) by the first order. The second-order differential field of the first analysis field stores a sequence of numbers obtained by differentiating the first feature quantity f(t) by the second order. The same applies to the second analysis field and subsequent fields. Note that the feature quantity data obtained by analyzing the temporal change of the music data, which is time-series data, and its time differential data are both time-series data.

[0044] The reaction field stores the reactions of the users who have appreciated the music. For example, for a binary-choice question such as "like" or "dislike", the reaction field stores numerical values such as "1" for "like" and "0" for "dislike". In the following description, it is assumed that the number of music pieces is N, that is, the number of records stored in the music database 36 is N, and the number of reactions recorded for each piece of music is M.

[0045] Note that the user may be a single person or a plurality of persons having common attributes. When the user is a plurality of persons, the reactions of the plurality of users are respectively recorded in one music database 36. The common attributes can be set to any attributes, such as "men in their 30s", "third-year high school students living in Chofu City", or "persons in their 40s who like *** (artist name)". These attributes can be converted by known methods and expressed as numerical values or vectors. In the following description, numerical values or vectors representing various attributes or combinations of various attributes related to the user are referred to as attribute feature quantities.

[0046] When the user is a plurality of persons, reactions may be collected from each of the plurality of users for the same piece of music. In such a case, the N records recorded in the music DB 36 include a plurality of records in which the same data is recorded in the music data field and the feature quantity field.

[0047] FIG. 5 is an explanatory diagram for explaining the explanatory data vector xn. In the following description, a case where two types of feature quantities, the first feature quantity f(t) and the second feature quantity g(t), are used will be described as an example. As described above, in the present embodiment, the first feature quantity f(t) and the second feature quantity g(t) are calculated at the same time interval. Further, the first feature quantity f(t) and the second feature quantity g(t) have the same number of elements.

[0048] That is, the first feature quantity f(t) and the second feature quantity g(t) are calculated for, for example, a portion cut out from the beginning of a piece of music for a predetermined time such as 30 seconds or 1 minute. The first feature quantity f(t) and the second feature quantity g(t) may be calculated for, for example, a portion cut out for a predetermined time from the vicinity of the so-called refrain in the music.

[0049] Details of the n-th record in the music DB 36 are shown on the left side of FIG. 5. That is, the table on the left side of FIG. 5 shows the details of part V in FIG. 4. The n-th first feature quantity fn(t) is a sequence of numbers "fn(t1), fn(t2), fn(t3) ······". Here, fn(t1) indicates the value of the first feature quantity fn(t) at time t1. Similarly, the first-order differential f'(t) of the first feature quantity fn(t) is a sequence of numbers "fn'(t1), fn'(t2), fn'(t3) ······", and the second-order differential fn"(t) is a sequence of numbers "fn"(t1), fn"(t2), fn"(t3) ······".

[0050] Similarly, the n-th second feature quantity gn(t) is a sequence of "gn(t1), gn(t2), gn(t3) ······". Here, gn(t1) represents the value of the second feature quantity gn(t) at time t1. Similarly, the first-order derivative g'(t) of the second feature quantity gn(t) is a sequence of "gn'(t1), gn'(t2), gn'(t3) ······", and the second-order derivative gn"(t) is a sequence of "gn"(t1), gn"(t2), gn"(t3) ······".

[0051] When calculating the first feature quantity f(t) and the second feature quantity g(t) respectively for L time points, the data obtained by analyzing one piece of music is represented by a two-dimensional matrix of L rows and 6 columns as shown in FIG. 5. One column of this two-dimensional matrix is denoted as the data vector Dn at time tn. The data vector D1(t1) at time t1 has six elements: fn(t1), fn'(t1), fn"(t1), gn(t1), gn'(t1), gn"(t1).

[0052] Incidentally, if sequences up to the second-order derivative of each of the k types of feature quantities are used, the data obtained by analyzing one piece of music is represented by a two-dimensional matrix of L rows and 3k columns. If sequences up to the m-th order derivative of each of the k types of feature quantities are used, the data obtained by analyzing one piece of music is represented by a two-dimensional matrix of L rows and (m + 1)k columns. The order of derivative up to which the sequence is used may be different for each feature quantity.

[0053] The explanatory data vector xn is defined as a vector obtained by arranging the data vectors at each time from t1 to tL in a vertical column. That is, the explanatory data vector xn in this embodiment is a vector having 6L elements.

[0054] Returning to FIG. 4, the description will be continued. The first feature quantity f(t) recorded in the feature quantity field of the first analysis field is an example of the first time-series feature quantity data related to the time-series change of the first feature quantity f(t), which is obtained from the original data that is time-series data. The second feature quantity g(t) recorded in the feature quantity field of the second analysis field is an example of the second time-series feature quantity data related to the time-series change of the second feature quantity g(t), which is obtained from the original data that is time-series data by a method different from the first feature quantity f(t).

[0055] By arranging a plurality of numerical sequences recorded in the music analysis field of one record of the music DB 36 in a row, an explanatory data vector xn is generated. By arranging the data recorded in the reaction field of the same record in a row, a target data vector yn is created. Therefore, in the music DB 36, the explanatory data vector xn created based on the music data and the target data vector yn representing the reaction of the user who listened to the music are recorded in association with each other.

[0056] FIG. 6 is an explanatory diagram for explaining a method of calculating the interpretation matrix A†. As described above, based on the reaction of the user to one piece of music, one target data vector yn is generated. In the following description, it will be described by taking as an example the case where M items of reactions are collected for one piece of music, that is, the case where one target data vector yn has M elements from Ob1n to ObMn. In FIG. 6, the explanatory data vector xn and the target data vector yn where n = 2 are shown surrounded by a broken line. As described above, the user is a kind of black box model that receives the input of the music represented by the explanatory data vector xn and outputs the reaction represented by the target data vector yn.

[0057] By arranging N explanatory database vectors xn for an N - curve piecewise linear curve in the row direction, an explanatory matrix X, which is a two - dimensional matrix, is created. The explanatory matrix X is a two - dimensional matrix with 6L rows and N columns. Similarly, by arranging N target database vectors yn in the row direction, a target matrix Y, which is a two - dimensional matrix, is generated. Here, the array order of the explanatory database vectors xn and the array order of the corresponding target database vectors yn are the same. That is, the target database vectors yn are arranged in the same order as the corresponding explanatory database vectors xn. The target matrix Y is a two - dimensional matrix with M rows and N columns.

[0058] Note that for the subsequent processes to be executed, the target database vectors yn need to be linearly independent. That is, the vector product of the target matrix Y and the transposed matrix of the target matrix Y needs to be a regular matrix.

[0059] Regarding the target matrix Y, the Moore - Penrose generalized inverse matrix Y† of Y is calculated. In the following explanations, Y† may be described as the target inverse matrix Y†. The target inverse matrix Y† is calculated by equation (2). The target inverse matrix Y† is a two - dimensional matrix with N rows and M columns. Y† = Y T (YY T ) -1 ‥‥‥ (2)

[0060] The interpretation matrix A† is the vector product of the explanatory matrix X and the target inverse matrix Y†. The equation for calculating the interpretation matrix A† is shown in equation (3). A† = XY† ‥‥‥ (3)

[0061] The interpretation matrix A† is a two - dimensional matrix with 6L rows and M columns, that is, the same number of rows as the number of elements of the explanatory database vectors xn and the same number of columns as the number of elements of the target database vectors yn. The element in the a - th row and b - th column of the interpretation matrix A† indicates the influence of the a - th element of the explanatory database vector xn on the b - th element of the target database vector yn.

[0062] As described above, the interpretation matrix A† can be generated based on the explanatory data vector xn calculated for each piece of music and the target data vector yn indicating the reaction of the user who listened to each piece of music.

[0063] For reference, the outline of the formula transformation for deriving Equation (3) from Equations (1) and (2) is shown below. First, multiply both sides of Equation (1) by the transpose matrix of the target matrix Y from the right to obtain Equation (4). XY T = A†YY T ‥‥‥ (4)

[0064] As described above, since the vector product of the target matrix Y and the transpose matrix of the target matrix Y is a regular matrix, the inverse matrix can be calculated. Multiply both sides of Equation (5) by this inverse matrix from the right to obtain Equation (5). XY T (YY T ) -1 = A†(YY T )(YY T ) -1 = A† ‥‥‥ (5)

[0065] After swapping the left and right sides of Equation (5) and then substituting Equation (2) into the right side, Equation (6) is obtained. Equation (3) is derived from both ends of Equation (6). A† = XY T (YY T ) -1 = XY† ‥‥‥ (6)

[0066] Figure 7 is an explanatory diagram for explaining the configuration of the interpretation matrix A†. In the following explanation, the case where the target data vector yn of two elements, "like" and "not like", is generated based on the user's answer to the question "Do you like this piece of music?" will be described as an example. The target data vector yn for the nth piece of music has two elements, Obn1 and Obn2. For example, when the user answers "like", Obn1 is "1" and Obn2 is "0". When the user answers "not like", Obn1 is "0" and Obn2 is "1".

[0067] When accepting intermediate answers such as "sort of like" for a question, Obn1 and Obn2 may be defined, for example, as "0.8" and "0.2". The sum of Obn1 and Obn2 is 1.

[0068] Since the number of elements M of the target data vector yn is 2, as shown on the left side of FIG. 7, the interpretation matrix A† is a matrix with 6L rows and 2 columns. The element in the a-th row and 1st column of the interpretation matrix A† indicates the influence of the a-th element of the explanatory data vector xn on the 1st element of the target data vector yn, that is, the answer "like". Similarly, the element in the a-th row and 2nd column of the interpretation matrix A† indicates the influence of the a-th element of the explanatory data vector xn on the 2nd element of the target data vector yn, that is, the answer "not like".

[0069] By grouping the elements of the interpretation matrix A† every 6 rows starting from the 1st row, a partial interpretation matrix Ap† regarding the first feature quantity f(t) can be generated as shown in the upper right of FIG. 7. Similarly, by grouping the elements of the interpretation matrix A† every 6 rows starting from the 2nd row, a partial interpretation matrix Ap† regarding the first-order derivative f'(t) of the first feature quantity f(t) can be created as shown in the lower right of FIG. 7. Similarly, partial interpretation matrices Ap† regarding the second-order derivative f"(t) of the first feature quantity f(t), the second feature quantity g(t), the first-order derivative g'(t) of the second feature quantity g(t), and the second-order derivative g"(t) of the second feature quantity g(t) can also be created. The partial interpretation matrices Ap† exemplified above are examples of the interpretation matrix A† related to specific feature quantities.

[0070] The data in the 1st column of the partial interpretation matrix Ap† regarding the first feature quantity f(t) indicates the influence of the first feature quantity f(t) at each time from time t1 to time tL on the answer "like". The data in the 2nd column of the partial interpretation matrix Ap† regarding the first feature quantity f(t) indicates the influence of the first feature quantity f(t) at each time from time t1 to time tL on the answer "not like". The influence of each element of the explanatory data vector xn on each element constituting the target data vector yn is called the global feature importance.

[0071] FIG. 8 is a flowchart for explaining the processing flow of the program. Prior to the execution of FIG. 8, music data to be appreciated by the user is recorded in the music data field of the music DB 36. The feature amounts of each piece of music are recorded in the feature amount field of the music DB 36. The combination of the music data and the feature amounts is previously recorded in a mass storage device connected via a network. No data is recorded in the reaction field of the music DB 36.

[0072] The control unit 11 may calculate the temporal change of the feature amount and its numerical differentiation based on the music data recorded in the music data field, and record them in the feature amount field. Since the method for calculating the temporal change of the feature amount and the method for calculating the numerical differentiation are well-known, the details thereof will be omitted.

[0073] The control unit 11 extracts one record from the music DB 36. The control unit 11 plays the music data recorded in the music data field (step S501). The music is output from the speaker 17. The control unit 11 presents a predetermined question to the user via the display unit 15, and acquires the user's answer via 16 (step S502). The control unit 11 records the answer acquired in step S502 in the reaction field of the extracted record (step S503).

[0074] The control unit 11 determines whether or not the processing of all the music data recorded in the music DB 36 has been completed (step S504). If it is determined that the processing has not been completed (NO in step S504), the control unit 11 returns to step S501. If it is determined that the processing has been completed (YES in step S504), the control unit 11 generates an explanatory matrix X based on the data recorded in the feature amount field of the music DB 36 (step S505). The control unit 11 generates a target matrix Y based on the data recorded in the reaction field of the music DB 36 (step S506).

[0075] The control unit 11 calculates a target inverse matrix Y†, which is the Moore-Penrose generalized inverse matrix, based on the target matrix Y (step S507). The control unit 11 calculates an interpretation matrix A†, which is the vector product of the explanatory matrix X and the target inverse matrix Y† (step S508). The control unit 11 ends the process.

[0076] [Usage Example] A usage example of the interpretation matrix A† calculated by the procedure described above will be described. In this specific example, the music recorded in the music DB 36 is forty Chopin piano solo pieces performed by various performers. Note that data regarding the same piece performed by different performers is treated as separate pieces. The data recorded in the feature field corresponds to the first thirty seconds of each piece. A single user with intermediate-level piano performance ability listened to the first thirty seconds of each piece and answered either "like" or "not like".

[0077] The first feature f(t) is the tempo graph of the music. The second feature g(t) is the dynamics of the music. The time-series data of each feature includes 2,582 data points from t1 to t2582.

[0078] As described above, the number N of the explanatory data vectors xn and the target data vectors yn in FIG. 6 is 40. The number L of the data of the time-series data f(t) and g(t) of the features is 2,582, the number of elements of the explanatory data vector xn is (2,582×6), and the number of elements of the target data vector yn is 2.

[0079] The explanatory matrix X is a two-dimensional matrix with (2,582×6) rows and 2 columns. As described with reference to FIG. 7, based on the explanatory matrix X, a partial interpretation matrix Ap† regarding the first feature f(t), a partial interpretation matrix Ap† regarding the first-order derivative f'(t) of the first feature f(t), a partial interpretation matrix Ap† regarding the second-order derivative f"(t) of the first feature f(t), and a partial interpretation matrix Ap† regarding the second feature g(t) can be created.

[0080] An example will be used to explain the partial interpretation matrix Ap† for the first feature quantity f(t). The data in the first column of this partial interpretation matrix Ap† indicates the influence of the first feature quantity f(t) at each time from time t1 to time tL for the response of "like". Similarly, the data in the second column indicates the influence of the first feature quantity f(t) at each time from time t1 to time tL for the response of "not like".

[0081] Figure 9 is an example of a global feature importance graph for the first feature quantity f(t). The horizontal axis represents time. For example, "1000" means time t1000. The vertical axis represents the value of the global feature importance for the first feature quantity f(t). The solid line represents the sequence of values in the first column of the components of the partial interpretation matrix Ap† for the first feature quantity f(1), that is, the influence of the first feature quantity f(t) for the response of "like". The dashed line represents the sequence of values in the second column of the partial interpretation matrix Ap†, that is, the influence of the first feature quantity f(t) for the response of "not like". Both the solid line and the dashed line represent the global feature importance, which is the influence of each element of the explanatory database vector xn on each element of the target database vector yn.

[0082] According to Figure 9, on average, the influence of the first feature quantity f(t) for the response of "like" is about 11 percent higher than the influence of the first feature quantity f(t) for the response of "dislike". This suggests that the tempo graph may affect the preference for music.

[0083] Figure 10 is an example of a global feature importance graph for the first-order derivative f'(t) of the first feature quantity f(t). The horizontal axis in Figure 10 indicates the time number. The vertical axis indicates the value of the global feature importance regarding the first-order derivative f'(t). The solid line represents the first column sequence of the component values of the partial interpretation matrix Ap† regarding the first-order derivative f'(t), that is, the influence of the first-order derivative f'(t) on the answer of "like". The dashed line represents the second column sequence of the partial interpretation matrix Ap† regarding the first-order derivative f'(t), that is, the influence of the first-order derivative f'(t) on the answer of "not like". Both the solid line and the dashed line indicate the global feature importance, which is the influence of each element of the explanatory database vector xn on each element of the target database vector yn.

[0084] As described above, the partial interpretation matrix Ap† is an exemplification of the interpretation matrix A† related to a specific feature quantity, and Figures 9 and 10 graphically representing the partial interpretation matrix Ap† are exemplifications of charts regarding the interpretation matrix A†.

[0085] Table 1 shows the statistics from t1 to t2582 regarding the data shown in Figure 10. Table 1 is an exemplification of a chart regarding the interpretation matrix A†.

[0086]

Table 1

[0087] According to Table 1, the value of "like" is larger than that of "not like". This data indicates that the change speed of the tempo graph may affect the preference for music. The large change range of the "not like" graph suggests that music with a rapidly fluctuating tempo graph may not be very popular.

[0088] Similarly, for the second-order derivative f"(t) of the first feature quantity f(t), the second feature quantity g(t), the first-order derivative g'(t) of the second feature quantity g(t), and the second-order derivative g"(t) of the second feature quantity g(t), the global feature quantities can also be charted.

[0089] Note that the above tempoogram and dynamics are examples of feature quantities related to music. By investigating various feature quantities, it may be possible to find feature quantities that can clearly distinguish between "liked" and "not liked".

[0090] As described above, the sound pressure data of a music piece is an example of the original data that is time-series data. The original data is not limited to sound pressure data. For example, the chord progression of a music piece may be used as the original data. For example, data recording a video, speech, reading, etc. may be used as the original data.

[0091] According to this embodiment, an information processing method for assisting the interpretation of the reactions of people who have listened to a music piece can be provided. For example, by comparing the interpretation matrices A† calculated for "people in their 20s" and "people in their 60s" respectively, it may be possible to interpret what parts of the music piece the preference for the music piece by age is due to. Here, the preference for the music piece includes not only the preference for the music piece itself expressed by the composer on the score, but also the preference for the performance expressed by the performer.

[0092] If it is possible to interpret what parts of a music piece are preferred by people with specific attribute feature quantities, it may be possible to provide a music piece that is more preferred by people with the same attribute feature quantities by producing a music piece that emphasizes the said parts. Here, the production of a music piece includes newly composing a music piece that emphasizes the preferred parts, arranging an existing music piece so as to emphasize the preferred parts, and performing so as to emphasize the preferred parts.

[0093] Similarly, if it is possible to interpret what parts are not preferred by people with specific attribute feature quantities, it may be possible to grasp the characteristics of a music piece that people with the said attribute feature quantities do not like. Furthermore, it is possible to infer whether an existing music piece is preferred by people with specific attribute feature quantities.

[0094] [Embodiment 2] This embodiment relates to a form in which respective time-series data are arranged continuously to create an explanatory database vector xn. For parts common to Embodiment 1, the description is omitted.

[0095] FIG. 11 is an explanatory diagram for explaining the explanatory database vector xn of Embodiment 2. The table in the upper left of FIG. 11 is the same as the table in the upper left of FIG. 5. The explanatory database vector xn of this embodiment is formed by arranging a sequence of numbers in the order of the first feature quantity f(t), the first-order derivative f'(t) of the first feature quantity f(t), the second-order derivative f"(t) of the first feature quantity f(t), the second feature quantity g(t), the first-order derivative g'(t) of the second feature quantity g(t), and the second-order derivative g"(t) of the second feature quantity g(t).

[0096] FIG. 12 is an explanatory diagram for explaining the configuration of the interpretation matrix A† of Embodiment 2. The interpretation matrix A† in FIG. 12 is created by the procedure described with reference to FIG. 6 using the explanatory database vector xn shown in FIG. 11. Since the parts related to the respective time-series data are continuous, partial interpretation matrices Ap† related to the first feature quantity f(t), partial interpretation matrices Ap† related to the first-order derivative of the first feature quantity f(t), etc. can be easily extracted.

[0097] According to this embodiment, for example, even when the number of data of the first feature quantity f(t) and the number of data of the second feature quantity g(t) are different, the interpretation matrix A† can be created. For example, the width of window 51 when calculating the first feature quantity f(t) and the width of window 51 when calculating the second feature quantity g(t) may be different. The length of the part cut out from the music when calculating the first feature quantity f(t) and the length of the part cut out from the music when calculating the second feature quantity g(t) may be different.

[0098] [Embodiment 3] This embodiment relates to a form using feature quantities in the frequency domain. For parts common to Embodiment 2, the description is omitted.

[0099] FIG. 13 is an explanatory diagram for explaining the explanatory data vector xn of Embodiment 3. The table in the upper left of FIG. 13 is the same as the left half of the table shown in the upper left of FIG. 11. That is, the first feature amount f(t) is the same as that in Embodiment 1.

[0100] The table in the lower left of FIG. 13 shows the feature amounts in the frequency domain. The second feature amount g(ω) represents, for example, the power spectrum of music data. The first derivative g'(ω) and the second derivative g"(ω) of the second feature amount g(ω) represent the frequency derivative of the second feature amount g(ω). The second feature amount g(ω) is an example of the frequency analysis data of the original data that is time-series data.

[0101] Similar to Embodiment 2, the explanatory data vector xn of the present embodiment is formed by arranging in sequence the first feature amount f(t), the first derivative f'(t) of the first feature amount f(t), the second derivative f"(t) of the first feature amount f(t), the second feature amount g(ω), the first derivative g'(ω) of the second feature amount g(ω), and the second derivative g"(ω) of the second feature amount g(ω).

[0102] According to the present embodiment, by using the second feature amount g(ω) which is in the frequency domain, an interpretation matrix A† that reflects periodic features that are difficult to extract only from time-axis data can be created.

[0103] Note that the explanatory data vector xn may be formed by arranging the first feature amount f(t), the first derivative f'(t) of the first feature amount f(t), the second derivative f"(t) of the first feature amount f(t), and the second feature amount g(ω) in a column. In such a case, it is not necessary to calculate the first derivative g'(ω) and the second derivative g"(ω) of the second feature amount g(ω).

[0104] The program is an example of a program product. The program may be provided on a recording medium or may be in a form distributed from an external computer. A computer program can be deployed to execute on a single computer, or on a single site, or be distributed across multiple sites and executed on multiple computers interconnected by a communication network.

[0105] The technical features (constituent elements) described in each embodiment can be combined with each other, and by combining them, new technical features can be formed. The embodiments disclosed this time should be considered illustrative in all respects and not restrictive. The scope of the present invention is indicated not by the above description, but by the claims, and is intended to include all modifications within the meaning and scope equivalent to the claims.

[0106] The independent claims and dependent claims described in the claims can be combined with each other in any combination regardless of the citation form. Furthermore, the claims use a form (multi-claim form) of describing claims that cite two or more other claims, but are not limited to this. It may be described using a form of describing a multi-claim (multi-multi-claim) that cites at least one multi-claim.

Explanation of Reference Numerals

[0107] 10 Information processing apparatus 11 Control unit 12 Main memory device 13 Auxiliary storage device 14 Communication unit 15 Display unit 16 Input unit 17 Speaker 19 Reading unit 36 Music DB 96 Portable recording medium 97 Program 98 Semiconductor memory

Claims

1. a plurality of sets of time-series feature amount data acquired from original data which is time-series data, explanation data vectors including first-order differential feature amount data of the time-series feature amount data and second-order differential feature amount data of the time-series feature amount data, and a target data vector representing a human response to the original data are recorded in association with each other; Calculate an interpretation matrix which is a vector product of an explanation matrix in which a plurality of sets of the explanation data vectors are arranged and a generalized inverse matrix of a target matrix in which the target data vectors are arranged in an order corresponding to the explanation data vectors; Output a chart related to the interpretation matrix An information processing method in which processing is performed by a computer.

2. The original data is music data. The information processing method according to claim 1 .

3. The objective matrix is ​​composed of objective data vectors that represent the responses of one person. The information processing method according to claim 1 .

4. The objective matrix is ​​composed of objective data vectors representing the responses of multiple people who have a common attribute. The information processing method according to claim 1 .

5. The time-series feature amount data is First time-series feature data acquired from the original data; and second time-series feature amount data obtained from the original data by a method different from that of the first time-series feature amount data. The information processing method according to claim 1 .

6. the original data is music data, the first time-series feature data is a time-series change in a tempogram; The second time-series feature data is a time-series change in dynamics. The information processing method according to claim 5.

7. The explanatory data vector is In addition to the time series feature amount data, the first-order differential feature amount data, and the second-order differential feature amount data, frequency analysis data of the time series data is included. The information processing method according to claim 1 .

8. a plurality of sets of time-series feature amount data acquired from original data which is time-series data, explanation data vectors including first-order differential feature amount data of the time-series feature amount data and second-order differential feature amount data of the time-series feature amount data, and a target data vector representing a human response to the original data are recorded in association with each other; Calculate an interpretation matrix which is a vector product of an explanation matrix in which a plurality of sets of the explanation data vectors are arranged and a generalized inverse matrix of a target matrix in which the target data vectors are arranged in an order corresponding to the explanation data vectors; Output a chart related to the interpretation matrix A program that causes a computer to carry out processing.

9. An information processing device including a control unit, The control unit is a plurality of sets of time-series feature amount data acquired from original data which is time-series data, explanation data vectors including first-order differential feature amount data of the time-series feature amount data and second-order differential feature amount data of the time-series feature amount data, and a target data vector representing a human response to the original data are recorded in association with each other; Calculate an interpretation matrix which is a vector product of an explanation matrix in which a plurality of sets of the explanation data vectors are arranged and a generalized inverse matrix of a target matrix in which the target data vectors are arranged in an order corresponding to the explanation data vectors; Output a chart related to the interpretation matrix Information processing device.

Citation Information

Patent Citations

  • Medical image processing apparatus, endoscope system, medical image processing system, method of operating medical image processing apparatus, program, and storage medium

    JP2023083555A