Information processing method, information processing device, and program

The method addresses the issue of unsuitable music selection by using periodicity and timbre features to align music with user biometrics, ensuring the induced psychological state matches the user's needs.

WO2026009785A1PCT designated stage Publication Date: 2026-01-08PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/022816
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-01
Filing Date
2025-06-25
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing music selection technologies do not utilize features related to the periodicity and timbre of music, leading to songs that may not be suitable for improving a user's biometric information, making it difficult to induce a desired psychological state.

Method used

An information processing method that selects songs based on features related to the periodicity and timbre of each song, using recurrence plots and Mel-frequency cepstrum coefficients, along with biometric information such as heart rate, to output music suited to the user's recent biometric information and induce a predetermined psychological state.

Benefits of technology

Enables the output of music that aligns with the user's recent biometric information, effectively inducing the desired psychological state by considering unique characteristics of each song, including periodicity, timbre, and loudness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025022816_08012026_PF_FP_ABST
    Figure JP2025022816_08012026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing method includes: acquiring a plurality of music pieces and first biological information that pertains to a user for a most-recent prescribed period; selecting one or more music pieces from among the plurality of music pieces on the basis of the first biological information and a plurality of first acoustic feature quantities indicating features of each of the plurality of music pieces; and outputting the selected one or more music pieces. The plurality of first acoustic feature quantities include a feature quantity relating to the periodicity of each music piece and a feature quantity relating to the timbre of each music piece.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing device, and program

[0001] The present disclosure relates to a technology for outputting music suited to biometric information of a user.

[0002] Patent Document 1 discloses selecting music to be provided to a user based on information about the bars or beats of the music so as to improve the user's biometric information.

[0003] The feature quantities related to the periodicity and timbre of a song can uniquely identify the song. However, the technology disclosed in Patent Literature 1 does not use the feature quantities related to the periodicity and timbre when selecting songs to provide to a user. As a result, songs may be provided that have bars or beats that improve the user's biometric information, but whose periodicity or timbre is not suitable for improving the user's biometric information.

[0004] Patent No. 7402577

[0005] An object of the present disclosure is to provide a technology that can output music that is suited to the user's most recent biometric information.

[0006] An information processing method in one aspect of the present disclosure is an information processing method in a computer, which includes acquiring a plurality of songs and first biometric information, which is a user's biometric information for a recent predetermined time period; selecting one or more songs from the plurality of songs based on a plurality of first acoustic features indicating characteristics of each of the plurality of songs and the first biometric information; and outputting the one or more selected songs, wherein the plurality of first acoustic features include a feature related to the periodicity of each song and a feature related to the timbre of each song.

[0007] 1 is a block diagram showing the configuration of an information processing system in an embodiment; FIG. 2 is a diagram showing an example of an image showing a recurrence plot of a song; FIG. 3 is a diagram showing an example of an image showing the Mel frequency cepstral coefficients of a song; FIG. 4 is a diagram showing an example of an image showing the dynamic range of a song; FIG. 5 is a diagram showing an example of an image showing a CVRR; FIG. 6 is a diagram showing an example of a questionnaire screen regarding subjective evaluation of songs; FIG. 7 is a diagram showing an example of a user's psychological state; FIG. 8 is a flowchart showing a learning process for a song selection model; FIG. 9 is a flowchart showing an automatic song selection process; FIG. 10 is a diagram showing an example of an image showing a recurrence plot of a part in a song; FIG. 11 is a diagram showing an example of an image showing the chroma vectors of a part in a song.

[0008] (Background to one aspect of the present disclosure) In recent years, the present inventors have been studying techniques for inducing a user's psychological state to a predetermined psychological state based on the user's biometric information and information about music. In the process, the present inventors have focused on the fact that feature quantities related to the periodicity and timbre of music can uniquely identify music, and have conducted extensive research. The term "timbre" as used herein refers to the texture or characteristics of a sound, and this also applies to the following description.

[0009] As described above, the technology disclosed in Patent Document 1 does not use features related to periodicity and timbre when selecting music to provide to a user. As a result, music may be provided that has periodicity or timbre that is not suitable for improving the user's biometric information. In this case, it is difficult to induce the user's psychological state to a desired psychological state by improving the user's biometric information.

[0010] Therefore, the inventors have conducted extensive research into technology that can output music that is appropriate for the user's most recent biometric information and that can also induce the user's psychological state into a predetermined psychological state, and have come up with the present disclosure described below.

[0011] In order to solve the above problems, the following techniques are disclosed.

[0012] (1) An information processing method in one aspect of the present disclosure is an information processing method in a computer, which includes acquiring a plurality of songs and first biometric information, which is a user's biometric information for a predetermined period of time in the most recent past; selecting one or more songs from the plurality of songs based on a plurality of first acoustic features indicating characteristics of each of the plurality of songs and the first biometric information; and outputting the one or more selected songs, wherein the plurality of first acoustic features include a feature related to the periodicity of each song and a feature related to the timbre of each song.

[0013] In this configuration, the plurality of first acoustic features indicating the characteristics of each of the plurality of songs include a feature related to the periodicity of each song and a feature related to the timbre of each song, which can uniquely identify each song. One or more songs are then selected based on the plurality of first acoustic features and first biometric information, which is biometric information of the user over a predetermined period of time. This makes it possible to output songs that are suited to the user's most recent biometric information, taking into account the unique characteristics of each song.

[0014] (2) In the information processing method described in (1) above, the feature related to periodicity may be information indicating a recurrence plot of each piece of music, and the feature related to timbre may be information indicating a Mel-frequency cepstrum coefficient of each piece of music.

[0015] In this case, by taking into consideration the recurrence plot and Mel-frequency cepstrum coefficients of each of the plurality of songs, it is possible to output a song that is suitable for the user's most recent biological information.

[0016] (3) In the information processing method described in (1) or (2) above, the plurality of first acoustic features may further include a feature related to loudness of each piece of music.

[0017] In this case, the feature amount relating to the loudness of each piece of music is further taken into consideration, and music suited to the most recent biometric information of the user can be output.

[0018] (4) In the information processing method described in (3) above, the feature amount related to loudness may be information indicating a dynamic range of each piece of music.

[0019] In this case, it is possible to take into consideration the dynamic range of each piece of music and output music that is suited to the user's most recent biological information.

[0020] (5) In the information processing method described in any one of (1) to (4) above, the biological information may include information based on the heart rate of the user.

[0021] In this case, it is possible to output music that is suited to the user's most recent heart rate, taking into account the unique characteristics of each piece of music.

[0022] (6) In the information processing method described in (5) above, the plurality of first acoustic features may further include features related to the tempo of each piece of music, and outputting the one or more pieces of music may include determining an order in which to output each piece of music based on the information based on the heart rate and the features related to the tempo, and outputting the one or more pieces of music in the determined order.

[0023] In this case, the tempo-related feature of each piece of music is further taken into consideration, and one or more pieces of music can be output in an order suited to the user's most recent heart rate. As a result, it is possible to output pieces of music suited to the user's most recent heart rate.

[0024] (7) In the information processing method described in any one of (1) to (6) above, the selection of the one or more songs may include estimating a psychological state that each song induces in the user based on the plurality of first acoustic features and the first biometric information, and selecting the one or more songs based on the result of the estimation.

[0025] In this case, one or more songs are selected based on the psychological state that each song induces in the user, which is estimated based on the plurality of first acoustic features and the first biometric information. Therefore, by outputting the one or more songs, the user's psychological state can be induced to the estimated psychological state.

[0026] (8) In the information processing method described in (7) above, the selection of the one or more songs may include selecting songs that are estimated to induce a predetermined psychological state in the user.

[0027] In this case, one or more pieces of music that are estimated to induce a predetermined psychological state in the user are selected, and by outputting the one or more pieces of music, the user's psychological state can be induced to the predetermined psychological state.

[0028] (9) In the information processing method described in (7) or (8) above, the estimation of the psychological state may include estimating whether the psychological state induced in the user by each piece of music is a calm, pleasant state, an awakened, pleasant state, or an unpleasant state.

[0029] In this case, one or more pieces of music are selected based on whether the psychological state induced in the user by each piece of music is a calm pleasant state, an awakened pleasant state, or an unpleasant state. Therefore, by outputting the one or more pieces of music, the user's psychological state can be induced to any one of an awakened pleasant state, a calm pleasant state, or an unpleasant state.

[0030] (10) In the information processing method described in any one of (7) to (9) above, the selection of the one or more songs may include selecting the one or more songs using a learning model that has learned the relationship between the plurality of first acoustic features of the songs, second biometric information that is the biometric information of the user for the predetermined period of time immediately before listening to the song, and information indicating the psychological state induced in the user by the song.

[0031] In this case, the learning model can be used to take into account the biometric information of the user immediately before listening to each piece of music, and music that is suitable for the user's most recent biometric information can be output.

[0032] (11) In the information processing method described in any one of (1) to (9) above, the selection of the one or more pieces of music may further include selecting the one or more pieces of music from the plurality of pieces of music based on a plurality of second acoustic features that indicate characteristics of parts that are the main melody or the tonic chord that constitute each of the plurality of pieces of music.

[0033] In this case, a plurality of second acoustic features indicating the characteristics of the main melody or tonic chord that constitute each of the plurality of pieces of music are further taken into consideration, and music suited to the user's most recent biometric information can be output.

[0034] (12) In the information processing method described in (11) above, the selection of the one or more songs may include selecting one or more songs from the plurality of songs using a learning model that has learned the relationship between the plurality of first acoustic features of the songs, second biometric information that is the user's biometric information for the predetermined time immediately before listening to the song, information indicating the psychological state induced in the user by the song, and the plurality of second acoustic features of the parts that make up the song.

[0035] In this case, the learning model can be used to take into account the biometric information of the user immediately before listening to each piece of music, and music that is suitable for the user's most recent biometric information can be output.

[0036] (13) In the information processing method described in (11) or (12) above, the plurality of second acoustic features may include information indicating a recurrence plot of the part in each piece of music.

[0037] In this case, a recurrence plot of the main melody or tonic chord that constitutes each of the multiple pieces of music is taken into consideration, and a piece of music that is suited to the user's most recent biological information can be output.

[0038] (14) In the information processing method described in (11) or (12) above, the plurality of second acoustic features may include a feature related to the intensity of sound for each scale of the part in each piece of music.

[0039] In this case, a feature amount relating to the intensity of each note of the main melody or tonic chord that constitutes each of the multiple pieces of music is taken into consideration, and music suited to the user's most recent biometric information can be output.

[0040] (15) In the information processing method described in (14) above, the feature amount relating to the intensity of the sound for each scale of the part may be information indicating a chroma vector of the part.

[0041] In this case, the chroma vectors of the main melodies or tonic chords that make up each of the multiple pieces of music are taken into consideration, and music that is suited to the user's most recent biometric information can be output.

[0042] (16) In the information processing method described in (10) above, the plurality of first acoustic features may include features represented by an image.

[0043] In this case, the first acoustic feature represented by the image can be used for training and inputting the training model.

[0044] (17) In the information processing method described in (12) above, at least one of the plurality of first acoustic features and the plurality of second acoustic features may include a feature represented by an image.

[0045] In this case, at least one of the first acoustic feature represented by an image and the second acoustic feature represented by an image can be used for training and input of the training model.

[0046] (18) In the information processing method described in any one of (1) to (17) above, the plurality of musical pieces may include a first musical piece that can be divided into a plurality of musical piece elements, the plurality of first acoustic features of the first musical piece include a plurality of third acoustic features that indicate characteristics of each of the plurality of musical piece elements, and the plurality of third acoustic features include a feature related to the periodicity of each musical piece element and a feature related to the timbre of each musical piece element, and the selection of the one or more musical pieces may further include selecting one or more musical piece elements from the plurality of musical piece elements as the one or more musical pieces based on the plurality of third acoustic features and the first biometric information.

[0047] In this case, it is possible to take into consideration the periodicity and timbre-related features specific to each musical element and output musical element suitable for the user's most recent biometric information.

[0048] The present disclosure can be realized not only as an information processing method that executes the characteristic processes described above, but also as an information processing device or the like that has a characteristic configuration corresponding to the characteristic processes executed by the information processing method. Furthermore, the present disclosure can also be realized as a computer program that causes a computer to execute the characteristic processes included in such an information processing method. Therefore, the same effects as those of the above information processing method can also be achieved in the following other aspects.

[0049] (19) In another aspect of the present disclosure, an information processing device includes an acquisition unit that acquires a plurality of songs and first biometric information, which is a user's biometric information for a predetermined period of time in the most recent past; a selection unit that selects one or more songs from the plurality of songs based on a plurality of first acoustic features that indicate characteristics of each of the plurality of songs and the first biometric information; and an output unit that outputs the one or more songs selected by the selection unit, wherein the plurality of first acoustic features include a feature related to the periodicity of each song and a feature related to the timbre of each song.

[0050] (20) In yet another aspect of the present disclosure, a program is a program for an information processing device, which causes the information processing device to function as an acquisition unit that acquires a plurality of songs and first biometric information, which is a user's biometric information for a specified period of time in the most recent past; a selection unit that selects one or more songs from the plurality of songs based on a plurality of first acoustic features that indicate characteristics of each of the plurality of songs and the first biometric information; and an output unit that outputs the one or more songs selected by the selection unit, wherein the plurality of first acoustic features include a feature related to the periodicity of each song and a feature related to the timbre of each song.

[0051] The present disclosure can also be realized as an information processing system operated by such an information processing program. Needless to say, such a computer program can be distributed on a computer-readable non-transitory recording medium such as a CD-ROM or via a communication network such as the Internet.

[0052] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in an independent claim that represents a superordinate concept will be described as optional components. Furthermore, in all embodiments, the respective contents can be combined.

[0053] 1 is a block diagram showing the configuration of an information processing system 100 according to an embodiment. The information processing system 100 is a system that induces a user's psychological state to a predetermined psychological state by outputting one or more pieces of music from among a plurality of pieces of music that are suited to the user's most recent biometric information. The information processing system 100 includes an information terminal 2, a speaker 3, and an information processing device 1.

[0054] The information terminal 2 is configured by an information processing device such as a personal computer or tablet device placed in a space where the user is present, or a mobile device such as a smartphone carried by the user. The space is, for example, a living room, a study or other workroom in the user's home, or an office room. The space is not limited to this, and may also be a room in a public facility such as a library, a coworking space, or a shared office. The information terminal 2 is connected to the information processing device 1 via a short-range wireless communication channel such as Bluetooth (registered trademark). The information terminal 2 transmits various instructions input by user operations to the information processing device 1.

[0055] The speaker 3 is connected to the information processing device 1 via a short-range wireless communication channel such as Bluetooth (registered trademark). The speaker 3 is placed in a space where the user is present, and outputs (plays) music indicated by sound source data received from the information processing device 1. Note that the speaker 3 may be an earphone that is worn on the user's ear and can be connected to the information processing device 1 via a short-range wireless communication channel.

[0056] The information processing device 1 is configured by a computer such as a cloud server or an edge server, etc. The information processing device 1 includes an interface unit 11, a communication unit 18, a processor 10, and a memory 17.

[0057] The interface unit 11 is a communication interface that connects the information processing device 1 to a short-range wireless communication channel such as Bluetooth (registered trademark). The interface unit 11 outputs information received from the information terminal 2 via the short-range wireless communication channel to the processor 10. The interface unit 11 transmits information input from the processor 10 to the information terminal 2 via the short-range wireless communication channel.

[0058] When the information processing device 1 is configured as an edge server, the interface unit 11 may include an interface circuit to which peripheral devices such as a keyboard, a mouse, a display, and a memory device are connected. In this case, when a peripheral device is connected to the interface circuit, the interface unit 11 inputs and outputs various information between the peripheral device and the processor 10 via the interface circuit.

[0059] The communication unit 18 is a communication interface that connects the information processing device 1 to a network such as the Internet or a local area network. The communication unit 18 transmits and receives various information to and from external devices via the network under the control of the processor 10. The external devices include the biometric server 4.

[0060] The biometric server 4 is configured with a computer such as a cloud server or an edge server. The biometric server 4 manages the user's biometric information periodically transmitted from various sensors used by the user. The biometric information includes information based on heartbeats. The heartbeat-based information includes the heart rate. The biometric information may further include body temperature, blood pressure, brain waves, etc.

[0061] The processor 10 is configured, for example, by a central processing unit (CPU). The processor 10 includes an acquisition unit 12, a calculation unit 13, an estimation unit 14, a configuration unit 15, and an output unit 16. The acquisition unit 12, the calculation unit 13, the estimation unit 14, the configuration unit 15, and the output unit 16 may be realized by the processor 10 executing an information processing program, or may be configured by a dedicated hardware circuit such as an ASIC. The information processing program may be recorded on a non-transitory computer-readable recording medium. Details of the acquisition unit 12, the calculation unit 13, the estimation unit 14, the configuration unit 15, and the output unit 16 will be described later.

[0062] The memory 17 is configured as a non-volatile rewritable storage device such as a hard disk drive or a solid state drive, and includes a customer information storage unit 171, a music information storage unit 172, a music selection model storage unit 173, and a sound source storage unit 174.

[0063] The customer information storage unit 171 stores information about users of the information processing system 100 (hereinafter referred to as customer information). The customer information includes a user ID for identifying the user and user attribute information. The attribute information includes the user's age, gender, and address. The attribute information also includes the user's client ID and token used when the user accesses the biometric server 4. The user's client ID is an ID for identifying the user on the biometric server 4. The token is authentication information used to authenticate that the user is a user who can access the biometric server 4.

[0064] The music information storage unit 172 stores information about each of a plurality of music pieces (hereinafter referred to as music information). The music information is also called meta information. The music information includes information indicating the genre of each music piece. The music information also includes a plurality of feature quantities indicating the characteristics of each music piece (hereinafter referred to as a plurality of first acoustic feature quantities). The plurality of first acoustic feature quantities include a feature quantity related to the periodicity of each music piece and a feature quantity related to the timbre of each music piece.

[0065] Specifically, the feature quantities related to the periodicity of a song include information indicating a recurrence plot of the song. The information indicating the recurrence plot of the song is information that visualizes the periodicity of the song in a two-dimensional image with the time axis as the second axis using a recurrence plot, which is one of the techniques of nonlinear time series analysis. Fig. 2 is a diagram showing an example of an image indicating the recurrence plot of the song. The information indicating the recurrence plot of the song may be a two-dimensional image that visualizes the periodicity of the song as shown in Fig. 2, or may be two-dimensional array data indicating the image.

[0066] The feature quantities related to the timbre of a song include information indicating the Mel frequency cepstral coefficients of the song. The information indicating the Mel frequency cepstral coefficients of the song is information obtained by quantifying the timbre of the song based on the Mel frequency scale using cepstrum, which is a signal analysis processing technique. The Mel frequency scale is a frequency scale designed based on the characteristics of human hearing. FIG. 3 is a diagram showing an example of an image indicating the Mel frequency cepstral coefficients of a song. The information indicating the Mel frequency cepstral coefficients of a song may be an image indicating time-series changes in each Mel frequency component of the song, as shown in FIG. 3, or may be array data of numerical values ​​indicating the time-series changes. A song can be uniquely identified by information indicating a recurrence plot and information indicating the Mel frequency cepstral coefficients.

[0067] The plurality of first acoustic features further includes a feature related to the loudness of each song. Specifically, the feature related to the loudness of each song includes information indicating the dynamic range of the song. The information indicating the dynamic range of the song is information that expresses the average range of the loudness (amplitude) of the sound of the song using RMS (Root Mean Square). FIG. 4 is a diagram showing an example of an image indicating the dynamic range of the song. The information indicating the dynamic range may be an image indicating a time-series change in the RMS of the song, or may be array data of numerical values ​​indicating the time-series change.

[0068] The plurality of first acoustic features further includes a feature related to the tempo of each piece of music. Specifically, the feature related to the tempo of each piece of music includes information indicating BPM (Beats Per Minute).

[0069] The music selection model storage unit 173 stores a music selection model (learning model) that has been machine-learned to determine the relationship between a plurality of first acoustic features of a song, biometric information (second biometric information) for a predetermined period of time (e.g., 24 hours) immediately before the user listens to the song, and information indicating the psychological state induced in the user by the song. When the music selection model receives the plurality of first acoustic features of a song and the biometric information for the predetermined period of time immediately before the user listens to the song, it outputs an estimation result of the psychological state induced in the user by the song. The music selection model is constructed by the processor 10.

[0070] The biological information used in the machine learning of the music selection model includes information indicating CVRR (Coefficient of Variation of R-R Intervals). CVRR is an index for evaluating the activity level of the autonomic nervous system, obtained by analyzing the heart rate fluctuation pattern. The heart rate fluctuation pattern indicates the temporal change in the interval between heartbeats (R-R interval). FIG. 5 is a diagram showing an example of an image indicating CVRR. The information indicating CVRR may be an image indicating the time series change in CVRR, or may be array data of numerical values ​​indicating the time series change.

[0071] The information indicating the psychological state induced in the user by a song, which is used for machine learning of the song selection model, is obtained by conducting a questionnaire survey of the user regarding the subjective evaluation of the song.

[0072] Specifically, the processor 10 obtains sound source data of a song for learning from the sound source storage unit 174 and transmits the sound source data to the speaker 3 using the interface unit 11. As a result, the speaker 3 outputs the song indicated by the sound source data received from the information processing device 1. The processor 10 obtains image data indicating a questionnaire screen from the memory 17 and transmits the image data to the information terminal 2 using the interface unit 11. As a result, the information terminal 2 displays the questionnaire screen indicated by the image data received from the information processing device 1 on a display provided in the information terminal 2.

[0073] An example of a questionnaire screen will now be described. FIG. 6 is a diagram showing an example of a questionnaire screen regarding subjective evaluation of music. The questionnaire screen includes a question Q1 regarding a pleasant or unpleasant psychological state, a question Q2 regarding an alert or calm psychological state, and one question Q3 regarding preferences. On the information terminal 2, for the first question Q1, the user performs an operation to select a number corresponding to the degree of pleasant or unpleasant psychological state when listening to music from nine numbers ranging from "-4" to "4" indicating the degree of pleasant or unpleasant psychological state. In this way, the user answers the first question Q1.

[0074] Here, among the nine numerical values, the negative numerical values ​​"-1" to "-4" indicate that the larger the absolute value, the greater the degree of unpleasant psychological state. In other words, "-4" indicates the highest degree of unpleasant psychological state. The positive numerical values ​​"1" to "4" indicate that the larger the absolute value, the greater the degree of pleasant psychological state. In other words, "4" indicates the highest degree of pleasant psychological state. "0" indicates a psychological state that is neither pleasant nor unpleasant.

[0075] In response to the second question Q2, the user selects a number from nine numbers ranging from "-4" to "4" that indicate the degree of alertness or calmness of the psychological state when listening to the music, thereby answering the second question Q2.

[0076] Here, among the nine numerical values, the negative numerical values ​​"-1" to "-4" indicate a greater degree of sedation as the absolute value increases. In other words, "-4" indicates the highest degree of sedation. The positive numerical values ​​"1" to "4" indicate a greater degree of awakening as the absolute value increases. In other words, "4" indicates the highest degree of awakening. "0" indicates a psychological state that is neither sedated nor awakening.

[0077] In response to the third question Q3, the user selects a number from nine numbers ranging from "-4" to "4" that indicate the degree of preference, which number corresponds to the degree of preference for the song. In this way, the user answers the third question Q3.

[0078] Here, among the nine numerical values, the negative numerical values ​​"-1" to "-4" indicate that the larger the absolute value, the greater the degree to which the user dislikes the song. In other words, "-4" indicates the greatest degree to which the user dislikes the song. The positive numerical values ​​"1" to "4" indicate that the larger the absolute value, the greater the degree to which the user likes the song. In other words, "4" indicates the greatest degree to which the user likes the song.

[0079] When the user performs a predetermined operation to end the questionnaire, the information terminal 2 transmits information including the numerical values ​​of the user's answers to the three questions Q1, Q2, and Q3 (hereinafter, referred to as evaluation information) to the information processing device 1. Note that the questionnaire screen shown in Fig. 6 is merely an example. The questionnaire screen may be appropriately configured to allow the user to input information indicating the psychological state and degree of preference when listening to the music.

[0080] In the information processing device 1, when the interface unit 11 receives evaluation information, the processor 10 acquires numerical values ​​indicating the psychological state and the degree of preference included in the evaluation information. Based on the acquired numerical values ​​indicating the psychological state and the degree of preference, the processor 10 identifies whether the psychological state induced in the user by the music is a calm pleasant state (hereinafter referred to as the first state), an awakened pleasant state (hereinafter referred to as the second state), or an unpleasant state (hereinafter referred to as the third state).

[0081] 7 is a diagram showing an example of a user's psychological state. Specifically, the processor 10 defines a pleasant and calmed region as a first state, a pleasant and pleasant region as a second state, and an unpleasant region as a third state in a two-dimensional coordinate system having two axes, similar to Russell's circular diagram, an arousal / calm axis indicating the activity level of the user's psychological state and a pleasant / unpleasant axis indicating the degree of comfort or discomfort of the user's psychological state.

[0082] If the numerical value indicating the degree of the pleasant or unpleasant psychological state is a negative numerical value, the processor 10 determines that the psychological state induced in the user by the music is the third state. If the numerical value indicating the degree of the pleasant or unpleasant psychological state is 0 or greater and the numerical value indicating the degree of the sedated or awakened psychological state is a negative numerical value, the processor 10 determines that the psychological state induced in the user by the music is the first state. If the numerical value indicating the degree of the pleasant or unpleasant psychological state is 0 or greater and the numerical value indicating the degree of the sedated or awakened psychological state is 0 or greater, the processor 10 determines that the psychological state induced in the user by the music is the second state.

[0083] The processor 10 may be configured to determine that the psychological state induced in the user by the music is the third state if the numerical value indicating the degree of the pleasant or unpleasant psychological state is equal to or less than 0. The processor 10 may be configured to determine that the psychological state induced in the user by the music is the first state if the numerical value indicating the degree of the pleasant or unpleasant psychological state is a positive numerical value and the numerical value indicating the degree of the sedated or awakened psychological state is equal to or less than 0.

[0084] The sound source storage unit 174 stores sound source data of a plurality of songs. The plurality of songs includes training songs used in machine learning of the song selection model and listening songs that the user is permitted to listen to.

[0085] When the interface unit 11 or the communication unit 18 receives information indicating a registration request for the songs to be learned and the songs to be listened to, along with the sound source data of the songs, the processor 10 (acquisition unit) acquires the sound source data received by the interface unit 11 or the communication unit 18. The processor 10 stores the acquired sound source data in the sound source storage unit 174. The processor 10 calculates a plurality of first acoustic feature quantities for the songs by performing acoustic analysis processing on the sound source data of the songs. The processor 10 stores the calculated plurality of first acoustic feature quantities for the songs in the song information storage unit 172.

[0086] Next, the learning process of the music selection model by the processor 10 will be described. Fig. 8 is a flowchart showing the learning process of the music selection model. The learning process of the music selection model is executed by the processor 10 for each of a plurality of music pieces for training stored in the music selection model storage unit 173 at times when it is expected that the user's biometric information will change significantly, such as once in the morning and once in the evening. For example, if 54 music pieces for training are stored in the music selection model storage unit 173, the learning process shown in Fig. 8 is executed 54 times for each of the 54 music pieces at times when it is expected that the user's biometric information will change significantly.

[0087] In step S21, the processor 10 acquires, from the song information storage unit 172, song information for one learning song to be processed (hereinafter, the target song).

[0088] In step S22, the processor 10 acquires biometric information (second biometric information) of the user to be used for training the music selection model.

[0089] Specifically, the processor 10 references the user's attribute information stored in the customer information storage unit 171 and acquires a client ID and a token used by the user when accessing the biometric server 4. The processor 10 controls the communication unit 18 to transmit, to the biometric server 4, information requesting the transmission of the user's biometric information for a predetermined period of time (e.g., 24 hours) together with the client ID and the token (hereinafter, "request information"). In response, when the communication unit 18 receives the user's biometric information for the predetermined period of time from the biometric server 4, the processor 10 calculates the CVRR by analyzing the heart rate fluctuation pattern included in the biometric information. The processor 10 acquires information indicating the calculated CVRR as the user's biometric information to be used for training the music selection model. Note that the processor 10 may execute the process of step S22 only in the first learning process of the learning process, which is performed the same number of times as the number of songs to be trained.

[0090] In step S23, the processor 10 causes the speaker 3 to output the target song. Specifically, the processor 10 acquires sound source data of the target song from the sound source storage unit 174. The processor 10 transmits the sound source data to the speaker 3 using the interface unit 11. As a result, the speaker 3 outputs the song indicated by the sound source data received from the information processing device 1.

[0091] In step S24, the processor 10 acquires evaluation information about the target song from the information terminal 2. Specifically, the processor 10 acquires image data representing a questionnaire screen ( FIG. 6 ) from the memory 17 and transmits the image data to the information terminal 2 using the interface unit 11. The information terminal 2 then displays the questionnaire screen represented by the image data received from the information processing device 1 on its display. The user uses the questionnaire screen to answer a questionnaire about the subjective evaluation of the target song and performs a predetermined operation to end the questionnaire. The information terminal 2 transmits evaluation information about the target song, including numerical values ​​representing the psychological state and preferences answered by the user, to the information processing device 1. As a result, when the interface unit 11 receives the evaluation information about the target song from the information terminal 2, the processor 10 acquires the evaluation information. The processor 10 stores the acquired evaluation information in the memory 17.

[0092] In step S25, based on the evaluation information acquired in step S24, the processor 10 determines whether the psychological state induced in the user by the target song is the first state (a calm, pleasant state), the second state (an awakened, pleasant state), or the third state (an unpleasant state).

[0093] In step S26, if no song selection model is stored in the song selection model storage unit 173, the processor 10 generates a song selection model. The song selection model is a learning model that learns, by machine learning, the relationship between the multiple first acoustic features of the target song included in the song information acquired in step S21, the information indicating the CVRR acquired in step S22, and the psychological state induced in the user by the target song identified in step S25. As the machine learning algorithm, a CNN (Convolutional Neural Network) or the like can be adopted. If a song selection model is stored in the song selection model storage unit 173, the processor 10 causes the song selection model to learn the above relationship by machine learning.

[0094] When the process of step S26 ends, the processor 10 ends the learning process for the target song, and executes the learning process shown in FIG. 8 with another learning song as the target song.

[0095] Next, the acquisition unit 12, the calculation unit 13, the estimation unit 14, the configuration unit 15, and the output unit 16 will be described in detail.

[0096] When the interface unit 11 receives instruction information from the information terminal 2 to automatically select and output a song along with a user ID, the acquisition unit 12 acquires biometric information (first biometric information) for the most recent specified period (e.g., 24 hours) of the user corresponding to the user ID from the biometric server 4.

[0097] Specifically, similar to the processing of step S22 ( FIG. 8 ) by the processor 10, the acquisition unit 12 acquires, from the customer information storage unit 171, a client ID and a token used by the user when accessing the biometric server 4. The acquisition unit 12 transmits, together with the client ID and the token, request information requesting transmission of the user's biometric information for the most recent predetermined time period to the biometric server 4. In response to this, when the communication unit 18 receives the user's biometric information for the most recent predetermined time period from the biometric server 4, the acquisition unit 12 acquires the biometric information.

[0098] Furthermore, when the interface unit 11 acquires preference information received from the information terminal 2, the acquisition unit 12 stores the preference information in the memory 17. The preference information includes information indicating a preferred genre of music. Specifically, when the user performs an operation on the information terminal 2 to input information indicating a preferred genre of music, the information terminal 2 transmits the preference information including the information indicating the preferred genre of music to the information processing device 1.

[0099] In addition, the acquisition unit 12 may refer to information indicating the degree of preference contained in the evaluation information of each song for learning stored in the memory 17 and information indicating the genre contained in the song information of each song for learning stored in the song information storage unit 172, and acquire, as preference information, information indicating the genre with the largest sum of numerical values ​​indicating the degree of preference.

[0100] The calculation unit 13 calculates information about the user's heart rate (hereinafter, heart rate information) based on the user's biometric information for the most recent predetermined time period acquired by the acquisition unit 12. Specifically, the calculation unit 13 calculates the CVRR by analyzing the heart rate fluctuation pattern included in the user's biometric information for the most recent predetermined time period, similar to the processing of step S22 ( FIG. 8 ) by the processor 10. The calculation unit 13 outputs information indicating the calculated CVRR as heart rate information.

[0101] The estimation unit 14 estimates, for each of the multiple songs for listening stored in the sound source memory unit 174, whether the psychological state induced in the user by each song is a first state (a calm, pleasant state), a second state (an awakened, pleasant state), or a third state (an unpleasant state), based on the multiple first acoustic features of each song and the heart rate information calculated by the calculation unit 13.

[0102] Specifically, the estimation unit 14 acquires song information for each of a plurality of songs to be listened to from the song information storage unit 172. The estimation unit 14 inputs the plurality of first acoustic features of each song included in the song information and the heart rate information calculated by the calculation unit 13 to the song selection model stored in the song selection model storage unit 173. As a result, the estimation unit 14 outputs to the song selection model an estimation result as to which of the first state, the second state, and the third state the psychological state induced in the user by each song is.

[0103] The composition unit 15 selects one or more pieces of music to be output from the speaker 3 based on the result of estimation by the estimation unit 14 .

[0104] Specifically, the composition unit 15 adds the songs that the estimation unit 14 estimates to induce the first psychological state in the user to a first playlist. A playlist is a list of one or more songs to be output from the speaker 3. The first playlist is a playlist made up of one or more songs that are estimated to induce the first psychological state in the user.

[0105] The composition unit 15 adds to a second playlist the songs that the estimation unit 14 estimates to induce the second psychological state in the user. The second playlist is a playlist that is composed of one or more songs that are estimated to induce the second psychological state in the user.

[0106] If the estimation unit 14 estimates that the psychological state induced in the user by a song is the third state, the composition unit 15 does not add the song to either the first playlist or the second playlist. In other words, the composition unit 15 does not select the song as one or more songs to be output from the speaker 3.

[0107] The configuration unit 15 acquires the playlist information received by the interface unit 11 from the information terminal 2. The playlist information is information that specifies a playlist containing one or more pieces of music to be output from the speaker 3.

[0108] Specifically, when a user performs a predetermined operation on the information terminal 2 to designate a first playlist as a playlist containing one or more songs to be output from the speaker 3, the information terminal 2 transmits playlist information indicating the first playlist to the information processing device 1. On the other hand, when a user performs a predetermined operation on the information terminal 2 to designate a second playlist as a playlist containing one or more songs to be output from the speaker 3, the information terminal 2 transmits playlist information indicating the second playlist to the information processing device 1.

[0109] The composition unit 15 selects one or more songs included in the first playlist or the second playlist indicated by the playlist information as one or more songs to be output from the speaker 3. In other words, when the playlist information indicates the first playlist, the composition unit 15 selects one or more songs included in the first playlist that are estimated to induce a first state (a predetermined psychological state) in the user as one or more songs to be output from the speaker 3. When the playlist information indicates the second playlist, the composition unit 15 selects one or more songs included in the second playlist that are estimated to induce a second state (a predetermined psychological state) in the user as one or more songs to be output from the speaker 3.

[0110] When the interface unit 11 has not acquired playlist information received from the information terminal 2, the configuration unit 15 selects one or more songs included in the first playlist as one or more songs to be output from the speaker 3. However, this is not limiting, and when the interface unit 11 has not acquired playlist information received from the information terminal 2, the configuration unit 15 may select one or more songs included in the second playlist as one or more songs to be output from the speaker 3.

[0111] The output unit 16 determines an order in which to output each of the one or more pieces of music based on the heart rate included in the biometric information acquired by the acquisition unit 12 and the feature amounts related to the tempo of the one or more pieces of music selected by the composition unit 15. The output unit 16 outputs each of the one or more pieces of music in the determined order.

[0112] Specifically, the output unit 16 refers to the tempo feature values ​​of the one or more songs selected by the composition unit 15, which are included in the song information stored in the song information storage unit 172. As described above, the tempo feature values ​​include information indicating BPM. When the composition unit 15 selects one or more songs included in the first playlist, the output unit 16 arranges the one or more songs in descending order of BPM. When the composition unit 15 selects one or more songs included in the second playlist, the output unit 16 arranges the one or more songs in ascending order of BPM. The output unit 16 determines to output the one or more songs arranged in descending or ascending order of BPM in order starting with the song with the BPM closest to the heart rate included in the biometric information.

[0113] The output unit 16 acquires sound source data of one or more songs arranged in ascending or descending order of BPM from the sound source storage unit 174. The output unit 16 transmits the sound source data of each of the one or more songs to the speaker 3 in the determined order using the interface unit 11. As a result, the speaker 3 outputs each of the one or more songs indicated by the sound source data received from the information processing device 1 in the order determined by the output unit 16. This makes it possible to quickly induce the user's psychological state to the first state or the second state.

[0114] When the composition unit 15 selects one or more songs included in the first playlist, the output unit 16 may determine the order in which to output the one or more songs so that, among songs with a BPM higher than the heart rate, the songs are output in order starting with the song with the BPM closest to the heart rate. Also, when the composition unit 15 selects one or more songs included in the second playlist, the output unit 16 may determine the order in which to output the one or more songs so that, among songs with a BPM lower than the heart rate, the songs are output in order starting with the song with the BPM closest to the heart rate.

[0115] Next, the automatic music selection process will be described. Fig. 9 is a flowchart showing the automatic music selection process. The automatic music selection process is started by the processor 10 when the interface unit 11 receives, from the information terminal 2, instruction information to play recommended music along with the user ID.

[0116] In step S1 , the acquisition unit 12 acquires, from the biometric server 4 , biometric information for the most recent predetermined period (for example, 24 hours) of the user corresponding to the user ID received by the interface unit 11 .

[0117] In step S2, the calculation unit 13 calculates the CVRR based on the user's biological information for the most recent predetermined time period acquired in step S1. The calculation unit 13 outputs information indicating the calculated CVRR as heart rate information.

[0118] In step S3 , the estimation unit 14 acquires, from the music information storage unit 172 , music information for each of the multiple music pieces to be listened to that are stored in the sound source storage unit 174 .

[0119] In step S4, the estimation unit 14 estimates, for each of the multiple songs to be listened to, which of the first state (a calm, pleasant state), the second state (an awakened, pleasant state), or the third state (an unpleasant state) the psychological state that each song will induce in the user, based on the multiple first acoustic features of each song included in the song information acquired in step S3 and the heart rate information output in step S2.

[0120] In step S5 , the composition unit 15 selects one or more pieces of music to be output from the speaker 3 based on the result of estimation by the estimation unit 14 .

[0121] In step S6, the configuration unit 15 determines whether preference information is stored in the memory 17. If preference information is stored in the memory 17 (YES in step S6), the process proceeds to step S7. If it is determined that preference information is not stored in the memory 17 (NO in step S6), the process proceeds to step S8.

[0122] In step S7, the composition unit 15 limits the one or more songs selected in step S5 to be output from the speaker 3 to one or more songs in the genre indicated by the preference information. Specifically, the composition unit 15 refers to the information indicating the genre included in the song information acquired in step S3, and deletes songs that are different from the songs in the genre indicated by the preference information from the one or more songs selected in step S5 to be output from the speaker 3. When the processing of step S7 ends, the processing proceeds to step S8.

[0123] In step S8, the output unit 16 determines the order in which to output each of the one or more songs selected in step S5 or the one or more songs restricted in step S7, based on the heart rate included in the biometric information acquired in step S1 and the features related to the tempo of each song included in the song information acquired in step S3.

[0124] In step S9, the output unit 16 causes the speaker 3 to output each of the one or more pieces of music selected in step S5 or the one or more pieces of music restricted in step S7 in the order determined in step S8.

[0125] In the above embodiment, the plurality of first acoustic features indicating the characteristics of each of the plurality of songs include a feature related to the periodicity of each song and a feature related to the timbre of each song, which can uniquely identify each song. One or more songs are then selected based on the plurality of first acoustic features and the user's biometric information for a predetermined period of time. This makes it possible to output songs that are suited to the user's most recent biometric information, taking into account the unique characteristics of each song.

[0126] Furthermore, a first playlist and a second playlist are created based on the results of estimating the psychological state induced in the user by each song, and one or more songs included in the first playlist or the second playlist indicated by the playlist information and restricted to the genre indicated by the preference information are output. Therefore, when one or more songs included in the first playlist are output, the user's psychological state can be induced to the first state. On the other hand, when one or more songs included in the second playlist are output, the user's psychological state can be induced to the second state. Furthermore, since each of the one or more songs is output in an order determined based on the user's heart rate and the tempo-related features of each song, the user's psychological state can be smoothly induced to the first state or the second state.

[0127] The present disclosure can employ the following modifications.

[0128] (1) In the above embodiment, an example has been described in which information indicating CVRR is used as biometric information to be input during machine learning of the music selection model and use. However, the biometric information to be input during machine learning of the music selection model and use is not limited to information indicating CVRR. For example, in addition to heart rate, the biometric information may be a heart rate-related index such as LF (Low Frequency) / HF (High Frequency) or CCVTP (Coefficient of Component Variance of Total Power).

[0129] (2) The composition unit 15 may further select one or more songs from the plurality of songs based on a plurality of acoustic features (hereinafter, a plurality of second acoustic features) that indicate the characteristics of the parts that are the main melody or tonic chord that constitute each of the plurality of songs.

[0130] Specifically, the plurality of second acoustic features indicating characteristics of parts of a musical piece include information indicating a recurrence plot of the parts in the musical piece. The information indicating the recurrence plot of the parts in the musical piece is information in which the periodicity of the parts in the musical piece is visualized by the recurrence plot in a two-dimensional image with the time axis as the second axis. Fig. 10 is a diagram showing an example of an image indicative of the recurrence plot of the parts in the musical piece. The information indicative of the recurrence plot of the parts in the musical piece may be a two-dimensional image indicative of the periodicity of the parts in the musical piece as shown in Fig. 10, or may be two-dimensional array data indicative of the image.

[0131] The plurality of second acoustic features indicating the characteristics of a part of a musical piece further include a feature related to the sound intensity for each scale of the part in the musical piece. The feature related to the sound intensity for each scale of the part includes information indicating a chroma vector of the part. The information indicating the chroma vector of the part is information in which one octave is divided into 12 semitones and the intensity corresponding to each semitone of the part is expressed as a vector element. FIG. 11 is a diagram showing an example of an image indicating the chroma vector of a part in a musical piece. The information indicating the chroma vector of the part may be an image indicating a time-series change in intensity corresponding to each semitone of the part, as shown in FIG. 11, or may be array data of numerical values ​​indicating the time-series change.

[0132] The processor 10 constructs a song selection model by machine learning the relationship between multiple first acoustic features of a song, biometric information for a predetermined period of time immediately before the user listens to the song, information indicating the psychological state induced in the user by the song, and multiple second acoustic features of a predetermined part that makes up the song, and stores the constructed song selection model in the song selection model memory unit 173.

[0133] Specifically, when storing sound source data of a song in the sound source storage unit 174, the processor 10 calculates the above-mentioned multiple first acoustic features of the song and multiple second acoustic features indicating characteristics of predetermined parts that make up the song by executing acoustic analysis processing on the sound source data of the song. The processor 10 stores the calculated multiple first acoustic features of the song and multiple second acoustic features indicating characteristics of predetermined parts that make up the song in the song information storage unit 172.

[0134] 8 , if no song selection model is stored in the song selection model storage unit 173, the processor 10 generates a song selection model that has been machine-learned to determine the relationship between the multiple first acoustic features of the target song and the multiple second acoustic features of the parts in the target song included in the song information acquired in step S21, the information indicating the CVRR acquired in step S22, and the psychological state induced in the user by the target song identified in step S25. If a song selection model is stored in the song selection model storage unit 173, the processor 10 causes the song selection model to learn the above relationship by machine learning.

[0135] The estimation unit 14 estimates, for each of a plurality of songs for listening stored in the sound source storage unit 174, whether the psychological state induced in the user by each song is a first state (a calm, pleasant state), a second state (an awakened, pleasant state), or a third state (an unpleasant state), based on the plurality of first acoustic features of each song, the plurality of second acoustic features of the parts that make up the song, and the heart rate information calculated by the calculation unit 13.

[0136] Specifically, the estimation unit 14 acquires song information for each of a plurality of songs to be listened to from the song information storage unit 172. The estimation unit 14 inputs the plurality of first acoustic features of each song and the plurality of second acoustic features of the parts that make up the song, which are included in the song information, and the heart rate information calculated by the calculation unit 13, to the song selection model stored in the song selection model storage unit 173. As a result, the estimation unit 14 outputs to the song selection model an estimation result as to which of the first state, the second state, and the third state the psychological state induced in the user by each song is.

[0137] The composition unit 15 selects one or more pieces of music to be output from the speaker 3 based on the estimation results by the estimation unit 14, which are based on the multiple first acoustic features of each piece of music, the multiple second acoustic features of the parts that make up each piece of music, and the heart rate information calculated by the calculation unit 13. In this way, the composition unit 15 selects one or more pieces of music to be output from the speaker 3, based on the multiple first acoustic features of each piece of music, the multiple second acoustic features of the parts that make up each piece of music, and the heart rate information calculated by the calculation unit 13.

[0138] In comparison with the above embodiment, this modification further takes into account a plurality of second acoustic features that indicate the characteristics of the parts that make up each of a plurality of pieces of music, thereby making it possible to output music that is suited to the user's most recent biometric information.

[0139] (3) In the above embodiment, it has been described that one piece of music information for one piece of music is stored in the music information storage unit 172. However, it is also possible to divide a long piece of music, such as a symphony, into multiple music elements, and store one piece of music information for each music element in the music information storage unit 172.

[0140] As a result, the estimation unit 14 may estimate, for each music element of each of the multiple music pieces to be listened to stored in the sound source storage unit 174, which psychological state each music element induces in the user is one of the first state (a calm, pleasant state), the second state (an awakened, pleasant state), and the third state (an unpleasant state), based on multiple acoustic features indicating the characteristics of each music element (hereinafter, multiple third acoustic features) and the heart rate information calculated by the calculation unit 13. As a result, the composition unit 15 may select one or more music elements to be output from the speaker 3 based on the multiple third acoustic features of each of the multiple music elements of each music piece and the heart rate information calculated by the calculation unit 13.

[0141] Furthermore, the estimation unit 14 may estimate, for each music element of each of a plurality of music pieces to be listened to stored in the sound source storage unit 174, which psychological state each music element induces in the user, among a first state (a calm, pleasant state), a second state (an awakened, pleasant state), and a third state (an unpleasant state), based on a plurality of third acoustic features indicating the characteristics of each music element, a plurality of acoustic features indicating the characteristics of the parts constituting each music element (hereinafter, a plurality of fourth acoustic features), and the heart rate information calculated by the calculation unit 13. As a result, the composition unit 15 may select one or more music elements to be output from the speaker 3, based on the plurality of third acoustic features of each of the plurality of music elements of each music piece, the plurality of fourth acoustic features of the parts constituting each music element, and the heart rate information calculated by the calculation unit 13.

[0142] (4) As described above, the biometric information includes heart rate. Therefore, in step S9 of the automatic music selection process shown in FIG. 9 , the acquisition unit 12 may acquire the user's biometric information for the most recent predetermined time period, similar to step S1, each time the output unit 16 finishes outputting one song. The output unit 16 may then output a song from among one or more songs arranged in ascending or descending order of BPM that has a heart rate closest to the heart rate included in the biometric information. This makes it possible to omit the output of one or more songs depending on the user's current heart rate. The heart rate may change in response to the user's level of alertness (sedation). Therefore, the user's psychological state can be induced to the first or second state more quickly than when one or more songs arranged in ascending or descending order of BPM are output one by one.

[0143] In this case, the calculation unit 13 may further calculate a CCVTP, which is an index of tension, based on the heartbeat interval indicated by the heart rate included in the biological information. If the CCVTP value is higher than the previous calculation, i.e., if the user's psychological state is approaching an alert (tension) state, the output unit 16 may next output a piece of music whose heart rate is closest to the heart rate. Similarly, if the CCVTP value is lower than the previous calculation, i.e., if the user's psychological state is approaching a sedated (relaxed) state, the output unit 16 may next output a piece of music whose heart rate is closest to the heart rate.

[0144] In this case, it is possible to omit output of one or more pieces of music depending on the user's current heart rate and level of alertness (tension), which allows for more appropriate and prompt induction of the user's psychological state based on the actual changes in the user's psychological state.

[0145] (5) In the above embodiment, an example has been described in which the configuration unit 15 selects one or more songs to be output from the speaker 3 from the first playlist or the second playlist indicated by the playlist information received from the information terminal 2. However, instead of this, the configuration unit 15 may select a playlist from which to select one or more songs in accordance with the user's current psychological state estimated based on the user's biometric information. This configuration can be realized, for example, as follows.

[0146] In step S5, the calculation unit 13 calculates the CCVTP based on the heartbeat interval indicated by the heart rate included in the user's biometric information for the most recent specified time period acquired in step S1, in the same manner as in variant example (4).

[0147] If the CCVTP value calculated by the calculation unit 13 is higher than the previous calculation, the composition unit 15 estimates that the user's current psychological state is a state of arousal (tension). In this case, the composition unit 15 selects one or more songs from a second playlist consisting of one or more songs that induce a second psychological state (a state of arousal and pleasure) in the user. On the other hand, if the CCVTP value calculated by the calculation unit 13 is lower than the previous calculation, the composition unit 15 estimates that the user's current psychological state is a state of sedation (relaxation). In this case, the composition unit 15 selects one or more songs from a first playlist consisting of one or more songs that induce a first psychological state (a state of arousal and pleasure) in the user.

[0148] Conversely, if the construction unit 15 estimates that the user's current psychological state is a state of arousal (tension), it may select one or more songs from the first playlist, and if it estimates that the user's current psychological state is a state of sedation (relaxation), it may select one or more songs from the second playlist. Furthermore, the construction unit 15 may estimate the user's current psychological state based on the user's biometric information using a method other than the above.

[0149] (6) In the automatic music selection process shown in FIG. 9, steps S6 and S7 may be omitted.

[0150] (7) The processor 10 may not perform the learning process shown in Figure 8, but may store a music selection model constructed by a process similar to the learning process shown in Figure 8 in the music selection model storage unit 173 in a device other than the information processing device 1.

[0151] (8) The plurality of first acoustic features may not include a feature related to the tempo of each piece of music. In this case, the output unit 16 may omit step S8 ( FIG. 9 ) and randomly output, in step S9 ( FIG. 9 ), one or more pieces of music selected in step S5 ( FIG. 9 ) or one or more pieces of music restricted in step S7 ( FIG. 9 ).

[0152] (9) The plurality of second acoustic features may not include information indicating a recurrence plot of the parts in each piece of music or features relating to the intensity of the sound for each scale of the parts in each piece of music.

[0153] (10) The estimation unit 14 may use generative artificial intelligence (AI) to estimate whether the psychological state induced in the user by each piece of music is the first state, the second state, or the third state. This configuration can be realized, for example, as follows.

[0154] The processor 10 performs the same processes as steps S21 to S25 of the music selection model learning process ( FIG. 8 ). Then, the processor 10 creates a database of relationships between the plurality of first acoustic features of the target music piece contained in the music information acquired in the same manner as step S21, the information indicating the CVRR acquired in the same manner as step S22, and the psychological state induced in the user by the target music piece identified in the same manner as step S25.

[0155] Specifically, the processor 10 generates information (hereinafter, database information) that associates multiple first acoustic features of the target song included in the song information, information indicating the CVRR, and the psychological state induced in the user by the target song. The processor 10 stores the generated database information in the song selection model storage unit 173. However, without being limited to this, the processor 10 may store the database information in a computer such as a cloud server or an edge server different from the information processing device 1 via the communication unit 18.

[0156] The estimation unit 14 inputs input information including a plurality of first acoustic features of each song and the heart rate information calculated by the calculation unit 13 to the generation AI. The generation AI references the database information using Retrieval Augmented Generation (RAG) or the like, and estimates whether the psychological state induced in the user by each song is the first state, the second state, or the third state. Note that the generation AI may be implemented as a function of the processor 10, or may be implemented by an external computer accessible by the processor 10 via the communication unit 18.

[0157] Note that the configuration unit 15 may select one or more pieces of music to be output from the speaker 3 based on the result of estimation by the estimation unit 14, and then use the generation AI to restrict the pieces of music based on the preference information. Also, the output unit 16 may use the generation AI to determine the order in which to output each of the one or more pieces of music.

[0158] The present disclosure is useful in the field of services that provide music to users who periodically acquire biometric information.

Claims

1. An information processing method on a computer, comprising: acquiring a plurality of pieces of music and first biometric information, which is a user's biometric information for a specified period of time in the most recent past; selecting one or more pieces of music from the plurality of pieces of music based on a plurality of first acoustic features indicating characteristics of each of the plurality of pieces of music and the first biometric information; and outputting the one or more selected pieces of music, wherein the plurality of first acoustic features include a feature related to the periodicity of each piece of music and a feature related to the timbre of each piece of music.

2. The information processing method according to claim 1, wherein the feature quantity relating to periodicity is information indicating a recurrence plot of each piece of music, and the feature quantity relating to timbre is information indicating a Mel-frequency cepstrum coefficient of each piece of music.

3. The information processing method according to claim 1, wherein the plurality of first acoustic features further includes a feature related to loudness of each piece of music.

4. The information processing method according to claim 3, wherein the loudness-related feature is information indicating the dynamic range of each piece of music.

5. The information processing method according to claim 1, wherein the biological information includes information based on the user's heart rate.

6. The information processing method according to claim 5, wherein the plurality of first acoustic features further include features related to the tempo of each piece of music, and outputting the one or more pieces of music includes: determining an order in which to output each of the one or more pieces of music based on the information based on the heart rate and the features related to the tempo; and outputting each of the one or more pieces of music in the determined order.

7. The information processing method of claim 1, wherein the selection of the one or more songs includes: estimating the psychological state that each song will induce in the user based on the plurality of first acoustic features and the first biometric information; and selecting the one or more songs based on the result of the estimation.

8. The information processing method according to claim 7, wherein the selection of one or more songs includes selecting songs that are estimated to induce a predetermined psychological state in the user.

9. An information processing method according to claim 7 or 8, wherein the estimation of the psychological state includes estimating whether the psychological state induced in the user by each piece of music is a calm, pleasant state, an awakened, pleasant state, or an unpleasant state.

10. The information processing method of claim 7, wherein the selection of the one or more songs includes selecting the one or more songs using a learning model that has learned the relationship between the plurality of first acoustic features of the songs, second biometric information which is the biometric information of the user for the predetermined period of time immediately before listening to the song, and information indicating the psychological state induced in the user by the song.

11. The information processing method according to claim 1, wherein the selection of the one or more pieces of music further includes selecting the one or more pieces of music from the plurality of pieces of music based on a plurality of second acoustic features indicating characteristics of the main melody or tonic chord parts that constitute each of the plurality of pieces of music.

12. The information processing method of claim 11, wherein the selection of the one or more songs includes selecting one or more songs from the plurality of songs using a learning model that has learned the relationship between the plurality of first acoustic features of the songs, second biometric information which is the user's biometric information for the predetermined period of time immediately before listening to the song, information indicating the psychological state induced in the user by the song, and the plurality of second acoustic features of the parts that make up the song.

13. The information processing method according to claim 11 or 12, wherein the plurality of second acoustic features include information indicating a recurrence plot of the part in each piece of music.

14. The information processing method according to claim 11 or 12, wherein the plurality of second acoustic features include features relating to the intensity of sound for each scale of the part in each piece of music.

15. The information processing method according to claim 14, wherein the feature amount relating to the intensity of the sound for each scale of the part is information indicating the chroma vector of the part.

16. The information processing method according to claim 10, wherein the plurality of first acoustic features include features represented by images.

17. The information processing method according to claim 12, wherein at least one of the plurality of first acoustic features and the plurality of second acoustic features includes a feature represented by an image.

18. The information processing method of claim 1, wherein the plurality of musical pieces include a first musical piece that can be divided into a plurality of musical piece elements, the plurality of first acoustic features of the first musical piece include a plurality of third acoustic features that indicate characteristics of each of the plurality of musical piece elements, the plurality of third acoustic features including a feature related to the periodicity of each musical piece element and a feature related to the timbre of each musical piece element, and the selection of the one or more musical pieces further includes selecting one or more musical piece elements from the plurality of musical piece elements as the one or more musical pieces based on the plurality of third acoustic features and the first biometric information.

19. An information processing device comprising: an acquisition unit that acquires a plurality of songs and first biometric information, which is a user's biometric information for a specified period of time in the most recent past; a selection unit that selects one or more songs from the plurality of songs based on a plurality of first acoustic features that indicate characteristics of each of the plurality of songs and the first biometric information; and an output unit that outputs the one or more songs selected by the selection unit, wherein the plurality of first acoustic features include a feature related to the periodicity of each song and a feature related to the timbre of each song.

20. A program for an information processing device, which causes the information processing device to function as: an acquisition unit that acquires a plurality of pieces of music and first biometric information, which is a user's biometric information for a specified period of time in the most recent past; a selection unit that selects one or more pieces of music from the plurality of pieces of music based on a plurality of first acoustic features that indicate characteristics of each of the plurality of pieces of music and the first biometric information; and an output unit that outputs the one or more pieces of music selected by the selection unit, wherein the plurality of first acoustic features include a feature related to the periodicity of each piece of music and a feature related to the timbre of each piece of music.

Citation Information

Patent Citations

  • Method and device for reducing mental stress of woman in the menopause with music

    JP2000140117A

  • Music piece retrieval method, music piece retrieval data registration method, music piece retrieval device and music piece retrieval data registration device

    JP2002278547A

  • Contents reproducing device and contents reproducing method

    JP2007250053A

  • Similar musical piece display system

    JP2011141406A

  • Determination device, determination method, determination program, terminal device, and music piece reproduction program

    JP2017041136A