Information processing method, information processing device, and information processing program

By calculating separate rhythmic and harmonic complexity values, the method provides a comprehensive evaluation of audio content, addressing the limitations of conventional methods and enhancing the accuracy and relevance of sound content assessment.

WO2026042649A1PCT designated stage Publication Date: 2026-02-26PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/028429
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-23
Filing Date
2025-08-12
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Conventional methods for evaluating audio content complexity fail to accurately reflect the correlation between rhythmic and harmonic complexities, and do not account for factors like beat changes and musical harmony, making it difficult to assess the true complexity of sound sources, especially in commercially available music.

Method used

Calculate a first complexity value indicating rhythmic complexity and a second complexity value indicating harmonic complexity, using features such as beat intervals and harmonic changes, and output these values separately to provide a comprehensive evaluation.

Benefits of technology

This approach allows for an accurate and multi-dimensional assessment of audio content complexity, enabling better understanding and selection of sound content based on user preferences and situational suitability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025028429_26022026_PF_FP_ABST
    Figure JP2025028429_26022026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing method comprises: extracting a feature amount of a sound content; calculating, on the basis of the extracted feature amount, a first complexity value and a second complexity value which are indices each indicating the complexity of the sound content, the first complexity value being an index indicating the complexity of the rhythm of the entire sound content and the second complexity value being an index indicating the complexity of the harmony of the entire sound content; and associating and outputting the sound content, and the first complexity value and the second complexity value.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing device, and information processing program

[0001] The present disclosure relates to techniques for evaluating audio content.

[0002] Patent Document 1 discloses a preference estimation device that extracts features of content, calculates a complexity value of the content based on the extracted features, calculates an emotional value evoked by viewing the content based on the features and the complexity value, and determines preferences based on the calculated emotional value.

[0003] However, in the conventional technology of Patent Document 1, a representative value of multiple complexity values ​​is calculated as a final complexity value, and the final complexity value is aggregated into one dimension. Therefore, this conventional technology does not reflect the correlation between each complexity value, and therefore the complexity of audio content cannot be accurately evaluated.

[0004] JP 2012-220653 A

[0005] The present disclosure has been made to solve such problems, and aims to provide a technology that can accurately evaluate the complexity of sound content.

[0006] An information processing method in one aspect of the present disclosure is an information processing method in a computer, which includes acquiring sound content, extracting features of the sound content, calculating a first complexity value and a second complexity value that are indicators of the complexity of the sound content based on the extracted features, wherein the first complexity value is an indicator of the rhythmic complexity of the entire sound content and the second complexity value is an indicator of the harmonic complexity of the entire sound content, and outputting the sound content in correspondence with the first complexity value and the second complexity value.

[0007] According to the present disclosure, the complexity of sound content can be accurately evaluated.

[0008] Fig. 1 is a block diagram showing an example of the overall configuration of an information processing system according to an embodiment of the present disclosure; Fig. 2 is a flowchart showing an example of processing when an information processing device according to an embodiment of the present disclosure generates relationship information; Fig. 3 is a diagram schematically showing a spiral array model; Fig. 4 is a diagram showing an example of a relationship information table; Fig. 5 is a diagram showing a scatter plot of a first example; Fig. 6 is a diagram showing a scatter plot of a second example; Fig. 7 is a flowchart showing an example of processing when an information processing device selects audio content; Fig. 8 is a diagram showing a graph used when determining attributes;

[0009] (Findings underlying the present disclosure) The prior art disclosed in Patent Document 1 analyzes the features of extracted sound content and calculates rhythmic redundancy, rhythmic typicality, melodic redundancy, and melodic typicality as complexity values.

[0010] Rhythmic redundancy is calculated as "1" if the rhythm pattern is the same as the previous section, and "0" if it is different. Rhythmic typicality is calculated as "1" if the pattern is the same as a frequently occurring rhythm pattern, and "0" if it is different. Melodic redundancy is calculated as "1" if the melody pattern is the same as the previous section, and "0" if it is different. Melodic typicality is compared with the results of machine learning, and is calculated as "1" if the melody progression is common, and "0" if the melody progression is rare.

[0011] However, in this conventional technology, the final output complexity value is the average or weighted average of rhythmic redundancy, rhythmic typicality, melodic redundancy, and melodic typicality. This results in a one-dimensional output complexity value, which makes it difficult to understand the correlation between each complexity value. Furthermore, this conventional technology only calculates the complexity by determining the redundancy and typicality of note information quantized into discrete notes for each measure, without taking into account factors such as beat changes within a measure and the degree of musical harmony, which are thought to affect complexity. Therefore, further improvement is required to accurately evaluate the complexity of sound content. Furthermore, as mentioned above, since this technology only deals with note information quantized into discrete notes, it is not practical for sound sources that contain a variety of sounds, such as commercially available music.

[0012] Therefore, the inventor discovered that sound complexity can be accurately evaluated by calculating a first complexity value indicating the rhythmic complexity of the entire sound content from the features of the sound content and a second complexity value indicating the harmonic complexity of the entire sound content, and outputting both values, which led to the present disclosure.

[0013] (1) An information processing method in one aspect of the present disclosure is an information processing method in a computer, which includes acquiring sound content, extracting features of the sound content, calculating a first complexity value and a second complexity value that are indicators of the complexity of the sound content based on the extracted features, wherein the first complexity value is an indicator of the rhythmic complexity of the entire sound content and the second complexity value is an indicator of the harmonic complexity of the entire sound content, and outputting the sound content in correspondence with the first complexity value and the second complexity value.

[0014] According to this configuration, first and second complexity values ​​are calculated from the feature quantities of the sound content, and the calculated first and second complexity values ​​are output. Thus, in this configuration, the first and second complexity values ​​are not integrated. Furthermore, this configuration outputs two indices, rhythmic complexity and harmonic complexity, in association with the sound content, making it possible to evaluate the sound content from multiple perspectives. Therefore, the complexity of the constituent sounds can be accurately evaluated.

[0015] (2) In the information processing method described in (1) above, the feature may include beat interval information indicating the beat interval of the beats included in the sound content and a flatness that specifies the degree to which the beats are buried in the sound content, and the first complexity value may be an index based on the beat interval information and the flatness.

[0016] According to this configuration, a first complexity value that can objectively evaluate the complexity of the beat can be calculated using the beat interval and the flatness that defines the degree to which the beat contained in the sound content is buried by other sounds.

[0017] (3) In the information processing method described in (2) above, calculating the first complexity value may include calculating a degree of variation indicating the variation of the beat interval from the beat interval information, and calculating the first complexity value by multiplying the degree of variation by the flatness.

[0018] According to this configuration, the first complexity value is calculated by multiplying the degree of variation in beat intervals by the degree of flatness, so that it is possible to calculate a first complexity value that can objectively evaluate the complexity of the beat.

[0019] (4) In the information processing method described in any one of (1) to (3) above, the feature may include information on chords in the sound content and a degree of pitch change indicating the degree of pitch change in the sound content, and the second complexity value may be an index based on the information on chords and the degree of pitch change.

[0020] According to this configuration, it is possible to calculate a second complexity value that can objectively evaluate the complexity of a harmony using chord information and the degree of pitch change.

[0021] (5) In the information processing method described in (4) above, the chord information may be a chromagram, and calculating the second complexity value may include arranging, for each unit time, a plurality of pitches included in the chromagram at a plurality of pitch positions defined by a spiral array model that expands tonetz, which represents the structure of a chord on a two-dimensional plane, into a three-dimensional spiral space; calculating, for each unit time, a first sum that is the sum of the distances between the arranged pitches; calculating a second sum by adding the first sum over the entire time of the chromagram; and calculating the second complexity value by multiplying the second sum by the degree of pitch change.

[0022] According to this configuration, a first sum is calculated, which is the sum of the distances between multiple pitches arranged in the spiral array model for multiple unit times, the second sum is calculated by adding the first sum over the entire time, and the second complexity value is calculated by multiplying the second sum by the degree of change in pitch, so that a second complexity value that can objectively evaluate the complexity of harmony can be calculated.

[0023] (6) In the information processing method described in any of (1) to (5) above, the method may further include outputting a scatter plot in which data points defined by the first complexity value and the second complexity value are mapped onto a coordinate space.

[0024] This configuration allows the complexity of the audio content to be presented in an easy-to-understand manner.

[0025] (7) In the information processing method described in (6) above, the data point may be defined by the feature amount in addition to the first complexity value and the second complexity value.

[0026] This configuration makes it possible to present the relationship between the complexity of sound content and the feature amount in an easy-to-understand manner.

[0027] (8) In the information processing method described in (6) or (7) above, the scatter diagram may include data points corresponding to each of a plurality of sound contents.

[0028] According to this configuration, data points corresponding to a plurality of sound contents are plotted on a scatter diagram, so that the differences in complexity of the plurality of sound contents can be presented in an easy-to-understand manner.

[0029] (9) In the information processing method described in any of (1) to (8) above, the sound content may include reference sound content for which the user's preference value is known, and the method may further include obtaining a preference value indicating the user's preference for the reference sound content, and generating relationship information indicating the correspondence between the preference value and a reference first complexity value and a reference second complexity value, which are the first complexity value and the second complexity value for the reference sound content.

[0030] According to this configuration, relationship information is generated that indicates the correspondence between the user's preference value and the reference first complexity value and the reference second complexity value.By using this relationship information, it is possible to estimate the user's preference value for sound content whose user preference value is unknown.

[0031] (10) In the information processing method described in (9) above, the sound content may include target sound content for which the user's preference value is unknown, and the method may further include obtaining a target first complexity value and a target second complexity value, which are the first complexity value and the second complexity value for the target sound content, and calculating an estimated preference value, which is the preference value of the target sound content, based on the target first complexity value and the target second complexity value and the relationship information.

[0032] According to this configuration, it is possible to estimate the user's preference value for the target sound content, the preference value of which is unknown.

[0033] (11) The information processing method according to (10) above may further include classifying the target sound content into one of a plurality of attributes according to the estimated preference value.

[0034] According to this configuration, the attribute of the target sound content can be determined using the estimated preference value.

[0035] (12) The information processing method described in (11) above may further include acquiring situation information indicating the user's situation, selecting target sound content from multiple target sound contents having attributes suitable for the situation indicated by the situation information, and outputting the selected target sound content.

[0036] According to this configuration, it is possible to output target sound content having attributes suitable for the user's environment.

[0037] (13) In the information processing method described in (1) above, the feature may include beat interval information indicating the beat interval of a beat included in the sound content and a flatness that specifies the degree to which the beat is buried in the sound content, the first complexity value may be calculated by inputting the beat interval information and the flatness into a first trained model, and the first trained model may be a trained model trained using a first dataset that includes the beat interval information, the flatness, and the first complexity value.

[0038] According to this configuration, the first complexity value can be calculated using the first trained model.

[0039] (14) In the information processing method described in (1) or (13) above, the feature may include a chromagram of the sound content and a pitch change degree indicating the degree of pitch change in the sound content, the second complexity value may be calculated by inputting the chromagram and the pitch change degree into a second trained model, and the second trained model may be a trained model trained using a second dataset including the chromagram, the pitch change degree, and the second complexity value.

[0040] According to this configuration, the second complexity value can be calculated using the second trained model.

[0041] (15) In another aspect of the present disclosure, an information processing device is an information processing device including a processor, wherein the processor performs the following operations: acquiring sound content; extracting features of the sound content; calculating a first complexity value and a second complexity value, which are indicators of the complexity of the sound content, based on the extracted features; the first complexity value is an indicator of the rhythmic complexity of the entire sound content, and the second complexity value is an indicator of the harmonic complexity of the entire sound content; and outputting the sound content in correspondence with the first complexity value and the second complexity value.

[0042] This configuration makes it possible to provide an information processing device that can accurately evaluate the complexity of sound.

[0043] (16) In another aspect of the present disclosure, an information processing program causes a computer to acquire sound content, extract features of the sound content, calculate a first complexity value and a second complexity value that are indicators of the complexity of the sound content based on the extracted features, the first complexity value being an indicator of the rhythmic complexity of the entire sound content and the second complexity value being an indicator of the harmonic complexity of the entire sound content, and output the sound content in correspondence with the first complexity value and the second complexity value.

[0044] According to this configuration, it is possible to provide an information processing program that can accurately evaluate the complexity of sound.

[0045] The present disclosure can also be realized as an information processing system operated by such an information processing program. Needless to say, such a computer program can be distributed on a non-transitory computer-readable recording medium such as a CD-ROM or via a communication network such as the Internet.

[0046] Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, and step orders shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components. Furthermore, in all of the embodiments, the respective contents can be combined.

[0047] 1 is a block diagram showing an example of the overall configuration of an information processing system 100 according to an embodiment of the present disclosure. The information processing system 100 is a system that analyzes audio content, calculates a first complexity value and a second complexity value, and presents an objective index of the complexity of the audio content. The information processing system 100 is also a system that selects audio content suitable for the user's situation based on the evaluation results.

[0048] The information processing system 100 includes an information processing device 1, a sound source server 2, a terminal device 3, and an audio device 4. The information processing device 1, the sound source server 2, the terminal device 3, and the audio device 4 are connected via a network 5 so as to be able to communicate with each other.

[0049] The information processing device 1 is configured as a computer such as a cloud server or an edge server. When the information processing device 1 is configured as a cloud server, the network 5 is configured as a wide area communication network including the Internet and a mobile phone communication network. When the information processing device 1 is configured as an edge server, the network 5 is configured as a local area network.

[0050] In the example of Fig. 1, the information processing device 1 is configured by separate devices, namely, the sound source server 2, the terminal device 3, and the audio device 4, but this is just an example. The information processing system 100 may be configured by the same device. Furthermore, the information processing device 1 and the sound source server 2 may be configured by the same device. Furthermore, the information processing device 1 may be implemented in the terminal device 3 or in the audio device 4. Furthermore, the information processing device 1, the terminal device 3, and the audio device 4 may be configured by a single device.

[0051] Sound content is, for example, content that stimulates the user's hearing through sound, and includes not only music but also environmental sounds, beeps, etc. Environmental sounds include, for example, sounds of the natural environment, such as the sound of a flowing river, the sound of leaves rustling in the wind, and the sound of a flowing waterfall. Furthermore, environmental sounds include, for example, sounds of the city environment, such as the sound of a hustle and bustle in an urban area and the sound of a running train. Sound content may also be content that includes video in addition to sound.

[0052] Audio content is composed of recorded sound data, and its data format may be, for example, WAV (Waveform Audio File Format), AIFF (Audio Interchange File Format), MP3 (MPEG-1 Audio Layer 3), AAC (Advanced Audio Coding), or FLAC (Free Lossless Audio Codec).

[0053] The sound source server 2 is a server that records sound content. It crawls sound content that can be viewed on the Internet and records the crawled sound content. When the sound source server 2 receives an acquisition request from the information processing device 1, it selects appropriate sound content from all the sound content that it records and transmits the selected sound content to the information processing device 1.

[0054] The terminal device 3 is owned by a user who uses a service provided by the information processing system 100. The terminal device 3 may be a desktop computer or a portable computer such as a smartphone or a tablet terminal. The terminal device 3 includes a display and an input device. The display displays a scatter diagram (described later) generated by the information processing device 1. The input device accepts input of preference values ​​(described later) and transmits the accepted preference values ​​to the information processing device 1.

[0055] The audio device 4 includes a playback device, a speaker, and an amplifier. The audio device 4 outputs the sound represented by the audio content from the speaker. The playback device is, for example, a compact disc player or a DVD player. The audio device 4 plays the audio content selected by the information processing device 1. The audio device 4 transmits the audio content to the information processing device 1.

[0056] The terminal device 3 and the audio device 4 may be configured as a single device, such as a portable computer and a desktop computer.

[0057] The information processing device 1 includes a processor 10, a communication unit 11, and a memory 12. The communication unit 11 is a communication circuit that connects the information processing device 1 to a network 5. The communication unit 11 transmits scatter plot data showing a scatter plot to the terminal device 3. The communication unit 11 receives sound content from the sound source server 2 and the audio device 4. The communication unit 11 transmits selected sound content to the audio device 4.

[0058] The processor 10 is composed of electrical circuits such as a CPU (Central Processing Unit). The processor 10 includes an acquisition unit 101, an extraction unit 102, an index calculation unit 103, a scatter diagram generation unit 104, a relationship information generation unit 105, and a music selection unit 106. The acquisition unit 101 to the music selection unit 106 are realized by the processor 10 executing an information processing program recorded in the memory 12. However, this is just an example, and the acquisition unit 101 to the music selection unit 106 may be composed of dedicated hardware circuits. Some of the components of the acquisition unit 101 and the music selection unit 106 may be distributed and located in a device different from the information processing device 1.

[0059] The acquisition unit 101 acquires audio content from the sound source server 2 and the audio device 4 via the communication unit 11. The acquisition unit 101 acquires audio content by, for example, transmitting an acquisition request to the sound source server 2 or the audio device 4.

[0060] The extraction unit 102 extracts feature quantities from the sound content acquired by the acquisition unit 101. The feature quantities include, for example, beat interval information and flatness. The beat interval information is information indicating the beat intervals of beats included in the sound content. The beat interval information is, for example, time-series data in which beat intervals are arranged in time series. The flatness specifies the degree to which the beats are buried in the sound content. The flatness has one value for one piece of sound content.

[0061] The extraction unit 102 may extract the beat interval using a method such as frequency domain analysis, autocorrelation, a learned model, square wave matching, etc. The extraction unit 102 may extract the flatness using equation (3) described later.

[0062] The feature quantity further includes a chromagram of the sound content and a pitch change degree. The chromagram is data indicating the intensity of 12 pitches (C, C#, D, D#, E, F, F#, G, G#, A, A#, B) of the sound content for each of a plurality of unit times.

[0063] The extraction unit 102 generates a chromagram, for example, by applying a short-time Fourier transform (STFT) to the audio content, calculating an amplitude spectrum from the STFT result, converting the amplitude spectrum to a decibel scale, and representing the converted amplitude spectrum in color.

[0064] The degree of pitch change indicates the degree of pitch change in the sound content.

[0065] The index calculation unit 103 calculates a first complexity value and a second complexity value, which are indices indicating the complexity of the audio content, based on the feature extracted by the extraction unit 102. The index calculation unit 103 outputs the calculated first and second complexity values. For example, the index calculation unit 103 may output the first and second complexity values ​​to the scatter diagram generation unit 104, or may output them to the terminal device 3 via the communication unit 11.

[0066] The first complexity value is an index indicating the rhythmic complexity of the entire sound content, and the second complexity value is an index indicating the harmonic complexity of the entire sound content.

[0067] The index calculation unit 103 may calculate the degree of variation that indicates the variation in beat intervals, and calculate the first complexity value by multiplying the calculated degree of variation by the degree of flatness.

[0068] The index calculation unit 103 arranges multiple pitches included in the chromagram at pitch positions of 12 pitches defined by the spiral array model for multiple unit times, and calculates a first sum, which is the sum of the distances between the arranged multiple pitches, for multiple unit times. The index calculation unit 103 calculates a second sum by adding the first sum over the entire time of the chromagram. The index calculation unit 103 calculates a second complexity value by multiplying the calculated second sum by the degree of pitch change. The spiral array model is a model in which tunnels are expanded into a three-dimensional spiral space. Details of the spiral array model will be described later using FIG. 3.

[0069] The scatter diagram generation unit 104 outputs a scatter diagram in which data points defined by the first complexity value and the second complexity value calculated by the index calculation unit 103 are mapped onto a coordinate space. In this case, the scatter diagram is configured in a two-dimensional coordinate space in which one axis indicates the first complexity value and the other axis indicates the second complexity value.

[0070] The data points may be defined by a feature in addition to the first and second complexity values. For example, the feature may be the time average value of the beat interval indicated by the beat interval information. In this case, the scatter diagram is configured in a coordinate space with three axes: the first complexity value, the second complexity value, and the time average value of the beat interval. The feature used in the scatter diagram may be the flatness and the degree of pitch change.

[0071] The scatter plot may include data points corresponding to each of a plurality of sound contents, in which case a scatter plot is generated on which a plurality of data points are plotted.

[0072] The relationship information generation unit 105 acquires preference values ​​indicating a user's preference for reference audio content. Here, the relationship information generation unit 105 acquires preference values ​​for each of multiple reference audio contents. The relationship information generation unit 105 generates relationship information indicating a correspondence between the acquired preference values ​​and multiple reference first complexity values ​​and multiple reference second complexity values, which are first complexity values ​​and second complexity values ​​for the multiple reference audio contents. The reference audio content refers to the audio content acquired by the acquisition unit 101 to generate the relationship information. The preference value is a numerical value indicating the preference of a user who listens to the reference content. The preference value is expressed, for example, as a numerical value in multiple levels (e.g., five levels). The relationship information generation unit 105 simply has the index calculation unit 103 calculate the reference first complexity value and the reference second complexity value.

[0073] The relationship information may be, for example, a multiple regression line or a relationship information table 400. The relationship information generating unit 105 may calculate the multiple regression line by performing a multiple regression analysis with a plurality of reference first complexity values ​​and a plurality of reference second complexity values ​​as explanatory variables and a plurality of preference values ​​as objective variables. The relationship information generating unit 105 may generate the relationship information table 400 by performing a variance analysis of the plurality of reference first complexity values, the plurality of reference second complexity values, and the plurality of preference values.

[0074] The relationship information generating unit 105 stores the generated relationship information in the memory 12 as relationship information 121 .

[0075] The music selection unit 106 acquires a target first complexity value and a target second complexity value, which are a first complexity value and a second complexity value for the target sound content. Here, the music selection unit 106 acquires a plurality of target first complexity values ​​and a plurality of target second complexity values ​​for a plurality of target sound contents. The music selection unit 106 acquires both complexity values ​​by having the index calculation unit 103 calculate the target first complexity value and the target second complexity value. The target sound content is sound content whose preference value is unknown. The target sound content is sound content acquired by the acquisition unit 101 for the purpose of providing the sound content to the user.

[0076] The music selection unit 106 calculates a plurality of estimated preference values, which are preference values ​​for each of the plurality of target sound contents, based on a plurality of first complexity values ​​and a plurality of second complexity values ​​for the plurality of target sound contents and the relationship information 121.

[0077] The music selection unit 106 classifies each of the plurality of target sound contents into one of a plurality of attributes according to the estimated preference value. The attributes include a first attribute indicating that the content is suitable as background music and a second attribute indicating that the content is suitable as non-background music.

[0078] The music selection unit 106 acquires situation information indicating the user's situation. For example, the music selection unit 106 may acquire situation information input using the terminal device 3 from the terminal device 3. The music selection unit 106 selects, from a plurality of target sound contents, target sound content having attributes suitable for the situation indicated by the situation information. The music selection unit 106 transmits the selected target sound content to the audio device 4 and causes the transmitted target sound content to be output from the audio device 4. This allows the user to listen to sound content suitable for their own situation.

[0079] The memory 12 is configured as a non-volatile rewritable storage device such as a solid state drive, etc. The memory 12 stores relationship information 121.

[0080] 2 is a flowchart illustrating an example of a process performed by the information processing device 1 according to the embodiment of the present disclosure to generate the relationship information 121. The flowchart in FIG. 2 is executed, for example, when a user uses a service provided by the information processing device 1 for the first time.

[0081] (Step S1) The acquisition unit 101 acquires reference sound content. Here, the acquisition unit 101 acquires multiple reference sound contents. For example, the acquisition unit 101 transmits an acquisition request to the sound source server 2 via the communication unit 11, and acquires the multiple reference sound contents transmitted from the sound source server 2 via the communication unit 11.

[0082] (Step S2) The acquisition unit 101 determines one reference sound content to be processed from the plurality of reference sound contents acquired in step S1.

[0083] (Step S3) The extraction unit 102 extracts the feature quantities of one reference sound content. Here, the feature quantities extracted include beat interval information, flatness, chromagram, and pitch change degree.

[0084] (Step S4) The index calculation unit 103 calculates a reference first complexity value and a reference second complexity value, which are first and second complexity values ​​of one reference audio content. The first complexity value is calculated using Equation (1).

[0085]

[0086] In formula (1), C 1 indicates the first complexity value, and nPVI (The normalized pairwise variability index) indicates the degree of variability of the beat interval indicated by the beat interval information. R indicates the flatness.

[0087] The degree of variation is calculated using equation (2).

[0088] In formula (2), d i indicates the i-th beat interval, and l indicates the number of beat intervals.

[0089] Flatness W R is calculated using, for example, equation (3) which is the flatness proposed by Dubnov.

[0090] In equation (3), x(n) represents the sample value of the n-th sample point of the sound signal represented by the sound content, and N represents the number of sample points included in the sound signal.

[0091] The second complexity value is calculated using equation (4).

[0092]

[0093] In formula (4), C 2 denotes the second complexity value, and W H indicates the degree of pitch change.

[0094] In equation (5), D represents the second sum. g(P j , P i ) are any two pitches P among the multiple pitches arranged in the spiral array model. j and pitch P i j and i are indices for specifying the pitch. jt is the pitch P on the chromagram in unit time t. j The value indicates whether or not the pitch P j If there is, B jt = 1, and the pitch P j If there is no B jt = 0. B itis the pitch P on the chromagram in unit time t. i The value indicates whether or not the pitch P i If there is, B it = 1, and the pitch P i If there is no B it In this embodiment, the pitches P j , P i Is B jt = 1, B it = 1, and the pitch P j , P i Is B jt = 0, B it = 0. In equation (5), Σ indicating the addition of j=1 to 12 and Σ indicating the addition of i=1 to 11, excluding Σ indicating the addition of t=1 to T, represent the first summation.

[0095] FIG. 3 is a diagram schematically illustrating a spiral array model 300. The spiral array model 300 is a three-dimensional extension of the Tonetz model proposed by Euler, which represents the structure of chords on a two-dimensional plane. The spiral array model 300 includes 12 pitch positions 301 corresponding to 12 pitches. The 12 pitch positions 301 are arranged on the spiral array model 300 in order, for example, starting from Bb, spaced at intervals of perfect fifths. Furthermore, the 12 pitch positions 301 are arranged at positions where the pitch angle of the spiral array model 300 is an integer multiple of 45 degrees. k is an index that identifies the pitch position 301. h is the height of the spiral array model 300 per unit pitch angle (45 degrees). r is the radius of the spiral array model 300. For example, if the pitch P included in the chromagram at pitch positions 301 where k=1, 5 E , P C When the pitch P E , P C The distance between g(P i , P j ) is calculated as

[0096] (Step S5) The scatter diagram generation unit 104 transmits the one reference sound content determined in step S2 to the audio device 4 via the communication unit 11, and plays back the one reference sound content. This allows the user to listen to the one reference sound content.

[0097] (Step S6) The scatter diagram generating unit 104 acquires the preference value for one reference sound content by the user from the terminal device 3.

[0098] (Step S7) The acquisition unit 101 determines whether or not processing for all reference sound contents has been completed. If processing for all reference sound contents has not been completed (NO in step S7), the process returns to step S2, and the next reference sound content is determined. On the other hand, if processing for all reference sound contents has been completed (YES in step S7), the process proceeds to step S8.

[0099] (Step S8) The relationship information generation unit 105 uses the plurality of reference first complexity values, the plurality of reference second complexity values, and the plurality of preference values ​​to generate relationship information 121. The relationship information generation unit 105 generates, for example, a multiple regression line or a relationship information table 400 as the relationship information 121. The multiple regression line is expressed by Equation (6).

[0100]

[0101] In equation (6), y represents the preference value, x1 represents the first complexity value, and x2 represents the second complexity value. 0 , β 1 , β 2 are partial regression coefficients. That is, when the reference first preference value and the reference second preference value are input to the formula (6), the relationship information generating unit 105 calculates the partial regression coefficient β such that preference values ​​corresponding to the reference first complexity value and the reference second complexity value are output. 0 , β 1 , β 2 Ask for.

[0102] FIG. 4 is a diagram showing an example of a relationship information table 400. The relationship information table 400 is a table defined by two axes corresponding to the first complexity value and the second complexity value. The first complexity value and the second complexity value are each divided into three classes: low, medium, and high. The first complexity value is defined as a class below 40 as low, a class between 40 and 50 as medium, and a class above 50 as high. The second complexity value is defined as a class below 12 as low, a class between 12 and 20 as medium, and a class above 20 as high. Note that the thresholds defining the classes of the first complexity value and the second complexity value are not limited to the above values, and other values ​​may be used.

[0103] Preference values ​​are registered in cells corresponding to the respective classes of the first and second complexity values. The scatter diagram generation unit 104 identifies the class to which the reference first complexity value belongs and the class to which the reference second complexity value belongs for a plurality of reference sound contents, and votes preference values ​​for cells where the identified two classes intersect. The scatter diagram generation unit 104 then registers a representative value (e.g., an average value, an added value, etc.) of the preference values ​​voted for in each cell as the final preference value for each cell in the relationship information table 400.

[0104] (Step S9) The relationship information generating unit 105 stores the relationship information 121 generated in step S8 in the memory 12.

[0105] (Step S10) The scatter diagram generating unit 104 generates a scatter diagram with the first complexity values ​​and the second complexity values ​​as data points for the plurality of reference sound contents.

[0106] 5 is a diagram showing a first example of a scatter diagram 500. The scatter diagram 500 is configured in a two-dimensional coordinate space with the second complexity value defined on the vertical axis and the first complexity value defined on the horizontal axis. One data point 501 corresponds to one reference sound content. This allows a user to grasp the complexity for multiple reference sound contents.

[0107] FIG. 6 shows a second example scatter diagram 600. The scatter diagram 600 is configured in a three-dimensional coordinate space with the first complexity value on the depth axis, the second complexity value on the horizontal axis, and the feature amount on the height axis. The feature amount is the time average value of the beat intervals of each reference content. One data point 601 corresponds to one reference sound content. This allows the user to understand the relationship between complexity and beat intervals for multiple reference sound contents.

[0108] (Step S11) The scatter diagram generation unit 104 outputs scatter diagram data for displaying the scatter diagram 500 or the scatter diagram 600 to the terminal device 3 via the communication unit 11. As a result, the terminal device 3 displays the scatter diagram 500 or the scatter diagram 600 on the display.

[0109] 7 is a flowchart showing an example of processing when the information processing device 1 selects audio content. Note that the flowchart of FIG. 7 is executed, for example, in response to a user's request for music selection or periodically after the relationship information 121 is generated.

[0110] (Step S21) The acquisition unit 101 acquires a plurality of target sound contents. Here, the acquisition unit 101 acquires a plurality of target sound contents. For example, the acquisition unit 101 may acquire a plurality of sound contents recorded on a CD-ROM read by the audio device 4 as the plurality of target sound contents. Note that the acquisition unit 101 may also acquire a plurality of target sound contents from the sound source server 2.

[0111] (Steps S22 to S24) Steps S22 to S24 are different in that the processing target is a plurality of target sound contents instead of a plurality of reference sound contents, and the detailed processing contents are the same as steps S2 to S4. Note that step S24 differs from step S4 in that a target first complexity value and a target second complexity value are calculated.

[0112] (Step S25) The music selection unit 106 calculates an estimated preference value by inputting the target first complexity value and the target second complexity value calculated in step S24 into a multiple regression line. Alternatively, the music selection unit 106 identifies cells corresponding to the target first complexity value and the target second complexity value calculated in step S24 from the relationship information table 400, and reads preference values ​​from the identified cells to calculate an estimated preference value.

[0113] (Step S26) The music selection unit 106 determines the attribute of the one target sound content determined in step S22 to be the attribute corresponding to the estimated attribute value calculated in step S26.

[0114] 8 is a diagram showing a graph 800 used in determining attributes. In the graph 800, the vertical axis represents preference value and the horizontal axis represents complexity.

[0115] The graph 800 is divided into three regions 801 to 803 according to preference values. Region 801 is a region with a high preference value. Region 803 is a region with a low preference value. Region 802 is an intermediate region with preference values ​​located between regions 801 and 802.

[0116] The music selection unit 106 determines the attribute of target sound content whose estimated preference value belongs to area 802 as the first attribute. The music selection unit 106 determines the attribute of target sound content whose estimated preference value belongs to area 801 or area 803 as the second attribute. The first attribute is an attribute indicating suitability as background music. Target sound content whose estimated preference value belongs to area 801 will attract the user's attention, so it is not suitable for use as background music while working. Therefore, the music selection unit 106 determines the attribute of target sound content whose estimated preference value belongs to area 801 as the second attribute.

[0117] The target sound content whose estimated preference value belongs to the region 803 is likely to cause discomfort to the user and interfere with work. Therefore, the music selection unit 106 determines the attribute of the target sound content whose estimated preference value belongs to the region 803 to be the second attribute.

[0118] The target sound content whose estimated preference value belongs to the region 802 provides the user with a moderate level of pleasure and is therefore suitable as background music. Therefore, the music selection unit 106 determines the attribute of the target sound content whose estimated preference value belongs to the region 802 to be the first attribute.

[0119] Note that the graph 800 shows that the preference value is high when the complexity is medium, and the preference value decreases as the complexity decreases and as the complexity increases, taking into account the inverted U-shaped characteristic that humans prefer sound content with medium complexity over simple sound content and overly complex sound content.

[0120] (Step S27) In step S27, the acquisition unit 101 determines whether or not the processing for all target sound contents has been completed. If the processing for all target sound contents has not been completed (NO in step S27), the processing returns to step S22. On the other hand, if the processing for all target sound contents has been completed (YES in step S27), the processing proceeds to step S28.

[0121] (Step S28) The music selection unit 106 acquires situation information input by the user from the terminal device 3 via the communication unit 11. The situation indicated by the situation information is, for example, at work, upon waking up, having a meal, relaxing, etc.

[0122] (Step S29) The music selection unit 106 selects target sound content having an attribute suitable for the situation indicated by the situation information from the plurality of target sound contents acquired in step S21. For example, if the situation information indicates that the user is working or eating, the music selection unit 106 may select target sound content having a first attribute. On the other hand, if the situation information indicates that the user has just woken up or is relaxing, the music selection unit 106 may select target sound content having a second attribute.

[0123] (Step S30) The music selection unit 106 outputs the selected target sound content to the audio device 4 via the communication unit 11. As a result, the user listens to the selected target sound content via the audio device 4.

[0124] (Step S31) The scatter diagram generation unit 104 generates a scatter diagram 500 or a scatter diagram 600. The scatter diagram generation unit 104 may plot data points 501 corresponding to each of the multiple target sound contents acquired in step S21 on the scatter diagram 500 or the scatter diagram 600. When generating the scatter diagram 500, the data points 501 are configured from the target first complexity value and the target second complexity value of the target sound content. On the other hand, when generating the scatter diagram 600, the data points 601 are configured from the target first complexity value, the target second complexity value, and the feature amount of the target sound content.

[0125] (Step S32) The scatter diagram generation unit 104 outputs scatter diagram data indicating the scatter diagram 500 or the scatter diagram 600 to the terminal device 3 via the communication unit 11. As a result, the scatter diagram 500 or the scatter diagram 600 relating to the target sound content is displayed on the display of the terminal device 3.

[0126] As described above, according to the information processing device 1 of this embodiment, first and second complexity values ​​are calculated from the feature quantities of sound content, and the calculated first and second complexity values ​​are output. In this way, the information processing device 1 does not integrate the first and second complexity values. Furthermore, the information processing device 1 uses two indices, rhythmic complexity and harmonic complexity, to evaluate sound content from multiple perspectives. Therefore, sound complexity can be accurately evaluated.

[0127] (Modifications) The present disclosure can employ the following modifications.

[0128] (1) The index calculation unit 103 may calculate the first complexity value by inputting the beat interval information and flatness extracted by the extraction unit 102 into a first trained model. Here, the first trained model is a trained model that has been pre-trained using a large number of first data sets in which the beat interval and flatness are explanatory variables and the first complexity value is a target variable. The first trained model is configured by a neural network. For example, the neural network may be a feedforward neural network, a convolutional neural network, or the like.

[0129] (2) The index calculation unit 103 may calculate the second complexity value by inputting the chromagram and the degree of pitch change extracted by the extraction unit 102 into a second trained model. Here, the second trained model is a trained model that has been pre-trained using a large number of second data sets, with the chromagram and the degree of pitch change as explanatory variables and the second complexity value as a target variable. The second trained model is composed of a neural network. For example, the neural network may be a feedforward neural network, a convolutional neural network, or the like.

[0130] The present disclosure is useful in the technical field of selecting music that suits a user's preferences.

Claims

1. An information processing method in a computer, comprising: acquiring sound content; extracting features of the sound content; calculating a first complexity value and a second complexity value, which are indices indicating the complexity of the sound content, based on the extracted features; the first complexity value is an index indicating the rhythmic complexity of the entire sound content, and the second complexity value is an index indicating the harmonic complexity of the entire sound content; and outputting the sound content in correspondence with the first complexity value and the second complexity value.

2. An information processing method as described in claim 1, wherein the feature includes beat interval information indicating the beat interval of beats contained in the sound content and a flatness that specifies the degree to which the beats are buried in the sound content, and the first complexity value is an index based on the beat interval information and the flatness.

3. An information processing method as described in claim 2, wherein calculating the first complexity value includes: calculating a degree of variation indicating the variation of the beat interval from the beat interval information; and calculating the first complexity value by multiplying the degree of variation by the flatness.

4. An information processing method as described in claim 1, wherein the feature includes chord information of the sound content and a degree of pitch change indicating the degree of pitch change in the sound content, and the second complexity value is an index based on the chord information and the degree of pitch change.

5. The information processing method according to claim 4, wherein the chord information is a chromagram, and calculating the second complexity value includes: arranging, for each unit time, a plurality of pitches included in the chromagram at a plurality of pitch positions defined by a spiral array model that expands tonetz, which represents the structure of a chord on a two-dimensional plane, into a three-dimensional spiral space; calculating, for each unit time, a first sum that is the sum of the distances between the plurality of arranged pitches; calculating a second sum by adding up the first sum over the entire time of the chromagram; and calculating the second complexity value by multiplying the second sum by the degree of pitch change.

6. The information processing method according to claim 1, further comprising outputting a scatter plot in which data points defined by said first complexity value and said second complexity value are mapped onto a coordinate space.

7. The information processing method according to claim 6, wherein the data point is defined by the feature amount in addition to the first complexity value and the second complexity value.

8. The information processing method according to claim 6 or 7, wherein the scatter plot includes data points corresponding to each of a plurality of sound contents.

9. The information processing method of claim 1, wherein the sound content includes reference sound content for which a user's preference value is known, and further comprising: acquiring a preference value indicating the user's preference for the reference sound content; and generating relationship information indicating a correspondence between the preference value and a reference first complexity value and a reference second complexity value, which are the first complexity value and the second complexity value for the reference sound content.

10. The information processing method of claim 9, wherein the sound content includes target sound content for which the user's preference value is unknown, and further comprising: acquiring a target first complexity value and a target second complexity value, which are the first complexity value and the second complexity value for the target sound content; and calculating an estimated preference value, which is the preference value of the target sound content, based on the target first complexity value and the target second complexity value and the relationship information.

11. The information processing method according to claim 10, further comprising classifying the target sound content into one of a plurality of attributes according to the estimated preference value.

12. An information processing method as described in claim 11, further comprising: acquiring situation information indicating the user's situation; selecting, from a plurality of target sound contents, target sound content having attributes suitable for the situation indicated by the situation information; and outputting the selected target sound content.

13. An information processing method as described in claim 1, wherein the feature includes beat interval information indicating the beat interval of beats included in the sound content and a flatness that specifies the degree to which the beat is buried in the sound content, the first complexity value is calculated by inputting the beat interval information and the flatness into a first trained model, and the first trained model is a trained model trained using a first dataset that includes the beat interval information, the flatness, and the first complexity value.

14. An information processing method as described in claim 1, wherein the feature includes a chromagram of the sound content and a pitch change degree indicating the degree of pitch change in the sound content, the second complexity value is calculated by inputting the chromagram and the pitch change degree into a second trained model, and the second trained model is a trained model trained using a second dataset including the chromagram, the pitch change degree, and the second complexity value.

15. An information processing device including a processor, wherein the processor performs the following operations: acquire sound content; extract features of the sound content; calculate a first complexity value and a second complexity value that are indices indicating the complexity of the sound content based on the extracted features; the first complexity value is an index indicating the rhythmic complexity of the entire sound content, and the second complexity value is an index indicating the harmonic complexity of the entire sound content; and output the sound content in association with the first complexity value and the second complexity value.

16. An information processing program that causes a computer to perform the following steps: acquire sound content; extract features of the sound content; calculate a first complexity value and a second complexity value that are indices indicating the complexity of the sound content based on the extracted features; the first complexity value is an index indicating the rhythmic complexity of the entire sound content, and the second complexity value is an index indicating the harmonic complexity of the entire sound content; and output the sound content in association with the first complexity value and the second complexity value.

Citation Information

Patent Citations

  • Change-responsive preference estimation device

    JP2012220653A