Feature output model generation system

The feature output model generation system processes singing data through machine learning to generate features that improve the accuracy of song and key recommendations by considering song similarities, addressing the limitations of existing systems.

JP7784420B2Active Publication Date: 2025-12-11NTT DOCOMO INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023517139
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-27
Filing Date
2022-03-17
Publication Date
2025-12-11
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

Existing systems fail to generate appropriate features from singing data for song and key recommendations during karaoke, limiting the accuracy of song and key suggestions based on user singing history.

Method used

A feature output model generation system that processes singing data through machine learning, dividing it into sections, and calculates feature distances to generate a model that considers song similarities, enabling accurate key and song recommendations.

Benefits of technology

The system effectively outputs features that enhance the accuracy of song and key recommendations by considering song similarities, allowing for more precise user suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007784420000001
    Figure 0007784420000001
  • Figure 0007784420000002
    Figure 0007784420000002
  • Figure 0007784420000003
    Figure 0007784420000003
Patent Text Reader

Abstract

This feature amount output model generation system generates a feature amount output model which suitably outputs a feature amount from singing data. A feature amount output model generation system 10 generates a feature amount output model that receives input of information based on singing data, which is time series voice data associated with singing of a song, and that outputs a feature amount of the singing data, said feature amount output model generation system comprising: a singing data acquisition unit 11 that acquires pieces of singing data for a respective plurality of songs; a division unit 12 that divides each piece of the singing data into a plurality of temporal sections; and a feature amount output model generation unit 13 that generates, by machine learning, the feature amount output model which receives input of information based on the singing data of a divided section and which outputs a feature amount of the divided section from the divided singing data, wherein the feature amount output model generation unit 13 performs machine learning by a reference based on the distance between feature amounts of the singing data associated with the same song and on the distance between feature amounts of the singing data associated with songs different from each other.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a feature output model generation system that receives information based on singing data, which is time-series data of voices related to the singing of a song, and generates a feature output model that outputs features of the singing data. [Background technology]

[0002] Conventionally, it has been proposed to recommend songs to a user based on the user's singing history at karaoke (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-78387 Summary of the Invention [Problem to be solved by the invention]

[0004] The key of a song is important when singing karaoke. Therefore, similar to the song recommendations described above, it is possible to recommend a key for the user to sing in. When recommending a key, it is possible to use singing data, which is data on the user's voice singing in the past. By using singing data, it is possible to recommend a more appropriate key. Furthermore, when recommending songs, it is possible to make more appropriate recommendations by using singing data.

[0005] A pre-prepared recommendation model or the like may be used to determine the key or song to be recommended. To make an appropriate recommendation, it may be possible to use the features of the singing data, rather than the singing data itself, as input to the recommendation model or the like. However, no method has been proposed to generate features from singing data for recommendation purposes. Therefore, it has not been possible to make an appropriate recommendation using singing data.

[0006] One embodiment of the present invention has been made in consideration of the above, and aims to provide a feature output model generation system that can generate a feature output model that appropriately outputs features from singing data. [Means for solving the problem]

[0007] In order to achieve the above-mentioned object, a feature output model generation system according to one embodiment of the present invention is a feature output model generation system that inputs information based on singing data, which is time-series data of audio related to the singing of a song, and generates a feature output model that outputs features of the singing data. The system includes: a singing data acquisition unit that acquires singing data for each of a plurality of songs to be used in generating the feature output model; a division unit that divides each piece of singing data acquired by the singing data acquisition unit into a plurality of temporal sections; and a feature output model generation unit that inputs information based on the singing data of the divided sections from the singing data divided by the division unit and generates a feature output model through machine learning that outputs features of the singing data of the divided sections. The feature output model generation unit performs machine learning using criteria based on the distance between features of singing data related to the same song and the distance between features of singing data related to different songs.

[0008] In a feature output model generation system according to one embodiment of the present invention, machine learning is performed using criteria based on the distance between features of singing data relating to the same song and the distance between features of singing data relating to different songs to generate a feature output model. The feature output model generated in this manner can output features appropriate for use in recommendations that take into account the similarity between songs. In other words, the feature output model generation system according to one embodiment of the present invention can generate a feature output model that appropriately outputs features from singing data. [Effects of the Invention]

[0009] A feature output model generated according to one embodiment of the present invention can output features suitable for use in recommendations that take into account the similarities between songs. That is, according to one embodiment of the present invention, a feature output model that appropriately outputs features from singing data can be generated. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating a functional configuration of a feature output model generation system according to an embodiment of the present invention. [Figure 2] FIG. 10 is a diagram illustrating an example of singing data used to generate a feature output model. [Figure 3] FIG. 10 is a diagram for explaining generation of a feature output model by machine learning. [Figure 4] 10 is a graph showing an example of features output by a feature output model. [Figure 5] 1 is a flowchart illustrating a process executed in a feature output model generation system according to an embodiment of the present invention. [Figure 6] FIG. 1 is a diagram illustrating a hardware configuration of a feature output model generation system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, an embodiment of a feature output model generation system according to the present invention will be described in detail with reference to the drawings. In the description of the drawings, the same elements are given the same reference numerals and duplicated explanations will be omitted.

[0012] FIG. 1 shows the functional configuration of a feature output model generation system 10 according to this embodiment. The feature output model generation system 10 is a system (device) that generates a feature output model. The feature output model (feature generation model) receives information based on singing data, which is time-series data of voices related to the singing of a song, and generates and outputs features of the singing data. The singing data is, for example, data based on a user's voice recorded when the user sings a song at a karaoke bar. Specific examples of singing data will be described later. Note that the singing data does not necessarily have to be based on the user's singing, as long as it is time-series data of voices related to the singing of a song. For example, the singing data may be data (model data) based on voices that are expected to earn a perfect score in a karaoke scoring system.

[0013] The features of the singing data output by the feature output model are used, for example, to change the key when a user sings a song at karaoke or to recommend a song. How the output features are specifically used will be described later. The feature output model generation system 10 generates (infers) the feature output model by machine learning. That is, the feature output model is a trained model (machine learning model) by machine learning.

[0014] The feature output model generation system 10 is realized by a computer such as a PC (personal computer) or a server device, etc. The feature output model generation system 10 may also be realized by a plurality of computers, i.e., a computer system.

[0015] Next, a description will be given of the functions of the feature output model generation system 10 according to this embodiment. As shown in Fig. 1, the feature output model generation system 10 includes a singing data acquisition unit 11, a division unit 12, and a feature output model generation unit 13.

[0016] The singing data acquisition unit 11 is a functional unit that acquires singing data for each of a plurality of songs to be used in generating a feature output model. The singing data acquisition unit 11 may acquire singing data including data indicating the duration of a time-series pitch.

[0017] Singing data is, for example, information indicating the duration of the same pitch of the voice (music note information), as shown in FIG. 2. Furthermore, singing data is information for each song, and can be associated with a song ID, for example, so that it is possible to identify which song the data pertains to. The duration is indicated, for example, by the elapsed time from the start time of playback of the song. The value in the "pitch" column shown in FIG. 2 is information indicating the pitch. Specifically, the value in the "pitch" column is a note number (MIDI number, MIDI key). For example, the value 62 in the "pitch" column corresponds to D4 in the international key. The information in the "time_from" and "time_to" columns shown in FIG. 2 indicates the timing when singing (vocalization) at the corresponding pitch started and ended, respectively. The units of "time_from" and "time_to" are seconds. The numbers on the left side of FIG. 2 are serial numbers for the information.

[0018] For example, the data in row 0 of Figure 2 indicates that the pitch of note number 62 continued for 0.332 seconds, from 8.005 seconds to 8.337 seconds, with the start time of the song as the base (0 seconds).

[0019] Note that, due to the presence of an introductory section, the user usually starts singing some time after the start of the song. Therefore, the first "time_from" is not 0 seconds. In the singing data shown in Figure 2, the first "time_from" is 8.005 seconds.

[0020] The singing data shown in FIG. 2 can be obtained, for example, by analyzing the voice sung by a conventional karaoke system (karaoke song file). The singing data acquisition unit 11 acquires singing data from the karaoke system, for example, obtained from the voice recorded when the user sings. Alternatively, the singing data acquisition unit 11 may acquire raw data of the voice related to the user's singing and generate and acquire singing data by a conventional method. The singing data acquisition unit 11 may also acquire singing data by a method other than the above. Note that the singing data acquired by the singing data acquisition unit 11 does not necessarily have to be based on the actual singing of the user, as long as it is data in the format shown in FIG. 2 that simulates actual singing. Note that the singing data does not necessarily have to be the format shown in FIG. 2, as long as it is time-series data of the voice related to the singing of a song.

[0021] The singing data acquisition unit 11 acquires multiple pieces of singing data for different songs. The user who sang the singing data associated with the singing data acquired by the singing data acquisition unit 11 may be any user, or may be multiple users. The user may be the same user as the user who sang the singing data associated with the singing data whose features are output by the feature output model (i.e., for example, the user who is the target of recommendation), or a different user. The singing data acquisition unit 11 acquires a sufficient amount of singing data to perform machine learning, which will be described later. The singing data acquisition unit 11 outputs the acquired singing data to the division unit 12.

[0022] The dividing unit 12 is a functional unit that divides each piece of singing data acquired by the singing data acquiring unit 11 into multiple time segments. As will be described later, machine learning is performed using the singing data divided into time segments. The dividing unit 12 performs division as follows.

[0023] The dividing unit 12 receives singing data from the singing data acquisition unit 11. Based on division rules stored in advance, the dividing unit 12 divides each input singing data into multiple sections. The dividing unit 12 divides each singing data into a certain number of sections. The dividing unit 12 divides the time period from the first "time_from" to the last "time_to" of the singing data, i.e., the time period related to the singing data, into equal parts so as to equalize the above-mentioned certain number of sections, and sets the divided timings (times) as timings to be used as divisions. If the timings to be used as divisions are included in a time period of consecutive identical pitches (i.e., a time period from "time_from" to "time_to" associated with one "pitch" value), the beginning or end of the time period, whichever is closer to the timing, is set as the division of the section. By dividing the data into sections in this way, consecutive identical pitches are not divided into multiple sections. For example, data 0 to 4 shown in FIG. 2 are set as data for one section (the first section).

[0024] The number of sections to be divided does not have to be a fixed number, but may be a number that is preset for each song. Furthermore, the timing used for dividing the data may not be equal as described above, but may be a timing that is preset for each song. The dividing unit 12 outputs the singing data to be divided and information indicating the divided sections to the feature output model generating unit 13.

[0025] The feature output model generation unit 13 is a functional unit that generates a feature output model by machine learning from the singing data divided by the division unit 12. The feature output model is a model that receives information based on the singing data of a divided section and outputs the feature of the singing data of that section. The feature output model generation unit 13 performs machine learning using criteria based on the distance between the feature of singing data related to the same song and the distance between the feature of singing data related to different songs.

[0026] The feature output model generation unit 13 may perform machine learning so that the distance between features of singing data relating to the same song is shorter than the distance between features of singing data relating to different songs.The feature output model generation unit 13 may determine the sections of singing data to be used for machine learning based on the distance between features output by the feature output model being generated.The feature output model generation unit 13 may convert data indicating pitch lengths included in the singing data divided by the division unit 12 into words, which are character strings corresponding to the pitch lengths for each consecutive identical pitch, and generate a feature output model that inputs information based on the converted words.

[0027] The features output by the feature output model are vectors with a preset number of dimensions (N-dimensional vectors). In other words, the feature output model is a model that performs embedding. As will be described later, singing data for a section is converted into words, and features are generated from the converted words. The feature output model is configured to include, for example, a neural network. More specifically, the feature output model is an LSTM (Long Short Term Memory) that is suitable for time-series data. However, the feature output model may be any model other than those described above, as long as it is generated by machine learning, inputs singing data for a section, and outputs features of the singing data for that section.

[0028] When recommending keys or songs, the similarity between songs can be a very important feature in determining a user's song preferences or ease of singing. The similarity between songs here refers to the extent to which the songs share the same rhythm, melody, and scale patterns, and the similarity of their pitches. As described above, the feature output model inputs information based on singing data. Furthermore, the feature output model is generated taking into account the distance between features of the singing data related to the songs. The features output by the feature output model generated in this way take into account the similarity between the songs. Music is expressed by rhythm, melody, and scale patterns, etc., but in the past, it was difficult to calculate the similarity between songs based on individual features such as rhythm, melody, and scale patterns.

[0029] The feature output model generation unit 13 generates a feature output model as follows: The feature output model generation unit 13 receives the singing data and information indicating the divided sections from the division unit 12. The feature output model generation unit 13 converts each consecutive identical pitch of the singing data (each line in FIG. 2) into a word, which is a character string corresponding to the length of the pitch. First, the feature output model generation unit 13 calculates the time length of the consecutive identical pitch. The length in seconds can be calculated by subtracting the value of "time_from" from the value of "time_to". Next, the feature output model generation unit 13 divides the calculated length by a preset unit time. The preset unit time is, for example, 0.1 seconds. The feature output model generation unit 13 rounds the calculated value to an integer. The rounding is performed using a preset method, for example, rounding up, rounding down, or rounding up. The feature output model generating unit 13 determines that a character string in which values ​​indicating pitch are consecutively arranged by the calculated integer is a consecutive word of the same pitch.

[0030] For example, for the data of pitch 0 in FIG. 2, the time length is first calculated as 8.337-8.005=0.332 seconds. Next, this calculated value is divided by a preset unit time to obtain 0.332 / 0.1=3.32. This value is rounded up or down to the nearest integer, 3. That number of consecutive 62 values ​​indicating the pitch are used as the word (note embedding) for that pitch data. The word indicates how long the pitch lasted. In the above example, the wording is performed based on how many tenths of a second the pitch lasted. The word represents the pitch and duration. The feature output model generation unit 13 converts all pitch data into words.

[0031] The feature output model generation unit 13 encodes the singing data based on the elapsed time of the same pitch (same note). This allows irregularly consecutive singing data (sound information) to be handled as information used to generate a feature output model. The encoding described above is expected to enable the handling of notes with the same pitch but slightly different lengths, and notes with similar pitches and the same length, as similar information.

[0032] As a preprocessing step for generating the feature output model, the feature output model generation unit 13 converts each converted word into an input feature to be input to the feature output model. In the following description, when simply referring to a feature, it refers to a feature that is the output of the feature output model, and a feature converted from a word is called an input feature. The input feature is a vector with a preset number of dimensions. The conversion from a word to an input feature is performed, for example, by a conventional natural language processing method. Specifically, the conversion is performed by a model generated by fastText. The above words may be used to generate a model by fastText, which is a machine learning method.

[0033] By using fastText, it is possible to group similar words into meaningful groups by taking into account the grouping of characters that make up a given word. Therefore, although "7171" and "717171" are different words, they are treated as being semantically close because they share the common component "71."

[0034] The feature output model generation unit 13 inputs the input features converted from the words included in the section, in the order in which the words appear in the singing data, to the feature output model for that section. For example, the feature output model has an input layer with neurons equal to the number of dimensions of the vector that is the input features. Also, an output layer with neurons equal to the number of dimensions of the vector that is the features (the above-mentioned N). The feature output model inputs the input features included in the section in order for each input feature. When all input features included in the section have been input, the feature output model outputs features.

[0035] The feature output model generation unit 13 performs machine learning for generating a feature output model using information on the three sections as one set. The three sections are a predetermined section (anchor), another section (positive) in the same song as the predetermined section, and another section (negative) in a different song from the predetermined section. The feature output model generation unit 13 selects (extracts) the above three sections.

[0036] The feature output model generation unit 13 performs machine learning using information on the three sections (sets of input features) selected as shown in FIG. 3 as input to the feature output model. The feature output model outputs features for each of the three sections. Specifically, as shown in FIG. 3, Anchor Embedding, which is an anchor feature, is obtained from Anchor Seq, which is input anchor information, by Embedding Net, which is the feature output model. Positive Embedding, which is a positive feature, is obtained from Positive Seq, which is positive information. Negative Embedding, which is a negative feature, is obtained from Negative Seq, which is negative information. Note that while FIG. 3 shows words as being input to the feature output model, in reality, input features are input to the feature output model.

[0037] The feature output model generation unit 13 performs machine learning based on the output features, i.e., the above-mentioned Anchor Embedding, Positive Embedding, and Negative Embedding. As shown on the right side of FIG. 3, the feature output model generation unit 13 performs machine learning so that the distance D1 between the Anchor Embedding and the Positive Embedding (the distance between the anchor and the positive) is shorter than the distance D2 between the Anchor Embedding and the Negative Embedding (the distance between the anchor and the negative). The above distance may be a Euclidean distance in an N-dimensional space, which is a vector space (feature space) of the features, or any other distance.

[0038] Specifically, the feature output model generation unit 13 performs machine learning using the following loss function Loss(): Loss(A,P,N)=Max(||f(A)-f(P)|| 2 -||f(A)-f(N)|| 2 +α,0) In the above equation, A is the Anchor Sequence. P is the Positive Sequence. N is the Negative Sequence. f(X) is the embedding vector obtained as the output when X is input to the Embedding Net. ||f(A)-f(P)|| is the distance between the anchor and the positive. ||f(A)-f(N)|| is the distance between the anchor and the negative. α is a preset hyperparameter that indicates the difference between the anchor-positive distance and the anchor-negative distance. Max(X,Y) is a function whose function value is the larger of X and Y. Machine learning based on the loss function itself can be performed in the same way as before.

[0039] The feature output model generation unit 13 selects three intervals, anchor, positive, and negative, from the intervals indicated by the information input from the division unit 12, and performs machine learning using the information of the selected intervals. The feature output model generation unit 13 generates a feature output model by repeatedly selecting the three intervals and performing machine learning. For example, the feature output model generation unit 13 generates a feature output model by repeating the above process until the generation of the feature output model converges based on preset conditions or a predetermined number of times, as in the conventional case.

[0040] The selection of the three sections, anchor, positive, and negative, is performed as follows. For example, the feature output model generation unit 13 selects the three sections by random sampling. In this case, the feature output model generation unit 13 randomly selects a song for each of the above repetitions, i.e., for each learning epoch, and randomly selects two sections included in the selected song to be the anchor and positive. Furthermore, the feature output model generation unit 13 randomly selects a song different from the selected song, and randomly selects one section included in the selected song to be the negative.

[0041] Alternatively, the feature output model generation unit 13 may select a negative using a feature output model that is currently being generated. In this case, the anchor and positive are selected in the same manner as above. The feature output model generation unit 13 randomly selects a song other than the song that includes the anchor and the positive. In this case, there may be multiple other songs. The feature output model generation unit 13 randomly selects a section included in the selected song to be a negative candidate. In this case, there are multiple negative candidates, for example, a preset number (N).

[0042] The feature output model generation unit 13 calculates the feature of the anchor and the feature of each negative candidate using the feature output model being generated. Next, the feature output model generation unit 13 calculates the distance between the anchor and the negative candidate for each negative candidate from the calculated feature. The feature output model generation unit 13 determines the negative candidate whose calculated distance is within a preset threshold as the negative candidate to be used in machine learning. By adopting negative candidates in this way, distance calculation is required for each sampling, but learning progresses intensively for negative candidates that should be far from the anchor.

[0043] Note that negative candidates may be all sections of all songs other than the anchor song, but instead of considering distance calculations to select all as candidates, negatives may be determined using a fixed number, N, of negative candidates as described above. This can speed up the machine learning processing. Note that the process of determining the sections to be used in machine learning may be performed as a mini-batch process.

[0044] The feature output model generation unit 13 outputs the generated feature output model. For example, the feature output model is transmitted or output to another device or module that uses the feature output model. Alternatively, the feature output model generation unit 13 may store the generated feature output model in the feature output model generation system 10 so that the generated feature output model can be used by another device or module that uses the feature output model.

[0045] FIG. 4 shows an example of features output by the feature output model generated by the feature output model generation system 10. In FIG. 4, one point corresponds to the feature of one section. Points of the same color (same intensity) correspond to points in the same song. In FIG. 4, the feature, which is a high-dimensional vector, is converted (dimensionality compressed) into a three-dimensional vector. The feature of sections included in the same song are feature of positions close to each other. In FIG. 4, the points included in the rectangular regions A1 and A2 correspond to the feature of sections included in the same song for each region A1 and A2. Note that, even if the songs are different, the feature of sections between songs with similar melody, genre, etc. is closer to each other than the feature of sections between songs with different melody, genre, etc. For example, even if the songs are different, the distance between the feature of sections between pop songs is closer than the distance between the feature of sections between pop and enka.

[0046] The feature output model, which is a trained model generated by the feature output model generation system 10, is expected to be used as a program module that is part of artificial intelligence software. The feature output model is used, for example, in a computer having a CPU (Central Processing Unit) and memory, and the CPU of the computer operates according to instructions from the feature output model stored in the memory. For example, the CPU of the computer operates in accordance with the instructions to input information to the feature output model, perform calculations according to the feature output model, and output results from the feature output model. Specifically, the CPU of the computer operates in accordance with the instructions to input information to the input layer of a neural network, perform calculations based on trained weighting coefficients in the neural network, and output results from the output layer of the neural network.

[0047] The feature output model generated as described above is used as follows. For example, the feature output model is used to recommend a key or a song when the song is sung at karaoke. Specifically, the feature output model is used when making the recommendation based on data on the past singing performance of the user to whom the recommendation is made. The singing performance data includes singing data similar to the singing data acquired by the singing data acquisition unit 11.

[0048] A feature output model is used to generate features (embedding) from the singing data. At this time, information to be input to the feature output model (for example, the above-mentioned input features) is generated from the singing data in the same way as when the feature output model is generated in the feature output model generation system 10. The information to be input to the feature output model is information for each section of the singing data. Division into sections is also performed in the same way as above.

[0049] The feature values ​​for each section may be arranged in the order of the sections for each song and concatenated to form feature values ​​for each song. In other words, the feature values ​​may be concatenated. In the case of simple concatenation, the dimension of the vector representing the feature values ​​for each song is the number of dimensions of the vector representing the feature values ​​output by the feature output model multiplied by the number of song sections. Furthermore, when concatenating, the vectors may be aggregated (averaged or added) to form feature values ​​for each song with lower dimensions.

[0050] Data on the past singing performance of the target user can be used for recommendations by converting it into features for each song related to singing as described above. When recommending a key, singing data that is model data for each key of the song the user intends to sing can be converted into features using the feature output model as described above and used. Note that to change the key of the model singing data, simply change the pitch value shown in Figure 2. For example, to raise the key by one, if the pitch value is 62, change it to 63. As a result, if the word corresponding to the pitch before the change was "626262," the word after the change will be "636363."

[0051] The recommendation is performed, for example, by using a recommendation model (Dense) that takes the above-mentioned feature quantities related to singing performance and feature quantities related to model data as inputs and outputs a value indicating how well the user's singing matches the key related to the model data. The above value can be calculated using model data in multiple different keys, and the key with the highest value can be recommended. The recommendation model can be created using conventional machine learning methods, etc.

[0052] By changing the key of a song as described above, even if the song being recommended to a woman is by a male singer, by changing the key it is possible to determine the similarity between the song and songs by female singers that the woman has sung in the past.

[0053] When recommending songs, model data of songs that are candidates for songs to be recommended to the user can be used. In this case, the recommendation is performed using a recommendation model that inputs, for example, the feature amounts related to the singing performance and the feature amounts related to the model data, and outputs a value indicating the degree to which the song related to the model data should be recommended. The value can be calculated using model data of multiple different songs, and the song with the highest value can be recommended.

[0054] The above is an example of a recommendation using features output by the feature output model, and the feature output model may be used for recommendations other than those described above. The feature output model may also be used for purposes other than recommendations. The above is the function of the feature output model generation system 10 according to this embodiment.

[0055] Next, the process executed by the feature output model generation system 10 according to this embodiment (the operation method performed by the feature output model generation system 10) will be described with reference to the flowchart of FIG.

[0056] In this process, the singing data acquisition unit 11 acquires singing data for each of a plurality of songs (S01). Next, the division unit 12 divides each singing data into a plurality of time segments (S02). Next, the feature output model generation unit 13 converts the singing data for each segment into an input format for the feature output model (S03). Specifically, the singing data is converted into words corresponding to the pitch length. Furthermore, the words are converted into input features.

[0057] Next, the feature output model generation unit 13 determines sections to be used in machine learning (S04). Specifically, the three sections, the anchor, positive, and negative sections, described above, are determined. Next, the feature output model generation unit 13 performs machine learning to generate a feature output model using information on the determined sections (S05). Specifically, as described above, machine learning is performed based on a criterion based on the distance between features of singing data relating to the same song and the distance between features of singing data relating to different songs. Next, the feature output model generation unit 13 determines whether to end the machine learning (S06).

[0058] If it is determined not to end the machine learning, the above-mentioned processes of S04 to S06 are performed again. If it is determined to end the machine learning, the generated feature output model is output from the feature output model generation unit 13 (S07). This process is executed by the feature output model generation system 10 according to this embodiment.

[0059] As described above, in this embodiment, machine learning is performed using criteria based on the distance between features of singing data relating to the same song and the distance between features of singing data relating to different songs, and a feature output model is generated. The feature output model generated in this manner can output features appropriate for use in recommendations that take into account the similarity between songs, as described with reference to FIG. 4 . That is, this embodiment can generate a feature output model that appropriately outputs features from singing data. As a result, similar songs can be extracted and keys or songs can be recommended with greater accuracy.

[0060] Furthermore, as in the above-described embodiment, machine learning may be performed so that the distance between features of singing data relating to the same song is shorter than the distance between features of singing data relating to different songs. This configuration makes it possible to generate a feature output model that reliably outputs features appropriately from singing data. However, machine learning does not necessarily have to be performed as described above, as long as it is performed based on criteria based on the distance between features of singing data relating to the same song and the distance between features of singing data relating to different songs.

[0061] Furthermore, as in the above-described embodiment, sections of singing data to be used for machine learning (for example, the above-described three sections: anchor, positive, and negative) may be determined based on the distances between features output by a feature output model in the process of being generated. With this configuration, machine learning can be performed efficiently as described above, and as a result, a feature output model with more advanced learning can be generated for the same learning process. Alternatively, the learning process in generating a feature output model can be reduced. However, the determination of sections to be used for machine learning does not have to be performed as described above.

[0062] Furthermore, as in the above-described embodiment, the singing data may include data indicating the length of a time-series pitch. This configuration allows for reliable and appropriate generation of a feature output model. Alternatively, the data indicating the length of pitches included in the divided singing data may be converted into words, which are character strings corresponding to the length of each consecutive identical pitch, to generate a feature output model that inputs information based on the converted words. This configuration allows for appropriate and easy handling of singing data using the above-described conventional natural language processing techniques to generate a feature output model. As a result, the feature output model can be generated appropriately and easily. However, the singing data does not have to be the above-described type; it may be time-series data of voices related to the singing of a piece of music. Furthermore, the singing data does not have to be treated as words as described above; it may be treated in any format that can be used as input to the feature output model.

[0063] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wires, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.

[0064] Functions include, but are not limited to, judgment, determination, judgment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, election, establishment, comparison, assumption, expectation, regard, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.

[0065] For example, the feature output model generation system 10 according to an embodiment of the present disclosure may function as a computer that performs information processing according to the present disclosure. Fig. 6 is a diagram illustrating an example of the hardware configuration of the feature output model generation system 10 according to an embodiment of the present disclosure. The feature output model generation system 10 described above may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.

[0066] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the feature output model generation system 10 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.

[0067] Each function in the feature output model generation system 10 is realized by loading predetermined software (programs) onto hardware such as a processor 1001 and a memory 1002, causing the processor 1001 to perform calculations, control communication via a communication device 1004, and control at least one of reading and writing of data in the memory 1002 and the storage 1003.

[0068] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured by a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, each function in the feature output model generation system 10 described above may be realized by the processor 1001.

[0069] Furthermore, the processor 1001 reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002, and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, each function of the feature output model generation system 10 may be implemented by a control program stored in the memory 1002 and running on the processor 1001. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.

[0070] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for performing information processing according to an embodiment of the present disclosure.

[0071] The storage 1003 is a computer-readable recording medium, and may be composed of at least one of an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. The storage 1003 may also be called an auxiliary storage device. The storage medium included in the feature output model generation system 10 may be, for example, a database, a server, or other appropriate medium including at least one of the memory 1002 and the storage 1003.

[0072] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also called, for example, a network device, a network controller, a network card, or a communication module.

[0073] The input device 1005 is an input device (for example, a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (for example, a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (for example, a touch panel).

[0074] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.

[0075] Furthermore, the feature output model generation system 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.

[0076] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0077] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.

[0078] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0079] Each aspect / embodiment described in this disclosure may be used alone, in combination, or switched depending on the implementation. Furthermore, notification of predetermined information (e.g., notification that "X is true") is not limited to being done explicitly, but may be done implicitly (e.g., by not notifying the predetermined information).

[0080] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.

[0081] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0082] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.

[0083] As used in this disclosure, the terms "system" and "network" are used interchangeably.

[0084] Furthermore, the information, parameters, etc. described in this disclosure may be expressed using absolute values, may be expressed using relative values ​​from a predetermined value, or may be expressed using other corresponding information.

[0085] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.

[0086] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.

[0087] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0088] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.

[0089] When used in this disclosure, the terms "include," "including," and variations thereof are intended to be inclusive, similar to the term "comprising." Furthermore, when used in this disclosure, the term "or" is not intended to be an exclusive or.

[0090] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0091] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different." [Explanation of symbols]

[0092] 10...feature output model generation system, 11...singing data acquisition unit, 12...division unit, 13...feature output model generation unit, 1001...processor, 1002...memory, 1003...storage, 1004...communication device, 1005...input device, 1006...output device, 1007...bus.

Claims

1. A feature output model generation system that receives information based on singing data, which is time-series data of voices related to singing of a song, and generates a feature output model that outputs features of the singing data, comprising: a singing data acquisition unit that acquires singing data for each of a plurality of songs to be used in generating a feature output model; a division unit that divides each of the singing data acquired by the singing data acquisition unit into a plurality of time segments; a feature output model generation unit that receives information based on the singing data of a divided section from the singing data divided by the division unit and generates, by machine learning, a feature output model that outputs the feature of the singing data of the section; The feature output model generation unit is a feature output model generation system that performs machine learning based on criteria based on the distance between features of singing data related to the same song and the distance between features of singing data related to different songs.

2. 2. The feature output model generation system according to claim 1, wherein the feature output model generation unit performs machine learning so that the distance between features of singing data relating to the same song is shorter than the distance between features of singing data relating to different songs.

3. 3. The feature output model generation system according to claim 1, wherein the feature output model generation unit determines a section of the singing data to be used for machine learning based on a distance between features output by the feature output model in the process of generation.

4. 4. The feature output model generation system according to claim 1, wherein the singing data acquisition unit acquires singing data including data indicating a time-series pitch length.

5. 5. The feature output model generation system according to claim 4, wherein the feature output model generation unit converts data indicating pitch lengths included in the singing data divided by the division unit into words, which are character strings corresponding to the pitch lengths for each consecutive identical pitch, and generates a feature output model that inputs information based on the converted words.

Citation Information

Patent Citations

  • USING A SYSTEM FOR PREDICTING MUSIC PREFERENCES TO DELIVER MUSIC CONTENT THROUGH A CELLULAR NETWORK

    JP2004510176A

  • Karaoke system

    JP2012078387A

  • Music selection and organization using rhythm, texture and pitch

    US20150220633A1

  • Musical piece recommendation model generation system and musical piece recommendation system

    WO2021005984A1

  • Musical composition recommendation model generation system, and musical composition recommendation system

    WO2021005985A1