Recommended Information Providing Device

The recommended information providing device addresses the limitation of conventional karaoke systems by employing a deep learning model to predict the next song based on user history, offering personalized and relevant recommendations that align with the user's singing order tendencies.

JP7685996B2Active Publication Date: 2025-05-30NTT DOCOMO INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022531652
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-17
Filing Date
2021-06-03
Publication Date
2025-05-30
Estimated Expiration
2041-06-03

AI Technical Summary

Technical Problem

Conventional karaoke devices do not consider the sequence or order of songs sung by users when recommending songs, leading to a lack of personalized and relevant recommendations.

Method used

A recommended information providing device that uses a deep learning model for multi-task learning to predict the next song a user will sing based on their past singing history, incorporating music and singer metadata to improve prediction accuracy.

Benefits of technology

The device effectively provides personalized song recommendations by grasping the singing order tendency of users, enhancing the user experience by suggesting songs that align with their sequential preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007685996000001
    Figure 0007685996000001
  • Figure 0007685996000002
    Figure 0007685996000002
  • Figure 0007685996000003
    Figure 0007685996000003
Patent Text Reader

Abstract

The purpose of the present invention is to provide useful information to a user by grasping a tendency in the order of songs to be sung. A recommendation information providing device 5: acquires song ID identifying each of a plurality of songs that a user sang in order in the past; creates a learning model M, which at least predicts the next song and a song after the next song the user is going to sing following an arbitrary song, from the song ID related to the arbitrary song, using the song ID related to the plurality of songs as training data; inputs the song ID related to a song to be sung by the user into the learning model M; and outputs recommendation information about songs recommended for the user to sing, on the basis of the next song predicted by the learning model M. The learning model M is a deep learning model of multi-task learning including a fully connected layer M31 which receives the song ID, and outputs an output value for predicting the next song, and a fully connected layer M32 which outputs an output value for predicting the song after the next song. An output of the fully connected layer M31 is connected to an input of the fully connected layer M32.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to a recommended information providing device that provides recommended information.

Background Art

[0002] Conventionally, in a karaoke device, a technique for presenting information for recommending a song suitable for a user's preference from among a plurality of songs to the user is known (see Patent Document 1 below). This device identifies the category of songs preferred by the user based on the user's singing history, extracts songs belonging to the identified category from all the songs sung by the preferred singer, and presents the extracted songs as recommended songs.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] With the above conventional device, recommended songs are presented based on the category preferred by the user, and the tendency of the order of the songs sung by the user is not considered. Therefore, it is desired to provide useful recommended information to the user by grasping the tendency of the order of the songs to be sung and predicting songs that match the tendency.

[0005] Therefore, in order to solve the above problems, an object is to provide a recommended information providing device capable of providing useful recommended information to the user by grasping the tendency of the order of the songs to be sung.

Means for Solving the Problems

[0006] The recommended information providing device of this embodiment is a recommended information providing device that provides recommended information, includes at least one processor, and the at least one processor acquires music identification information for identifying each of a plurality of pieces of music that the user has sung in order in the past as information capable of discriminating the singing order. Using the music identification information regarding the plurality of pieces of music as training data, a learning model is constructed to at least predict the piece of music one ahead that the user will sing next to an arbitrary piece of music and the piece of music two ahead that the user will sing next to the piece of music one ahead. The music identification information regarding the piece of music that is the singing target of the user is input into the learning model, and based on the piece of music one ahead predicted by the learning model, recommended information regarding the piece of music recommended for the user to sing is output. The learning model is a deep learning model of multi-task learning including a first fully connected layer that outputs an output value for predicting the piece of music one ahead when the music identification information is input, and a second fully connected layer that outputs an output value for predicting the piece of music two ahead, and the output of the first fully connected layer is connected to the input of the second fully connected layer.

[0007] According to this embodiment, music identification information regarding a plurality of pieces of music that the user has sung in the past is acquired as information capable of discriminating the singing order, and using that information as training data, a learning model is constructed to predict the piece of music one ahead that the user will sing next to an arbitrary piece of music and the piece of music two ahead that the user will sing next to that piece of music. Then, using the constructed learning model, the piece of music one ahead that the user will sing next to the target piece of music is predicted, and recommended information is output based on the predicted piece of music one ahead. As a result, after grasping the tendency of the pieces of music that the user sings continuously, the piece of music that the user will sing next to an arbitrary piece of music is predicted, and recommended information is output based on the prediction result, so that useful recommended information can be provided to the user. In particular, as the learning model, a deep learning model of multi-task learning including two fully connected layers whose inputs and outputs are connected to each other is used, so that by adjusting the parameters during the construction of the learning model, the accuracy of predicting the piece of music one ahead that the user will sing next to the target piece of music can be efficiently improved, and more useful information can be provided to the user. Note that the "user" referred to here is not limited to a single user, but also means a group including a plurality of users who sing at the same time.

Advantages of the Invention

[0008] According to one aspect of the present invention, useful recommendation information can be provided to a user by grasping the tendency of the order of the songs to be sung.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Modes for Carrying Out the Invention

[0010] The embodiments of the present invention will be described with reference to the accompanying drawings. Where possible, the same parts are denoted by the same reference numerals, and redundant descriptions are omitted.

[0011] FIG. 1 is a system configuration diagram showing the configuration of a karaoke system 1 according to the present embodiment. The karaoke system 1 is a device having a function of playing a music piece designated by a user, and a function of collecting singing voice of the user corresponding to the play and outputting the collected voice to a speaker or the like simultaneously with the playback voice of the music piece. The karaoke system 1 further has a function of providing recommendation information regarding a music piece recommended for singing next to the music piece designated immediately before to the user.

[0012] As shown in FIG. 1, the karaoke system 1 includes a karaoke device 2, a front server 3, a data management device 4, and a recommendation information providing device 5. The front server 3, the data management device 4, and the recommendation information providing device 5 are configured to be able to transmit and receive data to and from each other via a communication network such as a LAN (Local Area Network), a WAN (Wide Area Network), and a mobile communication network.

[0013] The karaoke device 2 provides a function of playing a music piece, a function of collecting singing voice of a user, and a function of outputting to a speaker or the like. The front server 3 is electrically connected to the karaoke device 2, and has a playback function of providing playback data for playing a music piece designated by a user to the karaoke device 2, a search function of a music piece according to a user operation, a designation function of a music piece to be played according to a user operation, a function of receiving information of a music piece sung by a user using the karaoke device 2 in response to the playback of the music piece and recording history information of the sung music piece, and the like. The front server 3 receives a user operation, provides a user interface for displaying information to the user, and includes a user terminal device (not shown) connected to the front server 3 by wire or wirelessly.

[0014] The data management device 4 is a data storage device (database device) that stores data processed by the front server 3 and the recommended information providing device 5. This data management device 4 includes a history information storage unit 101 that stores history information recording the history of a user's past singing of songs using the karaoke device 2, and a music information storage unit 102 that stores music information regarding songs that can be reproduced by the karaoke device 2 and singer information regarding the singers who provide those songs. Various information stored in the data management device 4 is updated at any time by the processing of the front server 3 or data obtained from the outside.

[0015] FIG. 2 shows an example of the data configuration of the history information stored in the data management device 4, and FIGS. 3, 4, and 5 show examples of the data configuration of the music information and singer information stored in the data management device 4.

[0016] As shown in FIG. 2, in one data record of the history information, a "session flag" that identifies a session in which a plurality of users sang songs together in a group using the karaoke device 2, a "terminal ID" that identifies the user terminal device used by the users in that group, a "singing time" indicating the date and time when the users in that group sang using the karaoke device 2, and a "music ID" (music identification information) that uniquely identifies the song sung are associated with each other. In the history information having such a configuration, a data record is accumulated and recorded each time a group of users specifies and sings a song using the karaoke device 2, and information on a plurality of songs sung in sequence by a group of users in one session in the past is recorded in a manner that allows the order of singing to be determined. The "session flag" that identifies a session is generated when the front server 3 determines that songs have been reproduced simultaneously using the same user terminal device.

[0017] FIG. 3 shows an example of the data configuration of music management information among music information. As such, in one data record of the music management information, a "music ID" (music identification information) for identifying a music playable using the karaoke system 1, a "singer ID" for identifying the singer who provides the music, a "singer name" indicating the name of the singer, and a "music name" indicating the name of the music are associated and stored. In the data management device 4, data records corresponding to each music playable by the front server 3 are stored including in the music management information.

[0018] FIG. 4 shows an example of the data configuration of music metadata recording the attributes of music among music information. As such, in one data record of the music metadata, a "music ID" (music identification information) for uniquely identifying a music playable using the karaoke system 1, a "lyricist" indicating the lyricist of the music, a "composer" indicating the composer of the music, a flag of "anime" indicating whether the music is related to an anime, a flag of "drama" indicating whether the music is related to a drama, and a flag of "western music" indicating whether the music is western music are associated and stored. In the data management device 4, data records corresponding to each music playable by the front server 3 are stored including in the music metadata.

[0019] FIG. 5 shows an example of the data configuration of singer metadata recording the attributes of a singer among music information. As such, in one data record of the singer metadata, a "singer ID" for uniquely identifying the singer of a music playable using the karaoke system 1, a "generation" indicating the generation of the singer, a flag of "gender" indicating the gender of the singer, a flag of "group" indicating whether the singer is a group, and a "release number" indicating the number of music released by the singer are associated and stored. In the data management device 4, data records corresponding to each singer of the music playable by the front server 3 are stored including in the singer metadata.

[0020] The recommended information providing device 5 is a device that provides recommended information regarding music to be recommended for singing to a group of users, and includes, as functional components, a data acquisition unit 201, a model construction unit 202, a prediction unit 203, and a recommended information generation unit 204. Hereinafter, the functions of each component will be described.

[0021] The data acquisition unit 201 acquires history information and music information from the data management device 4 prior to the construction process of the learning model for predicting the music to be recommended. Also, the data acquisition unit 201 also acquires history information and music information prior to the generation process of the recommended information. The data acquisition unit 201 delivers the acquired information to the model construction unit 202 or the prediction unit 203.

[0022] That is, at the time of the construction process of the learning model, the data acquisition unit 201 combines the information read from the history information storage unit 101 and the music information storage unit 102 of the data management device 4 to generate information regarding a plurality of music that the user group has sung in order in the past. Specifically, the data acquisition unit 201 reads out the data records of the history information having the same "session flag" in the order of singing indicated by the "singing time". Then, the data record of the music metadata corresponding to the "music ID" in the read data record and the data record of the singer metadata including the "singer ID" corresponding to the singer of the music indicated by the "music ID" are read out in order. The data acquisition unit 201 delivers the read data records of the music metadata and the data records of the singer metadata to the model construction unit 202 in order.

[0023] Also, at the time of the generation process of the recommended information, the data acquisition unit 201 reads out the "music ID" regarding the music to be predicted that the user group has sung immediately before from the history information, and acquires the data record of the music metadata regarding the music and the data record of the singer metadata regarding the singer of the music. The data acquisition unit 201 delivers the acquired data records to the prediction unit 203.

[0024] Returning to FIG. 1, the model construction unit 202 sequentially acquires the two data records read by the data acquisition unit 201, and uses the data records related to the music corresponding to the same "session flag" as training data. Based on the "music ID" related to any music, the machine learning learning model is constructed to predict the music that the user group will sing next after any music, and the music that the user group will sing second, ..., Nth (N is an integer greater than or equal to 2) after any music. In this embodiment, N = 3, but any integer can be selected as long as N is 2 or more.

[0025] Also, prior to constructing the learning model, the model construction unit 202 executes preprocessing to generate training data by processing the data records passed from the data acquisition unit 201. Specifically, the model construction unit 202 vectorizes the information of each item of the data record and converts it into a multi-dimensional vector, and concatenates the vectors of the information of each item horizontally to generate training data (concatenated vector). FIG. 6 shows an example of the vector of music metadata converted by the model construction unit 202, and FIG. 7 shows an example of the vector of singer metadata converted by the model construction unit 202. As shown in FIG. 6, for example, "music ID: AAA" in the data record of music metadata is converted into a 5-dimensional vector (+0.5, ..., +0.8, ...) and concatenated with the vectors of the information of other items to generate a concatenated vector. Also, as shown in FIG. 7, for example, "singer ID: aa" in the data record of singer metadata is converted into a 5-dimensional vector (... +0.4, ..., +0.8, ...) and concatenated with the vectors of the information of other items to generate a concatenated vector.

[0026] Then, the model construction unit 202 inputs a combination of the concatenated vector of the music metadata and the concatenated vector of the singer metadata into the learning model in the singing order specified by the corresponding history information, and optimizes the parameters of the learning model (trains the learning model) so that the prediction result of the learning model becomes the music from one song ahead to N songs ahead and the singer of the music from one song ahead to N songs ahead, which are specified by the history information and the music information. At this time, the model construction unit 202 uses a learning model of deep learning as the learning model.

[0027] FIG. 8 shows the configuration of the learning model M used by the model construction unit 202. As shown in FIG. 8, the learning model M includes an RNN (Recurrent Neural Network) layer M11 into which a concatenated vector of music metadata (music information) is input, an RNN layer M12 into which a concatenated vector of singer metadata (singer information) is input, a combined layer M20 in which the output layers of the RNN layers M11 and M12 are combined, and fully connected layers M31, M32, M33, M41, M42, M43 that are connected to the subsequent stage of the combined layer M20 and output prediction values. The number of fully connected layers can be changed according to how many songs ahead of the music to be predicted corresponding to the input concatenated vector are predicted. In this embodiment, it has six fully connected layers to predict up to three songs ahead.

[0028] The RNN layers M11 and M12 are each layers having two types of inputs and two types of outputs. Each of the RNN layers M11 and M12 has, as an input layer, a layer for inputting the concatenated vector in the current calculation step and a layer for inputting the output vector stacked by the RNN layers M11 and M12 in the previous calculation step, and, as an output layer, a layer for passing the output vector to the combined layer 20 and a layer for stacking the output vector for the next calculation step. These RNN layers M11 and M12 can combine and output information considering the present and the past by adding the layer stacked in the previous calculation to the layer calculated from the concatenated vector in the current calculation.

[0029] The combination layer M20 is a layer that horizontally concatenates the output layers of the RNN layer M11 and the RNN layer M12 to output a vector.

[0030] The fully-connected layers M31, M32, and M33 are respectively connected to the subsequent stage of the combination layer M20, and output predicted values (output vectors) of the songs sung by the user group for one song ahead, two songs ahead, and three songs ahead of the input concatenated vector. The fully-connected layer M31 has the same length as the number of candidate songs to be recommended, performs parameter calculations for the number obtained by multiplying the length of the combination layer M20 and the length of the fully-connected layer M31, and outputs an output vector obtained by horizontally concatenating the probability values for the candidate songs. Also, the fully-connected layer M32 has the same length as the fully-connected layer M31, its input is also connected to the output of the fully-connected layer M31, performs parameter calculations for the number obtained by adding the length of the combination layer M20 and the length of the fully-connected layer M31 and multiplying by the length of the fully-connected layer M32, and outputs an output vector obtained by horizontally concatenating the probability values for the candidate songs. Similarly, the fully-connected layer M33 has the same length as the fully-connected layer M31, its input is also connected to the output of the fully-connected layer M32, performs parameter calculations for the number obtained by adding the length of the combination layer M20 and the length of the fully-connected layer M32 and multiplying by the length of the fully-connected layer M33, and outputs an output vector obtained by horizontally concatenating the probability values for the candidate songs.

[0031] The fully-connected layers M41, M42, and M43 are each connected to the subsequent stage of the connection layer M20, and output predicted values (output vectors) of the singers of the songs sung by the user group for one song ahead, two songs ahead, and three songs ahead of the input concatenated vector. The fully-connected layer M41 has the same length as the number of candidate singers to be recommended, performs parameter calculation for the number obtained by multiplying the length of the connection layer M20 and the length of the fully-connected layer M41, and outputs an output vector in which the probability values for the candidate singers are concatenated horizontally. Also, the fully-connected layer M42 has the same length as the fully-connected layer M41, its input is also connected to the output of the fully-connected layer M41, performs parameter calculation for the number obtained by adding the length of the connection layer M20 and the length of the fully-connected layer M41 and multiplying by the length of the fully-connected layer M42, and outputs an output vector in which the probability values for the candidate singers are concatenated horizontally. Similarly, the fully-connected layer M43 has the same length as the fully-connected layer M41, its input is also connected to the output of the fully-connected layer M42, performs parameter calculation for the number obtained by adding the length of the connection layer M20 and the length of the fully-connected layer M42 and multiplying by the length of the fully-connected layer M43, and outputs an output vector in which the probability values for the candidate singers are concatenated horizontally.

[0032] The learning model M with the above configuration is a multi-task learning deep learning model that simultaneously solves multiple tasks of predicting the songs sung one song ahead, two songs ahead, and three songs ahead, and the singers of the songs sung one song ahead, two songs ahead, and three songs ahead by one model. According to such a model, it is possible to prevent useful information for predicting the song one song ahead and the singer of the song one song ahead, which are the main tasks, from falling out of the feature amounts. In other words, by sharing the intermediate layer in multiple tasks, it is possible to learn features common to multiple problems.

[0033] The model construction unit 202 uses the learning model M with the above configuration to simultaneously input the concatenated vector of music metadata and the connection vector of singer metadata into the learning model M, and the music one song ahead, two songs ahead, and three songs ahead predicted by the obtained output vector respectively match the music one song ahead, two songs ahead, and three songs ahead specified from the data records sequentially delivered from the data acquisition unit 201. The learning model M is trained. At this time, the music to be predicted is the music corresponding to the element with the largest probability value in the output vector. At the same time, the model construction unit 202 trains the learning model M so that each of the singers of the music one song ahead, two songs ahead, and three songs ahead predicted by the output vector of the learning model M is the singer of the music one song ahead, two songs ahead, and three songs ahead specified from the data records sequentially delivered from the data acquisition unit 201. At this time, the singer to be predicted is the singer corresponding to the element with the largest probability value in the output vector. As a result of the training, for example, the parameters of the weights (w) and biases (b) used in the parameter calculation in each layer of the learning model M are optimized.

[0034] In this embodiment, a deep learning model for multi-task learning is used, and training is performed to predict the music and singer several songs ahead together so as to obtain the correct answer, thereby improving the prediction accuracy of the music one song ahead and the singer of that music. That is, in the training of deep learning, the parameters closer to the output part are preferentially adjusted by the error backpropagation method. Therefore, in order to improve the prediction accuracy of the music two songs ahead, the parameters (w, b) in the fully connected layer M32 are adjusted so as to reduce the error of the prediction result by repeating the training, and the prediction accuracy is improved. At this time, if the error does not converge even after adjusting the parameters of the fully connected layer M32, the parameters at a location far from the output part are adjusted. For example, the parameters of the fully connected layer M31 are adjusted. The adjustment of the parameters of the fully connected layer M31 to improve the prediction accuracy of the music two songs ahead also affects the prediction result of the music one song ahead, and as a result, this adjustment also contributes to the improvement of the prediction accuracy of the music one song ahead. Similarly, the training effect for improving the prediction accuracy of the music three songs ahead, the singer two songs ahead, and the singer three songs ahead also contributes to the improvement of the prediction accuracy of the music one song ahead and the singer of the music one song ahead.

[0035] During the generation process of the recommendation information, the prediction unit 203 uses the learning model M constructed by the model construction unit 202 based on the data record of the music metadata regarding the music to be predicted and the data record of the singer metadata regarding the singer of that music, and obtains prediction values regarding the music one song ahead of the music to be predicted and the singer of the music one song ahead of the music to be predicted. Specifically, the prediction unit 203 performs the same preprocessing as the model construction unit 202 on the two data records to generate two concatenated vectors. Then, the prediction unit 203 obtains prediction values of the music one song ahead and the singer of the music one song ahead based on the output vector obtained by inputting the two generated concatenated vectors into the learning model M.

[0036] The recommended information generation unit 204 acquires, from the prediction unit 203, predicted values regarding the music one piece ahead and predicted values regarding the singer of the music one piece ahead, and based on these predicted values, generates and outputs recommended information on the music and the singer of the music to be recommended for singing to the user group. For example, the recommended information generation unit 204 generates recommended information regarding the music and singer corresponding to the element with the highest probability value among the respective output vectors, and the music and singer corresponding to the element with the second highest probability value among the respective output vectors. FIG. 9 shows an example of the data configuration of the recommended information generated by the recommended information generation unit 204. For example, the recommended information includes data in which the music "Song Z" and the singer "Singer Z" recommended first, and the music "Song B" and the singer "Singer B" recommended second are arranged in order. Then, the recommended information generation unit 204 outputs the generated recommended information to the terminal device of the front server 3 or the like.

[0037] Next, the processing of the recommended information providing apparatus 5 configured as described above will be described. FIG. 10 is a flowchart showing the procedure of the learning model construction process by the recommended information providing apparatus 5, and FIG. 11 is a flowchart showing the procedure of the recommended information generation process by the recommended information providing apparatus 5. The learning model construction process is started at a preset timing (for example, a regular timing), or at a timing when a certain amount of historical information is accumulated in the data management apparatus 4 or the like. The recommended information generation process is started at a preset timing, or at a timing when a music designation input is received from the user in the front server 3 or the like.

[0038] Referring to FIG. 10, when the learning model construction process is started, the data acquisition unit 201 sequentially acquires data records of historical information regarding the past music singing of the user group from the data management apparatus 4 in the order of singing (step S101). Also, the data acquisition unit 201 sequentially acquires data records of music metadata corresponding to the music recorded in the historical information and data records of singer metadata corresponding to the music from the data management apparatus 4 (step S102).

[0039] Next, preprocessing is executed by the model construction unit 202, and a pair of two data records is converted into two concatenated vectors in the order of acquisition (step S103). Then, the learning model M is trained by the model construction unit 202 so that the predicted value using the two concatenated vectors approaches the correct answer indicated by the history information and the music information, whereby the parameters of the learning model M are optimized (construction of the learning model, step S104). After being trained using a plurality of pairs of data records, the construction process of the learning model ends.

[0040] Next, referring to FIG. 11, when the generation process of the recommendation information is started, the data acquisition unit 201 acquires information for specifying a target music piece from the front server 3, and accordingly, from the data management device 4, data records of music metadata and data records of singer metadata regarding the target music piece are acquired (step S201). Then, preprocessing is executed by the prediction unit 203, and the two data records are each converted into a concatenated vector (step S202).

[0041] Next, the two concatenated vectors are input into the learning model M by the prediction unit 203, and based on the output vector of the learning model M, a predicted value regarding the music piece one ahead of the target music piece and a predicted value regarding the singer of the music piece one ahead of the target music piece are acquired (step S203). Then, based on these predicted values, the recommendation information generation unit 204 generates recommendation information regarding the music piece recommended for singing to the user group and the singer of that music piece (step S204). Finally, the recommendation information generation unit 204 outputs the recommendation information to the terminal device of the front server 3 or the like (step S205).

[0042] Next, the operation and effect of the recommended information providing apparatus 5 of the present embodiment will be described. According to this recommended information providing apparatus 5, music identification information regarding a plurality of pieces of music sung by a user in the past is acquired as information capable of discriminating the order in which they were sung, and a learning model M is constructed using this information as training data to predict the piece of music one piece ahead that the user will sing next to an arbitrary piece of music, and the piece of music two pieces ahead that the user will sing next to that. Then, using the constructed learning model M, the piece of music one piece ahead that the user will sing next to the target piece of music is predicted, and recommended information is output based on the predicted piece of music one piece ahead. As a result, after grasping the tendency of the pieces of music that the user sings continuously, the piece of music that the user will sing next to an arbitrary piece of music is predicted, and recommended information is output based on the prediction result, so that useful recommended information can be provided to the user. In particular, as the learning model M, since a deep learning model of multi-task learning including two fully connected layers whose inputs and outputs are connected to each other is used, the accuracy of predicting the piece of music one piece ahead that the user will sing next to the target piece of music can be efficiently improved by adjusting the parameters when constructing the learning model, and more useful information can be provided to the user.

[0043] Also, in the present embodiment, the learning model M is constructed such that each of the piece of music one piece ahead and the piece of music two pieces ahead predicted by inputting the music ID included in the history information is the piece of music one piece ahead next to the piece of music indicated by the music ID, and the piece of music two pieces ahead next to the piece of music one piece ahead, which are specified by the history information. By doing so, it is possible to construct the learning model M that grasps the tendency of three pieces of music that the user sings continuously, and it is possible to surely improve the prediction accuracy of the piece of music that the user will sing next to the target piece of music. As a result, more useful information can be provided to the user.

[0044] Also, in the present embodiment, the learning model M includes RNN layers M11 and M12 between the input and the fully connected layers M31, M32, M33, M41, M42, M43. According to such a learning model M, since the piece of music that the user will sing afterwards is predicted based on the information of two pieces of music that the user sings continuously, the prediction accuracy can be further improved. Therefore, more useful information can be provided to the user.

[0045] In addition, in this embodiment, a learning model M that predicts music three songs ahead to one song ahead from the music ID of any music is adopted. In the fully connected layers M32 and M33 of the learning model M, the respective inputs are connected to the outputs of the fully connected layers M31 and M32. With such a configuration, it is possible to construct the learning model M that grasps the tendency of four consecutive songs sung by the user, and it is possible to surely improve the prediction accuracy of the song to be sung next to the target song of the user. As a result, more beneficial information can be provided to the user.

[0046] Note that the block diagram used in the description of the above embodiment shows blocks of functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Also, the realization method of each functional block is not particularly limited. That is, each functional block may be realized using one physically or logically combined device, or two or more physically or logically separated devices may be directly or indirectly (for example, using wired, wireless, etc.) connected and realized using these multiple devices. The functional block may be realized by combining software with the above one device or the above multiple devices.

[0047] Functions include, but are not limited to, judgment, decision, determination, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, solution, selection, selection, establishment, comparison, assumption, expectation, regarded as, notification (broadcasting), notification (notifying), communication (communicating), transfer (forwarding), configuration (configuring), reconfiguration (reconfiguring), allocation (allocating, mapping), assignment (assigning), etc. For example, a functional block (component) that functions as transmission is called a transmission unit or a transmitter. In any case, as described above, the realization method is not particularly limited.

[0048] For example, the data management device 4 and the recommended information providing device 5 in one embodiment of the present disclosure may function as a computer that performs the processing of the present disclosure. FIG. 12 is a diagram showing an example of the hardware configuration of the data management device 4 and the recommended information providing device 5 according to one embodiment of the present disclosure. The above-described data management device 4 and recommended information providing device 5 may physically be configured as a computer device 100 including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.

[0049] In the following description, the term "device" can be read as a circuit, device, unit, etc. The hardware configuration of the data management device 4 and the recommended information providing device 5 may be configured to include one or more of each device shown in the figure, or may be configured without including some devices.

[0050] Each function in the data management device 4 and the recommended information providing device 5 is realized by causing a processor 1001 to perform calculations by loading a predetermined software (program) onto hardware such as the processor 1001 and the memory 1002, and controlling communication by the communication device 1004, or controlling at least one of reading and writing data in the memory 1002 and the storage 1003.

[0051] The processor 1001 controls the entire computer by operating an operating system, for example. The processor 1001 may be configured by a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic device, a register, and the like. For example, the above-described data acquisition unit 201, model construction unit 202, prediction unit 203, and recommended information generation unit 204 may be realized by the processor 1001.

[0052] Also, the processor 1001 reads a program (program code), software module, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002, and executes various processes according to these. As the program, a program that causes a computer to execute at least a part of the operations described in the above embodiments is used. For example, the data acquisition unit 201, the model construction unit 202, the prediction unit 203, and the recommended information generation unit 204 may be stored in the memory 1002 and realized by a control program operating in the processor 1001, and the same may be true for other functional blocks. Although it has been described that the above various processes are executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. Note that the program may be transmitted from a network via a telecommunication line.

[0053] The memory 1002 is a computer-readable recording medium and may be composed of at least one of, for example, ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. The memory 1002 may be referred to as a register, cache, main memory (main storage device), etc. The memory 1002 can store a program (program code), software module, etc. executable for performing the learning model construction process and the recommended information generation process according to an embodiment of the present disclosure.

[0054] Storage 1003 is a computer-readable recording medium, which may be composed of at least one of, for example, optical discs such as CD-ROM (Compact Disc ROM), hard disk drives, flexible disks, magneto-optical disks (e.g., compact discs, digital versatile discs, Blu-ray (registered trademark) discs), smart cards, flash memories (e.g., cards, sticks, key drives), floppy (registered trademark) disks, magnetic strips, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate media including at least one of memory 1002 and storage 1003.

[0055] Communication device 1004 is hardware (a transmission / reception device) for performing communication between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, a communication module, etc. Communication device 1004 may be configured to include, for example, a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. in order to implement at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, data acquisition unit 201 for receiving the above-mentioned information may be implemented by communication device 1004. This data acquisition unit 201 may be physically or logically separated into a transmission unit and a reception unit.

[0056] Input device 1005 is an input device for receiving external input (e.g., keyboard, mouse, microphone, switch, button, sensor, etc.). Output device 1006 is an output device for performing external output (e.g., display, speaker, LED lamp, etc.). For example, the above-mentioned recommended information generation unit 204, etc. may be implemented by output device 1006. Note that input device 1005 and output device 1006 may have an integrated configuration (e.g., a touch panel).

[0057] Also, each device such as the processor 1001 and the memory 1002 is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus or may be configured using different buses for each device.

[0058] Also, the data management device 4 and the recommended information providing device 5 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these hardware components.

[0059] The notification of information is not limited to the aspects / embodiments described in the present disclosure, and other methods may be used. For example, the notification of information may be implemented by physical layer signaling (e.g., downlink control information (DCI), uplink control information (UCI)), upper layer signaling (e.g., radio resource control (RRC) signaling, medium access control (MAC) signaling, notification information (master information block (MIB), system information block (SIB))), other signals, or combinations thereof. Also, the RRC signaling may be referred to as an RRC message and may be, for example, an RRC connection setup message, an RRC connection reconfiguration message, etc.

[0060] Each aspect / embodiment described in the present disclosure may be applied to at least one of systems using LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (new Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (Ultra-WideBand), Bluetooth (registered trademark), and other suitable systems, as well as next-generation systems extended based on these. Further, a plurality of systems may be combined and applied (for example, a combination of at least one of LTE and LTE-A and 5G, etc.).

[0061] The processing procedures, sequences, flowcharts, etc. of each aspect / embodiment described in the present disclosure may be rearranged as long as there is no contradiction. For example, for the methods described in the present disclosure, the elements of various steps are presented using an exemplary order and are not limited to the specific order presented.

[0062] Information, etc. may be output from an upper layer (or a lower layer) to a lower layer (or an upper layer). It may also be input and output via a plurality of network nodes.

[0063] The input and output information, etc. may be stored in a specific location (for example, a memory) or may be managed using a management table. The input and output information, etc. may be overwritten, updated, or appended. The output information, etc. may be deleted. The input information, etc. may be transmitted to other devices.

[0064] The determination may be made based on a value represented by 1 bit (either 0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0065] Each aspect / embodiment described in the present disclosure may be used alone, in combination, or switched for use during execution. Also, the notification of predetermined information (e.g., the notification of "being X") is not limited to being explicitly performed, and may be performed implicitly (e.g., by not performing the notification of the predetermined information).

[0066] As described in detail above regarding the present disclosure, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described in the present disclosure. The present disclosure can be implemented as modifications and variations without departing from the spirit and scope of the present disclosure defined by the claims. Therefore, the description of the present disclosure is for illustrative purposes and does not have any limiting meaning for the present disclosure.

[0067] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether called software, firmware, middleware, microcode, a hardware description language, or by any other name.

[0068] Also, software, instructions, information, etc. may be transmitted and received via a transmission medium. For example, when software is transmitted from a website, server, or other remote source using at least one of wired technologies (such as coaxial cables, optical fiber cables, twisted pairs, digital subscriber lines (DSL), etc.) and wireless technologies (such as infrared rays, microwaves, etc.), at least one of these wired and wireless technologies is included within the definition of the transmission medium.

[0069] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc., which may be referred to throughout the above description, may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0070] Note that for the terms described in this disclosure and the terms necessary for understanding this disclosure, they may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Also, a signal may be a message. Also, a component carrier (CC) may be referred to as a carrier frequency, a cell, a frequency carrier, etc.

[0071] The terms "system" and "network" used in this disclosure are used interchangeably.

[0072] Also, the information, parameters, etc. described in this disclosure may be represented using absolute values, relative values from a predetermined value, or relative to corresponding other information. For example, a radio resource may be indicated by an index.

[0073] The names used for the above-mentioned parameters are not limiting in any way. Furthermore, the mathematical formulas and the like using these parameters may be different from those explicitly disclosed in the present disclosure. Since various channels (e.g., PUCCH, PDCCH, etc.) and information elements can be identified by any suitable names, the various names assigned to these various channels and information elements are not limiting in any way.

[0074] The terms "determining" and "deciding" used in the present disclosure may include a wide variety of operations. "Determining" and "deciding" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up (searching, inquiring) (e.g., searching in a table, database, or another data structure), and considering something as having "determined" or "decided" what has been ascertained. Also, "determining" and "deciding" may include considering something as having "determined" or "decided" what has been received (e.g., receiving information), transmitted (e.g., transmitting information), input, output, accessed (e.g., accessing data in a memory). Further, "determining" and "deciding" may include considering something as having "determined" or "decided" what has been resolved, selected, chosen, established, compared, etc. That is, "determining" and "deciding" may include considering that some operation has been "determined" or "decided". Also, "determining (deciding)" may be read as "assuming", "expecting", "considering", etc.

[0075] The terms "connected" and "coupled," or any variations thereof, mean any direct or indirect connection or coupling between two or more elements, and can include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements can be physical, logical, or a combination thereof. For example, "connected" may be read as "accessed." As used in this disclosure, two elements can be considered to be "connected" or "coupled" to each other using at least one of one or more wires, cables, and printed electrical connections, and also, by way of some non-limiting and non-exhaustive examples, electromagnetic energy having wavelengths in the radio frequency region, microwave region, and optical (both visible and invisible) region, etc.

[0076] As used in this disclosure, the recitation "based on" does not mean "based only on" unless otherwise specified. In other words, the recitation "based on" means both "based only on" and "based at least in part on."

[0077] In this disclosure, when the terms "include," "including," and variations thereof are used, these terms are intended to be inclusive in the same manner as the term "comprising." Further, the term "or" as used in this disclosure is not intended to be exclusive.

[0078] In this disclosure, for example, when articles are added by translation, such as a, an, and the in English, this disclosure may include that the nouns following these articles are in the plural form.

[0079] In the present disclosure, the term "A and B are different" may mean that "A and B are different from each other". Note that the term may also mean that "A and B are each different from C". Terms such as "separate" and "coupled" may be interpreted in the same way as "different".

Industrial Applicability

[0080] One aspect of the present invention uses a recommendation information providing device that provides recommendation information, and is capable of providing useful recommendation information to a user by grasping a tendency of the order of songs to be sung.

Explanation of Signs

[0081] 5... Recommendation information providing device, 1001... Processor, 201... Data acquisition unit, 202... Model construction unit, 203... Prediction unit, 204... Recommendation information generation unit, M... Learning model, M11, M12... RNN layer, M31, M32, M33... Fully connected layer.

Claims

1. A recommended information providing device that provides recommended information, comprising at least one processor, wherein the at least one processor acquires, as information capable of discriminating the order of singing, music identification information for identifying each of a plurality of pieces of music that a user has sung in order in the past, using, as training data, a concatenated vector obtained by combining music attribute information and singer attribute information corresponding to each of the plurality of pieces of music with respect to the music identification information regarding the plurality of pieces of music, from a concatenated vector obtained by combining music attribute information and singer attribute information corresponding to an arbitrary piece of music with respect to the music identification information regarding the arbitrary piece of music, constructs a learning model for predicting at least a piece of music one piece ahead that the user will sing next to the arbitrary piece of music and a piece of music two pieces ahead that the user will sing next to the piece of music one piece ahead, inputs the concatenated vector regarding the piece of music to be sung by the user into the learning model, and outputs recommended information regarding a piece of music recommended for the user to sing based on the piece of music one piece ahead predicted by the learning model, the learning model is a deep learning model of multi-task learning including a first fully-connected layer that inputs the concatenated vector and outputs an output value for predicting the piece of music one piece ahead, and a second fully-connected layer that outputs an output value for predicting the piece of music two pieces ahead, the output of the first fully-connected layer is connected to the input of the second fully-connected layer, when constructing the learning model, the error backpropagation method is used, and when the prediction error of the second fully-connected layer does not converge as a result of adjusting the parameters of the second fully-connected layer, the parameters of the first fully-connected layer are adjusted, A recommended information providing device.

2. The at least one processor constructs the learning model such that each of the piece of music one piece ahead and the piece of music two pieces ahead predicted by inputting the concatenated vector included in the training data becomes the next piece of music of the piece of music indicated by the concatenated vector and the next two pieces of music of the next piece of music, which are specified by the training data, The recommended information providing device according to Claim 1.

3. The learning model includes a recurrent neural network (RNN) layer between the input and the first fully-connected layer and the second fully-connected layer, The recommended information providing device according to Claim 1 or 2.

4. The at least one processor Construct the learning model that further predicts, from the concatenated vector regarding any piece of music, the music three pieces ahead that the user will sing next after the music two pieces ahead. The learning model further includes a third fully-connected layer that outputs an output value for predicting the music three pieces ahead. The output of the second fully-connected layer is connected to the input of the third fully-connected layer. During the construction of the learning model, the error backpropagation method is used. When the prediction error of the third fully-connected layer does not converge as a result of adjusting the parameters of the second and third fully-connected layers, the parameters of the first fully-connected layer are adjusted. The recommended information providing apparatus according to any one of claims 1 to 3.

5. The learning model is a deep learning model of multi-task learning that further inputs the attribute information of the music specified by the music specifying information. The recommended information providing apparatus according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Communication karaoke (Orchestration without lyric) system and karaoke playing terminal

    JP1999052965A

  • Smart Audio Headphone System

    JP2018504719A

  • Karaoke device

    JP2019148769A