Recommended Information Providing Device

The recommended information providing device uses a learning model to predict scoring results from past singing data, addressing the challenge of providing suitable pitch settings for diverse music types by adjusting pitch settings based on user history, enhancing user-specific recommendations.

JP7714543B2Active Publication Date: 2025-07-29NTT DOCOMO INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022530471
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-09
Filing Date
2021-05-28
Publication Date
2025-07-29
Estimated Expiration
2041-05-28

AI Technical Summary

Technical Problem

Conventional karaoke devices struggle to provide recommended pitch settings for music with limited user singing history, making it difficult to match user preferences accurately.

Method used

A recommended information providing device that uses a learning model to predict scoring results based on past singing data, adjusting pitch settings to recommend suitable settings for various types of music by inputting pitch information into a constructed learning model and outputting recommended pitch settings based on predicted scoring results.

Benefits of technology

Enables the provision of recommended pitch settings tailored to individual user preferences, improving the accuracy of pitch recommendations for a wide variety of music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007714543000001
    Figure 0007714543000001
  • Figure 0007714543000002
    Figure 0007714543000002
  • Figure 0007714543000003
    Figure 0007714543000003
Patent Text Reader

Abstract

The purpose of the present invention is to provide recommendation information relating to setting appropriate for singing with respect to a wide variety of songs. A recommendation information provision device 5 is provided with at least one processor. The at least one processor acquires a scoring result relating to singing of each of songs by a user in the past with respect to each temporal section of the song, acquires pitch information indicating the pitches of sounds that constitute the song and are arranged on a time-series basis in the section, uses the scoring result and the pitch information as training data to build, from the pitch information, a learning model for predicting a scoring result relating to singing of a song by the user, acquires a scoring result relating to singing of a target song by the user on the basis of output of the learning model by inputting pitch information relating to the target song to the learning model while changing the pitches of sounds indicated by the pitch information into multiple kinds, and on the basis of the scoring results for multiple kinds of pitch information relating to the target song, outputs, as recommendation information, the contents of setting of the pitches of sounds to be recommended to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to a recommended information providing device that provides recommended information.

Background Art

[0002] Conventionally, in a karaoke device, every time a user sings, the user ID, the music ID, the scoring result, and the setting key information set in the karaoke device at the time of the user's singing are managed in association with each other. When the user makes a reservation to play a desired music, a technique of displaying information regarding a setting key having the highest average value of the scoring results in the setting key information on a display means is known (see Patent Document 1 below).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, depending on the above-described conventional device, since the setting information of the pitch of the recommended music is output using the history of the scoring results corresponding to the setting key information of the music that the user has already set, for a music with a small singing history of the user, it tends to be difficult for the user to obtain recommended information regarding the setting content of the recommended pitch. Therefore, conventionally, it has been desired to provide recommended information that matches the past singing tendency of the user for a wide variety of music.

[0005] Therefore, in order to solve the above problems, an object of the present invention is to provide a recommended information providing device that can provide recommended information regarding settings suitable for singing for a wide variety of music.

Means for Solving the Problems

[0006] The recommended information providing device according to this embodiment is a recommended information providing device that provides recommended information, includes at least one processor, and the at least one processor acquires a scoring result regarding the singing of the user's past music for each temporal section of the music, acquires pitch information indicating the pitch of the sounds that are the sounds constituting the music and are arranged in time series in the section, constructs a learning model that predicts the scoring result regarding the singing of the user's music from the pitch information using the scoring result and the pitch information as training data, inputs the pitch information regarding the target music into the learning model while changing the pitch of the sounds indicated by the pitch information into a plurality of types, acquires the scoring result regarding the singing of the user's target music based on the output of the learning model, and outputs, as recommended information, the setting content of the pitch to be recommended to the user based on the scoring results for a plurality of types of pitch information regarding the target music.

[0007] According to this embodiment, the scoring result for each section regarding the singing of the user's past music and the pitch information of the section are used as training data, and a learning model for predicting the scoring result is constructed. Then, the pitch information regarding the target music is input into the constructed learning model while the pitch of the sounds indicated by the pitch information is changed into a plurality of types, and based on the output thereof, the scoring result regarding the singing of the user's target music is acquired. Further, based on the scoring results for the pitch information changed into a plurality of types, recommended information regarding the setting content of the pitch is output. Thereby, based on the scoring tendency for the past pitch pattern of the user, it is possible to acquire the predicted values of the scoring results when the setting content of the pitch is changed in various ways when singing the target music. In addition, by outputting recommended information regarding the setting content of the pitch using those predicted values, it is possible to provide recommended information regarding settings suitable for singing for a wide variety of types of music.

Effect of the Invention

[0008] According to one aspect of the present invention, it is possible to provide recommended information regarding settings suitable for singing for a wide variety of types of music.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Mode for Carrying Out the Invention

[0010] Embodiments of the present invention will be described with reference to the accompanying drawings. When possible, the same parts are denoted by the same reference numerals, and redundant descriptions are omitted.

[0011] FIG. 1 is a system configuration diagram showing the configuration of a karaoke system 1 according to the present embodiment. The karaoke system 1 is a device having a known function of playing a music piece designated by a user, and a known function of collecting a singing voice of the user in response to the playback and evaluating and scoring the singing voice. The karaoke system 1 further has a function of providing recommendation information regarding a setting key for the pitch (scale) of the music piece to the user.

[0012] As shown in FIG. 1, the karaoke system 1 includes a karaoke device 2, a front server 3, a data management device 4, and a recommendation information providing device 5. The front server 3, the data management device 4, and the recommendation information providing device 5 are configured to be able to transmit and receive data to and from each other via a communication network such as a LAN (Local Area Network), a WAN (Wide Area Network), and a mobile communication network.

[0013] The karaoke device 2 provides a function for playing music and a function for collecting the singing voice of the user. The front server 3 is electrically connected to the karaoke device 2 and has a playback function for providing playback data for playing a music designated by the user to the karaoke device 2, a search function for music according to the user's operation, a scoring function for receiving data of the singing voice collected by the karaoke device 2 in response to the playback of the music and calculating a scoring result of the singing voice, etc. When providing the playback data of the music, the front server 3 has a function of providing playback data in which the pitch of the music is uniformly changed according to a setting key previously set by the user. For this setting key, for example, numerical values from -7 to +7 are assigned, and when the setting key increases by +1, the playback data is set so that the pitch of the music uniformly rises by a predetermined scale (for example, a semitone). In addition, the front server 3 also has a function of storing the scoring result of the singing voice by the user in the data management device 4 each time as history information. The front server 3 receives the operation of the user and provides a user interface for displaying information to the user, and includes a terminal device connected to the front server 3 by wire or wirelessly.

[0014] The data management device 4 is a data storage device (database device) for storing data processed by the front server 3 and the recommended information providing device 5. This data management device 4 includes a history information storage unit 101 for storing history information recording the scoring results regarding the singing of the user's past music, and a music information storage unit 102 for storing pitch information regarding the music playable on the karaoke device 2. Various information stored in the data management device 4 is updated at any time by the processing of the front server 3 or data acquired from the outside.

[0015] FIG. 2 shows an example of the data configuration of the history information stored in the data management device 4, and FIGS. 3 and 4 show an example of the data configuration of the music information stored in the data management device 4.

[0016] As shown in FIG. 2, the history information includes a "user identifier" for identifying a user, a "music identifier" for identifying a music that the user has sung in the past using the karaoke system 1, a "singing time" indicating the time when the music was sung in the past, a "total score" indicating a scoring result for the entire section of the music targeted by the function of the front server 3, a "section 1 score",..., a "section 24 score" indicating scoring results for each section of the music targeted by the function of the front server 3, which are associated and stored. In the scoring function of the front server 3, the time intervals of each music are divided into a predetermined number (for example, 24), the scoring results are calculated for each divided interval, and the overall scoring result "total score" of each music is calculated from the scoring results of all intervals. In the history information, the scoring results of each section and the overall scoring result calculated by the front server 3 are recorded for each user's singing of each music.

[0017] FIG. 3 shows an example of the data structure of the pitch information among the music information. Thus, the pitch information includes a "music identifier" for identifying a music that can be reproduced using the karaoke system 1, a "note start time (ms)" indicating the start time in the entire music of the sound (note) that constitutes the music, a "note end time (ms)" indicating the end time in the entire music of the sound, a "pitch" indicating the standard pitch (pitch) of the sound numerically, and a "strength" indicating the strength of the sound numerically, which are associated and stored. In the data management device 4, pitch information regarding all standard sounds (sounds before being changed by the setting key) arranged in time series in each music that can be reproduced by the front server 3 is stored.

[0018] FIG. 4 shows an example of the data structure of the section information among the music information. Thus, the section information includes a "music identifier" for identifying a music that can be reproduced using the karaoke system 1, a "section start time (ms)" indicating the start time in the entire music of the section of the music, and a "section end time (ms)" indicating the end time in the entire music of the section, which are associated and stored. In the data management device 4, section information regarding all sections that constitute each music that can be reproduced by the front server 3 is stored.

[0019] The recommended information providing device 5 is a device that provides recommended information regarding setting keys to a user, and includes, as functional components, a data acquisition unit 201, a model construction unit 202, a prediction unit 203, and a recommended information generation unit 204. Hereinafter, the functions of each component will be described.

[0020] The data acquisition unit 201 acquires history information and music information from the data management device 4 prior to the construction process of a learning model for predicting scoring results. Also, the data acquisition unit 201 also acquires music information prior to the prediction process of the scoring results. The data acquisition unit 201 delivers the acquired respective information to the model construction unit 202 or the prediction unit 203.

[0021] That is, at the time of the construction process of the learning model, the data acquisition unit 201 combines the information read from the history information storage unit 101 and the music information storage unit 102 of the data management device 4 to generate history information on the scoring results for each sound in each section of the music that the user has sung in the past. FIG. 5 shows an example of the data configuration of the history information generated by the data acquisition unit 201. As such, the history information is associated with a "user identifier" for identifying the user, a "music identifier" for identifying the music that the user has sung in the past, a "section" for identifying the section of the music, a "note start time (ms)" indicating the start time of the sound in that section, a "note end time (ms)" indicating the end time of the sound, a "pitch" numerically indicating the pitch of the sound, a "strength" numerically indicating the strength of the sound, and a "score" indicating the scoring result of the section in which the sound is included. The data acquisition unit 201 generates history information regarding all the sounds constituting each music that the user has sung in the past. Note that in the "pitch information" included in the history information, when the setting key has been changed from the standard key during the user's past singing, a numerical value corresponding to the changed pitch is recorded accordingly.

[0022] Also, during the prediction process of the scoring result, the data acquisition unit 201 acquires music information regarding the music to be predicted from the data management device 4. The data acquisition unit 201 delivers the acquired music information to the prediction unit 203.

[0023] Returning to FIG. 1, the model construction unit 202 constructs a machine learning model that predicts a scoring result regarding the singing of the target music by the user based on the pitch information of the target music, using the history information generated by the data acquisition unit 201 as training data. Prior to constructing the learning model, the model construction unit 202 executes preprocessing for processing the history information delivered from the data acquisition unit 201. Specifically, the model construction unit 202 converts each piece of information in the history information into a one-dimensional vector (sound vector) in which information on the pitch and intensity of each sound constituting the music sung by the user in the past is arranged. In addition, the model construction unit 202 converts each piece of information in the history information into a one-dimensional vector (score vector) in which the scoring results of the sections corresponding to each sound, which is a one-dimensional vector corresponding to the sound vector, are arranged, and a one-dimensional vector (user identification vector) in which user identification information regarding the singing user, which is a one-dimensional vector corresponding to the sound vector, is arranged.

[0024] FIG. 6 shows an example of the data configuration of the one-dimensional vector generated by the preprocessing of the model construction unit 202. In this way, the model construction unit 202 converts the history information into a sound vector V1, a score vector V2, and a user identification vector V3.

[0025] Then, the model construction unit 202 inputs the sound vector V1 and the user identification vector V3 into the learning model, and optimizes the parameters of the learning model (trains the learning model) so that the output result of the learning model approaches the score indicated by the score vector V2. At this time, the model construction unit 202 uses a deep learning model as the learning model.

[0026] FIG. 7 shows the configuration of the learning model M used by the model construction unit 202. As shown in FIG. 7, the learning model M is composed of a one-hot encoding unit M1, a GRU unit M2, a combining unit M3, and a dense unit M4.

[0027] The one-hot encoding unit M1 receives the user identification vector V3 and converts the user identification vector V3 into a two-dimensional vector. FIG. 8 shows an example of the data configuration of the two-dimensional vector converted by the one-hot encoding unit M1. Thus, in the two-dimensional vector, each row corresponds to the sound indicated by each element of the sound vector V1, and each column corresponds to each user indicated by each element of the user identification vector V3. For example, when one element "user identifier" in the user identification vector V3 is "A1", in the row corresponding to that element, the value of the column corresponding to "user identifier A1" is set to "1", and the values of the columns corresponding to other user identifiers are set to "0". The one-hot encoding unit M1 generates a two-dimensional vector for all the rows corresponding to all the elements included in the user identification vector V3.

[0028] The GRU unit M2 is a type of recurrent neural network (RNN). In addition to the normal output, it outputs a state. As inputs, in addition to the sound vector V1 as the normal input, the state output immediately before is input again. Thereby, the GRU unit M2 has a function of storing past input information and can process long-term time series information.

[0029] The combining unit M3 combines the output of the one-hot encoding unit M1 and the output of the GRU unit M2. The dense unit M4 is a fully-connected layer in deep learning, which multiplies the numerical sequence of a certain number of dimensions output from the combining unit M3 by weights (w) and adds a bias (b) to convert it into an output (Y) of an arbitrary number of dimensions. In this embodiment, the dense unit M4 converts it into a one-dimensional output vector Y in which the scoring results (scores) of each section of the music are arranged. FIG. 9 shows an example of the data configuration of the output vector converted by the dense unit M4. Thus, in the output vector (Y), each element represents the predicted value of the scoring result of each section composed of the sound corresponding to the element of the input sound vector V1.

[0030] The model construction unit 202 uses the learning model M with the above configuration to input the user identification vector V3 and the sound vector V1 into the learning model M, and trains the learning model M so that the resulting output vector (Y) approaches the scores of each section indicated by the score vector V2. As a result of the training, for example, the parameters of the weights (w) and the bias (b) in the dense unit M4 of the learning model M are optimized.

[0031] Returning to FIG. 1 again, the prediction unit 203 obtains the predicted values of the scoring results of each section regarding the singing of the target music of the user by using the learning model M constructed by the model construction unit 202 based on the music information regarding the target music. Specifically, the prediction unit 203 performs the same preprocessing as the model construction unit 202 on the music information to generate the sound vector V1 and the user identification vector V3 regarding the target music. Then, the prediction unit 203 obtains the predicted values of the scoring results of each section of the target music based on the output vector (Y) obtained by inputting the generated sound vector V1 and user identification vector V3 into the learning model M.

[0032] In this embodiment, the prediction unit 203 changes the numerical values of the pitch information for each section in the music information of the target music into multiple types, and based on the music information in which the pitch information is changed into multiple types, uses the learning model M to obtain the predicted values of the scoring results for each section. Specifically, in the music information of the target music, the prediction unit 203 uniformly increases or decreases the pitch information for all sections by a predetermined value from a standard pitch corresponding to the numerical value of the setting key set in the front server 3. For example, corresponding to the setting key "+1", the numerical value of the pitch information for all sections is set to increase by +1, and corresponding to the setting key "+2", the numerical value of the pitch information for all sections is set to increase by +2.

[0033] The recommended information generation unit 204 repeatedly obtains the predicted values of the scoring information for each section regarding the music in which the pitch information is changed into multiple types from the prediction unit 203, and calculates the predicted value of the overall scoring result for each music in which the pitch information is changed into multiple types. For example, as the predicted value of the overall scoring result, the average value of the predicted values of the scoring results for all sections is calculated. Then, the recommended information generation unit 204 selects the setting content (setting key) of the pitch to be recommended to the user based on the predicted values of the scoring results of the music in which the pitch information is changed into multiple types, and outputs the recommended information indicating the selected setting key together with the predicted value of the scoring result corresponding to the setting key.

[0034] For example, the recommended information generation unit 204 selects, as the setting key to be recommended to the user, those corresponding to the music with a relatively high predicted value of the scoring result, those corresponding to the music with a predicted value of the scoring result higher than a preset threshold, and the like. The recommended information and the information of the predicted value output by the recommended information generation unit 204 are output to the terminal device of the front server 3 and the like.

[0035] Next, the processing of the recommended information providing apparatus 5 configured as described above will be described. FIG. 10 is a flowchart showing the procedure of the learning model construction process by the recommended information providing apparatus 5, and FIG. 11 is a flowchart showing the procedure of the recommendation process regarding the setting key by the recommended information providing apparatus 5. The learning model construction process is started at a preset timing (for example, a regular timing), or at a timing when a certain amount of historical information is accumulated in the data management apparatus 4, etc. The recommendation process regarding the setting key is started at a preset timing, or at a timing when an instruction is received from the user in the front server 3, etc.

[0036] Referring to FIG. 10, when the learning model construction process is started, the data acquisition unit 201 acquires historical information regarding the scoring results of the user's past song singing from the data management apparatus 4 (step S101). Also, the data acquisition unit 201 acquires music information regarding the music recorded in the historical information from the data management apparatus 4 (step S102).

[0037] Next, preprocessing is executed by the model construction unit 202, and based on the historical information and the music information, a sound vector V1, a score vector V2, and a user identification vector V3 are generated (step S103). Thereafter, the model construction unit 202 trains the learning model M using the sound vector V1, the score vector V2, and the user identification vector V3, thereby optimizing the parameters of the learning model M (construction of the learning model, step S104), and the learning model construction process ends.

[0038] Next, referring to FIG. 11, when the recommendation process regarding the setting key is started, the data acquisition unit 201 acquires music information regarding the target music from the data management apparatus 4 (step S201). Thereafter, preprocessing is executed by the prediction unit 203, and based on the music information in which the pitch information is changed into a plurality of types, a sound vector V1 is generated, and a user identification vector V3 for identifying the user whose scoring result is to be predicted is generated for the elements corresponding to the sound vector V1 (step S202).

[0039] Next, the prediction unit 203 inputs the sound vector V1 and the user identification vector V3 into the learning model M, and based on the output vector of the learning model M, a predicted value of the scoring result for each section of the music with multiple types of setting keys changed is obtained (step S203). After that, based on the predicted values of the scoring results for each section of the music with multiple setting keys, the recommendation information generation unit 204 calculates the predicted value of the overall scoring result for each music with multiple setting keys (step S204). Finally, based on the predicted values of the scoring results for each music with multiple setting keys, the recommendation information generation unit 204 selects the setting key to be recommended to the user, and generates and outputs the recommendation information to the user (step S205).

[0040] FIG. 12 shows an example of the data configuration of the recommendation information output by the recommendation information providing apparatus 5. As described above, a plurality of records in which the item of "key setting content" indicating the type of the setting key and the item of "predicted score" indicating the predicted value of the overall scoring result are associated with each other are output. In the recommendation information having such a configuration, the recommended setting key is indicated by the "key setting content" corresponding to the "predicted score" indicating a relatively high numerical value.

[0041] Next, the operation and effect of the recommended information providing apparatus 5 of the present embodiment will be described. According to this recommended information providing apparatus 5, the scoring results for each section regarding the singing of the user's past music, and the pitch information of the section are used as training data, and a learning model M for predicting the scoring results is constructed. Then, the pitch information regarding the target music is input to the constructed learning model M while the pitch heights indicated by the pitch information are changed in multiple types, and based on the output thereof, the scoring results regarding the singing of the target music by the user are obtained. Further, based on the scoring results for the pitch information changed in multiple types, recommended information regarding the setting content of the pitch height is output. Thereby, based on the scoring tendency for the past pitch patterns of the user, it is possible to obtain the predicted values of the scoring results when the setting content of the pitch height is variously changed during the singing of the target music. In addition, by outputting the recommended information regarding the setting content of the pitch height using those predicted values, it is possible to provide the recommended information regarding the settings suitable for singing for a wide variety of music.

[0042] Also, in the present embodiment, a learning model M that takes the time-series pitch information as input and outputs the scoring results for each section of the music corresponding to the pitch information is used, and the learning model M is constructed so that the output of the learning model M approaches the scoring results for each section included in the training data. By doing so, it is possible to construct a learning model M that grasps the tendency of the scoring results for the pitch patterns of each section of the music, and it is possible to surely improve the prediction accuracy of the scoring results regarding the singing of the target music by the user. As a result, it is possible to provide recommended information suitable for the singing of the target music by the user.

[0043] Also, in the present embodiment, a learning model M that further inputs the identification information of the user is used. By doing so, it is possible to construct a learning model M that grasps the tendency of the scoring results for the pitch patterns of each user, and it is possible to surely improve the prediction accuracy of the scoring results for each individual user. As a result, it is possible to provide recommended information suitable for each individual user.

[0044] In addition, in the present embodiment, the scoring results for each section of the target music, which is the output of the learning model M, are averaged to obtain the scoring result regarding the singing of the target music by the user. In this way, it is possible to easily determine the user's proficiency in singing the target music.

[0045] Also, in the present embodiment, the pitch height indicated by the pitch information in all sections regarding the target music is uniformly changed by a predetermined value, and the pitch information is input to the learning model M. Based on the output of the learning model M, the scoring result regarding the singing of the target music by the user is obtained. With such a configuration, it is possible to maintain the prediction accuracy of the scoring result when the setting content of the pitch height is changed during the singing of the target music, and it is possible to provide useful recommendation information for the user during singing.

[0046] Note that the block diagrams used in the description of the above embodiment show the blocks of functional units. These functional blocks (components) are realized by an arbitrary combination of at least one of hardware and software. Also, the realization method of each functional block is not particularly limited. That is, each functional block may be realized using one physically or logically combined device, or may be realized using two or more physically or logically separated devices directly or indirectly (for example, using wired, wireless, etc.) connected, and these multiple devices. The functional block may be realized by combining software with the above one device or the above multiple devices.

[0047] Functions include, but are not limited to, judgment, decision-making, determination, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, solution, selection, selection determination, establishment, comparison, assumption, expectation, presumption, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), assigning, etc. For example, a functional block (component) that enables transmission is referred to as a transmitting unit or a transmitter. As described above, the implementation method is not particularly limited.

[0048] For example, the data management device 4 and the recommended information providing device 5 in one embodiment of the present disclosure may function as a computer that performs the processing of the present disclosure. FIG. 13 is a diagram showing an example of the hardware configuration of the data management device 4 and the recommended information providing device 5 according to one embodiment of the present disclosure. The above-described data management device 4 and recommended information providing device 5 may physically be configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.

[0049] In the following description, the term "device" can be read as a circuit, device, unit, etc. The hardware configuration of the data management device 4 and the recommended information providing device 5 may be configured to include one or more of each device shown in the figure, or may be configured without including some devices.

[0050] Each function in the data management device 4 and the recommended information providing device 5 is realized by causing the processor 1001 to perform operations by loading a predetermined software (program) onto hardware such as the processor 1001, the communication device 1004 to control communication, and at least one of reading and writing data in the memory 1002 and the storage 1003.

[0051] The processor 1001 controls the entire computer by operating, for example, an operating system. The processor 1001 may be constituted by a central processing unit (CPU: Central Processing Unit) including an interface with peripheral devices, a control device, an arithmetic device, registers, and the like. For example, the above-described data acquisition unit 201, model construction unit 202, prediction unit 203, recommendation information generation unit 204, and the like may be realized by the processor 1001.

[0052] Also, the processor 1001 reads a program (program code), software module, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002, and executes various processes according to these. As the program, a program for causing a computer to execute at least a part of the operations described in the above-described embodiments is used. For example, the data acquisition unit 201, model construction unit 202, prediction unit 203, and recommendation information generation unit 204 may be stored in the memory 1002 and realized by a control program operating in the processor 1001, and other functional blocks may be realized in the same manner. Although it has been described that the above-described various processes are executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. Note that the program may be transmitted from a network via a telecommunication line.

[0053] The memory 1002 is a computer-readable recording medium and may be constituted by at least one of, for example, ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. The memory 1002 may be called a register, a cache, a main memory (main storage device), etc. The memory 1002 can store a program (program code), a software module, etc. executable for carrying out the construction process and the recommendation process according to an embodiment of the present disclosure.

[0054] The storage 1003 is a computer-readable recording medium and may be constituted by at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (for example, a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (for example, a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. The storage 1003 may be called an auxiliary storage device. The above-described recording medium may be, for example, a database, a server, or other appropriate media including at least one of the memory 1002 and the storage 1003.

[0055] The communication device 1004 is hardware (a transceiver device) for performing communication between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, a communication module, etc. The communication device 1004 may include, for example, a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. in order to implement at least one of frequency-division duplexing (FDD) and time-division duplexing (TDD). For example, the data acquisition unit 201 that receives the above-described information may be implemented by the communication device 1004. This data acquisition unit 201 may be physically or logically separated into a transmission unit and a reception unit.

[0056] The input device 1005 is an input device (for example, a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives an external input. The output device 1006 is an output device (for example, a display, a speaker, an LED lamp, etc.) that performs an output to the outside. For example, the above-described recommended information generation unit 204, etc. may be implemented by the output device 1006. Note that the input device 1005 and the output device 1006 may have an integrated configuration (for example, a touch panel).

[0057] Also, each device such as the processor 1001 and the memory 1002 is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses for each device.

[0058] Further, the data management device 4 and the recommended information providing device 5 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these hardware components.

[0059] The notification of information is not limited to the aspects / embodiments described in the present disclosure, and other methods may be used. For example, the notification of information may be performed by physical layer signaling (e.g., downlink control information (DCI), uplink control information (UCI)), upper layer signaling (e.g., radio resource control (RRC) signaling, medium access control (MAC) signaling, notification information (master information block (MIB), system information block (SIB))), other signals, or a combination thereof. Further, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC connection setup message, an RRC connection reconfiguration message, or the like.

[0060] Each aspect / embodiment described in the present disclosure may be applied to at least one of systems using LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (new Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (Ultra-WideBand), Bluetooth (registered trademark), and other appropriate systems, as well as next-generation systems extended based on these. Also, multiple systems may be combined and applied (for example, a combination of at least one of LTE and LTE-A and 5G, etc.).

[0061] The processing procedures, sequences, flowcharts, etc. of each aspect / embodiment described in the present disclosure may be reordered as long as there is no contradiction. For example, for the methods described in the present disclosure, the elements of various steps are presented using an exemplary order and are not limited to the specific order presented.

[0062] Information, etc. may be output from an upper layer (or a lower layer) to a lower layer (or an upper layer). It may also be input and output via a plurality of network nodes.

[0063] The input and output information, etc. may be stored in a specific location (for example, a memory) or may be managed using a management table. The input and output information, etc. may be overwritten, updated, or appended. The output information, etc. may be deleted. The input information, etc. may be transmitted to other devices.

[0064] The determination may be made based on a value represented by 1 bit (either 0 or 1), may be made based on a boolean value (Boolean: true or false), or may be made based on a numerical comparison (e.g., comparison with a predetermined value).

[0065] Each aspect / embodiment described in the present disclosure may be used alone, may be used in combination, or may be switched and used during execution. Further, notification of predetermined information (e.g., notification of "being X") is not limited to being explicitly performed, and may be performed implicitly (e.g., by not performing the notification of the predetermined information).

[0066] As described above in detail regarding the present disclosure, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described in the present disclosure. The present disclosure can be implemented in modified and changed forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is for illustrative purposes and has no restrictive meaning for the present disclosure.

[0067] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether called software, firmware, middleware, microcode, a hardware description language, or by any other name.

[0068] Also, software, instructions, information, etc. may be transmitted and received via a transmission medium. For example, when software is transmitted from a website, server, or other remote source using at least one of wired technologies (such as coaxial cables, optical fiber cables, twisted pairs, digital subscriber line (DSL), etc.) and wireless technologies (such as infrared rays, microwaves, etc.), at least one of these wired and wireless technologies is included within the definition of the transmission medium.

[0069] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc., which may be referred to throughout the above description, may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0070] In addition, for the terms described in this disclosure and the terms necessary for understanding this disclosure, they may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Also, a signal may be a message. Also, a component carrier (CC) may be referred to as a carrier frequency, a cell, a frequency carrier, etc.

[0071] The terms "system" and "network" used in this disclosure are used interchangeably.

[0072] Also, the information, parameters, etc. described in this disclosure may be represented using absolute values, relative values from a predetermined value, or corresponding other information. For example, a radio resource may be indicated by an index.

[0073] The names used for the above-described parameters are not limiting in any way. Further, mathematical formulas and the like using these parameters may be different from those explicitly disclosed in this disclosure. Since various channels (e.g., PUCCH, PDCCH, etc.) and information elements can be identified by any suitable names, the various names assigned to these various channels and information elements are not limiting in any way.

[0074] The terms "determining" and "deciding" used in this disclosure may encompass a wide variety of operations. "Determining" and "deciding" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up (searching, inquiring) (e.g., searching in a table, database, or another data structure), and considering something as having "determined" or "decided" what has been ascertained. Also, "determining" and "deciding" may include considering something as having "determined" or "decided" what has been received (e.g., receiving information), transmitted (e.g., transmitting information), input, output, accessed (e.g., accessing data in a memory). Further, "determining" and "deciding" may include considering something as having "determined" or "decided" what has been resolved, selected, chosen, established, compared, etc. That is, "determining" and "deciding" may include considering that some operation has been "determined" or "decided". Also, "determining (deciding)" may be read as "assuming", "expecting", "considering", etc.

[0075] The terms "connected" and "coupled" and any variations thereof mean any direct or indirect connection or coupling between two or more elements, and can include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements can be physical, logical, or a combination thereof. For example, "connected" may be read as "accessed". As used in this disclosure, two elements can be considered to be "connected" or "coupled" to each other using at least one of one or more wires, cables, and printed electrical connections, and also using, as some non-limiting and non-exhaustive examples, electromagnetic energy having wavelengths in the radio frequency region, microwave region, and optical (both visible and invisible) region.

[0076] As used in this disclosure, the description "based on" does not mean "based only on" unless otherwise specified. In other words, the description "based on" means both "based only on" and "based at least in part on".

[0077] In this disclosure, when the terms "include", "including" and their variations are used, these terms are intended to be inclusive in the same manner as the term "comprising". Further, the term "or" as used in this disclosure is not intended to be exclusive.

[0078] In this disclosure, for example, when articles are added by translation, such as a, an, and the in English, this disclosure may include that the nouns following these articles are in the plural form.

[0079] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other". Note that the term may also mean "A and B are each different from C". Terms such as "separate" and "coupled" may also be interpreted in the same way as "different".

Industrial Applicability

[0080] One aspect of the present invention uses a recommendation information providing device that provides recommendation information, and enables the provision of recommendation information regarding settings suitable for singing with respect to a wide variety of music.

Explanation of Signs

[0081] 5... Recommendation information providing device, 1001... Processor, 201... Data acquisition unit, 202... Model construction unit, 203... Prediction unit, 204... Recommendation information generation unit, M... Learning model.

Claims

Claim 1 A recommended information providing apparatus for providing recommended information, comprising at least one processor, wherein the at least one processor acquires scoring results regarding the singing of the user's past songs for each temporal interval of the songs, acquires pitch information indicating the pitch of sounds that are arranged in time series in the interval and that constitute the songs, constructs a learning model for predicting the scoring results regarding the singing of the user's songs from the pitch information using the scoring results and the pitch information as training data, inputs the pitch information regarding the target song into the learning model while uniformly changing the pitch of the sounds indicated by the pitch information in all intervals regarding the target song by a plurality of types of predetermined numerical values, and acquires the scoring results regarding the singing of the target song of the user based on the output of the learning model, outputs the setting content of the pitch to be recommended to the user as the recommended information based on the scoring results for the plurality of types of the pitch information regarding the target song, Recommended information providing apparatus. Claim 2 The at least one processor uses a learning model that takes the time-series pitch information as input and outputs the scoring results for each interval of the song corresponding to the pitch information, and constructs the learning model such that the output of the learning model approaches the scoring results for each interval included in the training data, The recommended information providing apparatus according to claim 1. Claim 3 The at least one processor uses the learning model that further inputs the identification information of the user, The recommended information providing apparatus according to claim 1 or 2. Claim 4 The at least one processor averages the scoring results for each interval of the target song that is the output of the learning model to acquire the scoring results regarding the singing of the target song of the user, The recommended information providing apparatus according to claim 2.

Citation Information

Patent Citations

  • Suitable key recommendation system by user and by musical piece

    JP2007010922A

  • Karaoke system, control method of karaoke system, and control program of karaoke system and information recording medium thereof

    JP2011203479A

  • Karaoke device

    JP2016029429A

  • Manufacturing method of lithium ion conductive sulfide, lithium ion conductive sulfide manufactured thereby, solid electrolyte including lithium ion conductive sulfide, and total solid battery

    JP2017010922A

  • Karaoke system

    JP2018091982A