Prediction device

The prediction device addresses the challenge of accurately predicting karaoke scoring results for new songs by using a learning model constructed from past scoring data and pitch information, thereby improving prediction accuracy.

JP7682175B2Active Publication Date: 2025-05-23NTT DOCOMO INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022530470
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-09
Filing Date
2021-05-28
Publication Date
2025-05-23
Estimated Expiration
2041-05-28

AI Technical Summary

Technical Problem

Conventional karaoke devices struggle to accurately predict scoring results for new songs because they do not effectively incorporate the user's singing tendency based on musical pitch patterns.

Method used

A prediction device that uses scoring results from past songs and pitch information to construct a learning model, which then predicts scoring results for new songs based on their pitch information.

Benefits of technology

This approach improves the prediction accuracy of scoring results for new songs by reflecting the user's past scoring tendencies for pitch patterns, enhancing the reliability of the predicted scores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682175000001
    Figure 0007682175000001
  • Figure 0007682175000002
    Figure 0007682175000002
  • Figure 0007682175000003
    Figure 0007682175000003
Patent Text Reader

Abstract

The purpose of the present invention is to improve prediction accuracy of a scoring result for a new song sung by a user. A prediction device 5 is for predicting a scoring result. The prediction device: acquires a scoring result for a song sung by a user in the past for each temporal interval of the song; acquires musical pitch information which indicates the pitch of sounds constituting the song and arranged in time-series in the interval; uses the scoring result and the musical pitch information as training data and constructs a learning model for predicting a scoring result for a song sung by the user from the musical pitch information; and with the musical pitch information relating to the new song input to the learning model, the prediction device acquires a scoring result for the new song sung by the user on the basis of the output of the learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] One aspect of the present invention relates to a prediction device for predicting a scoring result. [Background technology]

[0002] Conventionally, karaoke devices that predict the score of a song that a user has not sung based on the data of the user's past singing history have been used. For example, a device is known that extracts the score of a song that has a difficulty level set to be the same or similar to that of a selected song from the user's singing history, and predicts the score of the selected song based on the extracted score (see Patent Document 1 below). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2018-91982 A Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the above conventional devices, the scoring result is predicted based on the scoring results of songs of similar difficulty, so the user's singing tendency based on the musical pitch pattern of the song is not easily reflected in the predicted result, which limits the accuracy of the prediction of the scoring result for the user's singing of a new song.

[0005] Therefore, in order to solve the above-mentioned problems, an object of the present invention is to provide a prediction device capable of improving the prediction accuracy of the scoring results regarding the singing of a new song by a user. [Means for solving the problem]

[0006] The prediction device of this embodiment is a prediction device that predicts scoring results, and is equipped with at least one processor, which obtains scoring results for the user's singing of past songs for each temporal section of the song, obtains pitch information that indicates the pitch of the sounds that make up the song and that are arranged in chronological order in the section, and uses the scoring results and pitch information as training data to construct a learning model that predicts the scoring results for the user's singing of the song from the pitch information, and inputs pitch information for a new song into the learning model to obtain scoring results for the user's singing of the new song based on the output of the learning model.

[0007] According to this embodiment, the scoring results for each section of the user's past singing of songs and the pitch information of the sections are used as training data to construct a learning model that predicts the scoring results. Then, the pitch information of the new song is input to the constructed learning model, and the scoring results for the user's singing of the new song are obtained based on the output. This makes it possible to obtain the scoring results for the singing of the new song that reflect the user's past scoring tendency for pitch patterns, and improves the prediction accuracy of the obtained scoring results. Effect of the Invention

[0008] According to one aspect of the present invention, it is possible to improve the prediction accuracy of the scoring result regarding the singing of a new song by a user. [Brief description of the drawings]

[0009] [Figure 1] 1 is a system configuration diagram showing the configuration of a karaoke system 1 according to the present embodiment. [Diagram 2] 4 is a diagram showing an example of a data configuration of history information stored in a data management device 4. FIG. [Diagram 3] 2 is a diagram showing an example of a data structure of music information stored in a data management device 4. FIG. [Figure 4] 2 is a diagram showing an example of a data structure of music information stored in a data management device 4. FIG. [Diagram 5] 4 is a diagram showing an example of a data configuration of history information generated by a prediction device 5. FIG. [Figure 6] FIG. 2 is a diagram showing an example of the data structure of a one-dimensional vector generated by a prediction device 5. [Figure 7] FIG. 2 is a diagram showing a configuration of a learning model used by a prediction device 5. [Figure 8] FIG. 2 is a diagram showing the data structure of a two-dimensional vector converted by a learning model used by a prediction device 5. [Figure 9] FIG. 2 is a diagram showing the data structure of an output vector output by a learning model used by a prediction device 5. [Figure 10] 11 is a flowchart showing the procedure of a learning model construction process performed by the prediction device 5. [Figure 11] 13 is a flowchart showing the procedure of a process for recommending music by the prediction device 5. [Figure 12] FIG. 2 is a diagram illustrating an example of a hardware configuration of a data management device 4 and a prediction device 5 according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will now be described with reference to the accompanying drawings. Whenever possible, the same reference numerals are used to designate the same parts, and redundant description will be omitted.

[0011] 1 is a system configuration diagram showing the configuration of a karaoke system 1 according to this embodiment. The karaoke system 1 is a device having a known function of playing a song designated by a user, and a known function of collecting the singing voice of the user in response to the playback, and evaluating and scoring the singing voice. The karaoke system 1 also has a function of predicting the scoring result for the singing of a new song by the user.

[0012] 1, the karaoke system 1 includes a karaoke device 2, a front server 3, a data management device 4, and a prediction device 5. The front server 3, the data management device 4, and the prediction device 5 are configured to be able to transmit and receive data to and from each other via a communication network such as a LAN (Local Area Network), a WAN (Wide Area Network), or a mobile communication network.

[0013] The karaoke device 2 provides a function for playing songs and a function for collecting the singing voice of the user. The front server 3 is electrically connected to the karaoke device 2 and has a playback function for providing the karaoke device 2 with playback data for playing a song designated by the user, a song search function in response to the user's operation, and a scoring function for receiving singing voice data collected by the karaoke device 2 in response to the playback of the song and calculating the scoring result of the singing voice. The front server 3 also has a function for storing the scoring result of the singing voice by the user as history information in the data management device 4 each time. The front server 3 provides a user interface for accepting user operations and displaying information to the user, and includes a terminal device connected to the front server 3 by wire or wirelessly.

[0014] The data management device 4 is a data storage device (database device) that stores data processed by the front server 3 and the prediction device 5. This data management device 4 includes a history information storage unit 101 that stores history information recording the scoring results regarding the user's past singing of songs, and a song information storage unit 102 that stores pitch information regarding songs that can be played on the karaoke device 2. The various information stored in the data management device 4 is updated as needed by the processing of the front server 3 or by data acquired from outside.

[0015] FIG. 2 shows an example of the data structure of history information stored in the data management device 4, and FIGS. 3 and 4 show an example of the data structure of music information stored in the data management device 4.

[0016] As shown in FIG. 2, the history information stores a "user identifier" for identifying a user, a "music identifier" for identifying a song that the user has previously sung using the karaoke system 1, a "singing time" indicating the time when the song was previously sung, a "total score" indicating the scoring result for the singing of all sections of the song by the function of the front server 3, and "score per section" ... "score per section 24" indicating the scoring result for the singing of each section of the song by the function of the front server 3, all of which are associated with each other. The scoring function of the front server 3 divides the time section of each song into a predetermined number (e.g., 24), calculates the scoring result for each divided section, and calculates the overall scoring result "total score" for each song from the scoring results for all sections. The history information records the scoring results for each section and the overall scoring result calculated by the front server 3 for each user's singing of each song.

[0017] 3 shows an example of the data structure of the pitch information in the music information. In this way, the pitch information stores a "music identifier" that identifies a music piece that can be played using the karaoke system 1, a "note start time (ms)" that indicates the start time of a note (note) that constitutes the music piece in the entire music piece, a "note end time (ms)" that indicates the end time of the note in the entire music piece, a "pitch" that numerically indicates the pitch (pitch) of the note, and a "strength" that numerically indicates the strength of the note, all of which are associated with each other. The data management device 4 stores pitch information on all notes that constitute each music piece that can be played by the front server 3 and are arranged in chronological order in each music piece.

[0018] 4 shows an example of the data configuration of the section information in the song information. In this way, the section information stores, in association with each other, a "song identifier" that identifies a song that can be played using the karaoke system 1, a "section start time (ms)" that indicates the start time of the section of the song in the entire song, and a "section end time (ms)" that indicates the end time of the section in the entire song. The data management device 4 stores section information on all sections that make up each song that can be played by the front server 3.

[0019] The prediction device 5 is a device that predicts a scoring result of the singing of a new song by the user by the front server 3, and includes, as functional components, a data acquisition unit 201, a model construction unit 202, a prediction unit 203, and a selection information generation unit 204. The function of each component will be described below.

[0020] Prior to the construction process of a learning model for predicting a scoring result, the data acquisition unit 201 acquires history information and music information from the data management device 4. Prior to the prediction process of a scoring result, the data acquisition unit 201 also acquires music information. The data acquisition unit 201 passes each piece of acquired information to the model construction unit 202 or the prediction unit 203.

[0021] That is, during the construction process of the learning model, the data acquisition unit 201 combines information read from the history information storage unit 101 and the music information storage unit 102 of the data management device 4 to generate history information of the scoring results for each note in each section of a song sung by the user in the past. FIG. 5 shows an example of the data configuration of the history information generated by the data acquisition unit 201. In this way, the history information is associated with a "user identifier" that identifies the user, a "music identifier" that identifies a song sung by the user in the past, a "section" that identifies a section of the song, a "note start time (ms)" that indicates the start time of a note in the section, a "note end time (ms)" that indicates the end time of the note, a "pitch" that numerically indicates the pitch of the note, a "strength" that numerically indicates the strength of the note, and a "score" that indicates the scoring result of the section including the note. The data acquisition unit 201 generates history information regarding all the notes that constitute each song sung by the user in the past.

[0022] Furthermore, during the process of predicting the scoring result, the data acquisition unit 201 acquires song information on the new song to be predicted from the data management device 4. The data acquisition unit 201 passes the acquired song information to the prediction unit 203.

[0023] Returning to FIG. 1, the model construction unit 202 uses the history information generated by the data acquisition unit 201 as training data to construct a machine learning learning model that predicts the scoring result of the user's singing of a new song based on the pitch information of the new song. Prior to constructing the learning model, the model construction unit 202 executes pre-processing to process the history information delivered from the data acquisition unit 201. In detail, the model construction unit 202 converts each piece of the history information into a one-dimensional vector (sound vector) in which the pitch and strength information of each note constituting the song sung by the user in the past is arranged. In addition, the model construction unit 202 converts each piece of the history information into a one-dimensional vector (score vector) corresponding to the sound vector in which the scoring results of the sections corresponding to each note are arranged, and a one-dimensional vector (user identification vector) corresponding to the sound vector in which user identification information of the user who sang is arranged.

[0024] 6 shows an example of the data configuration of a one-dimensional vector generated by pre-processing of the model construction unit 202. In this manner, the model construction unit 202 converts the history information into a sound vector V1, a score vector V2, and a user identification vector V3.

[0025] Then, the model construction unit 202 inputs the sound vector V1 and the user identification vector V3 into the learning model, and optimizes the parameters of the learning model so that the output result of the learning model approaches the score indicated by the score vector V2 (trains the learning model). At this time, the model construction unit 202 uses a deep learning learning model as the learning model.

[0026] Fig. 7 shows the configuration of the learning model M used by the model construction unit 202. As shown in Fig. 7, the learning model M is composed of a one-hot encoding unit M1, a GRU unit M2, a combination unit M3, and a dense unit M4.

[0027] The one-hot encoding unit M1 receives the user identification vector V3 and converts the user identification vector V3 into a two-dimensional vector. FIG. 8 shows an example of the data configuration of the two-dimensional vector converted by the one-hot encoding unit M1. In this way, in the two-dimensional vector, each row corresponds to a sound indicated by each element of the sound vector V1, and each column corresponds to each user indicated by each element of the user identification vector V3. For example, if the "user identifier" of one element included in the user identification vector V3 is "A1", in the row corresponding to that element, the value of the column corresponding to the "user identifier A1" is set to "1", and the values ​​of the columns corresponding to other user identifiers are set to "0". The one-hot encoding unit M1 generates two-dimensional vectors for rows corresponding to all elements included in the user identification vector V3.

[0028] The GRU unit M2 is a type of recurrent neural network (RNN), which outputs a state in addition to the normal output, and receives the previously output state again as input in addition to the sound vector V1 as the normal input. This allows the GRU unit M2 to have the function of storing past input information and to process long-term time-series information.

[0029] The combining unit M3 combines the output of the one-hot encoding unit M1 with the output of the GRU unit M2. The dense unit M4 is a fully connected layer in deep learning, and converts the numerical sequence of a certain number of dimensions output from the combining unit M3 into an output (Y) of an arbitrary number of dimensions by multiplying it by a weight (w) and adding a bias (b). In this embodiment, the dense unit M4 converts it into a one-dimensional output vector Y in which the scoring results (scores) of each section of the music are arranged. FIG. 9 shows an example of the data configuration of the output vector converted by the dense unit M4. Thus, in the output vector (Y), each element indicates a predicted value of the scoring result of each section composed of sounds corresponding to the elements of the input sound vector V1.

[0030] Using the learning model M configured as above, the model construction unit 202 inputs the user identification vector V3 and the sound vector V1 into the learning model M, and trains the learning model M so that the resulting output vector (Y) approaches the score for each section indicated by the score vector V2. As a result of the training, for example, the parameters of the weight (w) and bias (b) in the dense part M4 of the learning model M are optimized.

[0031] Returning to FIG. 1 again, the prediction unit 203 obtains a predicted value of the scoring result for each section of the user's singing of the new song, using the learning model M constructed by the model construction unit 202 based on the song information for the new song. Specifically, the prediction unit 203 performs preprocessing similar to that of the model construction unit 202 on the song information, and generates a sound vector V1 and a user identification vector V3 for the new song. Then, the prediction unit 203 obtains a predicted value of the scoring result for each section of the new song, based on the output vector (Y) obtained by inputting the generated sound vector V1 and user identification vector V3 into the learning model M.

[0032] The selection information generating unit 204 repeatedly obtains the predicted value of the scoring information for each section for the multiple songs from the prediction unit 203, and calculates a predicted value of the overall scoring result for each of the multiple songs. For example, as the predicted value of the overall scoring result, the average value of the predicted values ​​of the scoring results for all sections is calculated. Then, the selection information generating unit 204 selects a song to be recommended to the user from the multiple songs, and outputs selection information indicating the selected song together with the predicted value of the overall scoring result for each of the multiple songs. For example, the selection information generating unit 204 selects a song with a relatively high predicted value of the scoring result, a song with a predicted value of the scoring result higher than a preset threshold, etc., as a song to be recommended to the user. The selection information and the information of the predicted value output by the selection information generating unit 204 are output to a terminal device of the front server 3, etc.

[0033] Next, the processing of the prediction device 5 configured as above will be described. Fig. 10 is a flowchart showing the procedure of the learning model construction processing by the prediction device 5, and Fig. 11 is a flowchart showing the procedure of the music recommendation processing by the prediction device 5. The learning model construction processing is started at a preset timing (for example, at regular timing), or at a timing when a certain amount of history information is accumulated in the data management device 4. The music recommendation processing is started at a preset timing, or at a timing when an instruction is received from a user in the front server 3.

[0034] 10, when the process of constructing a learning model is started, the data acquisition unit 201 acquires history information on the scoring results of the user's past singing of songs from the data management device 4 (step S101). In addition, the data acquisition unit 201 acquires song information on songs recorded in the history information from the data management device 4 (step S102).

[0035] Next, the model construction unit 202 executes preprocessing, and generates a sound vector V1, a score vector V2, and a user identification vector V3 based on the history information and the music information (step S103). After that, the model construction unit 202 trains the learning model M using the sound vector V1, the score vector V2, and the user identification vector V3, thereby optimizing the parameters of the learning model M (construction of the learning model, step S104), and the construction process of the learning model is completed.

[0036] 11, when the music recommendation process is started, the data acquisition unit 201 acquires music information about a plurality of music pieces from the data management device 4 (step S201). Then, the prediction unit 203 executes preprocessing to generate a sound vector V1 based on the music information, and generates a user identification vector V3 for identifying a user whose score result is to be predicted, for the number of elements corresponding to the sound vector V1 (step S202).

[0037] Next, the prediction unit 203 inputs the sound vector V1 and the user identification vector V3 into the learning model M, and obtains a predicted value of the score result for each section of the multiple songs based on the output vector of the learning model M (step S203). After that, the selection information generation unit 204 calculates a predicted value of the overall score result for each of the multiple songs based on the predicted value of the score result for each section of the multiple songs (step S204). Finally, the selection information generation unit 204 selects songs to be recommended to the user to sing based on the predicted value of the score result for each of the multiple songs, and generates and outputs selection information for the user (step S205).

[0038] Next, the effect of the prediction device 5 of this embodiment will be described. According to this prediction device 5, the scoring results for each section of the user's past singing of songs and the pitch information for each section are used as training data to construct a learning model M that predicts the scoring results. Then, the pitch information for the new song is input to the constructed learning model M, and the scoring results for the user's singing of the new song are obtained based on the output. This makes it possible to obtain the scoring results for the singing of the new song that reflect the user's past scoring tendency for pitch patterns, and improve the prediction accuracy of the obtained scoring results.

[0039] In addition, in this embodiment, a learning model M is used that receives time-series pitch information and outputs a score for each section of a song corresponding to the pitch information, and the learning model M is constructed so that the output of the learning model M approaches the score for each section included in the training data. In this way, a learning model M that grasps the tendency of the score for the pitch pattern for each section of a song can be constructed, and the prediction accuracy of the score for the user's singing of a new song can be reliably improved.

[0040] In addition, in this embodiment, a learning model M is used that further inputs user identification information. In this way, a learning model M that grasps the tendency of the scoring results for the pitch patterns of each user can be constructed, and the prediction accuracy of the scoring results for each user can be reliably improved.

[0041] In this embodiment, the scoring results for the user's singing of the new song are obtained by averaging the scoring results for each section of the new song, which are the output of the learning model M. In this way, the user's strengths and weaknesses in singing the new song can be easily determined.

[0042] In addition, in this embodiment, the scoring results for the user's singing of multiple songs are repeatedly obtained, and songs to be recommended to the user are selected from the multiple songs based on the scoring results, and selection information is output. With this configuration, it is possible to select and output songs that are predicted to be good at singing for the user from the multiple songs, and to provide the user with information that is useful when singing.

[0043] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of at least one of hardware and software. The method of realizing each functional block is not particularly limited. That is, each functional block may be realized by using one device that is physically or logically combined, or may be realized by using two or more devices that are physically or logically separated and directly or indirectly connected (for example, by wire, wirelessly, etc.). The functional blocks may be realized by combining the one device or the multiple devices with software.

[0044] Functions include, but are not limited to, judgment, determination, judgment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocation, mapping, assignment, etc. For example, a functional block (component) that performs the function of transmission is called a transmitting unit or a transmitter. In either case, as described above, there is no particular limitation on the method of realization.

[0045] For example, the data management device 4 and the prediction device 5 according to an embodiment of the present disclosure may function as a computer that performs the processing of the present disclosure. Fig. 12 is a diagram showing an example of a hardware configuration of the data management device 4 and the prediction device 5 according to an embodiment of the present disclosure. The above-mentioned data management device 4 and the prediction device 5 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.

[0046] In the following description, the term "apparatus" can be replaced with a circuit, a device, a unit, etc. The hardware configuration of the data management apparatus 4 and the prediction apparatus 5 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.

[0047] Each function of the data management device 4 and the prediction device 5 is realized by loading a specified software (program) onto hardware such as a processor 1001 and a memory 1002, so that the processor 1001 performs calculations, controls communication via a communication device 1004, and controls at least one of reading and writing of data in the memory 1002 and the storage 1003.

[0048] The processor 1001 controls the entire computer by running an operating system, for example. The processor 1001 may be configured with a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the above-mentioned data acquisition unit 201, model construction unit 202, prediction unit 203, and selection information generation unit 204, etc. may be realized by the processor 1001.

[0049] Moreover, the processor 1001 reads out a program (program code), a software module, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002, and executes various processes according to the read out programs. As the program, a program that causes a computer to execute at least a part of the operations described in the above-mentioned embodiment is used. For example, the data acquisition unit 201, the model construction unit 202, the prediction unit 203, and the selection information generation unit 204 may be realized by a control program stored in the memory 1002 and operated in the processor 1001, and other functional blocks may be similarly realized. Although the above-mentioned various processes have been described as being executed by one processor 1001, they may be executed by two or more processors 1001 simultaneously or sequentially. The processor 1001 may be implemented by one or more chips. The program may be transmitted from a network via a telecommunication line.

[0050] The memory 1002 is a computer-readable recording medium, and may be configured by at least one of, for example, a read only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing the construction process and the recommendation process according to one embodiment of the present disclosure.

[0051] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, and the like. Storage 1003 may be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other suitable medium including at least one of memory 1002 and storage 1003.

[0052] The communication device 1004 is hardware (transmission / reception device) for performing communication between computers via at least one of a wired network and a wireless network, and is also called, for example, a network device, a network controller, a network card, a communication module, etc. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc., in order to realize at least one of Frequency Division Duplex (FDD) and Time Division Duplex (TDD). For example, the data acquisition unit 201 that receives the above-mentioned information may be realized by the communication device 1004. The data acquisition unit 201 may be implemented as a transmission unit and a reception unit that are physically or logically separated.

[0053] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that accepts an input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that performs output to the outside. For example, the above-mentioned selection information generation unit 204 and the like may be realized by the output device 1006. Note that the input device 1005 and the output device 1006 may be integrated into one configuration (e.g., a touch panel).

[0054] In addition, each device such as the processor 1001 and the memory 1002 is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.

[0055] Furthermore, the data management device 4 and the prediction device 5 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.

[0056] The notification of information is not limited to the aspects / embodiments described in the present disclosure, and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI), Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB), System Information Block (SIB))), other signals, or a combination thereof. In addition, the RRC signaling may be called an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.

[0057] Each aspect / embodiment described in the present disclosure may be applied to at least one of systems using LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (new Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (Ultra-WideBand), Bluetooth (registered trademark), and other suitable systems, and next-generation systems extended based on these. In addition, a combination of multiple systems (e.g., a combination of at least one of LTE and LTE-A and 5G, etc.) may be applied.

[0058] The order of the steps, sequences, flow charts, etc. of each aspect / embodiment described in this disclosure may be changed unless inconsistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0059] Information, etc. may be output from a higher layer (or a lower layer) to a lower layer (or a higher layer). Information may be input / output via multiple network nodes.

[0060] The input and output information may be stored in a specific location (e.g., memory) or may be managed using a management table. The input and output information may be overwritten, updated, or added to. The output information may be deleted. The input information may be transmitted to another device.

[0061] The determination may be based on a value represented by a single bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0062] Each aspect / embodiment described in this disclosure may be used alone, in combination, or switched according to execution. In addition, notification of predetermined information (e.g., notification that "X is the case") is not limited to being done explicitly, but may be done implicitly (e.g., not notifying the predetermined information).

[0063] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described in the present disclosure. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.

[0064] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0065] Additionally, software, instructions, information, etc. may be transmitted or received over a transmission medium. For example, if the software is transmitted from a website, server, or other remote source using wired and / or wireless technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave, etc.), then these wired and / or wireless technologies are included within the definition of transmission media.

[0066] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0067] In addition, the terms described in this disclosure and the terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Also, a signal may be a message. Also, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.

[0068] As used in this disclosure, the terms "system" and "network" are used interchangeably.

[0069] In addition, the information, parameters, etc. described in the present disclosure may be represented using absolute values, may be represented using relative values ​​from a predetermined value, or may be represented using other corresponding information. For example, a radio resource may be indicated by an index.

[0070] The names used for the above-mentioned parameters are not limiting in any way. Moreover, the formulas using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not limiting in any way.

[0071] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, search, inquiry (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in a memory), and the like. In addition, "judgment" and "decision" can include considering resolving, selecting, choosing, establishing, comparing, etc., to be a "judgment" or "decision." In other words, "judgment" and "decision" can include considering some action to be a "judgment" or "decision." In addition, "judgment" can be interpreted as "assuming," "expecting," "considering," etc.

[0072] The terms "connected" and "coupled", or any variations thereof, mean any direct or indirect connection or coupling between two or more elements, and can include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements can be physical, logical, or a combination thereof. For example, "connected" may be read as "accessed". As used in this disclosure, two elements can be considered to be "connected" or "coupled" to each other using at least one of one or more wires, cables, and printed electrical connections, and also, as some non-limiting and non-exhaustive examples, electromagnetic energy having wavelengths in the radio frequency region, microwave region, and optical (both visible and invisible) regions, etc.

[0073] As used in this disclosure, the recitation "based on" does not mean "based only on" unless otherwise specified. In other words, the recitation "based on" means both "based only on" and "based at least on".

[0074] In this disclosure, when the terms "include", "including", and their variations are used, these terms are intended to be inclusive, similar to the term "comprising". Further, the term "or" used in this disclosure is not intended to be exclusive.

[0075] In this disclosure, for example, when articles are added by translation, such as a, an, and the in English, this disclosure may include that the noun following these articles is in the plural form.

[0076] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different." [Industrial Applicability]

[0077] One aspect of the present invention is to use a prediction device that predicts a scoring result, thereby making it possible to improve the accuracy of predicting a scoring result regarding the singing of a new piece of music by a user. [Explanation of symbols]

[0078] 5...prediction device, 1001...processor, 201...data acquisition unit, 202...model construction unit, 203...prediction unit, 204...selection information generation unit, M...learning model.

Claims

1. A prediction device for predicting a scoring result, At least one processor; The at least one processor: Obtaining a score for the user's past singing of a song for each time section of the song; acquiring pitch information indicating pitches of sounds constituting the music piece that are arranged in time series in the section; Using the scoring result and the pitch information as training data, a learning model is constructed that predicts a scoring result regarding the singing of the song by the user from the pitch information; inputting the pitch information regarding the new musical piece into the learning model, and acquiring a score result regarding the singing of the new musical piece by the user based on an output of the learning model; Prediction device.

2. The at least one processor using a learning model that receives the time-series pitch information as an input and outputs a scoring result for each section of a piece of music corresponding to the pitch information, the learning model is constructed so that the output of the learning model approaches the scoring result for each section included in the training data; The prediction device according to claim 1 .

3. The at least one processor using the learned model further inputting an identity of the user; The prediction device according to claim 1 .

4. The at least one processor obtaining a score for the user's singing of the new song by averaging the score results for each section of the new song, which are the output of the learning model; The prediction device according to claim 2 .

5. The at least one processor repeatedly acquiring the scoring results regarding the singing of the plurality of musical pieces by the user, selecting musical pieces to be recommended to the user from the plurality of musical pieces based on the scoring results, and outputting selection information; The prediction device according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Karaoke system, control method of karaoke system, and control program of karaoke system and information recording medium thereof

    JP2011203479A

  • Karaoke device

    JP2016029429A

  • Karaoke system

    JP2018091982A

  • Server device and recommendation system

    JP2019148767A