Singing score method, device, equipment and storage medium

By preprocessing and multi-dimensionally scoring the singer's vocal audio data, the problem of the single scoring method in existing technologies is solved, and an accurate reflection of the singer's true level is achieved.

CN115691561BActive Publication Date: 2026-01-30BEIJING THUNDERSTONE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211266695.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-17
Publication Date
2026-01-30
Estimated Expiration
2042-10-17

AI Technical Summary

Technical Problem

The current singing evaluation method is too simplistic, which allows singers to perform in a perfunctory manner in response to the evaluation criteria, making it difficult to accurately reflect the singer's true level.

Method used

By acquiring the singer's vocal audio data, preprocessing it, and extracting multi-dimensional audio scores, including pitch, rhythm, breath control, emotion, and technique, a comprehensive score is obtained to obtain the singer's overall score.

Benefits of technology

This avoids singers performing perfunctorily in order to get high scores, accurately reflects the singer's true level, and improves the accuracy and multi-dimensionality of the scoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691561B_ABST
    Figure CN115691561B_ABST
Patent Text Reader

Abstract

This invention relates to the field of speech signal processing technology, and discloses a singing evaluation method, apparatus, device, and storage medium. The singing evaluation method includes: acquiring the singer's vocal audio data; preprocessing the vocal audio data to obtain target audio data; determining the multi-dimensional audio score corresponding to the target audio data; and determining the singer's total singing score based on the multi-dimensional audio score. Because this invention obtains the target audio data after preprocessing the singer's vocal audio data and then determines the multi-dimensional audio score corresponding to the target audio data, compared to the single singing evaluation method of the prior art, the above-mentioned singing evaluation method of this invention can avoid singers performing perfunctorily in order to obtain a high score, and accurately reflects the singer's true level.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech signal processing, and in particular to a singing scoring method, device, equipment and storage medium. BACKGROUND

[0002] At present, with the development of network technology and electronic technology, singing as a leisure way is no longer limited to traditional KTV rooms, but is also widely applied to mobile phone K singing, network K singing and home K singing scenes, and the form and location of singing are more diversified. The singing scoring corresponding to singing mostly adopts single variable scoring, such as using sound intensity method to score, and scoring according to the sound intensity; and a more advanced scoring method compares the human voice tone with the tone in the tone file to score.

[0003] However, the above scoring methods are single, and the singer can perform in a coping manner according to the scoring method to obtain a high score, such as the singer only needs to sing loudly in the sound intensity scoring to obtain a higher score, which makes it difficult to accurately reflect the real level of the singer.

[0004] The above content is only used to assist in understanding the technical solutions of the present application, and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0005] The main purpose of the present application is to provide a singing scoring method, device, equipment and storage medium, which aims to solve the technical problem that the existing singing scoring method is single, the singer can perform in a coping manner according to the scoring method to obtain a high score, and it is difficult to accurately reflect the real level of the singer.

[0006] To achieve the above purpose, the present application provides a singing scoring method, which comprises the following steps:

[0007] Obtaining human voice audio data of a singer;

[0008] Preprocessing the human voice audio data to obtain target audio data;

[0009] Determining a multi-dimensional audio score corresponding to the target audio data;

[0010] Determining a total singing score of the singer based on the multi-dimensional audio score.

[0011] Optionally, the step of preprocessing the human voice audio data to obtain target audio data comprises:

[0012] Filtering the human voice audio data, and windowing and framing the filtered human voice audio data;

[0013] Judging whether there is unvoiced frame data in each frame data after windowing and framing.

[0014] If existing, pitch extraction is performed on the non-plosive frame data to obtain target pitch data;

[0015] Target intensity data and target spectrum data corresponding to the frame data are determined, and the target pitch data, the target intensity data, and the target spectrum data are taken as target audio data.

[0016] Optionally, the step of, if existing, pitch extraction is performed on the non-plosive frame data to obtain target pitch data, comprises:

[0017] If existing, inverse frequency conversion is performed on the non-plosive frame data to obtain inverse frequency data;

[0018] Frequencies within a fundamental frequency search range in the inverse frequency data are extracted to obtain a vocal fundamental frequency;

[0019] Pitch conversion is performed on the vocal fundamental frequency, and median filtering is performed on the pitch data obtained by conversion to obtain target pitch data.

[0020] Optionally, the multi-dimensional audio score comprises a pitch score, a rhythm score, a breath score, an emotion score, and a skill score, and the step of determining a multi-dimensional audio score corresponding to the target audio data comprises:

[0021] A pitch matching degree is determined according to target pitch data and reference pitch data corresponding to each character in the current singing lyrics, and the pitch score is determined based on the pitch matching degree;

[0022] A matching time interval is determined according to the pitch matching degree, and the rhythm score is determined based on the matching time interval;

[0023] The number of times of singing pauses occurring in the singing process is determined according to each adjacent intensity value in the target intensity data, and the breath score is determined based on the number of times of singing pauses;

[0024] The emotion score is determined according to the variation amplitude of each intensity value in the target intensity data;

[0025] Harmonics corresponding to each spectrum are determined according to the target spectrum data, and the skill score is determined based on the number of harmonics;

[0026] Correspondingly, the step of determining a singing overall score of the singer based on the multi-dimensional audio score comprises:

[0027] The singing overall score of the singer is determined based on the pitch score, the rhythm score, the breath score, the emotion score, and the skill score.

[0028] Optionally, the step of determining the pitch matching degree according to the target pitch data corresponding to each character in the current singing lyrics and the reference pitch data comprises:

[0029] obtaining the frame data in the window corresponding to each character in the current singing lyrics and the reference pitch data corresponding to each character;

[0030] determining the pitch difference between the target pitch data in the frame data in the window and the reference pitch data;

[0031] determining the number of frames in which the pitch difference in each character is lower than a preset offset threshold;

[0032] determining the pitch matching degree based on the number of frames and the total number of frames in the window.

[0033] Optionally, the step of determining the matching time interval according to the pitch matching degree and determining the rhythm score based on the matching time interval comprises:

[0034] obtaining the actual time interval corresponding to each character and determining the offset time interval before and after the actual time interval;

[0035] determining the matching time interval with the highest pitch matching degree in the offset time interval;

[0036] determining the rhythm score based on the matching time interval and the actual time interval.

[0037] Optionally, the step of determining the number of times of occurrence of singing process interruption according to each adjacent sound intensity value in the target sound intensity data comprises:

[0038] determining whether each adjacent sound intensity value in the target sound intensity data is lower than the interruption threshold;

[0039] if so, determining that there is singing process interruption and determining the number of times of occurrence of singing process interruption.

[0040] In addition, in order to achieve the above-mentioned purpose, the present application further provides a singing score device, which comprises:

[0041] an audio data acquisition module for acquiring human voice audio data of a singer;

[0042] an audio data preprocessing module for preprocessing the human voice audio data to obtain target audio data;

[0043] a multi-dimensional audio score module for determining a multi-dimensional audio score corresponding to the target audio data;

[0044] a singing total score module for determining a singing total score of the singer based on the multi-dimensional audio score.

[0045] In addition, to achieve the above object, the present application further provides a singing scoring device, which comprises a memory, a processor and a singing scoring program stored in the memory and executable on the processor, and the singing scoring program is configured to implement the steps of the singing scoring method as described above.

[0046] In addition, to achieve the above object, the present application further provides a storage medium, which stores a singing scoring program, and the singing scoring program is executable on a processor to implement the steps of the singing scoring method as described above.

[0047] The present application obtains the human voice audio data of a singer, then pre-processes the human voice audio data to obtain target audio data, determines the multi-dimensional audio score corresponding to the target audio data, and finally determines the total singing score of the singer based on the multi-dimensional audio score. Compared with the single singing scoring method in the prior art, the singing scoring method of the present application can avoid the singer's coping singing for a high score, and accurately reflects the real level of the singer. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 The structure schematic diagram of the singing scoring device related to the hardware running environment of the embodiment of the present application;

[0049] Figure 2 The flowchart of the first embodiment of the singing scoring method of the present application;

[0050] Figure 3 The flowchart of the second embodiment of the singing scoring method of the present application;

[0051] Figure 4 The sliding window schematic diagram of the windowing and framing in the second embodiment of the singing scoring method of the present application;

[0052] Figure 5 The flowchart of the third embodiment of the singing scoring method of the present application;

[0053] Figure 6 The rhythm element calculation schematic diagram in the third embodiment of the singing scoring method of the present application;

[0054] Figure 7 The structure block diagram of the first embodiment of the singing scoring device of the present application.

[0055] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0056] It should be understood that the specific embodiments described herein are merely exemplary and do not limit the present application.

[0057] Referring to Figure 1 , Figure 1 The structure diagram of a singing scoring device related to the hardware running environment of the embodiment of the present application is shown.

[0058] As Figure 1 shown, the singing scoring device can include a processor 1001, for example, a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to realize the connection and communication between the components. The user interface 1003 can include a display screen, an input unit such as a keyboard, and optionally a standard wired interface, a wireless interface. The network interface 1004 can optionally include a standard wired interface, a wireless interface (such as a wireless fidelity (Wi-Fi) interface). The memory 1005 can be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk memory. The memory 1005 can also be a storage device independent of the aforementioned processor 1001.

[0059] Those skilled in the art can understand that Figure 1 the structure shown in the figure does not constitute a limitation on the singing scoring device, and can include more or fewer components than the figure, or combine certain components, or different component arrangements.

[0060] As Figure 1 shown, the memory 1005 as a storage medium can include an operating system, a network communication module, a user interface module, and a singing scoring program.

[0061] In Figure 1 the singing scoring device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the singing scoring device of the present application can be arranged in the singing scoring device, and the singing scoring device calls the singing scoring program stored in the memory 1005 through the processor 1001, and executes the singing scoring method provided by the embodiment of the present application.

[0062] The embodiment of the present application provides a singing scoring method, referring to Figure 2, Figure 2 A flowchart of a first embodiment of a singing score method.

[0063] In this embodiment, the singing score method comprises the following steps:

[0064] Step S10: Obtain the human voice audio data of the singer.

[0065] It should be noted that the execution subject of the method of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a mobile phone, a tablet computer, a personal computer, etc., and can also be other electronic devices that realize the same or similar functions. The following singing score device (referred to as score device) is used to describe this embodiment and each of the following embodiments.

[0066] It can be understood that the human voice audio data can be the corresponding human voice data when the singer sings a song.

[0067] In a specific implementation, the score device can collect the human voice audio data of the singer in real time when the singer sings, so as to perform subsequent singing scoring.

[0068] Step S20: Preprocess the human voice audio data to obtain target audio data.

[0069] It should be noted that the target audio data can be audio data used for multi-dimensional scoring. The target audio data can include pitch data, intensity data and spectrum data.

[0070] In a specific implementation, the score device can preprocess the human voice audio data of the singer, which includes filtering, windowing and framing, pitch calculation and spectrum calculation. After the above preprocessing, the target audio data including pitch data, intensity data and spectrum data can be obtained to perform subsequent multi-dimensional scoring.

[0071] Step S30: Determine the multi-dimensional audio score corresponding to the target audio data.

[0072] It should be noted that the multi-dimensions can be five dimensions of pitch, rhythm, breath, emotion and skill.

[0073] In a specific implementation, the score device can determine the scores in the five dimensions of pitch, rhythm, breath, emotion and skill, i.e., the pitch score, the rhythm score, the breath score, the emotion score and the skill score, based on the pitch data, the intensity data and the spectrum data in the target audio data, thereby avoiding too single scoring standard or dimension.

[0074] Step S40: Determine the total singing score of the singer based on the multi-dimensional audio score.

[0075] In a specific implementation, the scoring device can perform comprehensive scoring based on the audio scores in the above-mentioned multiple dimensions and the weighted values of actual requirements of each score to obtain the final singing total score of the singer.

[0076] The embodiment obtains the human voice audio data of the singer, then pre-processes the human voice audio data to obtain target audio data, determines the multi-dimensional audio scores corresponding to the target audio data, and finally determines the singing total score of the singer based on the multi-dimensional audio scores. Compared with the single singing scoring method in the prior art, the singing scoring method in the embodiment can avoid the singer's coping singing for high scores and accurately reflect the real level of the singer.

[0077] Reference Figure 3 , Figure 3 The flowchart of the second embodiment of the singing scoring method of the application is shown in FIG. 2.

[0078] Based on the first embodiment, in the embodiment, the step S20 includes:

[0079] Step S201: filtering the human voice audio data and windowing and framing the filtered human voice audio data.

[0080] It should be noted that when scoring based on human voice audio data, there are interference data such as noise and unvoiced data, so the embodiment is proposed to avoid the above-mentioned interference, improve the accuracy of obtaining the target audio data, and thus improve the accuracy of singing scoring.

[0081] It can be understood that the frequency range of human voice when speaking is between 300 and 3400 Hz, but when singing, the fluctuation of the tone is more rich, and the overtone in the high pitch part can even reach more than 14 KHz. In the subsequent tone extraction operation, in order to avoid high-frequency noise interference and the infection of the musical instruments not filtered out in the human voice, and at the same time, to retain more harmonic components of the human voice to improve the accuracy of tone extraction, it is necessary to filter the human voice audio data.

[0082] In a specific implementation, the scoring device can perform high-low pass filtering on the above-mentioned human voice audio data, the low pass filtering frequency point is 14000 Hz, and the high pass filtering frequency point is 70 Hz, so as to filter out the human voice audio data above 14000 Hz, and at the same time, filter out the low-frequency interference below 70 Hz.

[0083] It should be noted that the speech signal corresponding to the filtered human voice audio data is not a stationary signal, and its statistical properties change over time, so it is necessary to window and frame the above-mentioned filtered human voice audio data to improve the stability of the human voice audio time.

[0084] It can be understood that the scoring device can use 90-120 milliseconds of data amount at a preset sampling rate for windowing and framing, and the windows overlap by 50%. In order to improve the accuracy of the framing, a Hamming window can be used, and the Hamming window formula is as follows:

[0085] W(n) = 0.54 - 0.46 x cos(2 x pi x n / N), 0≤n≤N

[0086] Wherein, N is the window size, pi is 3.1415926, and W(n) is the window data.

[0087] It should be noted that the preset sampling rate can be a sampling rate set based on singing scoring, so as to improve the range of data acquisition and improve the accuracy of singing scoring. The preset sampling rate can be 48000 or 44100.

[0088] For the convenience of understanding, the following Figure 4 will be described, but the present solution is not limited thereto. Figure 4 FIG. 2 is a sliding window diagram for windowing and framing in the second embodiment of the singing scoring method of the present application. In the figure, each frame, i.e., frame 1 and frame 2, has a window size of 80-120 milliseconds of data amount, and the windows overlap by 50%.

[0089] It should be understood that Figure 4 Although two frames are shown in the figure, the frames corresponding to the filtered human voice audio data are all windowed and framed in the above Figure 4 manner, and the present embodiment will not be described again.

[0090] Step S202: determining whether there is non-pure tone frame data in each frame data after windowing and framing.

[0091] It should be noted that the above human voice audio data contains pure tone, and the pure tone cannot extract the pitch, so it is necessary to determine whether there is non-pure tone frame data in each frame data after windowing and framing, so as to avoid subsequent pitch extraction of pure tone data.

[0092] It can be understood that the non-pure tone frame data can be the frame data of the pure tone data in the above human voice audio data after windowing and framing.

[0093] In a specific implementation, the scoring device can detect whether there is pure tone frame data in each frame data after windowing and framing by using zero crossing rate, and the zero crossing rate calculation formula is as follows:

[0094]

[0095] Wherein, ZCR is the zero crossing rate, WL is the window length, sign(k) is the sign of the data in the frame, i.e., the sign is 1 when the data is greater than 0, the sign is -1 when the data is less than 0, and the sign is 0 when the data is equal to 0.

[0096] It should be understood that, under the preset sampling rate, such as 48K or 44.1K sampling rate, when the ZCR>0.1, the frame is determined as a unvoiced frame, the frame pitch is recorded as 0, and no subsequent pitch calculation is performed, otherwise, the frame is a voiced frame, and the subsequent pitch extraction calculation is performed.

[0097] Step S203: If existing, pitch extraction is performed on the unvoiced frame data to obtain target pitch data.

[0098] It should be noted that the target pitch data can be the pitch data corresponding to the singer's voice frequency data.

[0099] In a specific implementation, the scoring device can perform pitch extraction on the detected unvoiced frame data to obtain the pitch data corresponding to the voice frequency data when the unvoiced frame data exists in the frame data after windowing and framing, and record the pitch corresponding to the unvoiced frame data as 0, thereby avoiding the influence of the unvoiced frame data on the singing score.

[0100] Further, in order to improve the accuracy of pitch extraction, in the embodiment, the step S203 includes:

[0101] Step S2031: If existing, cepstrum conversion is performed on the unvoiced frame data to obtain cepstrum data.

[0102] In a specific implementation, the scoring device can perform pitch extraction on the unvoiced frame data in a cepstrum manner when the unvoiced frame data exists in the frame data after windowing and framing, that is, fast Fourier operation is performed on each frame data in the unvoiced frame data, the amplitude is taken as a logarithmic amplitude spectrum, and then fast Fourier operation is performed on the logarithmic amplitude spectrum again, thereby obtaining the cepstrum data.

[0103] Step S2032: Extracting the frequency in the cepstrum data within the fundamental frequency search range to obtain the voice fundamental frequency.

[0104] It should be noted that, since the audio range of a person when singing is between 70 and 650 Hz, the fundamental frequency search range can be determined as:

[0105]

[0106] Wherein, LowIndex is the minimum serial number of search data, HighIndex is the maximum serial number of search data, and Fs is the sampling rate.

[0107] In a specific implementation, the scoring device can perform frequency extraction on the cepstrum data based on the fundamental frequency search range formula, thereby obtaining the voice fundamental frequency.

[0108] Step S2033: pitch conversion is performed on the base frequency of the human voice, and median filtering is performed on the converted pitch data to obtain target pitch data.

[0109] In a specific implementation, the scoring device can perform pitch conversion on the base frequency of the human voice according to the twelve-tone equal temperament, and the pitch conversion formula is as follows:

[0110]

[0111] wherein Pitch is the twelve-tone equal temperament pitch, basefreq is the base frequency, and round is the rounding operation.

[0112] It should be understood that after the pitch data is extracted, the scoring device can perform median filtering on the pitch data calculated for each frame to remove outliers such as wild points, so as to improve the accuracy of the pitch data, and the median filtering can be 9-point median filtering.

[0113] Step S204: determining target intensity data and target spectrum data corresponding to the frame data, and taking the target pitch data, the target intensity data, and the target spectrum data as target audio data.

[0114] It should be noted that the target intensity data can be the average level of the current frame of audio data, and can be determined according to the following formula:

[0115]

[0116]

[0117] wherein INDATA is the data of the current frame.

[0118] It should be understood that when a person sings, the louder the voice is, the more difficult it is to sing, and the higher the score related to the intensity data is.

[0119] In a specific implementation, the scoring device can perform fast Fourier operation on the frame data to obtain target spectrum data corresponding to the current frame, and the target pitch data, the target intensity data, and the target spectrum data can be taken as target audio data.

[0120] Reference Figure 5 , Figure 5 is a flowchart of a third embodiment of the singing scoring method.

[0121] Based on the above embodiments, in this embodiment, the step S30 includes:

[0122] Step S301: determining a pitch matching degree according to the target pitch data corresponding to each character in the current singing lyrics and the reference pitch data, and determining the pitch score based on the pitch matching degree.

[0123] It should be noted that there is a deviation in determining the multi-dimensional audio score, which causes the multi-dimensional audio score to be unable to accurately reflect the actual level of the singer. Therefore, the embodiment is proposed to avoid the deviation in determining the multi-dimensional audio score, improve the accuracy of the score, and thus more accurately reflect the real level of the singer and improve the user experience.

[0124] It can be understood that the reference pitch data can be standard pitch data corresponding to the current singing song and can be stored in a pitch file.

[0125] It should be noted that the pitch matching degree can be a value corresponding to the degree of matching between the target pitch data and the reference pitch data.

[0126] It can be understood that the pitch score can be a score corresponding to the accuracy of the pitch data of each character in the current singing lyrics hitting the reference pitch data.

[0127] It should be noted that the minimum scoring unit of the scoring rule is a character, and therefore the pitch score is calculated in units of characters.

[0128] In a specific implementation, the scoring device can determine the pitch data corresponding to each character in the current singing lyrics in the target pitch data, obtain the standard pitch data corresponding to each character in the current lyrics, that is, the reference pitch data, from the pitch file, compare and match the pitch data corresponding to each character with the reference pitch data, obtain the pitch matching degree with the reference pitch data, determine the accuracy of the pitch data of each character in the current singing lyrics hitting the reference pitch data from the pitch matching degree, and thus determine the pitch score.

[0129] Further, in order to improve the accuracy of the pitch score, the step of determining the pitch matching degree according to the target pitch data corresponding to each character in the current singing lyrics and the reference pitch data comprises:

[0130] Step S3011: obtaining the frame data in the window corresponding to each character in the current singing lyrics and the reference pitch data corresponding to each character.

[0131] It should be noted that the frame data in the window can be the frame data in the current character.

[0132] In a specific implementation, the scoring device can read the frame data contained in the character when determining the character currently sung by the singer, and obtain the reference pitch data corresponding to the character from the pitch file.

[0133] Step S3012: determining a pitch difference between the target pitch data in the frame data in the window and the reference pitch data.

[0134] In a specific implementation, the scoring device can compare the target pitch data and the reference pitch data to determine a difference between the pitch data of each frame in the currently sung word and the reference pitch data corresponding to the word.

[0135] Step S3013: determining a number of frames in which the pitch difference in each word is lower than a preset offset threshold.

[0136] It should be noted that the preset offset threshold can be a measurement value for determining whether the pitch data deviates from a normal value.

[0137] In a specific implementation, the scoring device can determine whether the pitch difference in each word is lower than the preset offset threshold. If yes, it can be determined that the target pitch data is within a normal range. If no, it can be determined that the target pitch data deviates from a normal value, and the singer may have run out of tune when singing the word.

[0138] Step S3014: determining a pitch matching degree based on the number of frames and a total number of frame data in the window.

[0139] It should be noted that the frame data in the window can be all data contained in the currently sung word, and accordingly, the total number of frames can be the number of all frames contained in the word.

[0140] In a specific implementation, the scoring device can determine the pitch matching degree according to the following matching degree formula:

[0141]

[0142] wherein PitchMatch is the pitch matching degree, HitFrameNum is the number of frames in which the pitch data of each frame in the current word and the pitch corresponding to the current word has a difference lower than a preset offset threshold, and WordFrameNum is the total number of frames in the current word.

[0143] It should be understood that, in a period of time in which the energy of the beginning and end of the word is small and the unvoiced component is not counted into the pitch, the pitch calculation is also not counted into the pitch, and therefore the pitch score can be determined according to the following formula:

[0144]

[0145] wherein PitchScore is the pitch score, and FullMatch is a matching degree optimal value.

[0146] It should be understood that, if the matching degree optimal value is 0.5, the pitch matching degree is compared with 0.5 to determine the pitch score.

[0147] Step S302: determining a matching time interval according to the pitch matching degree, and determining the rhythm score based on the matching time interval.

[0148] It should be noted that the matching time interval can be a time interval corresponding to a part in which the target pitch data and the reference pitch data are matched in the actual time length of the currently sung word.

[0149] In a specific implementation, the scoring device can obtain the actual time length of the currently sung word, determine the matching time interval based on the point with the highest pitch matching degree in the time length, and determine whether each word sung is on the beat based on the matching time interval and the actual time length, thereby determining the rhythm score.

[0150] Further, in order to improve the accuracy of the rhythm score, the step S302 comprises:

[0151] Step S3021: obtaining the actual time interval corresponding to each word, and determining the offset time interval before and after the actual time interval.

[0152] It should be noted that the actual time interval corresponding to each word can be obtained from the played accompaniment.

[0153] It can be understood that, in order to avoid the singer starting to sing at a certain time before or after the song file plays the current word, the real-time singing score of the singer cannot be accurately determined, and therefore the actual time interval is offset to more accurately determine the rhythm score of the singer.

[0154] It should be noted that the offset time interval can be a pre-set offset interval, which can be adjusted in real time based on actual conditions.

[0155] Step S3022: determining the matching time interval with the highest pitch matching degree in the offset time interval.

[0156] In a specific implementation, the scoring device can determine the pitch matching degree based on the above-mentioned manner in the offset time interval, and determine the position with the highest pitch matching degree, to determine the matching time interval.

[0157] Step S303: determining the rhythm score based on the matching time interval and the actual time interval.

[0158] In a specific implementation, the scoring device can take the time length of each word in the song as a window, take the word start time coordinate in the pitch file as a reference, search for the position with the highest pitch matching degree in the forward and backward sliding offset fixed time range in the pitch buffer, and thereby determine the rhythm score of the word.

[0159] For ease of understanding, reference can be made toFigure 6 The following will be explained but not limited to the present solution. Figure 6 The following is a schematic diagram for calculating the rhythm element in the third embodiment of the singing score method. In the diagram, the pitch buffer stores the pitch of each frame in the window, the arrow direction indicates that the time axis is from left to right, WORDTIME is the duration of a single word in the song, BASEPOS is the starting time coordinate in the pitch file, DELAY is the offset duration, RIGHTPOS is the coordinate with the highest pitch matching degree, BASEPOS+WORDTIME+DELAY is the time coordinate of the offset duration on the right side of the diagram, BASEPOS-DELAY is the time coordinate of the offset duration on the left side of the diagram, and BEAT is the rhythm element of a single word. Thus, the rhythm element BEAT of the word can be obtained as follows: BEAT=|BASEPOS-RIGHTPOS|.

[0160] It should be noted that the duration of each frame in the pitch buffer can be determined by the following formula:

[0161]

[0162] where FrameTime is the duration of each frame in the pitch buffer, WL is the duration of each word in the song, and Fs is the sampling rate.

[0163] It can be understood that when the time of the current frame is the end time of the current word plus the offset duration, the rhythm element of the word can be determined in the above manner.

[0164] It should be noted that the rhythm score is determined by the rhythm element, and the rhythm score calculation formula is as follows:

[0165]

[0166] where BeatScore is the rhythm score, and BEAT is the rhythm element.

[0167] It can be understood that when the rhythm element is greater than 200 milliseconds, it can be determined that there is a serious rhythm error, and thus the rhythm score is 0.

[0168] Step S303: determining the number of times of singing breaks in the singing process according to the adjacent sound intensity values in the target sound intensity data, and determining the breath score based on the number of times of singing breaks.

[0169] It should be noted that the breath stability can be judged according to whether there is a break. For example, when the number of times of breaks exceeds 3, it can be determined that the breath is unstable, and the breath score is 0.

[0170] In a specific implementation, the scoring device can obtain the adjacent sound intensity values in the target sound intensity data, thereby determining whether there is a continuous break in a single word, and determining the breath score according to the number of breaks.

[0171] Further, in order to improve the accuracy of the breath score, the step of determining the number of times of occurrence of the break in singing according to the adjacent sound intensity values in the target sound intensity data comprises:

[0172] Step S3031: determining whether the adjacent sound intensity values in the target sound intensity data are all lower than the break threshold.

[0173] It should be noted that the break threshold can be a standard value for determining the occurrence of the break.

[0174] In a specific implementation, the scoring device can determine whether the sound intensity in a single word is continuously lower than the break threshold according to whether the adjacent sound intensity values in the target sound intensity data are lower than the break threshold.

[0175] Step S3032: if yes, determining that the break occurs in the singing process and determining the number of times of occurrence of the break.

[0176] In a specific implementation, when the scoring device detects that the adjacent sound intensity values in the target sound intensity data are all lower than the break threshold, it can be determined that the break occurs once and the number of times of occurrence of the break in a word is determined, and the breath score calculation formula is as follows:

[0177]

[0178] wherein, BreathScore is the breath score, and STACCATO is the number of times of occurrence of the break.

[0179] It should be understood that when the number of times of occurrence of the break in a word exceeds a set threshold number of times, it can be determined that the breath is unstable, for example, when the number of times of occurrence of the break in a word exceeds 3, it can be determined that the breath is unstable, and the breath score is 0, thereby improving the accuracy of the breath score.

[0180] Step S304: determining the emotion score according to the variation amplitude of the sound intensity values in the target sound intensity data.

[0181] In a specific implementation, the scoring device can improve the variation amplitude of the sound intensity values in the target sound intensity data in a sentence of lyrics to perform emotion scoring for the singer, and the variation amplitude of the sound intensity values can be determined by the sound intensity mean square deviation and the maximum and minimum absolute value of the sound intensity, and the emotion score calculation formula is as follows:

[0182]

[0183] wherein, EmotionScore is the emotion score, E(k) is each frame sound intensity value of a sentence of lyrics, N is the number of frames of a sentence of lyrics, MAXE is the maximum sound intensity, and MINE is the minimum sound intensity.

[0184] Step S305: determining the harmonics corresponding to each frequency spectrum according to the target frequency spectrum data, and determining the skill score based on the number of the harmonics.

[0185] It should be noted that the singing skill is reflected by the vocal cavity resonance, and the better the vocal cavity resonance is, the higher the skill score is, and the more the harmonics are from the frequency spectrum.

[0186] In a specific implementation, the scoring device can determine the harmonics corresponding to each frequency spectrum according to the target frequency spectrum data, and calculate the skill score by taking the mean square deviation of each frame of frequency spectrum in a word as the harmonic characteristic, that is, taking the average value of the mean square deviation of each frame of frequency spectrum in a word as the skill score, and the skill score calculation formula is as follows:

[0187]

[0188] wherein, SkillScore is the skill score, AAvg is the average value of the frequency spectrum, N is the number of frequency spectrum data, and SN is the number of frames of a sentence.

[0189] Correspondingly, in the embodiment, the step S40 includes:

[0190] Step S401: determining the total score of the singer based on the pitch score, the beat score, the breath score, the emotion score and the skill score.

[0191] In a specific implementation, after obtaining the pitch score, the beat score, the breath score, the emotion score and the skill score, the scoring device can determine the total score of the singer based on the actual demand weighting value, so as to obtain a more standard total score of the singer, and the total score calculation formula is as follows:

[0192] TotalScore=A1*PitchScore+A2*BeatScore+A3*BreathScore+A4*EmotionScore+A5*SkillScore, wherein, TotalScore is the total score of the singer, A1 is the actual demand weighting value of the pitch score, A2 is the actual demand weighting value of the beat score, A3 is the actual demand weighting value of the breath score, A4 is the actual demand weighting value of the emotion score, A5 is the actual demand weighting value of the skill score, PitchScore is the pitch score, BeatScore is the beat score, BreathScore is the breath score, EmotionScore is the emotion score, and SkillScore is the skill score.

[0193] In addition, the embodiment of the present application also provides a storage medium, and the storage medium stores a singing score program, and the singing score program is executed by a processor to realize the steps of the singing score method as described above. In addition, the embodiment of the present application also provides a storage medium, and the storage medium stores a singing score program, and the singing score program is executed by a processor to realize the steps of the singing score method as described above.

[0194] Refer to Figure 7 , Figure 7 is a structural block diagram of a first embodiment of the singing score device.

[0195] As Figure 7 shown, the singing score device provided by the embodiment of the present application comprises:

[0196] An audio data acquisition module 501 is configured to acquire human voice audio data of a singer.

[0197] An audio data preprocessing module 502 is configured to preprocess the human voice audio data to obtain target audio data.

[0198] A multi-dimension audio score module 503 is configured to determine a multi-dimension audio score corresponding to the target audio data.

[0199] A singing total score module 504 is configured to determine a singing total score of the singer based on the multi-dimension audio score.

[0200] The embodiment acquires human voice audio data of a singer, then preprocesses the human voice audio data to obtain target audio data, determines a multi-dimension audio score corresponding to the target audio data, and finally determines a singing total score of the singer based on the multi-dimension audio score. Since the embodiment obtains target audio data by preprocessing human voice audio data of a singer, and then determines a multi-dimension audio score corresponding to the target audio data, compared with a single singing score method in the prior art, the above singing score method of the embodiment can avoid singers from performing in a fake manner to obtain a high score, and accurately reflect the real level of the singer.

[0201] Other embodiments or specific implementations of the singing score device of the present application can refer to the above-mentioned method embodiments, which will not be described here.

[0202] It should be noted that, in this document, the terms “comprising”, “including”, or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles, or systems that include a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent to such processes, methods, articles, or systems. Without more limitations, the element defined by the statement “comprising a” does not exclude the presence of other identical elements in the process, method, article, or system that includes the element.

[0203] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0204] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and the necessary general hardware platform, of course, also can be through hardware, but in many cases the former is the better embodiment. Based on such understanding, the technical solutions of the present application essentially or say the part of the prior art contribution can be embodied in the form of software products, the computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disc), including a number of instructions to make a terminal device (may be a mobile phone, computer, server, or network equipment, etc.) executes the method described in various embodiments of the present application.

[0205] The above is only the preferred embodiment of the present application, not therefore limit the patent scope of the present application, any equivalent structure or equivalent flow transformation made by using the content of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A singing score method, characterized by, The singing score method comprises the following steps: Obtaining human voice audio data of a singer; Preprocessing the human voice audio data to obtain target audio data; Determining a multi-dimensional audio score corresponding to the target audio data; Determining a total singing score of the singer based on the multi-dimensional audio score; The step of preprocessing the human voice audio data to obtain target audio data comprises: Filtering the human voice audio data and windowing and framing the filtered human voice audio data; Determining whether there is non-pure tone frame data in each frame of data after windowing and framing; If there is, extracting the pitch of the non-pure tone frame data to obtain target pitch data; Determining target intensity data and target spectrum data corresponding to each frame of data, and taking the target pitch data, the target intensity data, and the target spectrum data as target audio data; The multi-dimensional audio score comprises a pitch score, a rhythm score, a breath score, an emotion score, and a skill score, and the step of determining the multi-dimensional audio score corresponding to the target audio data comprises: Determining a pitch matching degree according to target pitch data corresponding to each character in the current singing lyrics and reference pitch data, and determining the pitch score based on the pitch matching degree; Determining a matching time interval according to the pitch matching degree, and determining the rhythm score based on the matching time interval; Determining the number of times of singing pauses according to adjacent intensity values in the target intensity data, and determining the breath score based on the number of times of singing pauses; Determining the emotion score according to the amplitude of change of each intensity value in the target intensity data; Determining harmonics corresponding to each spectrum according to the target spectrum data, and determining the skill score based on the number of harmonics; Correspondingly, the step of determining the total singing score of the singer based on the multi-dimensional audio score comprises: Determining the total singing score of the singer based on the pitch score, the rhythm score, the breath score, the emotion score, and the skill score. The step of determining the number of times of singing pauses according to adjacent intensity values in the target intensity data comprises: Determining whether adjacent intensity values in the target intensity data are all lower than a pause threshold; If so, determining that a singing pause occurs, and determining the number of times of singing pauses.

2. The singing score method of claim 1, wherein, The step of extracting the pitch of the non-pure tone frame data to obtain target pitch data if there is non-pure tone frame data comprises: If so, performing inverse frequency conversion on the non-pure tone frame data to obtain inverse frequency data; Extracting a frequency within a fundamental frequency search range from the inverse frequency data to obtain a human voice fundamental frequency; Performing pitch conversion on the human voice fundamental frequency, and performing median filtering on the converted pitch data to obtain target pitch data.

3. The singing score method of claim 1, wherein, The step of determining a pitch matching degree according to target pitch data corresponding to each character in the current singing lyrics and reference pitch data comprises: Obtaining windowed frame data corresponding to each character in the current singing lyrics and reference pitch data corresponding to each character; Determining a pitch difference value between the target pitch data in the windowed frame data and the reference pitch data; determine a number of frames in which the pitch difference in each character is lower than a preset offset threshold value; determine a pitch matching degree based on the number of frames and a total number of frame data in the window.

4. The singing score method of claim 1, wherein, The step of determining a matching time interval according to the pitch matching degree and determining the rhythm score based on the matching time interval, comprises: obtaining an actual time interval corresponding to each character and determining offset time intervals before and after the actual time interval; determining a matching time interval with the highest pitch matching degree in the offset time intervals; determining the rhythm score based on the matching time interval and the actual time interval.

5. A singing score device, characterized by The device comprises: an audio data acquisition module configured to acquire human voice audio data of a singer; an audio data preprocessing module configured to preprocess the human voice audio data to obtain target audio data; a multi-dimensional audio score module configured to determine a multi-dimensional audio score corresponding to the target audio data; a singing total score module configured to determine a singing total score of the singer based on the multi-dimensional audio score; The audio data preprocessing module is further configured to filter the human voice audio data and window and frame the filtered human voice audio data. The audio data preprocessing module is further configured to determine whether there is non-pure tone frame data in each frame of the windowed and framed data. The audio data preprocessing module is further configured to extract the pitch of the non-pure tone frame data to obtain target pitch data if there is non-pure tone frame data. The audio data preprocessing module is further configured to determine target sound intensity data and target frequency spectrum data corresponding to each frame of data, and take the target pitch data, the target sound intensity data, and the target frequency spectrum data as target audio data. The singing total score module is further configured to determine a pitch matching degree according to target pitch data and reference pitch data corresponding to each character in the current singing lyrics, and determine a pitch score based on the pitch matching degree. The singing total score module is further configured to determine a matching time interval according to the pitch matching degree, and determine a rhythm score based on the matching time interval. The singing total score module is further configured to determine the number of times of occurrence of a singing pause according to each adjacent sound intensity value in the target sound intensity data, and determine a breath score based on the number of times of occurrence of the singing pause. The singing total score module is further configured to determine an emotion score according to the variation amplitude of each sound intensity value in the target sound intensity data. The singing total score module is further configured to determine harmonics corresponding to each frequency spectrum according to the target frequency spectrum data, and determine a skill score based on the number of harmonics. The singing total score module is further configured to determine the singing total score of the singer based on the pitch score, the rhythm score, the breath score, the emotion score, and the skill score. The singing total score module is further configured to determine whether each adjacent sound intensity value in the target sound intensity data is lower than a singing pause threshold value. The singing total score module is further configured to determine that a singing pause occurs during singing and determine the number of times of occurrence of the singing pause if each adjacent sound intensity value in the target sound intensity data is lower than a singing pause threshold value.

6. A singing score device, characterized by The device comprises a memory, a processor, and a singing score program stored on the memory and executable on the processor, the singing score program being configured to implement the steps of the singing score method according to any one of claims 1 to 4.

7. A storage medium, characterized by The storage medium has a singing score program stored thereon, the singing score program being executable by the processor to implement the steps of the singing score method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-dimensional singing scoring system

    CN109448754A

  • Real-time singing scoring method and system

    CN109903778A

  • Singing breath scoring method and device

    CN112992183A