Song matching method and device, electronic equipment and storage medium
By using a timbre feature matching method, target songs with similar timbre to the singer's voice are identified, solving the problem of low song matching accuracy and achieving higher matching accuracy and a better singing experience.
Patent Information
- Application Number
- CN202211175487.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2042-09-26
AI Technical Summary
In existing technologies, song matching methods have low accuracy and cannot effectively match songs suitable for the singer, resulting in a poor singing experience.
By matching the timbre characteristics of the singer with the timbre characteristics of songs in the music library, a target song is determined, making it highly similar to the singer's timbre and suitable for the singer to perform.
It improves the accuracy of song matching, making the matched songs more suitable for the singer's timbre and enhancing the singing experience.
Smart Images

Figure CN115565508B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of Internet, and in particular, to a song matching method and device, electronic equipment and storage medium. BACKGROUND
[0002] With the development of computer technology, using singing software to sing has become a popular trend. However, the number of popular songs has exceeded one million, so how to match a suitable song for a singing object from a large number of songs for singing is a problem to be solved.
[0003] Currently, the registration information and behavior records of the singing object are obtained, such as the song listening records and song singing records authorized by the singing object, and the songs in the song library are labeled, such as the release year and the style type. According to the association between the registration information and behavior records of the singing object and the song labels, the songs for the singing object are matched.
[0004] The problem of the above technical solution is that there are many songs associated with the registration information and behavior records of the singing object, most of which do not match the timbre of the singing object and are not suitable for the singing object to sing, which leads to low matching accuracy. SUMMARY
[0005] The present disclosure provides a song matching method, device, electronic equipment and storage medium, which determines the target song for the singing object to sing by the timbre characteristics of the singing object and the timbre characteristics of the songs in the song library, so that the timbre of the target song has high similarity with the timbre of the singing object and is suitable for the singing object to sing, thereby improving the accuracy of song matching. The technical solution of the present disclosure is as follows:
[0006] According to an aspect of an embodiment of the present disclosure, a song matching method is provided, comprising:
[0007] In response to a song matching request of a singing object, first timbre information is determined based on a voice signal of the singing object, the first timbre information being used to represent at least one timbre category of the voice signal;
[0008] At least one target song is determined from a song library based on the first timbre information and a plurality of second timbre information, the second timbre information being used to represent at least one timbre category of the songs in the song library, and the similarity between the timbre category of the target song and the timbre category of the singing object being greater than a similarity threshold value;
[0009] The at least one target song is returned to the singing object, so that the singing object matches the at least one target song.
[0010] According to another aspect of the embodiments of the present disclosure, a song matching device is provided, comprising:
[0011] A first determination unit configured to, in response to a song matching request of a singing object, determine first timbre information based on a voice signal of the singing object, the first timbre information being used to represent at least one timbre category of the voice signal;
[0012] A second determination unit configured to determine at least one target song from a song library based on the first timbre information and a plurality of second timbre information, the second timbre information being used to represent at least one timbre category of songs in the song library, a similarity between a timbre category of the target song and a timbre category of the singing object being greater than a similarity threshold;
[0013] A matching unit configured to return the at least one target song to the singing object, so that the singing object matches with the at least one target song.
[0014] In some embodiments, the first determination unit comprises:
[0015] A dividing sub-unit configured to divide the voice signal of the singing object into a plurality of voice segments;
[0016] An extracting sub-unit configured to, for any voice segment, perform feature extraction on the voice segment to obtain a timbre feature of the voice segment, the timbre feature being used to represent a timbre category of the voice segment;
[0017] A clustering sub-unit configured to cluster the timbre features of the plurality of voice segments to obtain at least one timbre category of the voice signal and at least one category feature, the category feature being used to represent a cluster center.
[0018] In some embodiments, the dividing sub-unit is configured to divide the voice signal into a plurality of voice segments equally according to a target time length, or divide the voice signal into a plurality of voice segments according to a plurality of sentences included in the voice signal, one voice segment including one sentence.
[0019] In some embodiments, the extracting sub-unit is configured to perform spectral feature extraction on the voice segment to obtain a mel-frequency cepstrum feature of the voice segment, determine a plurality of voice frame features of the voice segment based on the mel-frequency cepstrum feature, and determine the timbre feature of the voice segment based on the plurality of voice frame features.
[0020] In some embodiments, the feature extraction on the voice segment to obtain the timbre feature of the voice segment is implemented based on an audio feature extractor.
[0021] The training step of the audio feature extractor includes:
[0022] Based on the spectrum feature extraction layer in the audio feature extractor, spectrum feature extraction is performed on a sample audio signal to obtain a sample mel-frequency cepstrum feature of the sample audio signal.
[0023] Based on the timbre feature extraction layer in the audio feature extractor, the sample mel-frequency cepstrum feature is processed to obtain a sample timbre feature of the sample audio signal.
[0024] Based on the object discriminator, the pitch discriminator, and the audio type discriminator in the audio feature extractor, the sample timbre feature is processed to obtain an object loss, a pitch loss, and an audio type loss.
[0025] Based on the object loss, the pitch loss, and the audio type loss, the audio feature extractor is trained.
[0026] In some embodiments, the clustering subunit is configured to cluster the timbre features of the plurality of speech segments based on a plurality of clustering information to obtain clustering results of the plurality of clustering information, the clustering information being used to indicate a number of classes during clustering, and the clustering result being used to represent an inter-class distance and an intra-class distance; determine target clustering information based on the clustering results of the plurality of clustering information, the target clustering information being the clustering information with the largest ratio of average inter-class distance to average intra-class distance; and determine the at least one timbre class and the at least one class feature of the speech signal based on the clustering result of the target clustering information.
[0027] In some embodiments, the device further includes:
[0028] The signal acquisition unit is configured to acquire the speech signal input by the singing object in a historical time period, or return prompt information to the singing object and acquire the speech signal input by the singing object based on the prompt information.
[0029] In some embodiments, the second determination unit is configured to acquire at least one third timbre information from the plurality of second timbre information, the number of timbre classes of the at least one third timbre information being not greater than the number of timbre classes indicated by the first timbre information; acquire at least one timbre information with a similarity greater than a similarity threshold from the at least one third timbre information; and determine the at least one target song corresponding to the at least one timbre information from the song library.
[0030] In some embodiments, the second determining unit is further configured to determine, for any third timbre information, a first similarity between at least one first timbre category in the first timbre information and at least one second timbre category in the third timbre information; determine, for any first timbre category in the first timbre information, a second similarity of the first timbre category, the second similarity being based on a minimum value among the first similarities between the first timbre category and the at least one second timbre category; and determine a sum value of the at least one second similarity of the at least one first timbre category as a similarity between the first timbre information and the third timbre information.
[0031] In some embodiments, the apparatus further comprises:
[0032] a song dividing unit configured to divide, for any song in the song library, the song into a plurality of song segments;
[0033] a feature extracting unit configured to perform feature extraction on any song segment to obtain a timbre feature of the song segment, the timbre feature being used to represent a timbre category of the song segment;
[0034] a clustering unit configured to cluster the timbre features of the plurality of song segments to obtain at least one timbre category of the song and at least one category feature, the category feature being used to represent a clustering center.
[0035] According to another aspect of embodiments of the present disclosure, an electronic device is provided, comprising:
[0036] one or more processors;
[0037] a memory for storing program code executable by the processor;
[0038] wherein the processor is configured to execute the program code to implement the above-described song matching method.
[0039] According to another aspect of embodiments of the present disclosure, a computer-readable storage medium is provided, which, when program code stored therein is executed by a processor of an electronic device, enables the electronic device to perform the above-described song matching method.
[0040] According to another aspect of embodiments of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the above-described song matching method.
[0041] The embodiment of the present disclosure provides a song matching method, which determines at least one timbre category of a voice signal of a singing object based on the voice signal, so that at least one target song similar to the timbre category of the voice signal can be determined from a song library based on the at least one timbre category, and then the at least one target song is returned to the singing object, so that the singing object matches the at least one target song, so that the timbre category of the target song matched with the singing object is highly similar to the timbre category of the singing object, and the singing object is suitable for singing, and the accuracy of song matching is improved.
[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0043] The accompanying drawings incorporated in the specification and forming a part of it, illustrate embodiments consistent with the present disclosure, and together with the description, serve to explain the principles of the present disclosure, and do not limit the present disclosure.
[0044] Figure 1 is a schematic diagram of an implementation environment of a song matching method according to an exemplary embodiment.
[0045] Figure 2 is a flowchart of a song matching method according to an exemplary embodiment.
[0046] Figure 3 is a flowchart of another song matching method according to an exemplary embodiment.
[0047] Figure 4 is a training flowchart of an audio feature extractor according to an exemplary embodiment.
[0048] Figure 5 is a flowchart of extracting timbre features of a singing object according to an exemplary embodiment.
[0049] Figure 6 is a flowchart of extracting timbre features of a song library according to an exemplary embodiment.
[0050] Figure 7 is a block diagram of a song matching method according to an exemplary embodiment.
[0051] Figure 8 is a block diagram of a song matching device according to an exemplary embodiment.
[0052] Figure 9 is a block diagram of another song matching device according to an exemplary embodiment.
[0053] Figure 10 is a block diagram of an electronic device according to an exemplary embodiment.
[0054] Figure 11 is a structural schematic diagram of a server according to an exemplary embodiment. DETAILED DESCRIPTION
[0055] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings.
[0056] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or a chronological sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0057] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant national and regional laws, regulations and standards. For example, the voice signals involved in the present application are obtained under sufficient authorization.
[0058] Figure 1 is a schematic diagram of an implementation environment of a song matching method according to an exemplary embodiment. Taking the electronic device provided as a server as an example, referring to Figure 1 , the implementation environment specifically includes: a terminal 101 and a server 102.
[0059] The terminal 101 can be at least one of a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player, an MP4 player, and a laptop computer. The terminal 101 can be installed and run with an application program for song matching. The terminal 101 can be connected to the server 102 through a wireless network or a wired network.
[0060] The terminal 101 can be at least one of a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player, an MP4 player, and a laptop computer. The terminal 101 can be installed and run with an application program for song matching. The terminal 101 can be connected to the server 102 through a wireless network or a wired network.
[0061] The server 102 can be at least one of a server, a plurality of servers, a cloud computing platform, and a virtualization center. The server 102 can be connected to the terminal 101 and other terminals through a wireless network or a wired network. Alternatively, the number of servers can be more or less, which is not limited in the embodiments of the present disclosure. Of course, the server 102 can also include other functional servers to provide more comprehensive and diversified services.
[0062] Figure 2 A flowchart of a song matching method according to an exemplary embodiment is shown in FIG. 2. As shown in FIG. 2, the method is performed by a server and includes the following steps. Figure 2
[0063] In step S201, in response to a song matching request of a singing object, the server determines first timbre information based on a voice signal of the singing object, the first timbre information being used to represent at least one timbre category of the voice signal.
[0064] In the embodiments of the present disclosure, when the singing object uses an application program for singing or any application program containing a song recording function, the terminal can send a song matching request to the server by triggering a point singing operation on the application program. The server is used to provide background service for the application program. The server can obtain the voice signal of the singing object collected in a historical time period or the voice signal of the singing object collected in real time in response to the song matching request. The server can determine at least one timbre category of the voice signal according to the extracted features by performing feature extraction on the voice signal of the singing object. The at least one timbre category can reflect the timbre of the singing object.
[0065] In step S202, the server determines at least one target song from the song library based on the first timbre information and a plurality of second timbre information, the second timbre information is used to represent at least one timbre category of the songs in the song library, and the similarity between the timbre category of the target song and the timbre category of the singing object is greater than a similarity threshold.
[0066] In the embodiments of the present disclosure, the song library is a collection of songs, and each song in the song library has a corresponding timbre category. For any song, the song can correspond to one or more timbre categories. For example, the beginning part of the song is bass, the middle part is tenor, and the end part uses a certain special singing style, so that the song corresponds to three different timbre categories. Optionally, the server stores a plurality of second timbre information of a plurality of songs in the song library, and if the song library is updated, the server can obtain the second timbre information of the newly added song and save it.
[0067] After the server obtains the first timbre information, the server can determine a target song similar in timbre to the singing object from a plurality of songs based on the similarity between at least one timbre category of the voice signal and at least one timbre category of each song in the song library. Since the similarity between the timbre category of the target song and the voice signal of the singing object is greater than the similarity threshold, the timbre of the target song is not much different from the timbre of the singing object, and the target song is suitable for the singing object to sing. The similarity threshold is used to judge the similarity between the timbre category of the target song and the timbre category of the voice signal of the singing object.
[0068] In step S203, the server returns at least one target song to the singing object, so that the singing object matches with at least one target song.
[0069] In the embodiments of the present disclosure, after the server determines at least one target song from a plurality of songs in the song library, the server can display the target song on the application interface, so that the singing object can select a song from the target song returned by the server to sing.
[0070] The embodiments of the present disclosure provide a song matching method, which determines at least one timbre category of a voice signal of a singing object based on the voice signal, so that at least one target song similar in timbre to the voice signal can be determined from a song library based on the at least one timbre category, and then the at least one target song is returned to the singing object, so that the singing object matches with at least one target song. The similarity between the timbre category of the target song matched with the singing object and the timbre category of the singing object is high, and the target song is suitable for the singing object to sing, thereby improving the accuracy of song matching.
[0071] In some embodiments, the first timbre information is determined based on the singing object's voice signal, including:
[0072] The singing object's voice signal is divided into a plurality of voice segments;
[0073] For any voice segment, feature extraction is performed on the voice segment to obtain a timbre feature of the voice segment, the timbre feature being used to represent a timbre category of the voice segment;
[0074] The timbre features of the plurality of voice segments are clustered to obtain at least one timbre category of the voice signal and at least one category feature, the category feature being used to represent a cluster center.
[0075] By clustering the timbre features of the singing object, the timbre of the singing object can be more accurately reflected.
[0076] In some embodiments, the singing object's voice signal is divided into a plurality of voice segments, including any of the following:
[0077] The voice signal is equally divided into a plurality of voice segments according to a target time length;
[0078] The voice signal is divided into a plurality of voice segments according to a plurality of sentences included in the voice signal, one voice segment including one sentence.
[0079] By dividing the singing object's voice signal into a plurality of voice segments in two ways, the singing object's voice signal can be more accurately processed.
[0080] In some embodiments, feature extraction is performed on the voice segment to obtain a timbre feature of the voice segment, including:
[0081] Spectral feature extraction is performed on the voice segment to obtain a mel-frequency cepstrum feature of the voice segment;
[0082] Based on the mel-frequency cepstrum feature, a plurality of voice frame features of the voice segment are determined;
[0083] Based on the plurality of voice frame features, the timbre feature of the voice segment is determined.
[0084] By extracting the timbre feature of the voice segment, the timbre feature of the singing object can be obtained, and the accuracy of song matching is improved.
[0085] In some embodiments, feature extraction is performed on the voice segment to obtain a timbre feature of the voice segment, which is implemented based on an audio feature extractor;
[0086] The training steps of the audio feature extractor include:
[0087] The sample audio signal is subjected to spectral feature extraction based on a spectral feature extraction layer in the audio feature extractor, to obtain sample mel-frequency cepstral features of the sample audio signal.
[0088] The sample mel-frequency cepstral features are processed based on a timbre feature extraction layer in the audio feature extractor, to obtain sample timbre features of the sample audio signal.
[0089] The sample timbre features are processed based on an object discriminator, a pitch discriminator, and an audio type discriminator in the audio feature extractor, to obtain an object loss, a pitch loss, and an audio type loss.
[0090] The audio feature extractor is trained based on the object loss, the pitch loss, and the audio type loss.
[0091] Through training of the audio feature extractor, the features of the singing object timbre can be accurately extracted, and the accuracy of song matching is improved.
[0092] In some embodiments, the timbre features of the plurality of voice segments are clustered to obtain at least one timbre category and at least one category feature of the voice signal, including:
[0093] The timbre features of the plurality of voice segments are respectively clustered based on a plurality of clustering information, to obtain clustering results of the plurality of clustering information, the clustering information being used to indicate a category number during clustering, and the clustering result being used to represent an inter-class distance and an intra-class distance.
[0094] Based on the clustering results of the plurality of clustering information, a target clustering information is determined, the target clustering information being a clustering information with a maximum ratio of an average inter-class distance to an average intra-class distance.
[0095] Based on the clustering result of the target clustering information, at least one timbre category and at least one category feature of the voice signal are determined.
[0096] By clustering the timbre features of the singing object based on a plurality of clustering information, a plurality of different clustering results can be obtained, so that the target clustering information determined based on different clustering results can accurately reflect the timbre of the singing object.
[0097] In some embodiments, the method further includes:
[0098] Obtaining voice signals input by the singing object in a historical time period; or,
[0099] Returning prompt information to the singing object, and obtaining voice signals input by the singing object based on the prompt information.
[0100] By collecting the singing object's recorded voice signal in real time, the timbre information determined based on the voice signal can reflect the current timbre of the singing object.
[0101] In some embodiments, based on the first timbre information and the plurality of second timbre information, at least one target song is determined from a song library, including:
[0102] From the plurality of second timbre information, at least one third timbre information is obtained, and the number of timbre categories of the at least one third timbre information is not greater than the number of timbre categories indicated by the first timbre information;
[0103] From the at least one third timbre information, at least one timbre information is obtained, and the similarity between the at least one timbre information and the first timbre information is greater than a similarity threshold;
[0104] At least one target song corresponding to the at least one timbre information is determined from the song library.
[0105] By obtaining the third timbre information based on the category number, the songs in the song library are filtered, and then the target song satisfying the similarity condition is determined from the filtered songs based on the similarity, so that the timbre of the determined target song has a high similarity with the timbre of the singing object, and the accuracy of matching the song for the singing object is improved.
[0106] In some embodiments, the method further includes:
[0107] For any third timbre information, a first similarity between at least one first timbre category in the first timbre information and at least one second timbre category in the third timbre information is determined;
[0108] For any first timbre category in the first timbre information, a second similarity of the first timbre category is determined, and the second similarity is based on the minimum value of the first similarities between the first timbre category and the at least one second timbre category;
[0109] The sum of at least one second similarity of at least one first timbre category is determined as the similarity between the first timbre information and the third timbre information.
[0110] By determining the target song based on the similarity between the first timbre information and the third timbre information, the target song has a high similarity with the singing object, and the accuracy of song matching is improved.
[0111] In some embodiments, the method further includes:
[0112] For any song in the song library, the song is divided into a plurality of song segments;
[0113] For any song segment, feature extraction is performed on the song segment to obtain a timbre feature of the song segment, and the timbre feature is used to represent a timbre category of the song segment.
[0114] The timbre features of the plurality of song segments are clustered to obtain at least one timbre category of the song and at least one category feature, and the category feature is used to represent a clustering center.
[0115] Through feature extraction and clustering on the song segments, timbre information of the songs in the song library can be obtained, so that a song similar to the timbre of the singing object can be selected, and the accuracy of song matching is improved.
[0116] Figure 3 is a flowchart of another song matching method according to an example embodiment, as shown in Figure 3 The method is performed by a server and includes the following steps.
[0117] In step S301, in response to a song matching request of a singing object, the server obtains a voice signal of the singing object.
[0118] In the embodiments of the present disclosure, the singing object sends a song matching request to the server through a terminal when ordering a song, and the server can obtain the voice signal of the singing object through different ways in response to the song matching request, and then determine the timbre of the singing object based on the voice signal to match a suitable song for the singing object.
[0119] In some embodiments, the server can obtain the voice signal of the singing object in the following two ways.
[0120] In the first way, the server obtains the voice signal previously saved by the singing object as the voice signal of the singing object. Correspondingly, the server obtains the voice signal input by the singing object in a historical time period. The historical time period can be one day, three days or seven days, which is not limited in the embodiments of the present disclosure. It should be noted that the voice signal input by the singing object in the historical time period is the voice signal without the mute part. By obtaining the voice signal input in the historical time period, the singing object does not need to frequently input the voice signal, and the server can match the song for the singing object, thereby improving the efficiency of song matching.
[0121] In a second mode, the server collects the voice signal recorded by the singing object in real time. Accordingly, after receiving the song matching request, the server returns prompt information to the singing object, and obtains the voice signal input by the singing object based on the prompt information. The prompt information can be a prompt to sing a high note, a prompt to sing a low note, or a prompt to sing a singing style that the singing object wants to use, and the embodiments of the present disclosure do not limit the content of the prompt information. By collecting the voice signal recorded by the singing object in real time, the timbre information determined based on the voice signal can reflect the current timbre of the singing object.
[0122] For example, after receiving the song matching request of the singing object, the server sends prompt information to the terminal of the singing object. The terminal displays the prompt information on the application interface, and the singing object records according to the prompt information displayed on the application interface. If the prompt information is a prompt to sing a high note, the singing object can record a high note audio; if the prompt information is a prompt to sing a low note, the singing object can record a low note audio; if the prompt information is a prompt to sing a singing style that the singing object wants to use this time, the singing object can record a singing style audio. The server obtains all audios recorded by the singing object according to the prompt information, removes the mute part, and takes the audios as the voice signal used by the singing object in this time of singing. The server matches a song for the singing object according to the voice signal.
[0123] In step S302, the server divides the voice signal of the singing object into a plurality of voice segments.
[0124] In the embodiments of the present disclosure, the server can divide the voice signal in multiple ways. For any division method, the server can divide the voice signal of the singing object into a plurality of voice segments.
[0125] In some embodiments, the server can divide the voice signal according to the time length. Accordingly, the server divides the voice signal into a plurality of voice segments according to a target time length. The target time length can be 1 second, 5 seconds, or 10 seconds, and the embodiments of the present disclosure do not limit this.
[0126] For example, the total time length of the voice signal of the singing object is 30 seconds, the server divides it every 5 seconds to obtain 6 voice segments, or the server divides it every 10 seconds to obtain 3 voice segments.
[0127] In some embodiments, the server can divide the voice signal according to the sentences in the voice signal. Accordingly, the voice signal is divided into a plurality of voice segments according to a plurality of sentences included in the voice signal, and one voice segment includes one sentence.
[0128] For example, the singing object's voice signal includes three sentences, and the server divides the voice signal into three voice segments. Each voice segment includes one sentence.
[0129] In some embodiments, the server can first divide the voice signal according to the sentences included in the voice signal, and then divide the voice segment corresponding to each sentence according to the time length to obtain multiple voice segments.
[0130] For example, the singing object's voice signal includes three sentences, and the server can divide the voice signal into three voice segments. For a voice segment with a time length of 10 seconds, the server divides it every 5 seconds to obtain 2 voice segments, or divides it every 2 seconds to obtain 5 voice segments.
[0131] In step S303, for any voice segment, the server extracts the voice segment to obtain the timbre feature of the voice segment, and the timbre feature is used to represent the timbre category of the voice segment.
[0132] In the embodiments of the present disclosure, the server can first extract the spectral feature of the voice segment, and then determine the timbre feature of the voice segment based on the extracted spectral feature. Correspondingly, the server extracts the spectral feature of the voice segment to obtain the mel-frequency cepstrum feature of the voice segment. Then, the server determines multiple voice frame features of the voice segment based on the mel-frequency cepstrum feature. Finally, the server determines the timbre feature of the voice segment based on the multiple voice frame features. The voice frame feature represents the timbre feature of each frame of the voice segment. The timbre feature of the voice segment is determined based on the timbre features of all voice frames of the voice segment. Through the extraction of the timbre feature of the voice segment, the timbre feature of the singing object can be obtained, and the accuracy of song matching is improved.
[0133] In some embodiments, the server can extract the features of the voice segment based on an audio feature extractor. The audio feature extractor can be trained by the server or directly obtained by the server. Taking the case that the audio feature extractor is trained by the server as an example, the training steps of the audio feature extractor include the following four steps.
[0134] Step one, the server extracts the spectral feature of the sample audio signal based on the spectral feature extraction layer in the audio feature extractor to obtain the sample mel-frequency cepstrum feature of the sample audio signal.
[0135] The sample audio signal is the original singing audio of multiple singers or the voice audio of multiple objects saved in the database.
[0136] In some embodiments, the server frames the sample audio signal by using a sliding window to obtain a plurality of sample audio frames. Then, the sample audio signal after framing can be converted to a time-frequency domain by formula (1) below, and the sample audio signal in the time-frequency domain can be subjected to spectral feature extraction by formula (2) below to obtain sample mel-frequency cepstral coefficients of the sample audio signal.
[0137] S(n, f) = STFT(s(t)) (1)
[0138] wherein S(n, f) represents the sample audio signal in the time-frequency domain; n represents the nth frame of the sample audio signal after framing, the total number of frames is N, and the value of n is 0 < n ≤ N; f represents the center frequency, F is the maximum value of the center frequency, and the value of f is 0 < f ≤ F; t represents time, the total duration of the sample audio signal is T, and the value of t is 0 < t ≤ T; s(t) represents the sample audio signal in the time domain; and STFT() represents a short-time Fourier transform function.
[0139] M(n, k) = Mel(|S(n, f)|) (2)
[0140] wherein M(n, k) represents the sample mel-frequency cepstral coefficients of the sample audio signal; k represents the dimension number of the mel-frequency cepstral coefficients; |S(n, f)| represents the amplitude value of the sample audio signal in the time-frequency domain; and Mel() represents a mel-frequency cepstral coefficient extraction function.
[0141] Step two, the server processes the sample mel-frequency cepstral coefficients based on the timbre feature extraction layer in the audio feature extractor to obtain sample timbre features of the sample audio signal.
[0142] Wherein the server can process the sample mel-frequency cepstral coefficients into frame-level sample timbre features, and then perform statistical characteristic pooling on the frame-level sample timbre features to obtain the sample timbre features. The process of statistical characteristic pooling is that, based on the frame-level sample timbre features, the mean and variance of all sample audio frames included in a sentence are calculated, and the mean and variance are determined as the sample timbre features, i.e., the sentence-level sample timbre features.
[0143] In some embodiments, the server processes the sample mel-frequency cepstral coefficients by formula (3) below to obtain frame-level sample timbre features. The frame-level sample timbre features are subjected to statistical characteristic pooling by formula (4) below to obtain the sample timbre features.
[0144] R(n, l) = g(M(n, k)) (3)
[0145] wherein R(n, l) represents a feature of an audio frame, n represents an nth frame, and l represents a feature dimension number of the nth frame, and g() represents a function of processing the above-mentioned sample mel-frequency cepstrum feature by using a deep neural network.
[0146] v = statistic pooling(R(n, l)) (4)
[0147] wherein v represents a vector representation of a sample timbre feature at a sentence level, statistic pooling() represents a statistical characteristic pooling function, and R(n, l) represents a feature of an audio frame.
[0148] In step three, the server processes the sample timbre feature based on an object discriminator, a pitch discriminator, and an audio type discriminator in the audio feature extractor to obtain an object loss, a pitch loss, and an audio type loss.
[0149] wherein the object discriminator is configured to discriminate whether the sample audio signal is a singer or an object, the pitch discriminator is configured to discriminate whether the sample audio signal is bass, medium, or treble, and the audio type discriminator is configured to discriminate whether the sample audio signal is a karaoke audio or a speech audio based on the above-mentioned timbre feature information.
[0150] For any sample audio signal, the server can obtain a label of the sample audio signal, which is configured to represent that the sample audio signal includes three different dimensions, a first dimension is a singer or an object, a second dimension is bass, medium, or treble, and a third dimension is a karaoke audio or a speech audio.
[0151] For example, for any sample audio signal, if the sample audio signal is a song sung by a singer A, the pitch belongs to bass, and the audio type belongs to a karaoke audio, then the label of the sample audio signal is singer A, bass, and a karaoke audio.
[0152] In some embodiments, the server can obtain a prediction probability of the discriminator based on the object discriminator, the pitch discriminator, and the audio type discriminator in the audio feature extractor. Then, the server obtains the object loss, the pitch loss, and the audio type loss based on the label of the above-mentioned sample audio signal and the prediction probability of the above-mentioned discriminator.
[0153] For example, the prediction probability of the sample audio signal being a singer or an object obtained by the object discriminator, the prediction probability of the sample audio signal being bass, medium or treble obtained by the pitch discriminator, and the prediction probability of the sample audio signal being original audio or speech audio obtained by the audio type discriminator. Taking the object discriminator as an example, for any sample audio signal, the object discriminator discriminates which singer or object the sample audio signal belongs to, and obtains the prediction result of the object discriminator. Based on the prediction result and the label of the sample audio signal, the prediction probability of the object discriminator is obtained.
[0154] In some embodiments, the object loss of the object discriminator is calculated by the following formula (5). Similarly, the server can obtain the pitch loss of the pitch discriminator and the audio type loss of the audio type discriminator.
[0155]
[0156] wherein J1 represents the cross-entropy loss of the object discriminator; C represents the total number of singers and objects in the sample audio signal; s represents the sample audio signal; P(c|s) represents the probability of the sample audio signal being predicted as a singer or an object by the object discriminator. c P(c|s) is 0 if the sample audio signal is neither a singer nor an object. c
[0157] Step four, the server trains the audio feature extractor based on the object loss, the pitch loss and the audio type loss.
[0158] wherein the training loss of the discriminator is calculated by the following formula (6), and the audio feature extractor is trained based on the training loss.
[0159] J = (1 - a - b) * J1 + a * J2 + b * J3 (6)
[0160] wherein J represents the loss of the discriminator; a and b represent weights, which are not limited by the present disclosure; J1 represents the cross-entropy loss of the object discriminator; J2 represents the cross-entropy loss of the pitch discriminator; and J3 represents the cross-entropy loss of the audio type discriminator. The loss of the discriminator represents the error between the prediction probability and the label, and the smaller the loss, the closer the prediction probability is to the label, i.e., the closer the prediction probability is to the true value.
[0161] For example, Figure 4 is a training flowchart of an audio feature extractor according to an exemplary embodiment. Referring to FIG. 7, the training flowchart of the audio feature extractor includes the following steps. Figure 4 As shown, the audio feature extractor includes a spectral feature extraction layer, a timbre feature extraction layer, an object discriminator, a pitch discriminator, and an audio type discriminator. The server inputs a sample audio signal in the database into the audio feature extractor, performs spectral feature extraction on the sample audio signal based on the spectral feature extraction layer of the audio feature extractor, and obtains spectral features of the sample audio signal. Then, the server performs timbre feature extraction on the spectral features of the sample audio signal based on the timbre feature extraction layer of the audio feature extractor, and obtains timbre features of the sample audio signal. Then, the server discriminates the timbre features of the sample audio signal based on the object discriminator, the pitch discriminator, and the audio type discriminator, and obtains discrimination results of the three discriminators. Finally, the server obtains an object loss, a pitch loss, and an audio type loss based on the discrimination results and a label of the sample audio signal, and adjusts parameters of the audio feature extractor based on the above losses.
[0162] In step S304, the server clusters the timbre features of the plurality of voice segments to obtain first timbre information of the voice signal, the first timbre information including at least one timbre category of the voice signal and at least one category feature, the category feature being used to represent a clustering center.
[0163] In the embodiments of the present disclosure, the server clusters the timbre features of the plurality of voice segments based on distances between the timbre features of the voice segments. The distance can be an Euclidean distance, a Mahalanobis distance, or a cosine distance, which is not limited in the embodiments of the present disclosure. The clustering process refers to a process of classifying similar timbre features of voice segments into one category and classifying different timbre features of voice segments into different categories.
[0164] wherein, for any timbre feature, an array with a dimension of M*1 can be used to represent the timbre feature, i.e., the timbre feature is represented as an array with M rows and 1 column, and M has the same meaning as the feature dimension 1 described above. If the number of timbre features of the singing object is R, the timbre features of the voice segments can be represented by R arrays with a dimension of M*1. Wherein, R is a positive integer.
[0165] In some embodiments, the server can perform clustering based on multiple clustering information to obtain multiple different clustering results. The clustering information is used to indicate the number of classes when clustering, i.e., the number of classes in which the timbre features of the voice segments are divided. Correspondingly, the server performs clustering on the timbre features of the multiple voice segments based on the multiple clustering information to obtain the clustering results of the multiple clustering information. Then, the server determines target clustering information based on the clustering results of the multiple clustering information. Finally, the server determines at least one timbre class and at least one class feature of the voice signal based on the clustering result of the target clustering information. The class feature represents the average of all timbre features in the class corresponding to the timbre class, and one timbre class corresponds to one class feature. The clustering result is used to represent the inter-class distance and the intra-class distance. The inter-class distance represents the distance between the class features of two different classes. The intra-class distance represents the distance between the timbre features in each class and the class feature in this class. The target clustering information is the clustering information with the largest ratio of the average inter-class distance to the average intra-class distance. By performing clustering on the timbre features of the singing object based on multiple clustering information, multiple different clustering results can be obtained, so that the target clustering information determined based on different clustering results can more accurately reflect the timbre of the singing object.
[0166] For example, the plurality of clustering information indicates that the number of timbre categories in the clustering process increases from 2, and the maximum value of the timbre categories is the number of timbre features of the voice segment. The timbre categories are denoted by K_user, and the category features are denoted by an array with a dimension of M*1. Assuming that the voice segment has 4 timbre features, which are represented by 4 arrays with a dimension of M*1. The server can cluster the above-mentioned timbre features based on the clustering information of the timbre category 2, the clustering information of the timbre category 3, and the clustering information of the timbre category 4. If K_user is 2, the server clusters the 4 timbre features to obtain 2 timbre categories and 2 category features, i.e., 2 arrays with a dimension of M*1; if K_user is 3, the server clusters the 4 timbre features to obtain 3 timbre categories and 3 category features, i.e., 3 arrays with a dimension of M*1; if K_user is 4, the server clusters the 4 timbre features to obtain 4 timbre categories and 4 category features, i.e., 4 arrays with a dimension of M*1. Taking the timbre category 3 as an example, the category features represent the average value of the timbre features included in each class, which are represented by 3 arrays with a dimension of M*1. The inter-class distance is represented as the distance between the category features corresponding to the 3 timbre categories. The intra-class distance represents the variance of the distance between the timbre features in each class and the category features of the class. The average inter-class distance represents the average value of the inter-class distances between the 3 timbre categories. The average intra-class distance represents the average value of the intra-class distances corresponding to the 3 timbre categories. The server can obtain the ratio of the average inter-class distance to the average intra-class distance when the timbre category is 3. Similarly, the server can obtain the clustering results of the timbre category 2 and the timbre category 4, and the ratio of the average inter-class distance to the average intra-class distance when the timbre category is 2 and the ratio of the average inter-class distance to the average intra-class distance when the timbre category is 4. The server determines the maximum ratio from the above-mentioned 3 ratios, and determines the clustering information corresponding to the maximum ratio as the target clustering information.
[0167] In order to make the process described in steps S301 to S304 easier to understand, Figure 5 is a flowchart of extracting the timbre features of the singing object according to an exemplary embodiment. Referring to Figure 5As shown, the method comprises the following steps: 501, the singing object sends a song matching request to the server through the terminal. 502, the server determines whether there is an audio signal of the singing object, if yes, step 503 is performed, if not, step 504 is performed. 503, the server obtains the audio signal saved in the historical time period of the singing object. 504, the server sends prompt information to the terminal of the singing object, prompting the singing object to record. 505, whether the singing object wants to record multiple timbres, if yes, step 504 is performed, if not, step 506 is performed. 506, remove the mute part as the voice signal of the singing object. 507, the server divides the voice signal of the singing object to obtain multiple voice segments. 508, the server extracts the timbre features of the voice segments to obtain multiple timbre features. 509, the server clusters the multiple timbre features to obtain at least one timbre category. 510, the server obtains the category features of each timbre category.
[0168] In step S305, the server determines at least one target song from the song library based on the first timbre information and multiple second timbre information, the second timbre information is used to represent at least one timbre category of the songs in the song library, and the similarity between the timbre category of the target song and the timbre category of the singing object is greater than the similarity threshold.
[0169] In the embodiments of the present disclosure, the server determines at least one target song based on the similarity between the multiple second timbre information and the first timbre information after obtaining the first timbre information of the singing object in response to the song matching request of the singing object. Since the similarity between the target song and the voice signal of the singing object is greater than the similarity threshold, the target song is similar to the timbre of the singing object and is suitable for the singing object to sing.
[0170] In some embodiments, the server can first filter at least one third timbre information from the plurality of second timbre information based on the first timbre information, and then determine at least one target song from the at least one third timbre information. Accordingly, the server obtains at least one third timbre information from the plurality of second timbre information. Then, the server obtains at least one timbre information with a similarity greater than a similarity threshold to the first timbre information from the at least one third timbre information. Finally, the server determines at least one target song corresponding to the at least one timbre information from the song library. Wherein the number of timbre categories of the at least one third timbre information is not greater than the number of timbre categories indicated by the first timbre information. It should be noted that the number of target songs can be the number requested by the singing object to match, or the number randomly matched by the server for the singing object, and the present embodiment does not limit it. By obtaining the third timbre information based on the category number, the song in the song library is filtered, and then the target song satisfying the similarity condition is determined from the filtered song based on the similarity, so that the timbre of the determined target song has a high similarity with the timbre of the singing object, and the accuracy of matching the song for the singing object is improved.
[0171] In some embodiments, the server determines the similarity between the first timbre information and the third timbre information based on the similarity between each timbre category. Accordingly, for any third timbre information, the server determines a first similarity between at least one first timbre category in the first timbre information and at least one second timbre category in the third timbre information. Then, for any first timbre category in the first timbre information, the server determines a second similarity of the first timbre category. Finally, the server determines the sum of at least one second similarity of at least one first timbre category as the similarity between the first timbre information and the third timbre information. Wherein the second similarity is based on the minimum value in the first similarity between the first timbre category and the at least one second timbre category. By determining the target song based on the similarity between the first timbre information and the third timbre information, the target song has a high similarity with the singing object, and the accuracy of song matching is improved.
[0172] In some embodiments, for any song in the song library, the server obtains the audio signal of the song. If the server stores the a cappella audio of the song, the server obtains the a cappella audio, removes the silent part, and takes it as the audio signal of the song. If the server does not store the a cappella audio of the song, the server extracts the audio of the song from the audio with accompaniment, removes the silent part, and takes it as the audio signal of the song. Then, the server extracts the features of the audio signals of all the songs in the song library based on the audio feature extractor obtained by training or directly obtained in step S303, obtains the timbre features of the audio signals of all the songs in the song library, and saves the second timbre information obtained by performing step S304 on the timbre features. If the song library is updated, the server obtains the second timbre information of the added songs and saves them. One song corresponds to one second timbre information. For the audio signal of any song, the server divides the audio signal into multiple song segments, and obtains the timbre features of the song segments. Correspondingly, for the audio signal of any song in the song library, the server divides the audio signal into multiple song segments. Then, for any song segment, the server extracts the features of the song segment to obtain the timbre features of the song segment. Finally, the server clusters the timbre features of the multiple song segments to obtain at least one timbre category of the song and at least one category feature. The timbre features are used to represent the timbre category of the song segment, and the category features are used to represent the cluster center. By extracting the features of the song segments and clustering, the timbre information of the songs in the song library can be obtained, so that the songs similar to the timbre of the singing object can be selected, and the accuracy of song matching is improved.
[0173] For example, assuming that the first timbre information of the singing object is four arrays with M*1 dimensions, the first timbre category is four, and the server determines the songs with a timbre category number not greater than four from the song library as candidate songs. For the audio signal of any candidate song, assuming that the timbre category number of the audio signal of the candidate song is three, i.e., the second timbre category is three, the server obtains the third timbre information of the candidate song, i.e., three arrays with M*1 dimensions. For any array corresponding to the first timbre category, the distances of the category features between the array and the three arrays of the audio signal of the candidate song are calculated respectively, three distances are obtained, and the three distances are determined as the first similarity. The minimum value of the three distances is the second similarity. Similarly, the server can obtain the second similarity between the remaining three arrays of the first timbre information of the singing object and the three arrays of the audio signal of the candidate song, and obtain three second similarities. The sum of the four second similarities is the similarity between the timbre of the singing object and the timbre of the candidate song. Similarly, the server can obtain the similarity between the timbre of all the songs in the song library and the timbre of the singing object. The server determines at least one target song by sorting the above similarities from small to large.
[0174] In order to make the process of the server obtaining the second timbre information of each song in the song library more easily understood, Figure 6 is a flowchart of extracting the timbre features of songs in a song library according to an example embodiment. Referring to Figure 6 As shown, the process includes the following steps: 601, the server selects a song from the song library. 602, the server determines whether there is an accompaniment-free audio of the song, and if so, step 603 is performed, and if not, step 604 is performed. 603, the server obtains the accompaniment-free audio. 604, the server extracts the audio of the song from the audio with accompaniment. 605, the silent part is removed as the audio signal of the song. 606, the server divides the audio signal of the song to obtain a plurality of song segments. 607, the server extracts the timbre features of the song segments to obtain a plurality of timbre features. 608, the server clusters the plurality of timbre features to obtain at least one timbre category. 609, the server obtains the category features of each timbre category.
[0175] In step S306, the server returns at least one target song to the singing object, so that the singing object matches with at least one target song.
[0176] In the embodiments of the present disclosure, the server determines at least one target song from a plurality of songs in the song library, and the target song is similar to the timbre of the singing object and suitable for the singing object to sing. The server displays the target song on the application interface of the singing object, so that the singing object can select a song from the target song returned by the server to sing.
[0177] The above steps S301 to S306 exemplarily show an implementation of the song matching method provided by the present application. In order to make the song matching method more easily understood, Figure 7 is a block diagram of a song matching method according to an example embodiment. Referring to Figure 7 As shown, the server extracts the timbre features of the audio signals of all songs in the song library, and saves the timbre features of all songs in the song library in the server. The server extracts the timbre features of the singing object based on the song matching request of the singing object. Based on the timbre features of the singing object and the timbre features of the songs in the song library, the server obtains the similarity between the singing object and the songs in the song library. Based on the similarity, the server matches the songs suitable for singing for the singing object.
[0178] The embodiment of the present disclosure provides a song matching method, which comprises the following steps: determining at least one timbre category of a voice signal of a singing object based on the voice signal, determining at least one target song similar to the timbre category of the voice signal from a song library based on the at least one timbre category, and returning the at least one target song to the singing object, so that the singing object matches the at least one target song, and the timbre category of the target song matched with the singing object is highly similar to the timbre category of the singing object, and the singing object is suitable for singing, thereby improving the accuracy of song matching.
[0179] Figure 8 is a block diagram of a song matching device according to an exemplary embodiment. Referring to Figure 8 , the device comprises a first determination unit 801, a second determination unit 802 and a matching unit 803.
[0180] The first determination unit 801 is configured to, in response to a song matching request of a singing object, determine first timbre information based on a voice signal of the singing object, the first timbre information being used to represent at least one timbre category of the voice signal.
[0181] The second determination unit 802 is configured to determine at least one target song from a song library based on the first timbre information and a plurality of second timbre information, the second timbre information being used to represent at least one timbre category of a song in the song library, and a similarity between the timbre category of the target song and the timbre category of the singing object is greater than a similarity threshold.
[0182] The matching unit 803 is configured to return the at least one target song to the singing object, so that the singing object matches the at least one target song.
[0183] In some embodiments, Figure 9 is a block diagram of another song matching device according to an exemplary embodiment. Referring to Figure 9 , the first determination unit 801 comprises:
[0184] The division sub-unit 8011 is configured to divide the voice signal of the singing object into a plurality of voice segments.
[0185] The extraction sub-unit 8012 is configured to, for any voice segment, extract features of the voice segment to obtain timbre features of the voice segment, the timbre features being used to represent a timbre category of the voice segment.
[0186] The clustering sub-unit 8013 is configured to cluster the timbre features of the plurality of voice segments to obtain at least one timbre category of the voice signal and at least one category feature, the category feature being used to represent a clustering center.
[0187] In some embodiments, the dividing sub-unit 8011 is configured to divide the speech signal into a plurality of speech segments according to a target time length, or divide the speech signal into a plurality of speech segments according to a plurality of sentences included in the speech signal, and each speech segment includes one sentence.
[0188] In some embodiments, the extracting sub-unit 8012 is configured to perform spectral feature extraction on the speech segment to obtain a mel-frequency cepstrum feature of the speech segment, determine a plurality of speech frame features of the speech segment based on the mel-frequency cepstrum feature, and determine a timbre feature of the speech segment based on the plurality of speech frame features.
[0189] In some embodiments, the clustering sub-unit 8013 is configured to respectively cluster the timbre features of the plurality of speech segments based on a plurality of clustering information to obtain clustering results of the plurality of clustering information, the clustering information is used to indicate a number of categories during clustering, and the clustering result is used to represent an inter-class distance and an intra-class distance, determine a target clustering information based on the clustering results of the plurality of clustering information, the target clustering information is a clustering information with a maximum ratio of an average inter-class distance to an average intra-class distance, and determine at least one timbre category and at least one category feature of the speech signal based on the clustering result of the target clustering information.
[0190] In some embodiments, referring to FIG. 8, Figure 9 The apparatus further includes:
[0191] The signal obtaining unit 804 is configured to obtain a speech signal input by the singing object in a historical time period, or return prompt information to the singing object and obtain a speech signal input by the singing object based on the prompt information.
[0192] In some embodiments, the second determining unit 802 is configured to obtain at least one third timbre information from the plurality of second timbre information, the number of timbre categories of the at least one third timbre information is not greater than the number of timbre categories indicated by the first timbre information, obtain at least one timbre information with a similarity greater than a similarity threshold from the at least one third timbre information, and determine at least one target song corresponding to the at least one timbre information from the song library.
[0193] In some embodiments, the second determining unit 802 is further configured to determine a first similarity between at least one first timbre category in the first timbre information and at least one second timbre category in the third timbre information for any third timbre information, determine a second similarity of any first timbre category in the first timbre information, the second similarity is based on a minimum value in the first similarities between the first timbre category and the at least one second timbre category, and determine a sum value of at least one second similarity of the at least one first timbre category as a similarity between the first timbre information and the third timbre information.
[0194] In some embodiments, referring to Figure 9 As shown, the apparatus further includes:
[0195] The song dividing unit 805 is configured to divide any song in the song library into a plurality of song segments;
[0196] The feature extraction unit 806 is configured to, for any song segment, perform feature extraction on the song segment to obtain a timbre feature of the song segment, the timbre feature being used to represent a timbre category of the song segment.
[0197] The clustering unit 807 is configured to cluster the timbre features of the plurality of song segments to obtain at least one timbre category of the song and at least one category feature, the category feature being used to represent a clustering center.
[0198] The embodiments of the present disclosure provide a song matching apparatus. The at least one timbre category of the voice signal of the singing object is determined based on the voice signal of the singing object, so that at least one target song similar to the timbre category of the voice signal can be determined from the song library based on the at least one timbre category, and then the at least one target song is returned to the singing object, so that the singing object matches the at least one target song, and the timbre category of the target song matched with the singing object is highly similar to the timbre category of the singing object, which is suitable for the singing object to sing, and the accuracy of song matching is improved.
[0199] Figure 10 is a block diagram of an electronic device 1000 according to an exemplary embodiment. Generally, the electronic device 1000 includes a processor 1001 and a memory 1002.
[0200] The processor 1001 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 1001 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 1001 can also include a main processor and a co-processor, the main processor being a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit), and the co-processor being a low-power processor for processing data in a standby state. In some embodiments, the processor 1001 can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing of content to be displayed on a display screen. In some embodiments, the processor 1001 can further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.
[0201] The memory 1002 can include one or more computer-readable storage media that can be non-transitory. The memory 1002 can also include high-speed random access memory and nonvolatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1002 is used to store at least one program code for being executed by the processor 1001 to implement a song matching method provided by the method embodiments in the present disclosure.
[0202] In some embodiments, the electronic device 1000 can also optionally include a peripheral device interface 1003 and at least one peripheral device. The processor 1001, the memory 1002, and the peripheral device interface 1003 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 1003 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1008.
[0203] The peripheral interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002 and the peripheral interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002 and the peripheral interface 1003 can be implemented on a separate chip or circuit board, and the present embodiments are not limited in this regard.
[0204] The radio frequency circuit 1004 is configured to receive and send RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1004 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts electromagnetic signals received into electrical signals. Optionally, the radio frequency circuit 1004 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 1004 can communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1004 can also include NFC (Near Field Communication) related circuitry, and the present disclosure is not limited in this regard.
[0205] The display screen 1005 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 1005 is a touch display screen, the display screen 1005 is further configured to capture touch signals on or above the surface of the display screen 1005. The touch signals can be input to the processor 1001 as control signals for processing. In this case, the display screen 1005 can also be configured to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 1005 can be one, arranged on the front panel of the electronic device 1000; in other embodiments, the display screen 1005 can be at least two, arranged on different surfaces of the electronic device 1000 or in a folding design; in some embodiments, the display screen 1005 can be a flexible display screen, arranged on a curved surface or a folding surface of the electronic device 1000. Even, the display screen 1005 can also be arranged in an irregular shape other than a rectangle, i.e., a special-shaped screen. The display screen 1005 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc.
[0206] The camera assembly 1006 is configured to capture images or videos. Optionally, the camera assembly 1006 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is arranged on the front panel of the electronic device, and the rear-facing camera is arranged on the back of the electronic device. In some embodiments, the rear-facing camera is at least two, which is any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to realize the background blur function by fusing the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function by fusing the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 1006 can further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0207] The audio circuit 1007 can include a microphone and a speaker. The microphone is used to collect sound waves of a user and an environment, and convert the sound waves into an electrical signal input to the processor 1001 for processing, or input to the radio frequency circuit 1004 to realize voice communication. For the purpose of stereo sound collection or noise reduction, the microphone can be multiple, respectively arranged at different parts of the electronic device 1000. The microphone can also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert an electrical signal from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker can be a traditional diaphragm speaker, or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, not only can the electrical signal be converted into a sound wave audible to humans, but also can be converted into a sound wave inaudible to humans for ranging purposes. In some embodiments, the audio circuit 1007 can also include a headphone jack.
[0208] The power supply 1008 is used to supply power to each component in the electronic device 1000. The power supply 1008 can be alternating current, direct current, disposable battery or rechargeable battery. When the power supply 1008 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0209] Those skilled in the art can understand that the structure shown in the figure does not constitute a limitation on the electronic device 1000, and can include more or fewer components than the figure, or combine certain components, or use different component arrangements. Figure 10
[0210] When the computer device is configured as a server, Figure 11 is a structural schematic diagram of a server according to an exemplary embodiment. The server 1100 can have a large difference due to different configurations or performances, and can include one or more processors (Central Processing Units, CPU) 1101 and one or more memories 1102, wherein the memory 1102 stores at least one computer program, the at least one computer program is loaded and executed by the processor 1101 to realize the song matching method provided by each method embodiment. Of course, the server can also have a wired or wireless network interface, a keyboard and an input / output interface, etc. to perform input / output, and the server can also include other components for realizing the function of the device, which will not be described here.
[0211] In an example embodiment, a computer readable storage medium, for example, a memory 1002 including instructions, is also provided, which can be executed by the processor 1001 of the electronic device 1000 to complete the song matching method described above. Optionally, the computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0212] In an example embodiment, a computer program product is also provided, which includes a computer program, which, when executed by a processor, implements the song matching method described above.
[0213] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the present disclosure cover any and all variations of the present disclosure that come within the scope of the following claims and their equivalents. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0214] It should be understood that the present disclosure is not limited to the precise structures as herein described and illustrated in the drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present disclosure. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A song matching method, characterized by, The method comprises: in response to a song matching request of a singing object, dividing a voice signal of the singing object into a plurality of voice segments; for any voice segment, performing feature extraction on the voice segment to obtain a timbre feature of the voice segment, the timbre feature being used to represent a timbre category of the voice segment; based on a plurality of clustering information, clustering the timbre features of the plurality of voice segments respectively to obtain clustering results of the plurality of clustering information, the clustering information being used to indicate a category number during clustering, and the clustering result being used to represent an inter-class distance and an intra-class distance; based on the clustering results of the plurality of clustering information, determining target clustering information, the target clustering information being the clustering information with the largest ratio of average inter-class distance to average intra-class distance; based on the clustering result of the target clustering information, determining first timbre information of the voice signal, the first timbre information being used to represent at least one timbre category of the voice signal; based on the first timbre information and a plurality of second timbre information, determining at least one target song from a song library, the second timbre information being used to represent at least one timbre category of the songs in the song library, and a similarity between the timbre category of the target song and the timbre category of the singing object being greater than a similarity threshold; returning the at least one target song to the singing object to enable the singing object to match with the at least one target song.
2. The song matching method of claim 1, wherein, The dividing of the voice signal of the singing object into a plurality of voice segments comprises any of the following: equally dividing the voice signal into a plurality of voice segments according to a target time length; dividing the voice signal into a plurality of voice segments according to a plurality of sentences included in the voice signal, one voice segment including one sentence.
3. The song matching method of claim 1, wherein, The feature extraction on the voice segment to obtain the timbre feature of the voice segment comprises: performing spectral feature extraction on the voice segment to obtain a mel-frequency cepstrum feature of the voice segment; based on the mel-frequency cepstrum feature, determining a plurality of voice frame features of the voice segment; based on the plurality of voice frame features, determining the timbre feature of the voice segment.
4. The song matching method of claim 1, wherein, The feature extraction on the voice segment to obtain the timbre feature of the voice segment is implemented based on an audio feature extractor; The training step of the audio feature extractor comprises: based on a spectral feature extraction layer in the audio feature extractor, performing spectral feature extraction on a sample audio signal to obtain a sample mel-frequency cepstrum feature of the sample audio signal; based on a timbre feature extraction layer in the audio feature extractor, processing the sample mel-frequency cepstrum feature to obtain a sample timbre feature of the sample audio signal; based on an object discriminator, a pitch discriminator, and an audio type discriminator in the audio feature extractor, processing the sample timbre feature to obtain an object loss, a pitch loss, and an audio type loss; based on the object loss, the pitch loss, and the audio type loss, training the audio feature extractor.
5. The song matching method of claim 1, wherein, The method further comprises: obtaining the voice signal input by the singing object in a historical time period; or, Return prompt information to the singing object, and acquire the voice signal input by the singing object based on the prompt information.
6. The song matching method of claim 1, wherein, The method further comprises: The method further comprises: The method further comprises: The method further comprises:
7. The song matching method of claim 6, wherein, The method further comprises: The method further comprises: The method further comprises: The method further comprises:
8. The song matching method of claim 1, wherein, The method further comprises: The method further comprises: The method further comprises: The method further comprises:
9. A song matching apparatus characterized by comprising: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises A second determining unit is configured to determine at least one target song from a song library based on the first timbre information and a plurality of second timbre information, the second timbre information being used to represent at least one timbre category of songs in the song library, and a similarity between a timbre category of the target song and a timbre category of the singing object being greater than a similarity threshold value; A matching unit is configured to return the at least one target song to the singing object, so that the singing object matches the at least one target song.
10. An electronic device, comprising: The electronic device comprises: one or more processors; a memory for storing program codes executable by the processors; wherein the processors are configured to execute the program codes to implement the song matching method according to any one of claims 1 to 8.
11. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to perform the song matching method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Song recommendation method based on singer voice characteristics
CN106991163A
Audio processing method and device, electronic equipment and storage medium
CN115101094A