Audio Matching Detection Method and Device, Electronic Device, Storage Medium
By filtering and matching the start and end times and pitch of audio note sequences, the problem of low accuracy of existing audio matching detection methods is solved, and higher pitch matching detection accuracy is achieved.
Patent Information
- Application Number
- CN202210082795.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-01-24
AI Technical Summary
The existing audio matching detection methods have low accuracy, making it difficult to effectively evaluate the degree of matching between user imitation audio and standard audio.
By obtaining the sequence of notes of standard audio and audio to be detected, the notes with a duration greater than or equal to the threshold are filtered out, and the pitch is matched according to the start and end time and pitch, the pitch is determined, the impact of vibrato is reduced, and the detection accuracy is improved.
Improves the accuracy of audio matching detection and enables more accurate evaluation of the pitch matching parameters of user imitation audio and standard audio.
Smart Images

Figure CN114491140B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly, to an audio matching detection method and apparatus, an electronic device, a storage medium, and a program product. Background Art
[0002] With the development of technology and economy, people's lives are becoming increasingly rich. They can not only enjoy audio such as songs, instrumental music, and movies, but also imitate this audio by singing, playing musical instruments, etc. In order to enable users to know whether the audio obtained by their imitation matches the standard audio, it is necessary to detect the audio. However, the current audio matching detection methods have low accuracy. Summary of the Invention
[0003] To solve the above technical problems, embodiments of the present application provide an audio matching detection method and apparatus, an electronic device, a storage medium, and a program product.
[0004] According to one aspect of the embodiments of the present application, an audio matching detection method is provided. The method includes:
[0005] Obtaining a first note sequence corresponding to a standard audio and a second note sequence corresponding to an audio to be detected;
[0006] Screening out first notes with a duration greater than or equal to a first threshold from the first note sequence, and finding multiple second notes in the second note sequence whose start and end times match the start and end times of the first notes;
[0007] Screening out first target notes with pitches matching the pitches of the first notes from the multiple second notes, and determining the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target notes and the duration of the first notes;
[0008] Determining a pitch matching parameter of the audio to be detected according to the pitch similarity.
[0009] According to one aspect of the embodiments of the present application, an audio matching detection apparatus is provided. The apparatus includes:
[0010] An obtaining module configured to obtain a first note sequence corresponding to a standard audio and a second note sequence corresponding to an audio to be detected;
[0011] A searching module configured to screen out first notes with a duration greater than or equal to a first threshold from the first note sequence, and find multiple second notes in the second note sequence whose start and end times match the start and end times of the first notes;
[0012] A similarity determination module, configured to screen out a first target note from the multiple second notes whose pitch matches the pitch of the first note, and determine the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the duration of the first note;
[0013] A matching detection module, configured to determine the pitch matching parameter of the audio to be detected according to the pitch similarity.
[0014] According to one aspect of the embodiments of the present application, an electronic device is provided, including:
[0015] One or more processors;
[0016] A storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the audio matching detection method as described above.
[0017] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor of an electronic device, the electronic device is caused to execute the audio matching detection method as described above.
[0018] According to one aspect of the embodiments of the present application, a computer program product is provided, including a computer program, and when the computer instructions are executed by a processor, the audio matching detection method as described above is implemented.
[0019] In the technical solution provided by the embodiments of the present application, after obtaining the first note sequence corresponding to the standard audio and the second note sequence corresponding to the monitored audio, a first note with a duration greater than or equal to a first threshold is screened out from the first note sequence, and multiple second notes whose start and end times match the start and end times of the first note are found from the second note sequence. A first target note whose pitch matches the pitch of the first note is screened out from the multiple second notes, and the pitch similarity between the audio to be detected and the standard audio is determined according to the duration of the first target note and the duration of the first note. The pitch matching parameter of the audio to be detected is determined according to the pitch similarity. That is to say, when detecting the pitch matching parameter of the audio to be detected, if the corresponding standard audio contains a first note with a relatively long time, the pitch matching parameter is determined according to the duration of the first note and the duration of the first target note whose start and end times and pitch match those of the first note in the audio to be detected, so as to reduce the influence of "vibrato" on the pitch matching detection and improve the accuracy of the pitch matching detection.
[0020] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings
[0021] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments in line with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts. In the accompanying drawings:
[0022] Figure 1 is a schematic diagram of an implementation environment related to this application;
[0023] Figure 2 is a flowchart of an audio matching detection method shown in an exemplary embodiment of this application;
[0024] Figure 3 is a schematic diagram of a note sequence shown in an exemplary embodiment of this application;
[0025] Figure 4 is Figure 2 a flowchart of step S110 in the shown embodiment in an exemplary embodiment;
[0026] Figure 5 is a schematic diagram of an audio signal shown in an exemplary embodiment of this application;
[0027] Figure 6 is a schematic diagram of obtaining notes by quantizing the fundamental frequency shown in an exemplary embodiment of this application;
[0028] Figure 7 is Figure 2 a flowchart of step S130 in the shown embodiment in an exemplary embodiment;
[0029] Figure 8 is a flowchart of determining a rhythm matching parameter shown in an exemplary embodiment of this application;
[0030] Figure 9 is Figure 8 a flowchart of step S220 in the shown embodiment in an exemplary embodiment;
[0031] Figure 10 is a schematic diagram of an audio frame mapping relationship shown in an exemplary embodiment of this application;
[0032] Figure 11 is Figure 2 a flowchart of step S110 in the shown embodiment in an exemplary embodiment;
[0033] Figure 12 is a process diagram of determining a pitch matching parameter shown in an exemplary embodiment of this application;
[0034] Figure 13 It is a process diagram for determining rhythm matching parameters shown in an exemplary embodiment of the present application;
[0035] Figure 14 It is a schematic structural diagram of an audio matching detection device shown in an exemplary embodiment of the present application;
[0036] Figure 15 It shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0037] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0038] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0039] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0040] It should also be noted that: "a plurality of" mentioned in the present application refers to two or more. " / " describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the front and rear associated objects.
[0041] In order to enable users to know whether the audio obtained by their own imitation matches the standard audio, it is necessary to detect the audio. Currently, usually taking the standard audio as the standard, the pitch matching parameter is determined according to the pitch deviation degree of the audio to be detected. However, the accuracy of this method is relatively low. Based on this, the embodiments of the present application provide an audio matching detection method and device, an electronic device, a storage medium, and a program product, enriching the content of video events and improving the video event generation efficiency.
[0042] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an implementation environment involved in this application. The implementation environment includes a terminal device 100 and a server 200. The terminal device 100 and the server 200 communicate with each other through a wired or wireless network. The terminal device 100 can upload its own data to the server 200 or obtain data from the server 200.
[0043] It should be understood that Figure 1 the numbers of the terminal device 100 and the server 200 in
[0044] are merely illustrative. According to actual needs, there can be any number of terminal devices 100 and servers 200.
[0045] The terminal device 100 can include, but is not limited to, a smart phone, a tablet, a laptop computer, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, and so on.
[0046] In an exemplary embodiment, the audio matching detection method provided by the embodiments of this application can be executed by the terminal device 100. Correspondingly, the audio matching detection device can be placed in the terminal device 100. Among them, the terminal device 100 can obtain a first note sequence corresponding to a standard audio and a second note sequence corresponding to the audio to be detected. Then, the first notes with a duration greater than or equal to a first threshold are screened out from the first note sequence, and multiple second notes whose start and end times match the start and end times of the first notes are found from the second note sequence. Furthermore, the first target notes with a pitch matching the pitch of the first notes are screened out from the multiple second notes, and the pitch similarity between the audio to be detected and the standard audio is determined according to the duration of the first target notes and the duration of the first notes. The pitch matching parameter of the audio to be detected is determined according to the pitch similarity. In this way, determining the pitch matching parameter of the audio to be detected according to the duration of the first target notes and the second target notes can reduce the influence of "vibrato" on the pitch matching detection and improve the accuracy of the pitch matching detection.
[0047] In another exemplary embodiment, the server 200 may have a function similar to that of the terminal device 100 to execute the audio matching detection method provided in the embodiments of the present application. Accordingly, the audio matching detection device may be placed in the server 200. Among them, the terminal device 100 may upload the audio to be detected to the server 200. After receiving the audio to be detected uploaded by the terminal device 100, the server 200 obtains the standard audio corresponding to the audio to be detected, and obtains the first note sequence corresponding to the standard audio and the second note sequence corresponding to the audio to be detected. The server 200 screens out the first notes with a duration greater than or equal to the first threshold from the first note sequence, and finds multiple second notes in the second note sequence whose start and end times match the start and end times of the first notes. Then, the server 200 finds the first target note in the second note sequence whose start and end times match the start and end times of the first notes, so as to determine the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the duration of the first note, and determine the pitch matching parameter of the audio to be detected according to the pitch similarity.
[0048] In another exemplary embodiment, the terminal device 100 and the server 200 may also jointly execute the audio matching detection method provided in the embodiments of the present application. For example, the terminal device 100 may obtain the second note sequence corresponding to the audio to be detected and upload it to the server 200. The server 200 obtains the first note sequence corresponding to the standard audio, screens out the first notes with a duration greater than or equal to the first threshold from the first note sequence, finds multiple second notes in the second note sequence whose start and end times match the start and end times of the first notes, screens out the first target notes whose pitch matches the pitch of the first notes from the second notes, determines the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the duration of the first note, determines the pitch matching parameter of the audio to be detected according to the pitch similarity, and sends the pitch matching parameter to the terminal device 100.
[0049] It should be noted that, in addition to the application scenarios involved above, the embodiments of the present application can also be applied to various application scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc. In actual applications, corresponding adjustments can be made according to specific application scenarios. For example, if it is applied to the cloud technology scenario, the steps corresponding to the pitch matching detection method can be performed in the cloud; if it is applied to the intelligent transportation or assisted driving scenario, the terminal device 100 may be an in-vehicle terminal, a navigation terminal, etc., and the audio matching detection method can be applied to perform matching detection on the audio to be detected obtained by the in-vehicle terminal.
[0050] It should be noted that in this application, when dealing with user-related data such as the audio to be detected, when the method of this application is applied to specific products or technologies, all of them are obtained with the permission or consent of the user, and the extraction, use, and processing of the relevant data comply with the local safety standards and local laws and regulations.
[0051] See Figure 2 , Figure 2 which is a flowchart of an audio matching detection method shown in an exemplary embodiment of this application. This method can be applied to Figure 1 the implementation environment shown in Figure 1 and can be executed by the terminal device 100 in the shown implementation environment, or can be executed by the server 200, or can be jointly executed by the terminal device 100 and the server 200.
[0052] As Figure 2 shown, in an exemplary embodiment, this audio matching detection method may include steps S110 to S140, which are introduced in detail as follows:
[0053] Step S110, obtain the first note sequence corresponding to the standard audio and the second note sequence corresponding to the audio to be detected.
[0054] It should be noted that the audio to be detected is the audio whose matching degree with the standard audio is to be detected, and it can be an audio uploaded by the user. For example, it includes but is not limited to a song sung by the user, a passage spoken by the user, or a piece of music played by the user, etc.
[0055] The standard audio is the standard for detecting the audio to be detected, and it can be an audio with the same content as the audio to be detected. For example, the standard audio and the audio to be detected can be audio obtained by singing the same song, playing the same piece of music, or getting the same passage. Depending on the different content expressed, the audio matching detection method can be applied to different application scenarios. For example, the audio matching detection method can be applied to the song scoring scenario. Correspondingly, the audio to be detected can be the audio obtained by the user singing a certain song, and the standard audio can be the original singer's audio of this song; the audio matching detection method can be applied to the performance scoring scenario. Correspondingly, the audio to be detected can be the audio obtained by the user playing a certain piece of music, and the standard audio can be the audio played by a professional for this piece of music; the audio matching detection method can be applied to the dubbing scoring scenario. Correspondingly, the standard audio can be the original audio of a film or television work, and the audio to be detected can be the audio obtained by the user dubbing this film or television work. It should be noted that the application scenarios listed here are only exemplary. According to actual needs, the audio matching detection method can also be applied to other application scenarios, and this embodiment does not limit the application scenarios of the audio detection method.
[0056] The first note sequence is a note sequence obtained by processing a standard audio, which includes multiple notes. Among them, each note corresponds to a pitch and start and end times. It should be understood that the pitch is the height of the sound, and the essence of sound is a mechanical wave. The pitch of the sound is determined by the frequency of the mechanical wave; the start and end times include the start time and the end time.
[0057] The second note sequence is a note sequence obtained by processing the audio to be detected, which includes multiple notes, and each note corresponds to a pitch and start and end times.
[0058] In order to determine the matching degree between the audio to be detected and the standard audio, in this embodiment, the audio to be detected can be obtained and processed to obtain the second note sequence; it is also necessary to determine the standard audio corresponding to the audio to be detected and obtain the first note sequence corresponding to the standard audio.
[0059] The specific method for obtaining the first note sequence can be flexibly set according to actual needs. For example, in one example, the first note sequence corresponding to the standard audio can be found from the corresponding storage location. That is to say, the standard audio can be processed in advance to obtain the first note sequence, and the obtained first note sequence can be stored in the corresponding storage location. After obtaining the audio to be detected and determining the standard audio according to the audio to be detected, the first note sequence can be directly found from the corresponding storage location, so as to improve the response speed; in another example, after obtaining the audio to be detected and determining the standard audio according to the audio to be detected, the standard audio can be processed to obtain the first note sequence. Among them, the method for processing the standard audio to obtain the first note sequence includes but is not limited to converting the format of the standard audio to the MIDI (Musical Instrument Digital Interface) format to obtain the first note sequence.
[0060] The specific method for obtaining the second note sequence can be flexibly set according to actual needs. Among them, in order to make the comparison between the first note sequence and the second note sequence more referenceable, the method for processing the audio to be detected to obtain the second note sequence can be the same as the method for processing the standard audio to obtain the first note sequence. For example, if the method for processing the standard audio to obtain the first note sequence is to convert the format of the standard audio to the MIDI format to obtain the first note sequence; then the method for processing the audio to be detected to obtain the second note sequence can be to convert the format of the audio to be detected to the MIDI format to obtain the second note sequence.
[0061] In some embodiments, to further improve the accuracy of the pitch matching parameter, the audio to be detected can also be aligned with the standard audio in terms of time. The specific alignment method can be flexibly set according to actual needs. For example, the DTW (Dynamic Time Warping) algorithm can be used to adjust the audio to be detected so that it is aligned with the standard audio. After the alignment process, the second note sequence corresponding to the audio to be detected is obtained.
[0062] Step S120: Select the first notes with a duration greater than or equal to the first threshold from the first note sequence, and find multiple second notes in the second note sequence whose start and end times match those of the first notes.
[0063] The first threshold is used to determine the time of the first note, and its specific value can be flexibly set according to actual needs. For example, it can be 1.5 seconds, etc.
[0064] Due to the instability of the fundamental frequency of vocalization, if there is a vibrato in the audio provided by the user, it will cause the same note to be divided into multiple parts during the process of obtaining the second note sequence corresponding to the audio to be detected. For example, as shown in Figure 3 In the first note sequence 31 corresponding to the standard audio, there is a first note 311 with a relatively long time. In the audio to be detected, due to the existence of "vibrato", there are multiple second notes 321 with relatively short times at the corresponding position of the first note 311 in the second note sequence 32 corresponding to the audio to be detected. Moreover, if the duration of a certain note (i.e., the time difference between the start time and the end time of the note) is long, it is easy to have a vibrato. Therefore, to reduce the influence of "vibrato" on the pitch matching parameter, in this embodiment, the first notes with a duration greater than or equal to the first threshold are selected from the first note sequence, that is, the first notes with a relatively long duration are selected from the first note sequence, and then multiple second notes whose start and end times match those of the first notes are found from the second note sequence.
[0065] Among them, the start and end times of the second note matching those of the first note include at least one of the following situations:
[0066] First, the second note whose start and end times are within the start and end time range of the first note, that is, the start time of the second note is greater than or equal to the start time of the first note, and the end time of the second note is less than or equal to the end time of the first note. For example, if the start and end time range of the first note is from 2 minutes and 05 seconds to 5 minutes and 20 seconds, then both the start time and the end time of the second note are within the range from 2 minutes and 05 seconds to 5 minutes and 20 seconds;
[0067] Second, notes where the time difference between the start time of the second note and the start time of the first note is less than a preset first duration threshold;
[0068] Third, notes where the time difference between the end time of the second note and the end time of the first note is less than a preset first duration threshold. Among them, the specific value of the preset first duration threshold can be flexibly set according to actual needs. For example, it can be set to 0.1 seconds, etc. To improve the accuracy, the first duration threshold is less than the first threshold.
[0069] Step S130, screen out the first target notes from multiple second notes whose pitches match the pitch of the first note, and determine the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the duration of the first note.
[0070] Among them, the pitch of the second note matching the pitch of the first note can be that the difference between the pitch of the first note and the pitch of the first note is less than or equal to the pitch threshold, and the pitch threshold can be flexibly set according to actual needs. For example, it can be set to one semitone.
[0071] In this embodiment, after determining the first note and multiple second notes corresponding to the first note, screen out the notes from the multiple second notes whose pitches match the pitch of the first note, use the screened notes as the first target notes, and determine the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the first note.
[0072] Step S140, determine the pitch matching parameter of the audio to be detected according to the pitch similarity.
[0073] After determining the pitch similarity, in this embodiment, the pitch matching parameter of the audio to be detected can also be determined according to the pitch similarity. Among them, the pitch similarity of the audio to be detected can be directly used as the pitch matching parameter of the audio to be detected. Of course, the pitch similarity of the audio to be detected can also be processed to obtain the pitch matching parameter of the audio to be detected. The specific processing method can be flexibly set according to actual needs. For example, the pitch matching parameter can be in percentage system (that is, the full score is 100 points). Correspondingly, the percentage of the pitch similarity can be determined, and the numerator of the percentage can be used as the pitch matching parameter.
[0074] In this embodiment, when detecting the pitch matching degree between the audio to be detected and the corresponding standard audio, if the standard audio contains a first note with a relatively long time, the pitch matching parameter is determined according to the duration of the first note and the duration of the first target note in the audio to be detected whose start and end times and pitches all match the first note, so as to reduce the influence of "vibrato" and the like on the pitch matching detection and improve the accuracy of the pitch matching detection.
[0075] See Figure 4 ,Figure 4 For Figure 2 the schematic diagram of step S110 in the illustrated embodiment in an exemplary embodiment. As Figure 4 shown, the process of obtaining the first note sequence corresponding to the standard audio and the second note sequence corresponding to the audio to be detected may include steps S111 - S112, which are introduced in detail as follows:
[0076] Step S111, if the type of the audio to be detected is a mixed audio, then perform sound source separation on the audio to be detected to obtain a dry audio.
[0077] It should be noted that dry sound is pure human voice without music.
[0078] In some embodiments, the audio matching detection method can be applied to detect human voice audio. For example, the audio to be detected can be a song sung by a user, a passage spoken, etc. When recording the audio to be detected, background music and other noises may be recorded. To avoid the influence of noises on the pitch matching detection, in this embodiment, it can also be determined whether the type of the audio to be detected is a mixed audio. If so, perform sound source separation on the audio to be detected to obtain a dry audio.
[0079] Step S112, extract the fundamental frequency of the dry audio, and quantize the extracted fundamental frequency to obtain the second note sequence.
[0080] After obtaining the dry audio, the fundamental frequency of the dry audio can be extracted, and the extracted fundamental frequency can be quantized to obtain the second note sequence.
[0081] Among them, the method of extracting the fundamental frequency of the dry audio can be flexibly set according to actual needs. For example, the pYIN algorithm can be used to extract the fundamental frequency from the dry audio. Among them, the pYIN algorithm is an algorithm for extracting the fundamental frequency of audio.
[0082] In one example, the dry audio of the audio to be detected can be as Figure 5 shown. The process of extracting the fundamental frequency from the dry audio of the audio to be detected and quantizing the extracted fundamental frequency to obtain the second note sequence can be referred to Figure 6 shown, Figure 6 where the curve is the fundamental frequency and the straight line is the quantized note.
[0083] It should be noted that in some embodiments, in order to further improve the accuracy of the pitch matching parameter, sound source separation can also be performed on the standard audio to obtain the dry audio corresponding to the standard audio, extract the fundamental frequency of the dry audio corresponding to the standard audio, and quantize the extracted fundamental frequency to obtain the first note sequence. Among them, this process can be pre - processed or performed after determining the audio to be detected.
[0084] In this embodiment, if the type of the audio to be detected is a mixed audio, the audio to be detected is subjected to sound source separation to obtain a dry audio, the fundamental frequency of the dry audio is extracted, and the extracted fundamental frequency is quantized to obtain a second note sequence, thereby avoiding the influence of noises such as accompaniment on the pitch matching detection and improving the accuracy of the pitch matching parameter.
[0085] See Figure 7 , Figure 7 For Figure 2 the schematic diagram of step S130 in the embodiment shown in Figure 7 As shown, the process of determining the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the duration of the first note may include steps S131 - S133, which are introduced in detail as follows:
[0086] Step S131: Determine a third note other than the first note from the first note sequence.
[0087] In this embodiment, in addition to comparing the first note included in the first note sequence with the notes in the second note sequence, it is also necessary to compare the notes other than the first note in the first note sequence with the notes in the second note sequence. Therefore, the notes other than the first note can be determined from the first note sequence first, and the determined notes are used as the third notes.
[0088] Step S132: Determine a second target note in the second note sequence whose start time matches the start time of the third note and whose pitch matches the pitch of the third note.
[0089] Among them, the matching of the start times may mean that the time difference between the start times is less than a preset second duration threshold, and the matching of the pitches may mean that the difference between the pitches is less than or equal to a pitch threshold. The second duration threshold can be flexibly set according to actual needs. For example, it can be set to 0.5 seconds.
[0090] In this embodiment, after determining the third note, a second target note in the second note sequence whose start time matches the start time of the third note and whose pitch matches the pitch of the third note is determined. In one example, a note whose difference from the start time of the third note does not exceed 0.5 seconds and whose pitch difference from the third note does not exceed one and a half tones can be found from the second note sequence, and the found note is used as the second target note.
[0091] To improve accuracy, in some embodiments, multiple fourth notes other than the first note may also be determined from the second note sequence, and the fourth note uniquely corresponding to each third note is determined. Among them, the fourth note uniquely corresponding to the third note may be: among the multiple fourth notes, the fourth note that matches the start time of the third note and has the smallest start time difference; then, it is determined whether the pitch of the fourth note matches the corresponding third note. If they match, the fourth note is used as the second target note; that is, for the third notes in the first note sequence other than the first note and the fourth notes in the second note sequence other than the second note, they are compared according to a relatively strict one-to-one correspondence relationship to determine whether the fourth note is the second target note.
[0092] Step S133, determine the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note, the duration of the first note, the duration of the second target note, and the duration of the third note.
[0093] After determining the first target note, the first note, the second target note, and the third note, the pitch similarity between the audio to be detected and the standard audio can be determined according to the duration of the first target note, the duration of the first note, the duration of the second target note, and the duration of the third note. Among them, the specific determination method can be flexibly set according to actual needs.
[0094] In one embodiment, step S133 may include: obtaining a first ratio of the duration of the first target note to the duration of the first note, and a second ratio of the duration of the second target note to the duration of the third note; performing a weighted sum on the first ratio and the second ratio to obtain the pitch similarity between the audio to be detected and the standard audio.
[0095] Among them, if the number of the first target note, the first note, the second target note, and the third note is all multiple, the first ratio is the ratio of the sum of the durations of multiple first target notes to the sum of the durations of multiple first notes; the second ratio is the ratio of the sum of the durations of multiple second target notes to the sum of the durations of multiple third notes.
[0096] The weights corresponding to the first ratio and the second ratio can be flexibly set according to actual needs. For example, they can be determined according to the durations occupied by the first note and the third note. For example, the longer the duration, the larger the corresponding weight value.
[0097] In another embodiment, step S133 may include: taking the sum of the duration of the first target note and the duration of the second target note as a first value, and taking the sum of the duration of the first note and the duration of the third note as a second value, and taking the ratio of the first value to the second value as the pitch similarity between the audio to be detected and the standard audio.
[0098] In this embodiment, a third note other than the first note is determined from the first note sequence, and a second target note is determined from the second note sequence, where the start time of the second target note matches the start time of the third note and the pitch of the second target note matches the pitch of the third note. The pitch similarity between the audio to be detected and the standard audio is determined based on the duration of the first target note, the duration of the first note, the duration of the second target note, and the duration of the third note, thereby improving the accuracy of pitch matching detection.
[0099] See Figure 8 , Figure 8 which is a flowchart of obtaining the rhythm matching parameters of the audio to be detected shown in an exemplary embodiment. As Figure 8 shown, the audio matching detection method may further include step S210-step S230, which are introduced in detail as follows:
[0100] Step S210, obtain the first cepstrum corresponding to the standard audio and the second cepstrum corresponding to the audio to be detected.
[0101] It should be noted that the cepstrum is a signal spectrum obtained by performing the inverse Fourier transform on the Fourier transform spectrum of a signal after logarithmic operation.
[0102] In this embodiment, the first cepstrum corresponding to the standard audio and the second cepstrum corresponding to the audio to be detected can be obtained. Among them, the specific obtaining method can be flexibly set according to actual needs.
[0103] In some embodiments, the specific process of obtaining the second cepstrum corresponding to the audio to be detected may include step 211-step 213, which are introduced in detail as follows:
[0104] Step 211, perform Fourier transform on the audio to be detected to obtain a frequency spectrum.
[0105] In this embodiment, in order to determine the rhythm matching degree of the audio to be detected, the signal of the obtained audio to be detected can be subjected to Fourier transform to obtain the frequency spectrum of the audio to be detected.
[0106] Step 212, filter the obtained frequency spectrum according to the filtering information, and perform logarithmic operation on the filtered frequency spectrum to obtain a logarithmic spectrum; where the filtering information includes a variety of filtering parameters, and the filtering parameters corresponding to different frequencies are different.
[0107] Since the human auditory system has different sensitivities to audio signals of different frequencies and only focuses on certain specific frequency components, that is, the human auditory system is selective about frequencies. Therefore, in order to improve the accuracy of rhythm matching detection, in this embodiment, the obtained spectrum can be filtered according to the filtering information, and a logarithmic operation is performed on the filtered spectrum to obtain a log spectrum. Among them, the filtering information includes various filtering parameters, and the filtering parameters corresponding to different frequencies are different, so that the spectrum obtained after filtering is closer to the audio signal received by the human auditory system.
[0108] It should be noted that the filtering information can be flexibly set according to actual needs. In one example, since the Mel-Frequency Cepstral Coefficients (MFCC) takes into account the human auditory characteristics, the filtering parameters corresponding to Mel can be used to filter the obtained spectrum, thereby mapping the linear spectrum into the Mel non-linear spectrum based on auditory perception.
[0109] Step 213: Perform an inverse Fourier transform on the log spectrum to obtain a second cepstrum.
[0110] Performing an inverse Fourier transform on the obtained log spectrum can obtain a second cepstrum.
[0111] The method for obtaining the first cepstrum corresponding to the standard audio can be flexibly set according to actual needs. In one example, the first cepstrum corresponding to the standard audio can be found from the corresponding storage location. That is to say, the standard audio can be pre-processed to obtain the first cepstrum and stored in the corresponding storage location, and then the first cepstrum corresponding to the standard audio can be directly found from the corresponding storage location; or, in another example, after obtaining the audio to be detected, the standard audio can be determined, and then the standard audio can be processed to obtain the first cepstrum. Among them, the method for processing the standard audio to obtain the first cepstrum is similar to the method for processing the audio to be detected to obtain the second cepstrum. For example, the method for processing the standard audio to obtain the first cepstrum can be similar to steps 211 - 212, that is, the standard audio can be Fourier-transformed to obtain the corresponding spectrum, the spectrum of the standard audio can be filtered according to the filtering information to obtain the log spectrum of the standard audio, and then an inverse Fourier transform is performed on the log spectrum of the standard audio to obtain the first cepstrum.
[0112] Step S220: Determine the rhythm similarity between the audio to be detected and the standard audio according to the similarity between the first cepstrum and the second cepstrum.
[0113] After obtaining the first cepstrum and the second cepstrum, the rhythm similarity between the audio to be detected and the standard audio can be determined according to the similarity between the first cepstrum and the second cepstrum. Among them, the similarity between the first cepstrum and the second cepstrum can be directly used as the rhythm similarity between the audio to be detected and the standard audio, or the similarity between the first cepstrum and the second cepstrum can be processed and then used as the rhythm similarity between the audio to be detected and the standard audio. The specific processing method can be flexibly set according to actual needs.
[0114] In some embodiments, in order to improve the accuracy of rhythm matching detection, the first cepstrum and the second cepstrum can be aligned in time first, and then the rhythm similarity between the audio to be detected and the standard audio can be determined according to the similarity between the aligned first cepstrum and the second cepstrum. Among them, the specific alignment method can be flexibly set according to actual needs. For example, the DTW algorithm can be used to align the first cepstrum and the second cepstrum.
[0115] Step S230, determine the rhythm matching parameter of the audio to be detected according to the rhythm similarity.
[0116] After determining the rhythm similarity between the audio to be detected and the standard audio, the rhythm matching parameter of the audio to be detected is determined according to the rhythm similarity between the audio to be detected and the standard audio. Among them, the rhythm similarity between the audio to be detected and the standard audio can be directly used as the rhythm matching parameter of the audio to be detected, or the rhythm similarity between the audio to be detected and the standard audio can be processed to obtain the rhythm matching parameter of the audio to be detected. The specific processing method can be flexibly set according to actual needs. For example, the rhythm matching parameter can be in percentage system (that is, the highest score is 100 points). Correspondingly, the percentage of the rhythm similarity can be determined, and the numerator of the percentage can be used as the rhythm matching parameter.
[0117] In this embodiment, the first cepstrum corresponding to the standard audio and the second cepstrum corresponding to the audio to be detected are obtained. The rhythm similarity between the audio to be detected and the standard audio is determined according to the similarity between the first cepstrum and the second cepstrum, and the rhythm matching parameter of the audio to be detected is determined according to the rhythm similarity. Thus, the rhythm matching degree of the audio to be detected can be determined. Moreover, the rhythm matching degree is determined based on the similarity between the cepstrums corresponding to the audio to be detected and the standard audio respectively, which can reduce the influence of the pitch deviation on the rhythm matching degree and improve the accuracy of rhythm matching detection.
[0118] See Figure 9 , Figure 9 is Figure 8 the flowchart of step S220 in the embodiment shown in an exemplary embodiment. As Figure 9 shown, determining the rhythm similarity between the audio to be detected and the standard audio according to the similarity between the first cepstrum and the second cepstrum may include step S221-step S223, which are introduced in detail as follows:
[0119] Step S221: Obtain various mapping relationships between the first audio frames included in the first cepstrum and the second audio frames included in the second cepstrum, and calculate the differences between the first cepstrum and the second cepstrum under different mapping relationships; wherein, the differences include the differences between the first audio frames and the corresponding second audio frames.
[0120] It should be noted that the differences between the first cepstrum and the second cepstrum include the energy spectrum differences between the first audio frames and the corresponding second audio frames; if the number of the first audio frames and the number of the second audio frames are multiple, the differences between each second cepstrum and the corresponding first cepstrum can be determined first, and then the determined differences can be summed or averaged to obtain the differences between the first cepstrum and the second cepstrum.
[0121] In some embodiments, if the first cepstrum and the second cepstrum are obtained by filtering based on Mel corresponding filter parameters, the MFCC features can be respectively extracted from the first cepstrum and the second cepstrum to obtain the first MFCC feature sequence corresponding to the first cepstrum and the second MFCC feature sequence corresponding to the second cepstrum. An MFCC feature is the feature corresponding to an audio frame, and an MFCC feature includes multiple feature vectors; then, according to the sum of the squares of the differences between the feature vectors of the MFCC features, the differences between the corresponding first audio frame and the second audio frame are determined.
[0122] Since the mapping relationships between the first audio frames and the second audio frames are different, the differences between the first cepstrum and the second cepstrum are also different. In order to determine the minimum difference between the first cepstrum and the second cepstrum, in this embodiment, various mapping relationships between the first audio frames and the second audio frames can be obtained, and the differences between the first cepstrum and the second cepstrum under different mapping relationships can be calculated.
[0123] Step S222: Determine the minimum difference from the calculated differences, and select the target mapping relationship corresponding to the minimum difference from the various mapping relationships.
[0124] Determine the minimum difference from the calculated differences, and select the mapping relationship corresponding to the minimum difference from the various mapping relationships, and use the selected mapping relationship as the target mapping relationship.
[0125] It should be noted that for steps S221 - S222, the DTW algorithm can be used to determine the target mapping relationship between the first audio frame and the second audio frame.
[0126] Step S223: Determine the similarity between the first cepstrum and the second cepstrum according to the target mapping relationship, and use the determined similarity as the rhythm similarity between the audio to be detected and the standard audio.
[0127] See Figure 10 as shown Figure 10Among them, the abscissa is the number of frames of the first audio frame, and the ordinate is the number of frames of the second audio frame. If the rhythm of the audio to be detected is exactly the same as that of the standard audio, the mapping relationship between the first audio frame and the second audio frame is the diagonal line 1001 in the figure, that is, the nth frame in the audio to be detected matches the nth frame in the standard audio, where n is an integer greater than or equal to 1; if the mapping relationship between the first audio frame and the second audio frame is Figure 10 the solid line 1002 in
[0128] it indicates that there is a difference between the rhythm of the audio to be detected and that of the standard audio. In order to determine the degree of difference, after determining the target mapping relationship, the similarity between the first cepstrum and the second cepstrum can be determined according to the target mapping relationship, and the determined similarity can be used as the rhythm similarity between the audio to be detected and the standard audio. It should be noted that the number of frames represents the position of the audio frame in the audio. For example, the 1st frame, the 2nd frame, the 3rd frame, etc.
[0129] Among them, the specific method for determining the similarity between the first cepstrum and the second cepstrum according to the target mapping relationship can be flexibly set according to actual needs.
[0130] Step 310, screen out target audio frames with a time difference less than the second threshold from multiple second audio frames according to the target mapping relationship.
[0131] Among them, the second threshold can be flexibly set according to actual needs. For example, it can be 3, etc.
[0132] If in the target mapping relationship, the time difference between the second audio frame and the corresponding first audio frame is large, it indicates that the rhythm of the audio to be detected is inconsistent with that of the standard audio. Therefore, target audio frames with a time difference less than the second threshold can be screened out from multiple second audio frames according to the target mapping relationship.
[0133] Among them, the time difference between the second audio frame and the corresponding first audio frame can be the difference between the start times of the second audio frame and the corresponding first audio frame.
[0134] Alternatively, the time difference between the second audio frame and the corresponding first audio frame can be the difference in the number of frames between the second audio frame and the corresponding first audio frame. The difference in the number of frames can characterize the time difference between audio frames. For example, assuming that in the target mapping relationship, the first frame in the second cepstrum corresponds to the tenth frame in the first cepstrum, the second frame in the second cepstrum corresponds to the twentieth frame in the first cepstrum, and the second threshold is 15 frames. Since the time difference between the first frame in the second cepstrum and the tenth frame in the first cepstrum is 9 frames, and the time difference between the second frame in the second cepstrum and the twentieth frame in the first cepstrum is 18 frames, therefore, the first frame in the second cepstrum is the target audio frame, and the second frame in the second cepstrum is not the target audio frame.
[0135] Step 320, determine the similarity between the first cepstrum and the second cepstrum according to the number of target audio frames and the number of first audio frames.
[0136] If the number of target audio frames is larger, it indicates that the similarity between the first cepstrum and the second cepstrum is higher. Therefore, the similarity between the first cepstrum and the second cepstrum can be determined according to the number of target audio frames and the number of first audio frames. Among them, the ratio of the number of target audio frames to the number of first audio frames can be used as the similarity between the first cepstrum and the second cepstrum.
[0137] In another implementation, under the condition that the number of first audio frames and the number of second audio frames are both multiple, the process of determining the similarity between the first cepstrum and the second cepstrum according to the target mapping relationship may include Step 410 - Step 430, which are introduced in detail as follows:
[0138] Step 410, respectively obtain the time differences between multiple first audio frames and the corresponding second audio frames according to the target mapping relationship.
[0139] Among them, the calculation method of the time difference between the second audio frame and the corresponding first audio frame can refer to the foregoing description and will not be elaborated here.
[0140] Step 420, sum up the obtained time differences to obtain the total time difference.
[0141] After obtaining the time differences corresponding to multiple second audio frames respectively, the obtained time differences can be summed up to obtain the total time difference.
[0142] Step 430, determine the similarity between the first cepstrum and the second cepstrum according to the total time difference.
[0143] Among them, the smaller the total time difference, the more similar the first cepstrum and the second cepstrum. Therefore, the similarity between the first cepstrum and the second cepstrum can be determined according to the total time difference. Among them, the total time difference and the similarity can be inversely proportional.
[0144] In this embodiment, multiple mapping relationships between the first audio frames included in the first cepstrum and the second audio frames included in the second cepstrum are obtained, and the differences between the first cepstrum and the second cepstrum under different mapping relationships are calculated; wherein, the differences include the differences between the first audio frames and the corresponding second audio frames; the minimum difference is determined from the calculated differences, and the target mapping relationship corresponding to the minimum difference is selected from the multiple mapping relationships; the similarity between the first cepstrum and the second cepstrum is determined according to the target mapping relationship, and the determined similarity is used as the rhythm similarity between the audio to be detected and the standard audio, so as to improve the accuracy of subsequent rhythm matching parameters.
[0145] In an exemplary embodiment, after Figure 8 the step S230 shown, the audio matching detection method may further include: performing weighted summation on the rhythm matching parameter and the pitch matching parameter to obtain the comprehensive matching parameter of the audio to be detected. Wherein, the weight values corresponding to the rhythm matching parameter and the pitch matching parameter can be flexibly set according to actual needs.
[0146] In an exemplary embodiment, referring to Figure 11 shown, Figure 11 is Figure 2 the flowchart of step S110 in the embodiment shown in an exemplary embodiment. As Figure 11 shown, the process of obtaining the first note sequence corresponding to the standard audio and the second note sequence corresponding to the audio to be detected may include steps S510 - S530, which are introduced in detail as follows:
[0147] Step S510, obtaining the first note sequence corresponding to each of the multiple first sub-audios included in the standard audio; wherein, the multiple first sub-audios include multiple sub-audios obtained by segmenting the standard audio according to a preset segmentation method.
[0148] It should be noted that the preset segmentation method can be flexibly set according to actual needs. For example, it can be divided according to time periods, and the standard audio is divided into multiple sub-audios with the same duration; or, if the standard audio is a human voice audio, which usually includes a passage spoken by the user, the standard audio can be segmented according to the pause time of the speech, so as to obtain sub-audios containing different sentences. For example, one sentence can correspond to one sub-audio.
[0149] In this embodiment, the specific method for obtaining the first note sequences corresponding to the multiple first sub-audios included in the standard audio can be flexibly set according to actual needs. For example, in one example, the first note sequences corresponding to the multiple first sub-audios can be obtained from the corresponding storage locations; that is, the standard audio is segmented into multiple first sub-audios in advance according to a preset segmentation method, the multiple first sub-audios are respectively processed to obtain the first note sequences corresponding to each of them, and then stored, so as to facilitate directly obtaining the first note sequences from the corresponding storage locations subsequently. In another example, the standard audio can be segmented into multiple first sub-audios according to a preset segmentation method, and the multiple first sub-audios are respectively processed to obtain the first note sequences corresponding to the multiple first sub-audios.
[0150] Step S520: Segment the audio to be detected according to a preset segmentation method to obtain multiple second sub-audios.
[0151] In this embodiment, the audio to be detected is segmented using the same segmentation method as the standard audio, thereby obtaining multiple second sub-audios.
[0152] Step S530: Process the multiple second sub-audios respectively to obtain the second note sequences corresponding to the multiple second sub-audios.
[0153] In this embodiment, after obtaining the multiple second sub-audios, the multiple second sub-audios can be processed respectively to obtain the second note sequences corresponding to the multiple second sub-audios. Among them, the method for processing each second sub-audio to obtain the corresponding second note sequence can refer to the foregoing description (for example, the foregoing steps S111 - step S112), and will not be elaborated here.
[0154] In this embodiment, the first note sequences corresponding to the multiple first sub-audios included in the standard audio can be obtained; among them, the multiple first sub-audios include multiple sub-audios obtained by segmenting the standard audio according to a preset segmentation method; the audio to be detected is segmented according to the preset segmentation method to obtain multiple second sub-audios; the multiple second sub-audios are processed respectively to obtain the second note sequences corresponding to the multiple second sub-audios, thereby facilitating subsequent processing of the sub-audios and improving the processing speed.
[0155] In some embodiments, if the number of first sub-audios and the number of second sub-audios are both multiple, the pitch similarity between the audio to be detected and the standard audio may include the audio similarities corresponding to multiple audio combinations respectively, where each audio combination includes a first sub-audio included in the audio to be detected and a second sub-audio included in the standard audio, and the start time of the first sub-audio matches the start time of the second sub-audio. In one example, each note combination may include a first sub-audio and a second sub-audio, and the second sub-audio is the sub-audio with the smallest start time difference from the first sub-audio in the audio to be detected, so as to compare the first sub-audio and the second sub-audio one by one to improve the accuracy.
[0156] To determine the audio similarities corresponding to multiple audio combinations respectively, Figure 2 In step S120 in the illustrated embodiment, the process of screening out the first target note whose pitch matches the pitch of the first note from multiple second notes and determining the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the duration of the first note may include: for each audio combination, screening out the first target note whose pitch matches the pitch of the first note corresponding to the audio combination from the multiple second notes corresponding to the audio combination, and determining the pitch similarity between the first sub-audio and the second sub-audio in the audio combination according to the duration of the first target note corresponding to the audio combination and the duration of the first note, and using the obtained pitch similarity as the similarity of the corresponding audio combination.
[0157] Under the condition that the number of first sub-audios and the number of second sub-audios are both multiple and the pitch similarity between the audio to be detected and the standard audio includes the audio similarities corresponding to multiple audio combinations respectively, Figure 2 In step S130 in the illustrated embodiment, the process of determining the pitch matching parameter of the audio to be detected according to the pitch similarity may include: determining the pitch matching parameter of the corresponding second sub-audio according to the pitch similarity of each audio combination; performing weighted summation on the determined pitch matching parameters to obtain the pitch matching parameter of the audio to be detected.
[0158] Among them, the specific process of determining the pitch matching parameter of the corresponding second sub-audio according to the pitch similarity of the audio combination can be referred to the foregoing description and will not be elaborated here. For example, the pitch similarity of the audio combination can be directly used as the pitch matching parameter of the corresponding second sub-audio.
[0159] In some embodiments, when the number of the first sub-audios and the number of the second sub-audios are both multiple, in the foregoing steps S210-S230, the first cepstrum may include first sub-cepstrums respectively corresponding to the multiple first sub-audios, and the second cepstrum may include second sub-cepstrums respectively corresponding to the multiple second sub-audios. Thus, the rhythm similarity between the corresponding first sub-audio and the second sub-audio is determined according to the similarity between the first sub-cepstrum and the corresponding second sub-cepstrum, the rhythm matching parameter of the corresponding second sub-audio is determined according to the rhythm similarity between the first sub-audio and the corresponding second sub-audio, and then the rhythm matching parameters respectively corresponding to the multiple second sub-audios are weighted and summed to obtain the rhythm matching parameter of the audio to be detected. In this way, not only can the rhythm matching parameter of the audio to be detected be obtained, but also the rhythm matching parameters of each segmented audio can be obtained.
[0160] In this embodiment, not only can the pitch matching parameter of the audio to be detected be determined, but also the pitch matching parameters of each second sub-audio included in the audio to be detected can be determined, so that the user can understand the pitch matching degree of each segment.
[0161] The following takes the audio matching detection method of the present application applied to the song scoring scenario as an example for description. Among them, the process of determining the pitch matching parameter can be referred to Figure 12 as shown, including:
[0162] Obtain the first dry audio and the second dry audio. Among them, the audio to be detected may be an audio obtained by a user's singing of a certain song, and the standard audio may be the original audio of the song; the source separation can be performed on the audio to be detected and the standard audio respectively to obtain the first dry audio corresponding to the standard audio and the second dry audio corresponding to the audio to be detected.
[0163] Fundamental frequency extraction: The fundamental frequency can be extracted from the first dry audio and the second dry audio respectively through the pYIN algorithm to obtain the first fundamental frequency corresponding to the first dry audio and the second fundamental frequency corresponding to the second dry audio.
[0164] Audio transcription: The first dry audio and the second dry audio can be quantified respectively through the Tony algorithm to obtain the first note sequence corresponding to the first dry audio and the second note sequence corresponding to the second dry audio. Among them, the formats corresponding to the first note sequence and the second note sequence may be MIDI.
[0165] Audio calibration: Select the first notes with a duration greater than or equal to the first threshold from the first note sequence, find multiple second notes in the second note sequence whose start and end times match the start and end times of the first note, select the notes with the pitch matching the pitch of the first note from the multiple second notes, and obtain the first ratio of the duration of the first target note to the duration of the first note.
[0166] Determine the pitch matching parameter: Determine the third note other than the first note from the first note sequence, and determine the second target note that meets the preset conditions from the second note sequence; Obtain the second ratio of the duration of the second target note to the duration of the third note, and perform weighted summation on the first ratio and the second ratio to obtain the pitch matching parameter. The preset condition may be that the pitch difference from the third note does not exceed one and a half tones and the start time difference does not exceed 0.5 s.
[0167] Among them, the second target note can be determined according to the maximum matching algorithm of the bipartite graph. For example, the fourth note other than the first note sequence can be determined from the second note sequence, and the third note is used as a node in one subset of the bipartite graph, and the fourth note is used as a node in another subset of the bipartite graph to form a bipartite graph. Then, the maximum matching algorithm is executed. During the execution of the maximum matching algorithm, the nodes in the two subsets are compared. If the pitch difference does not exceed one and a half tones and the start time difference does not exceed 0.5 s, it is determined that the two nodes match. Thus, the second target note is determined.
[0168] The process of determining the rhythm matching parameter can be referred to Figure 13 as shown, including:
[0169] Extract the MFCC feature sequence: respectively obtain the first MFCC feature sequence corresponding to the first dry audio, and the second MFCC feature sequence corresponding to the second dry audio.
[0170] Perform dynamic time warping adjustment: Based on the DTW algorithm, when the difference between the first MFCC feature sequence and the second MFCC feature sequence is the smallest, determine the target mapping relationship between the first audio frame in the first MFCC feature sequence and the second audio frame in the second MFCC feature sequence.
[0171] Determine the rhythm matching parameter: Calculate the difference in the number of frames between each first audio frame and the corresponding second audio frame according to the target mapping relationship. If the difference in the number of frames is less than or equal to the second threshold, the corresponding second audio frame is used as the target audio frame, and the ratio of the number of target audio frames to the number in the first audio is used as the rhythm matching parameter.
[0172] By determining the rhythm matching parameter and the pitch matching parameter in the above manner, the accuracy can be improved.
[0173] Refer to Figure 14 , Figure 14 is a block diagram of an audio matching detection device shown in an exemplary embodiment of the present application. As Figure 14 shown, the device includes:
[0174] An acquisition module 1401, configured to acquire a first note sequence corresponding to a standard audio and a second note sequence corresponding to an audio to be detected;
[0175] A search module 1402, configured to filter out first notes with a duration greater than or equal to a first threshold from a first note sequence, and find a plurality of second notes in a second note sequence whose start and end times match the start and end times of the first notes;
[0176] A similarity determination module 1403, configured to filter out first target notes from the plurality of second notes whose pitches match the pitch of the first note, and determine the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the duration of the first note;
[0177] A matching detection module 1404, configured to determine the pitch matching parameter of the audio to be detected according to the pitch similarity.
[0178] In another exemplary embodiment, the apparatus further includes:
[0179] An inverse cepstrum acquisition module, configured to acquire a first inverse cepstrum corresponding to the standard audio and a second inverse cepstrum corresponding to the audio to be detected;
[0180] A first determination module, configured to determine the rhythm similarity between the audio to be detected and the standard audio according to the similarity between the first inverse cepstrum and the second inverse cepstrum;
[0181] A second determination module, configured to determine the rhythm matching parameter of the audio to be detected according to the rhythm similarity.
[0182] In another exemplary embodiment, the inverse cepstrum acquisition module includes:
[0183] A spectrum determination module, configured to perform a Fourier transform on the audio to be detected to obtain a spectrum;
[0184] A logarithmic spectrum determination module, configured to filter the obtained spectrum according to filtering information and perform a logarithmic operation on the filtered spectrum to obtain a logarithmic spectrum; wherein, the filtering information includes a variety of filtering parameters, and the filtering parameters corresponding to different frequencies are different;
[0185] An inverse cepstrum determination module, configured to perform an inverse Fourier transform on the logarithmic spectrum to obtain a second inverse cepstrum.
[0186] In another exemplary embodiment, the first determination module includes:
[0187] A difference determination module, configured to obtain a variety of mapping relationships between a first audio frame included in the first inverse cepstrum and a second audio frame included in the second inverse cepstrum, and calculate the differences between the first inverse cepstrum and the second inverse cepstrum under different mapping relationships; wherein, the differences include the differences between the first audio frame and the corresponding second audio frame;
[0188] A mapping relationship determination module, configured to determine the minimum difference from the calculated differences and select a target mapping relationship corresponding to the minimum difference from multiple mapping relationships;
[0189] A third determination module, configured to determine the similarity between the first cepstrum and the second cepstrum according to the target mapping relationship, and use the determined similarity as the rhythm similarity between the audio to be detected and the standard audio.
[0190] In another exemplary embodiment, under the condition that the numbers of the first audio frames and the second audio frames are both multiple, the third determination module includes:
[0191] A screening module, configured to screen out target audio frames from multiple second audio frames whose time difference from the corresponding first audio frames is less than a second threshold according to the target mapping relationship;
[0192] A fourth determination module, configured to determine the similarity between the first cepstrum and the second cepstrum according to the number of the target audio frames and the number of the first audio frames.
[0193] In another exemplary embodiment, under the condition that the numbers of the first audio frames and the second audio frames are both multiple, the third determination module includes:
[0194] A time difference acquisition module, configured to acquire the time differences between multiple first audio frames and the corresponding second audio frames respectively according to the target mapping relationship;
[0195] A total time difference acquisition module, configured to sum up the acquired time differences to obtain a total time difference;
[0196] A fourth determination module, configured to determine the similarity between the first cepstrum and the second cepstrum according to the total time difference.
[0197] In another exemplary embodiment, the apparatus further includes:
[0198] A comprehensive matching detection module, configured to perform weighted summation on the rhythm matching parameter and the pitch matching parameter to obtain a comprehensive matching parameter of the audio to be detected.
[0199] In another exemplary embodiment, the similarity determination module 1403 includes:
[0200] A note determination module, configured to determine a third note other than the first note from the first note sequence;
[0201] A target note determination module, configured to determine a second target note from the second note sequence whose start time matches the start time of the third note and whose pitch matches the pitch of the third note;
[0202] A pitch similarity determination module, configured to determine the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note, the duration of the first note, the duration of the second target note, and the duration of the third note.
[0203] In another exemplary embodiment, the pitch similarity determination module includes:
[0204] A ratio determination module, configured to obtain a first ratio of the duration of the first target note to the duration of the first note, and a second ratio of the duration of the second target note to the duration of the third note;
[0205] A weighted summation module, configured to perform weighted summation on the first ratio and the second ratio to obtain the pitch similarity between the audio to be detected and the standard audio.
[0206] In another exemplary embodiment, the acquisition module 1401 includes:
[0207] A separation module, configured to perform sound source separation on the audio to be detected to obtain a dry audio if the type of the audio to be detected is a mixed audio;
[0208] A transcription module, configured to extract the fundamental frequency of the dry audio and quantize the extracted fundamental frequency to obtain a second note sequence.
[0209] In another exemplary embodiment, the acquisition module 1401 includes:
[0210] A sub-audio acquisition module, configured to obtain first note sequences corresponding to multiple first sub-audios included in the standard audio; wherein, the multiple first sub-audios include multiple sub-audios obtained by segmenting the standard audio according to a preset segmentation method;
[0211] A segmentation module, configured to segment the audio to be detected according to a preset segmentation method to obtain multiple second sub-audios;
[0212] A note sequence determination module, configured to process multiple second sub-audios respectively to obtain second note sequences corresponding to the multiple second sub-audios.
[0213] In another exemplary embodiment, when the quasi-similarity includes audio similarities corresponding to multiple audio combinations, each audio combination includes a first sub-audio included in the audio to be detected and a second sub-audio included in the standard audio, and the start time of the first sub-audio matches the start time of the second sub-audio, the matching detection module 1404 includes:
[0214] A sub-audio matching detection module, configured to determine the pitch matching parameter corresponding to the second sub-audio according to the pitch similarity of each audio combination;
[0215] The audio matching detection module is configured to perform weighted summation on the determined pitch matching parameters to obtain the pitch matching parameters of the audio to be detected.
[0216] It should be noted that the audio matching detection device provided in the above embodiment and the audio matching detection method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be elaborated here.
[0217] An embodiment of the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the audio matching detection method provided in each of the above embodiments.
[0218] Figure 15 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.
[0219] It should be noted that Figure 15 The computer system 1500 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0220] As Figure 15 shown, the computer system 1500 includes a central processing unit (CPU) 1501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1502 or the program loaded from the storage section 1508 into the random access memory (RAM) 1503, such as executing the method described in the above embodiment. In the RAM 1503, various programs and data required for system operation are also stored. The CPU 1501, ROM 1502, and RAM 1503 are connected to each other through a bus 1504. The input / output (I / O) interface 1505 is also connected to the bus 1504.
[0221] The following components are connected to the I / O interface 1505: an input section 1506 including a keyboard, a mouse, etc.; an output section 1507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1508 including a hard disk, etc.; and a communication section 1509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1509 performs communication processing via a network such as the Internet. A drive 1510 is also connected to the I / O interface 1505 as required. A removable medium 1511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1510 as required so that a computer program read from it can be installed into the storage section 1508 as required.
[0222] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 1509, and / or installed from the removable medium 1511. When the computer program is executed by a central processing unit (CPU) 1501, various functions defined in the system of the present application are executed.
[0223] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0224] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0225] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the units themselves in some cases.
[0226] Another aspect of this application also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of an electronic device, the electronic device implements the method as described above. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device.
[0227] Another aspect of this application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and when the computer instructions are executed by a processor, the methods provided in the above various embodiments are implemented. Among them, the computer instructions can be stored in a computer-readable storage medium; the processor of the electronic device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in the above various embodiments.
[0228] The above content is only a preferred exemplary embodiment of this application and is not used to limit the implementation of this application. Those of ordinary skill in the art can easily make corresponding adaptations or modifications according to the main concept and spirit of this application. Therefore, the protection scope of this application should be subject to the protection scope required by the claims.
Claims
1. An audio matching detection method, characterized in that, The method includes: Obtaining a first note sequence corresponding to a standard audio and a second note sequence corresponding to an audio to be detected; Screening out first notes in the first note sequence with a duration greater than or equal to a first threshold, and finding multiple second notes in the second note sequence whose start and end times match the start and end times of the first notes; Screening out first target notes in the multiple second notes whose pitches match the pitches of the first notes, and determining the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target notes and the duration of the first notes; Determining a pitch matching parameter of the audio to be detected according to the pitch similarity; Wherein, the determining the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target notes and the duration of the first notes includes: Determining third notes other than the first notes in the first note sequence; Determining a fourth note in the second note sequence whose start time matches the start time of the third note. If the pitch of the fourth note matches the pitch of the third note, then using the fourth note as the second target note corresponding to the third note; Determining the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target notes, the duration of the first notes, the duration of the second target notes, and the duration of the third notes.
2. The method according to claim 1, characterized in that, The method further includes: Obtaining a first cepstrum corresponding to the standard audio and a second cepstrum corresponding to the audio to be detected; Determining the rhythm similarity between the audio to be detected and the standard audio according to the similarity between the first cepstrum and the second cepstrum; Determining a rhythm matching parameter of the audio to be detected according to the rhythm similarity.
3. The method according to claim 2, wherein The obtaining the first cepstrum corresponding to the standard audio and the second cepstrum corresponding to the audio to be detected includes: Performing a Fourier transform on the audio to be detected to obtain a frequency spectrum; Filtering the obtained frequency spectrum according to filtering information, and performing a logarithmic operation on the filtered frequency spectrum to obtain a log spectrum; wherein, the filtering information includes multiple filtering parameters, and the filtering parameters corresponding to different frequencies are different; Performing an inverse Fourier transform on the log spectrum to obtain the second cepstrum.
4. The method according to claim 2, wherein The determining the rhythm similarity between the audio to be detected and the standard audio according to the similarity between the first cepstrum and the second cepstrum includes: Obtaining multiple mapping relationships between a first audio frame included in the first cepstrum and a second audio frame included in the second cepstrum, and calculating the differences between the first cepstrum and the second cepstrum under different mapping relationships; wherein, the differences include the differences between the first audio frame and the corresponding second audio frame; Determining the minimum difference from the calculated differences, and selecting the target mapping relationship corresponding to the minimum difference from the multiple mapping relationships; Determining the similarity between the first cepstrum and the second cepstrum according to the target mapping relationship, and using the determined similarity as the rhythm similarity between the audio to be detected and the standard audio.
5. The method according to claim 4, wherein The number of the first audio frames and the number of the second audio frames are respectively multiple; determining the similarity between the first cepstrum and the second cepstrum according to the target mapping relationship includes: Filtering out target audio frames from multiple second audio frames according to the target mapping relationship, where the time difference between the target audio frames and the corresponding first audio frames is less than a second threshold; Determining the similarity between the first cepstrum and the second cepstrum according to the number of the target audio frames and the number of the first audio frames.
6. The method according to claim 4, wherein The number of the first audio frames and the number of the second audio frames are respectively multiple; determining the similarity between the first cepstrum and the second cepstrum according to the target mapping relationship includes: Obtaining the time differences between multiple first audio frames and the corresponding second audio frames respectively according to the target mapping relationship; Summing up the obtained time differences to obtain a total time difference; Determining the similarity between the first cepstrum and the second cepstrum according to the total time difference.
7. The method according to claim 2, wherein After determining the rhythm matching parameter of the audio to be detected according to the rhythm similarity, the method further includes: Performing weighted summation on the rhythm matching parameter and the pitch matching parameter to obtain a comprehensive matching parameter of the audio to be detected.
8. The method according to claim 1, characterized in that, Determining the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note, the duration of the first note, the duration of the second target note, and the duration of the third note includes: Obtaining a first ratio of the duration of the first target note to the duration of the first note, and a second ratio of the duration of the second target note to the duration of the third note; Performing weighted summation on the first ratio and the second ratio to obtain the pitch similarity between the audio to be detected and the standard audio.
9. The method according to any one of claims 1 to 8, characterized in that Obtaining the first note sequence corresponding to the standard audio and the second note sequence corresponding to the audio to be detected includes: If the type of the audio to be detected is a mixed audio, performing sound source separation on the audio to be detected to obtain a dry audio; Extracting the fundamental frequency of the dry audio and quantifying the extracted fundamental frequency to obtain the second note sequence.
10. The method according to any one of claims 1 to 8, characterized in that, Obtaining the first note sequence corresponding to the standard audio and the second note sequence corresponding to the audio to be detected includes: Obtaining the first note sequence corresponding to each of the multiple first sub-audios included in the standard audio; where the multiple first sub-audios include multiple sub-audios obtained by segmenting the standard audio according to a preset segmentation method; Segmenting the audio to be detected according to the preset segmentation method to obtain multiple second sub-audios; Processing each of the multiple second sub-audios respectively to obtain the second note sequence corresponding to each of the multiple second sub-audios.
11. The method according to claim 10, wherein The pitch similarity includes the audio similarity corresponding to multiple audio combinations respectively, and each audio combination includes a first sub-audio included in the audio to be detected and a second sub-audio included in the standard audio, and the start time of the first sub-audio matches the start time of the second sub-audio; Determining the pitch matching parameter of the audio to be detected according to the pitch similarity includes: Determining the pitch matching parameter of the corresponding second sub-audio according to the pitch similarity of each audio combination; Perform weighted summation on the determined pitch matching parameters to obtain the pitch matching parameter of the audio to be detected.
12. An audio matching detection device, characterized in that, The device includes: An acquisition module configured to acquire a first note sequence corresponding to a standard audio and a second note sequence corresponding to the audio to be detected; A search module configured to screen out first notes with a duration greater than or equal to a first threshold from the first note sequence, and search for a plurality of second notes in the second note sequence whose start and end times match the start and end times of the first notes; A similarity determination module configured to screen out first target notes whose pitches match the pitch of the first note from the plurality of second notes, and determine the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the duration of the first note; A matching detection module configured to determine the pitch matching parameter of the audio to be detected according to the pitch similarity; Wherein, the determining the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note and the duration of the first note includes: Determining a third note other than the first note from the first note sequence; Determining a fourth note in the second note sequence whose start time matches the start time of the third note, and if the pitch of the fourth note matches the pitch of the third note, using the fourth note as the second target note corresponding to the third note; Determining the pitch similarity between the audio to be detected and the standard audio according to the duration of the first target note, the duration of the first note, the duration of the second target note, and the duration of the third note.
13. An electronic device, characterized in that, Includes: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors cause the electronic device to implement the method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, Having computer-readable instructions stored thereon, which when executed by a processor of the computer cause the computer to execute the method according to any one of claims 1-11.
15. A computer program product, characterized in that, Including a computer program, which when executed by a processor implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Instrumental partner-training system
CN107424476A
Audio recognition method and device, and storage medium
CN110880329A
Multi-dimensional evaluation method for music
CN112382256A