Determination Method, Device, Equipment and Storage Medium for Audition Music
By calculating and analyzing the average loudness difference of multiple vocal clips in music, the starting point of the audition music is automatically determined, which solves the problems of poor listening effect and low manual setting efficiency in the prior art, and realizes a more efficient and attractive audition music determination method.
Patent Information
- Application Number
- CN202211435476.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-11-16
AI Technical Summary
The starting part of the audition music in the prior art is not necessarily the most attractive part, which leads to poor listening effects. At the same time, manual setting of audition music requires a lot of human resources and is inefficient.
By obtaining multiple vocal clips of the target music, the average loudness difference value of each vocal clip is calculated, and whether there is a chorus part is determined based on the first preset condition, and the starting point of the listening music is determined.
It improves the attractiveness and user experience of listening music, reduces human resources investment, and improves the determination efficiency of listening music.
Smart Images

Figure CN115757858B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and in particular, to a method, apparatus, device, and storage medium for determining audition music. Background Art
[0002] Currently, music audition functions are open to non-member users in various music applications (APPs), and some segments of paid music can be played as audition music for non-member users. Specifically, the above audition music can be determined by the following two methods: First, the starting part of the paid music is used as the audition music; Second, the chorus part of the paid music is manually set as the audition music by a human.
[0003] However, since the starting part of the music is not necessarily the most attractive part, the audition effect is poor. At the same time, manually setting the audition music requires a large amount of human resources, resulting in low efficiency. Therefore, how to determine the audition music is a technical problem that urgently needs to be solved. Summary of the Invention
[0004] The present application provides a method, apparatus, device, and storage medium for determining audition music to at least solve the problems of poor audition effect and inability to efficiently set audition music in the prior art. The technical solution of the present application is as follows:
[0005] According to a first aspect of the present application, there is provided a method for determining audition music, the method including: obtaining multiple human voice segments of a target music; determining an average loudness difference of each human voice segment; the average loudness difference is used to indicate the difference between the average loudness of the human voice segment and the average loudness of the previous adjacent human voice segment of the human voice segment in the target music; in the case where there is at least one first human voice segment among the multiple human voice segments, determining a target human voice segment from the at least one first human voice segment and the first human voice segment among the multiple human voice segments; the first human voice segment satisfies a first preset condition, and the first preset condition includes: the average loudness difference of the first human voice segment is greater than or equal to a first threshold, and the average loudness differences of the first preset number of human voice segments before the first human voice segment and the second preset number of human voice segments after the first human voice segment in the target music are both less than a second threshold, and the second threshold is less than the first threshold; determining the audition music of the target music according to the target human voice segment; the starting point of the audition music is the starting point of the target human voice segment.
[0006] In a possible implementation manner, determining the target voice segment from at least one first voice segment and the first voice segment among multiple voice segments includes: for any one of the first voice segments, determining whether there is an ending voice segment of the first voice segment in the voice segments after the first voice segment; the time interval between the ending voice segment of the first voice segment and the next adjacent voice segment after the ending voice segment of the first voice segment is greater than or equal to a third threshold; in the case where there is an ending voice segment of the first voice segment, determining the target voice segment from at least one first voice segment and the first voice segment among multiple voice segments; in the case where there is no ending voice segment of the first voice segment, determining the first voice segment among multiple voice segments as the target voice segment.
[0007] In a possible implementation manner, in the case where there is an ending voice segment of the first voice segment, determining the target voice segment from at least one first voice segment and the first voice segment among multiple voice segments includes: in response to the first voice segment satisfying a second preset condition, determining the first voice segment as a second voice segment, obtaining at least one second voice segment, and determining the target voice segment from at least one second voice segment; the second preset condition includes: the number of third voice segments in the voice segment sequence of the first voice segment is greater than or equal to a fourth threshold, and the repetition times of the third voice segment among multiple voice segments is greater than or equal to a fifth threshold; the voice segment sequence of the first voice segment includes the first voice segment, the ending voice segment of the first voice segment, and the voice segments between the first voice segment and the ending voice segment of the first voice segment; if none of the at least one first voice segment satisfies the second preset condition, determining the first voice segment among multiple voice segments as the target voice segment.
[0008] In a possible implementation manner, determining the target voice segment from at least one first voice segment and the first voice segment among multiple voice segments includes: for any one of the first voice segments, determining whether there is an ending voice segment of the first voice segment in the voice segments after the first voice segment; the time interval between the ending voice segment of the first voice segment and the next adjacent voice segment after the ending voice segment of the first voice segment is greater than or equal to a third threshold; in the case where there is an ending voice segment of the first voice segment, determining the first voice segment as the target voice segment.
[0009] In a possible implementation manner, the above method further includes: in the case where there is no first voice segment among multiple voice segments, determining the first voice segment among multiple voice segments as the target voice segment.
[0010] In a possible implementation manner, determining the trial listening music of the target music according to the target human voice segment includes: determining whether there is an ending human voice segment of the target human voice segment in the human voice segments after the target human voice segment; the time interval between the ending human voice segment of the target human voice segment and the next adjacent human voice segment after the ending human voice segment of the target human voice segment is greater than or equal to a third threshold; in the case where there is an ending human voice segment of the target human voice segment, determining the music starting from the target human voice segment and ending at the end of the ending human voice segment of the target human voice segment as the trial listening music; in the case where there is no ending human voice segment of the target human voice segment, determining the music with a preset duration after the start of the target human voice segment as the trial listening music.
[0011] According to a second aspect of the present application, there is provided a device for determining trial listening music, the device including a determining unit and an obtaining unit; the obtaining unit is used to obtain multiple human voice segments of the target music; the determining unit is used to determine the average loudness difference of each human voice segment; the average loudness difference is used to indicate the difference between the average loudness of the human voice segment and the average loudness of the previous adjacent human voice segment of the human voice segment in the target music; the determining unit is further used to, in the case where there is at least one first human voice segment among the multiple human voice segments, determine the target human voice segment from at least one first human voice segment and the first human voice segment among the multiple human voice segments; the first human voice segment satisfies a first preset condition, and the first preset condition includes: the average loudness difference of the first human voice segment is greater than or equal to a first threshold, and the average loudness differences of the first preset number of human voice segments before the first human voice segment and the second preset number of human voice segments after the first human voice segment in the target music are both less than a second threshold, and the second threshold is less than the first threshold; the determining unit is further used to determine the trial listening music of the target music according to the target human voice segment; the starting point of the trial listening music is the starting point of the target human voice segment.
[0012] In a possible implementation manner, the above determining unit is specifically used for: for any one of the first human voice segments, determining whether there is an ending human voice segment of the first human voice segment in the human voice segments after the first human voice segment; the time interval between the ending human voice segment of the first human voice segment and the next adjacent human voice segment after the ending human voice segment of the first human voice segment is greater than or equal to a third threshold; in the case where there is an ending human voice segment of the first human voice segment, determining the target human voice segment from at least one first human voice segment and the first human voice segment among the multiple human voice segments; in the case where there is no ending human voice segment of the first human voice segment, determining the first human voice segment among the multiple human voice segments as the target human voice segment.
[0013] In a possible implementation, the above-mentioned determining unit is specifically configured to: in response to the first voice segment satisfying a second preset condition, determine the first voice segment as a second voice segment, obtain at least one second voice segment, and determine a target voice segment from the at least one second voice segment; the second preset condition includes: the number of third voice segments in the voice segment sequence of the first voice segment is greater than or equal to a fourth threshold, and the number of repetitions of the third voice segment among multiple voice segments is greater than or equal to a fifth threshold; the voice segment sequence of the first voice segment includes the first voice segment, the ending voice segment of the first voice segment, and the voice segments between the first voice segment and the ending voice segment of the first voice segment; if none of the at least one first voice segment satisfies the second preset condition, determine the first voice segment among the multiple voice segments as the target voice segment.
[0014] In a possible implementation, the above-mentioned determining unit is specifically configured to: for any first voice segment, determine whether there is an ending voice segment of the first voice segment among the voice segments after the first voice segment; the time interval between the ending voice segment of the first voice segment and the next adjacent voice segment after the ending voice segment of the first voice segment is greater than or equal to a third threshold; in the case where there is an ending voice segment of the first voice segment, determine the first voice segment as the target voice segment.
[0015] In a possible implementation, the above-mentioned determining unit is specifically configured to: in the case where there is no first voice segment among the multiple voice segments, determine the first voice segment among the multiple voice segments as the target voice segment.
[0016] In a possible implementation, the above-mentioned determining unit is specifically configured to: determine whether there is an ending voice segment of the target voice segment among the voice segments after the target voice segment; the time interval between the ending voice segment of the target voice segment and the next adjacent voice segment after the ending voice segment of the target voice segment is greater than or equal to a third threshold; in the case where there is an ending voice segment of the target voice segment, determine the music starting from the target voice segment and ending at the ending voice segment of the target voice segment as the audition music; in the case where there is no ending voice segment of the target voice segment, determine the music with a preset duration after the target voice segment as the audition music.
[0017] According to the third aspect of the present application, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the method according to the first aspect and any of its possible implementations.
[0018] According to a fourth aspect of the present application, there is provided a computer-readable storage medium. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to the first aspect and any possible implementation manner thereof above.
[0019] According to a fifth aspect of the present application, there is provided a computer program product. The computer program product includes computer instructions. When the computer instructions run on an electronic device, the electronic device is enabled to execute the method according to the first aspect and any possible implementation manner thereof above.
[0020] The technical solution provided in the first aspect of the present application at least brings the following beneficial effects: In the prior art, the audition music of the target music is usually determined by using the starting part of the target music as the audition music or manually setting the chorus part as the audition music, which may lead to poor audition effect and low efficiency. The present application determines whether there is at least one first human voice segment among multiple human voice segments based on a first preset condition, and determines the starting point of the audition music based on the judgment result. The first preset condition includes that the average loudness difference of the first human voice segment is greater than or equal to a first threshold, and the average loudness differences of the first preset number of human voice segments before the first human voice segment and the second preset number of human voice segments after the first human voice segment in the target music are both less than a second threshold. Therefore, the first preset condition characterizes the characteristics that the loudness of the chorus part in the music becomes larger or the rhythm becomes faster. In this way, the determined target human voice segment is close to the starting part of the chorus part in the target music, and the audition music determined based on the target human voice segment is also closer to the chorus part of the target music, which can improve the user's audition experience. At the same time, by using the method of extracting the audition music from the target music by the electronic device, it is also possible to reduce human resources and improve the efficiency of determining the audition music.
[0021] At the same time, in the case where there is no first human voice segment that meets the first preset condition, the present application determines the first human voice segment in the target music as the target human voice segment. In this way, the audition music determined based on the first human voice segment is the main song part of the target music, which can improve the user's audition experience.
[0022] It should be noted that the technical effects brought by any implementation manner in the second aspect to the fifth aspect can refer to the technical effects brought by the corresponding implementation manner in the first aspect, which will not be elaborated here.
[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings
[0024] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application, and do not constitute an undue limitation on the present application.
[0025] Figure 1 is a schematic diagram of an implementation architecture shown according to an exemplary embodiment;
[0026] Figure 2 is a flowchart of a method for determining audition music shown according to an exemplary embodiment;
[0027] Figure 3 is a flowchart of another method for determining audition music shown according to an exemplary embodiment;
[0028] Figure 4 is a flowchart of another method for determining audition music shown according to an exemplary embodiment;
[0029] Figure 5 is a flowchart of another method for determining audition music shown according to an exemplary embodiment;
[0030] Figure 6 is a flowchart of another method for determining audition music shown according to an exemplary embodiment;
[0031] Figure 7 is a block diagram of a device for determining audition music shown according to an exemplary embodiment;
[0032] Figure 8 is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0033] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0035] Before introducing the method for determining the audition music provided by this application in detail, a brief introduction to the implementation environment (implementation architecture) involved in this application will be given first.
[0036] The method for determining the audition music provided by the embodiments of the present application can be applied to a system for determining audition music. Figure 1 A schematic structural diagram of the system for determining the audition music is shown. As Figure 1 shown, the system 10 for determining the audition music includes a device 11 for determining the audition music and an electronic device 12. The device 11 for determining the audition music is connected to the electronic device 12. The connection between the device 11 for determining the audition music and the electronic device 12 can be wired or wireless, and the embodiments of the present application do not limit this.
[0037] The device 11 for determining the audition music can be used to interact with the electronic device 12. For example, it can obtain the target music from the electronic device 12 and send the determined audition music of the target music to the electronic device 12.
[0038] The device 11 for determining the audition music can also be used to process the obtained target music. For example, it can obtain multiple human voice segments of the target music, determine the average loudness difference of each human voice segment. In the case where there is at least one first human voice segment among the multiple human voice segments, determine the target human voice segment from the at least one first human voice segment and the first human voice segment among the multiple human voice segments. In the case where there is no first human voice segment among the multiple human voice segments, determine the first human voice segment among the multiple human voice segments as the target human voice segment. Determine the audition music of the target music according to the target human voice segment.
[0039] The electronic device 12 has a lyrics prompter that stores the lyrics of the target song and highlights the lyrics in response to the singing operation of the singer singing the target song. The electronic device 12 also has a channel extractor that can extract the pure human voice part in the target song to obtain the human voice music.
[0040] The electronic device 12 can be used to interact with the device 11 for determining the audition music. For example, it can send the target music to the device 11 for determining the audition music and receive the audition music of the target music sent by the device 11 for determining the audition music.
[0041] Optionally, the electronic device may be a physical machine, such as: a desktop computer, also known as a desktop or desktop computer (desktop computer), a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The electronic device may also be a server, or a server cluster composed of multiple servers.
[0042] Optionally, the above-mentioned determining device 11 for audition music may also implement the functions to be implemented by the above-mentioned determining device 11 for audition music through a virtual machine (VM) deployed on a physical machine.
[0043] It should be noted that the determining device 11 for audition music and the electronic device 12 may be independent devices, or may be integrated into the same device. The present disclosure does not make specific limitations on this.
[0044] When the determining device 11 for audition music and the electronic device 12 are integrated into the same device, the communication method between the determining device 11 for audition music and the electronic device 12 is the communication between internal modules of the device. In this case, the communication process between the two is the same as the "communication process between the determining device 11 for audition music and the electronic device 12 when they are independent of each other".
[0045] In the following embodiments provided by the present disclosure, the present disclosure takes the determining device 11 for audition music and the electronic device 12 being independently arranged as an example for description.
[0046] For ease of understanding, the following specifically introduces the method for determining audition music provided by the present application with reference to the accompanying drawings.
[0047] Figure 2 It is a flowchart of a method for determining audition music shown according to an exemplary embodiment. This method can be applied to an electronic device, or can be applied to a determining device for audition music connected to the electronic device. At the same time, this method can also be applied to devices similar to the electronic device or the determining device for audition music. Hereinafter, taking this method being applied to an electronic device as an example, this method will be described, as Figure 2 shown, the method for determining audition music includes the following steps:
[0048] S201. Obtain multiple vocal segments of the target music.
[0049] As a possible implementation, the electronic device obtains the target music and extracts the vocal music from the target music based on audio processing technology. Further, the electronic device divides the vocal music into multiple vocal segments based on a preset lyric prompter.
[0050] It should be noted that during the process of extracting the vocal music, the electronic device can also remove the background music in the target music based on audio processing technology to obtain the vocal music.
[0051] The electronic device can determine the multiple vocal segments included in the vocal music according to the highlighting operation of each lyric by the lyric prompter.
[0052] Exemplarily, taking a certain target music including multiple lyric segments a1, a2... am as an example, during the process of the lyric prompter highlighting the lyric segment a1, the electronic device determines the vocal segment b1 corresponding to the lyric segment a1 from the vocal music based on the highlighting operation of the lyric segment a1 by the lyric prompter. As the lyric prompter sequentially highlights the subsequent lyric segments, the electronic device also determines the vocal segment corresponding to each lyric segment based on the lyric prompter, and thus obtains multiple vocal segments b1, b2... bm corresponding one by one to the multiple lyric segments a1, a2... am. At the same time, the order of the multiple vocal segments in the target music is the same as the order of the multiple lyric segments in the target music. Where m is the total number of vocal segments in the target music.
[0053] S202. Determine the average loudness difference of each vocal segment.
[0054] Among them, the average loudness difference is used to indicate the difference between the average loudness of the vocal segment and the average loudness of the previous adjacent vocal segment of the vocal segment in the target music.
[0055] As a possible implementation, the electronic device obtains the amplitude data of each vocal segment in the vocal music and obtains the average loudness of each vocal segment according to the amplitude data of each vocal segment and the conversion software.
[0056] Among them, the amplitude data of each vocal segment includes the amplitude corresponding to the vocal of each vocal segment at each moment.
[0057] Further, the electronic device determines the difference between the average loudness of each vocal segment and the average loudness of the previous adjacent vocal segment of the vocal segment as the average loudness difference of each vocal segment.
[0058] It should be noted that when the vocal segment is the first vocal segment in the target music, the electronic device determines that the average loudness difference of the vocal segment is a preset value. For example, the preset value can be 0.
[0059] When the human voice segment is not the first human voice segment, taking the nth human voice segment in the target music as an example, the electronic device determines that the difference between the average loudness Ln of the nth human voice segment and the average loudness Ln-1 of the (n-1)th human voice segment is the average loudness difference Vn of the nth human voice segment.
[0060] S203. Determine whether there is at least one first human voice segment among multiple human voice segments.
[0061] Among them, the first human voice segment satisfies the first preset condition, and the first preset condition includes: the average loudness difference of the first human voice segment is greater than or equal to the first threshold, and the average loudness differences of the first preset number of human voice segments before the first human voice segment and the average loudness differences of the second preset number of human voice segments after the first human voice segment in the target music are both less than the second threshold, and the second threshold is less than the first threshold.
[0062] As a possible implementation, the electronic device traverses each human voice segment and determines whether each human voice segment satisfies the above first preset condition. When there is any human voice segment that satisfies the above first preset condition, the electronic device determines that this human voice segment is the first human voice segment.
[0063] It should be noted that the first threshold, the second threshold, the first preset number, and the second preset number can all be set in the electronic device by the operation and maintenance personnel in advance.
[0064] Exemplarily, for the above nth human voice segment, when the electronic device determines that the average loudness difference of the nth human voice segment is greater than or equal to the first threshold, it determines whether the average loudness differences of the (n-1)th human voice segment and the (n-2)th human voice segment are less than the second threshold (taking the first preset number equal to 2 as an example), and determines whether the average loudness differences of the (n+1)th human voice segment and the (n+2)th human voice segment are less than the second threshold (taking the second preset number equal to 2 as an example).
[0065] If the average loudness differences of the (n-1)th human voice segment, the (n-2)th human voice segment, the (n+1)th human voice segment, and the (n+2)th human voice segment are all less than the second threshold, the electronic device determines that the nth human voice segment is the first human voice segment. If the average loudness difference of any one of the (n-1)th human voice segment, the (n-2)th human voice segment, the (n+1)th human voice segment, and the (n+2)th human voice segment is greater than or equal to the second threshold, the electronic device determines that the nth human voice segment is not the first human voice segment.
[0066] For the above nth human voice segment, if the electronic device determines that the average loudness difference of the nth human voice segment is less than the first threshold, it determines that the nth human voice segment is not the first human voice segment.
[0067] It should be noted that the first preset quantity and the second preset quantity may be the same value or different values, and the present application does not limit this.
[0068] S204. When there is at least one first human voice segment among multiple human voice segments, determine a target human voice segment from the at least one first human voice segment and the first human voice segment among the multiple human voice segments.
[0069] As a possible implementation, when there is at least one first human voice segment among multiple human voice segments, the electronic device determines whether the at least one first human voice segment includes the target human voice segment, and when the at least one first human voice segment does not include the target human voice segment, determines that the first human voice segment among the multiple human voice segments is the target human voice segment.
[0070] For the specific implementation of this step, reference can be made to the subsequent description of the embodiments of the present application, and details will not be elaborated here.
[0071] S205. When there is no first human voice segment among multiple human voice segments, determine the first human voice segment among the multiple human voice segments as the target human voice segment.
[0072] S206. Determine a trial listening music of the target music according to the target human voice segment.
[0073] Among them, the starting point of the trial listening music is the starting point of the target human voice segment.
[0074] As a possible implementation, the electronic device determines the target music after a preset duration starting from the target human voice segment as the trial listening music.
[0075] It should be noted that the preset duration can be set by the operation and maintenance personnel in the electronic device in advance.
[0076] Exemplarily, when it is determined that the human voice segment b3 is the target human voice segment and the preset duration is 30 seconds, the electronic device uses the target music 30 seconds after the human voice segment b3 as the trial listening music of the target music.
[0077] It can be understood that in the prior art, the audition music of the target music is usually determined by using the starting part of the target music as the audition music or manually setting the chorus part as the audition music, which may lead to poor audition effect and low efficiency. In this application, it is determined whether there is at least one first vocal segment among multiple vocal segments based on a first preset condition, and the starting point of the audition music is determined based on the judgment result. The first preset condition includes that the average loudness difference of the first vocal segment is greater than or equal to a first threshold value, and the average loudness differences of a first preset number of vocal segments before the first vocal segment and a second preset number of vocal segments after the first vocal segment in the target music are both less than a second threshold value. Therefore, the first preset condition characterizes the characteristics of the chorus part in the music where the loudness increases or the rhythm becomes faster. In this way, the target vocal segment determined based on the first preset condition is close to the starting part of the chorus part in the target music, and the audition music determined based on the target vocal segment is also closer to the chorus part of the target music, which can improve the user's audition experience. At the same time, by using the method of extracting the audition music from the target music by an electronic device, it is also possible to reduce human resources and improve the efficiency of determining the audition music.
[0078] Meanwhile, in the case where there is no first vocal segment that meets the first preset condition, this application determines the first vocal segment in the target music as the target vocal segment. In this way, the audition music determined based on the first vocal segment is the main song part of the target music, which can improve the user's audition experience.
[0079] In some embodiments, in order to be able to determine the target vocal segment from at least one first vocal segment and the first vocal segment among multiple vocal segments, as Figure 3 shown, S204 provided in the embodiments of this application may include the following steps:
[0080] S2041. For any first vocal segment, determine whether there is an ending vocal segment of the first vocal segment among the vocal segments after the first vocal segment.
[0081] Among them, the time interval between the ending vocal segment of the first vocal segment and the next adjacent vocal segment of the ending vocal segment of the first vocal segment is greater than or equal to a third threshold value.
[0082] As a possible implementation manner, for any first vocal segment, the electronic device obtains the time interval between each vocal segment after the first vocal segment and the next adjacent vocal segment of each vocal segment, and determines it as the time interval of each vocal segment after the first vocal segment.
[0083] Further, the electronic device traverses the time intervals of each vocal segment after the first vocal segment, and determines whether the time interval of each vocal segment after the first vocal segment is less than the third threshold value.
[0084] When the time interval of any human voice segment is less than a third threshold, the electronic device determines that the human voice segment after the first human voice segment is not the ending human voice segment of the first human voice segment. When the human voice segment is greater than or equal to the third threshold, the electronic device determines that the human voice segment after the first human voice segment is the ending human voice segment of the first human voice segment.
[0085] It should be noted that the third threshold can be set in the electronic device by the operation and maintenance personnel in advance.
[0086] Exemplarily, for the i-th human voice segment after the first human voice segment, the electronic device determines whether the time interval Ti of the i-th human voice segment after the first human voice segment is less than the third threshold. When Ti is less than the third threshold, the electronic device determines that the i-th human voice segment after the first human voice segment is not the ending human voice segment of the first human voice segment.
[0087] When Ti is greater than or equal to the third threshold, the electronic device determines that the i-th human voice segment after the first human voice segment is the ending human voice segment of the first human voice segment.
[0088] S2042: When there is an ending human voice segment of the first human voice segment, determine a target human voice segment from the first human voice segment among at least one first human voice segment and multiple human voice segments.
[0089] As a possible implementation manner, when there is an ending human voice segment of the first human voice segment, the electronic device determines whether the at least one first human voice segment includes the target human voice segment, and when the at least one first human voice segment does not include the target human voice segment, determines that the first human voice segment among the multiple human voice segments is the target human voice segment.
[0090] For the specific implementation manner of this step, reference can be made to the subsequent description of the embodiments of the present application, and details will not be elaborated here.
[0091] S2043: When there is no ending human voice segment of the first human voice segment, determine the first human voice segment among the multiple human voice segments as the target human voice segment.
[0092] It can be understood that in the present application, the ending human voice segment of the first human voice segment is determined based on the time interval of each human voice segment after the first human voice segment. Since the time interval characterizes the characteristic that the time interval between the ending human voice segment of the chorus part in the music and the next adjacent human voice segment is large, in this way, the ending human voice segment determined based on the time interval is close to the ending segment of the chorus part.
[0093] In some embodiments, in the presence of an ending voice segment of the first voice segment, in order to determine a target voice segment from at least one first voice segment and the first voice segment among multiple voice segments, as Figure 4 shown, S2042 provided by the embodiments of the present application may include the following steps:
[0094] S301. Determine whether the first voice segment meets a second preset condition.
[0095] Among them, the second preset condition includes: the number of third voice segments in the voice segment sequence of the first voice segment is greater than or equal to a fourth threshold, and the number of repetitions of the third voice segment in multiple voice segments is greater than or equal to a fifth threshold. The voice segment sequence of the first voice segment includes the first voice segment, the ending voice segment of the first voice segment, and the voice segments between the first voice segment and the ending voice segment of the first voice segment.
[0096] As a possible implementation, the electronic device determines the voice segment sequence of the first voice segment according to the determined first voice segment, the ending voice segment of the first voice segment, and the voice segments between the first voice segment and the ending voice segment of the first voice segment.
[0097] The electronic device obtains the number of repetitions of each voice segment in the voice segment sequence of the first voice segment in the target music, and determines the voice segments with the number of repetitions greater than or equal to the fifth threshold as the third voice segments.
[0098] Further, the electronic device determines whether the number of third voice segments is greater than or equal to the fourth threshold.
[0099] When the number of third voice segments is greater than or equal to the fourth threshold, the electronic device determines that the above-mentioned first voice segment meets the second preset condition.
[0100] When the number of third voice segments is less than the fourth threshold, the electronic device determines that the above-mentioned first voice segment does not meet the second preset condition.
[0101] Subsequently, the electronic device can traverse each first voice segment to determine whether each first voice segment meets the second preset condition.
[0102] It should be noted that the fourth threshold and the fifth threshold can be set in the electronic device by the operation and maintenance personnel in advance. The number of repetitions of the above-mentioned voice segment in the target music is the same as the number of repetitions of the text content of the voice segment in the target music.
[0103] Exemplarily, the electronic device determines that the sequence of voice segments b7, b8, b9, b10, b11 is the sequence of voice segments of the first voice segment b7, and obtains that the repetition times of the voice segments b7, b8, b9, b10, b11 in the target music are p7, p8, p9, p10, p11 respectively. Further, the electronic device determines that the voice segments b7, b8, b11 with repetition times greater than the fifth threshold are the third voice segments. When the number 3 (b7, b8, b11) of the third voice segments is greater than or equal to the fourth threshold, the electronic device determines that the first voice segment b7 meets the second preset condition. When the number 3 (b7, b8, b11) of the third voice segments is less than the fourth threshold, the electronic device determines that the first voice segment does not meet the second preset condition.
[0104] Similarly, the electronic device also makes the same judgment on the first voice segments b21, b35, and determines whether the first voice segments b21, b35 meet the second preset condition.
[0105] S302. In response to the first voice segment meeting the second preset condition, determine the first voice segment as the second voice segment to obtain at least one second voice segment.
[0106] S303. Determine the target voice segment from at least one second voice segment.
[0107] As a possible implementation manner, in the case of the ending voice segment of the first voice segment, the electronic device randomly determines a second voice segment from at least one second voice segment as the target voice segment according to the determined at least one second voice segment.
[0108] Exemplarily, in the case of determining that the voice segments b8, b21 are the second voice segments, the electronic device randomly determines the voice segment b8 as the target voice segment.
[0109] S304. If none of the at least one first voice segment meets the second preset condition, determine the first voice segment among the multiple voice segments as the target voice segment.
[0110] It can be understood that in the present application, the second human voice segment is determined based on a second preset condition, and the second preset condition includes: the number of third human voice segments in the human voice segment sequence of the first human voice segment is greater than or equal to a fourth threshold, and the number of repetitions of the third human voice segment among multiple human voice segments is greater than or equal to a fifth threshold. The human voice segment sequence of the first human voice segment includes the first human voice segment, the ending human voice segment of the first human voice segment, and the human voice segments between the first human voice segment and the ending human voice segment of the first human voice segment. Since the number of repetitions characterizes the characteristic that the lyrics in the chorus part repeat in the target music, in this way, the human voice segment sequence of the target human voice segment determined based on the number of repetitions is closer to the chorus part of the target music, thereby enabling the subsequent determined audition music to improve the user experience.
[0111] In some embodiments, in order to be able to determine the target human voice segment from at least one first human voice segment and the first human voice segment among multiple human voice segments, as Figure 5 shown, S204 provided by the embodiment of the present application may further include the following steps:
[0112] S2044. For any one first human voice segment, determine whether there is an ending human voice segment of the first human voice segment among the human voice segments after the first human voice segment.
[0113] Among them, the time interval between the ending human voice segment of the first human voice segment and the next adjacent human voice segment of the ending human voice segment of the first human voice segment is greater than or equal to a third threshold.
[0114] For the specific implementation manner of this step, reference may be made to the description in S2041 of the above embodiment of the present application, and details will not be elaborated here.
[0115] S2045. In the case where there is an ending human voice segment of the first human voice segment, determine the first human voice segment as the target human voice segment.
[0116] As a possible implementation manner, in the case where there is an ending human voice segment of the first human voice segment, the electronic device determines the first human voice segment as the target human voice segment.
[0117] Exemplarily, when the electronic device determines that the human voice segment b9 is the first human voice segment and the human voice segment b14 is the ending human voice segment of the first human voice segment, the human voice segment b9 is used as the target human voice segment.
[0118] S2046. In the case where there is no ending human voice segment of the first human voice segment, determine the first human voice segment among multiple human voice segments as the target human voice segment.
[0119] It can be understood that the present application determines whether there is an ending vocal segment of the first vocal segment. In the case where there is an ending vocal segment of the first vocal segment, the determined first vocal segment is used as the target vocal segment, and the audition music determined based on the target vocal segment is also closer to the chorus part of the target music, which can improve the user's audition experience. In the case where there is no ending vocal segment of the first vocal segment, the first vocal segment among multiple vocal segments is directly determined as the target vocal segment, which improves the efficiency of determining the target vocal segment. At the same time, the main verse part of the target music can be directly used as the audition music, which can improve the user experience.
[0120] In some embodiments, in order to determine the audition music of the target music according to the target vocal segment, as Figure 6 shown, S206 provided by the embodiment of the present application may further include the following steps:
[0121] S2061. Determine whether there is an ending vocal segment of the target vocal segment in the vocal segments after the target vocal segment.
[0122] Wherein, the time interval between the ending vocal segment of the target vocal segment and the next adjacent vocal segment after the ending vocal segment of the target vocal segment is greater than or equal to the third threshold.
[0123] For the specific implementation manner of this step, reference may be made to the description in S2041 of the above embodiment of the present application. The difference is that the determined ending vocal segment in this step is the target vocal segment, which can be replaced with reference in actual application, and will not be elaborated here.
[0124] S2062. In the case where there is an ending vocal segment of the target vocal segment, determine the music from the start of the target vocal segment to the end of the ending vocal segment of the target vocal segment as the audition music.
[0125] As a possible implementation manner, in the case where there is an ending vocal segment of the target vocal segment, the electronic device determines the music from the start of the target vocal segment to the end of the ending vocal segment of the target vocal segment as the audition music.
[0126] Exemplarily, if the electronic device determines that the vocal segment b3 is the target vocal segment and determines that the vocal segment b8 is the ending vocal segment of the target vocal segment b3, then the music from the start of the target vocal segment b3 to the end of the ending vocal segment b8 is used as the audition music.
[0127] S2063. In the case where there is no ending vocal segment of the target vocal segment, determine the music for a preset duration after the start of the target vocal segment as the audition music.
[0128] As a possible implementation, the electronic device determines the music within a preset duration starting from the target human voice segment as the trial listening music.
[0129] It should be noted that the preset duration can be pre-set in the electronic device by the operation and maintenance personnel.
[0130] It can be understood that in this application, it is determined whether there is an ending human voice segment of the target human voice segment. In the case where there is an ending human voice segment of the target human voice segment, the music starting from the target human voice segment and ending at the ending human voice segment of the target human voice segment is determined as the trial listening music, so that the user can trial listen to the complete or continuous chorus part, thereby improving the user experience. In the case where there is no ending human voice segment of the target human voice segment, the music within a preset duration starting from the target human voice segment is directly determined as the trial listening music, which can improve the efficiency of determining the trial listening music.
[0131] Figure 7 A determining device 500 for trial listening music shown according to an exemplary embodiment is as Figure 7 shown. The determining device 500 for trial listening music includes an obtaining unit 501 and a determining unit 502.
[0132] The obtaining unit is configured to obtain multiple human voice segments of the target music.
[0133] The determining unit is configured to determine the average loudness difference of each human voice segment. The average loudness difference is used to indicate the difference between the average loudness of the human voice segment and the average loudness of the previous adjacent human voice segment of the human voice segment in the target music.
[0134] The determining unit is further configured to, in the case where there is at least one first human voice segment among the multiple human voice segments, determine the target human voice segment from the at least one first human voice segment and the first human voice segment among the multiple human voice segments. The first human voice segment satisfies a first preset condition, and the first preset condition includes: the average loudness difference of the first human voice segment is greater than or equal to a first threshold, and the average loudness differences of the first preset number of human voice segments before the first human voice segment and the second preset number of human voice segments after the first human voice segment in the target music are both less than a second threshold, and the second threshold is less than the first threshold.
[0135] The determining unit is further configured to determine the trial listening music of the target music according to the target human voice segment. The starting point of the trial listening music is the starting point of the target human voice segment.
[0136] Optionally, as Figure 7 shown, the determining unit 502 provided in the embodiment of the present application is specifically configured to:
[0137] For any first vocal segment, determine whether there is an ending vocal segment of the first vocal segment among the vocal segments after the first vocal segment. The time interval between the ending vocal segment of the first vocal segment and the next adjacent vocal segment after the ending vocal segment of the first vocal segment is greater than or equal to a third threshold value.
[0138] In the case where there is an ending vocal segment of the first vocal segment, determine a target vocal segment from at least one first vocal segment and the first vocal segment among multiple vocal segments.
[0139] In the case where there is no ending vocal segment of the first vocal segment, determine the first vocal segment among multiple vocal segments as the target vocal segment.
[0140] Optionally, as Figure 7 shown, the determining unit 502 provided in the embodiment of the present application is specifically configured to:
[0141] In response to the first vocal segment satisfying a second preset condition, determine the first vocal segment as a second vocal segment, obtain at least one second vocal segment, and determine a target vocal segment from at least one second vocal segment. The second preset condition includes: the number of third vocal segments in the vocal segment sequence of the first vocal segment is greater than or equal to a fourth threshold value, and the number of repetitions of the third vocal segment among multiple vocal segments is greater than or equal to a fifth threshold value. The vocal segment sequence of the first vocal segment includes the first vocal segment, the ending vocal segment of the first vocal segment, and the vocal segments between the first vocal segment and the ending vocal segment of the first vocal segment.
[0142] If none of the at least one first vocal segment satisfies the second preset condition, determine the first vocal segment among multiple vocal segments as the target vocal segment.
[0143] Optionally, as Figure 7 shown, the determining unit 502 provided in the embodiment of the present application is specifically configured to:
[0144] For any first vocal segment, determine whether there is an ending vocal segment of the first vocal segment among the vocal segments after the first vocal segment. The time interval between the ending vocal segment of the first vocal segment and the next adjacent vocal segment after the ending vocal segment of the first vocal segment is greater than or equal to a third threshold value.
[0145] In the case where there is an ending vocal segment of the first vocal segment, determine the first vocal segment as the target vocal segment.
[0146] Optionally, as Figure 7 shown, the determining unit 502 provided in the embodiment of the present application is specifically configured to:
[0147] In the case where the first voice segment does not exist in multiple voice segments, the first voice segment among the multiple voice segments is determined as the target voice segment.
[0148] Optionally, as Figure 7 shown, the determination unit 502 provided in the embodiment of the present application is specifically configured to:
[0149] Determine whether there is an ending voice segment of the target voice segment in the voice segments after the target voice segment. The time interval between the ending voice segment of the target voice segment and the next adjacent voice segment after the ending voice segment of the target voice segment is greater than or equal to the third threshold.
[0150] In the case where there is an ending voice segment of the target voice segment, determine the music starting from the target voice segment and ending at the end of the ending voice segment of the target voice segment as the audition music.
[0151] In the case where there is no ending voice segment of the target voice segment, determine the music with a preset duration after the start of the target voice segment as the audition music.
[0152] Figure 8 It is a block diagram of an electronic device shown according to an exemplary embodiment. As Figure 8 shown, the electronic device 600 includes but is not limited to: a processor 601 and a memory 602.
[0153] Among them, the above-mentioned memory 602 is used to store the executable instructions of the above-mentioned processor 601. It can be understood that the above-mentioned processor 601 is configured to execute instructions to implement the method for determining audition music in the above-mentioned embodiment.
[0154] It should be noted that those skilled in the art can understand that Figure 8 the structure of the electronic device shown in Figure 8 does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than
[0155] shown, or combine certain components, or have different component arrangements. The processor 601 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 602, and calling the data stored in the memory 602, it executes various functions of the electronic device and processes data, thereby monitoring the entire electronic device. The processor 601 may include one or more processing units. Optionally, the processor 601 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 601 either.
[0156] The memory 602 can be used to store software programs and various data. The memory 602 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required by at least one functional module (such as an acquisition unit, a determination unit), etc. In addition, the memory 602 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0157] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions. The above instructions can be executed by a processor of an electronic device to implement the method for determining audition music in the above embodiment.
[0158] In actual implementation, the functions of the acquisition unit 501 and the determination unit 502 can both be Figure 8 implemented by the processor 601 in the above calling a computer program stored in the memory 602. The specific execution process can refer to the description of the method for determining audition music in the above embodiment, and will not be elaborated here.
[0159] Optionally, the computer-readable storage medium can be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0160] In an exemplary embodiment, an embodiment of the present application also provides a computer program product including one or more instructions. The one or more instructions can be executed by a processor of an electronic device to complete the method in the above embodiment.
[0161] It should be noted that when the instructions in the above computer-readable storage medium or the one or more instructions in the computer program product are executed by a processor of an electronic device, each process of the above method embodiment is implemented, and the same technical effects as the above method can be achieved. To avoid repetition, it will not be elaborated here.
[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0163] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0164] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0166] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks or optical discs that can store program codes.
[0167] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for determining audition music, characterized in that Including: Obtaining multiple vocal segments of the target music; Determining the average loudness difference of each vocal segment; The average loudness difference is used to indicate the difference between the average loudness of the vocal segment and the average loudness of the previous adjacent vocal segment of the vocal segment in the target music; When there is at least one first vocal segment among the multiple vocal segments, determining a target vocal segment from the at least one first vocal segment and the first vocal segment among the multiple vocal segments; The first vocal segment satisfies a first preset condition, and the first preset condition includes: the average loudness difference of the first vocal segment is greater than or equal to a first threshold, and the average loudness differences of the first preset number of vocal segments before the first vocal segment and the second preset number of vocal segments after the first vocal segment in the target music are both less than a second threshold, and the second threshold is less than the first threshold; Determining a trial listening music of the target music according to the target vocal segment; the starting point of the trial listening music is the starting point of the target vocal segment; The determining the target vocal segment from the at least one first vocal segment and the first vocal segment among the multiple vocal segments includes: For any first vocal segment, determining whether there is an ending vocal segment of the first vocal segment among the vocal segments after the first vocal segment; the time interval between the ending vocal segment of the first vocal segment and the next adjacent vocal segment of the ending vocal segment of the first vocal segment is greater than or equal to a third threshold; When there is an ending vocal segment of the first vocal segment, determining the target vocal segment from the at least one first vocal segment and the first vocal segment among the multiple vocal segments; When there is no ending vocal segment of the first vocal segment, determining the first vocal segment among the multiple vocal segments as the target vocal segment.
2. The determination method according to claim 1, characterized in that, The determining the target vocal segment from the at least one first vocal segment and the first vocal segment among the multiple vocal segments when there is an ending vocal segment of the first vocal segment includes: In response to the first vocal segment satisfying a second preset condition, determining the first vocal segment as a second vocal segment, obtaining at least one second vocal segment, and determining the target vocal segment from the at least one second vocal segment; the second preset condition includes: the number of third vocal segments in the vocal segment sequence of the first vocal segment is greater than or equal to a fourth threshold, and the repetition times of the third vocal segment in the multiple vocal segments is greater than or equal to a fifth threshold; the vocal segment sequence of the first vocal segment includes the first vocal segment, the ending vocal segment of the first vocal segment, and the vocal segments between the first vocal segment and the ending vocal segment of the first vocal segment; If none of the at least one first vocal segment satisfies the second preset condition, determining the first vocal segment among the multiple vocal segments as the target vocal segment.
3. The determination method according to claim 1, wherein Determining a target voice segment from the at least one first voice segment and the first voice segment among the multiple voice segments includes: For any one of the first voice segments, determining whether there is an ending voice segment of the first voice segment among the voice segments after the first voice segment; the time interval between the ending voice segment of the first voice segment and the next adjacent voice segment after the ending voice segment of the first voice segment is greater than or equal to a third threshold; When there is an ending voice segment of the first voice segment, determining the first voice segment as the target voice segment.
4. The determination method according to claim 1, wherein The method further includes: When the first voice segment does not exist among the multiple voice segments, determining the first voice segment among the multiple voice segments as the target voice segment.
5. The determination method according to any one of claims 1-4, characterized in that, Determining a trial listening music of the target music according to the target voice segment includes: Determining whether there is an ending voice segment of the target voice segment among the voice segments after the target voice segment; the time interval between the ending voice segment of the target voice segment and the next adjacent voice segment after the ending voice segment of the target voice segment is greater than or equal to a third threshold; When there is an ending voice segment of the target voice segment, determining the music starting from the target voice segment and ending at the ending voice segment of the target voice segment as the trial listening music; When there is no ending voice segment of the target voice segment, determining the music with a preset duration after the start of the target voice segment as the trial listening music.
6. A determining device for audition music, characterized in that, Including an acquisition unit and a determination unit; The acquisition unit is configured to acquire multiple voice segments of a target music; The determination unit is configured to determine an average loudness difference of each voice segment; the average loudness difference is used to indicate the difference between the average loudness of the voice segment and the average loudness of the previous adjacent voice segment of the voice segment in the target music; The determination unit is further configured to, when there is at least one first voice segment among the multiple voice segments, determine a target voice segment from the at least one first voice segment and the first voice segment among the multiple voice segments; The first voice segment satisfies a first preset condition, and the first preset condition includes: the average loudness difference of the first voice segment is greater than or equal to a first threshold, and the average loudness differences of the first preset number of voice segments before the first voice segment and the average loudness differences of the second preset number of voice segments after the first voice segment in the target music are both less than a second threshold, and the second threshold is less than the first threshold; The determination unit is further configured to determine a trial listening music of the target music according to the target voice segment; the starting point of the trial listening music is the starting point of the target voice segment; The determination unit is specifically configured to: For any first human voice segment, determine whether there is an ending human voice segment of the first human voice segment in the human voice segments after the first human voice segment; the time interval between the ending human voice segment of the first human voice segment and the next adjacent human voice segment after the ending human voice segment of the first human voice segment is greater than or equal to a third threshold; In the case where there is an ending human voice segment of the first human voice segment, determine the target human voice segment from the first human voice segment among the at least one first human voice segment and the first human voice segment among the multiple human voice segments; In the case where there is no ending human voice segment of the first human voice segment, determine the first human voice segment among the multiple human voice segments as the target human voice segment.
7. The determination device according to claim 6, characterized in that, The determining unit is specifically configured to: In response to the first human voice segment satisfying a second preset condition, determine the first human voice segment as a second human voice segment, obtain at least one second human voice segment, and determine the target human voice segment from the at least one second human voice segment; The second preset condition includes: the number of third human voice segments in the human voice segment sequence of the first human voice segment is greater than or equal to a fourth threshold, and the number of repetitions of the third human voice segment in the multiple human voice segments is greater than or equal to a fifth threshold; the human voice segment sequence of the first human voice segment includes the first human voice segment, the ending human voice segment of the first human voice segment, and the human voice segments between the first human voice segment and the ending human voice segment of the first human voice segment; If none of the at least one first human voice segments satisfy the second preset condition, determine the first human voice segment among the multiple human voice segments as the target human voice segment.
8. The determination device according to claim 6, characterized in that The determining unit is specifically configured to: For any first human voice segment, determine whether there is an ending human voice segment of the first human voice segment in the human voice segments after the first human voice segment; the time interval between the ending human voice segment of the first human voice segment and the next adjacent human voice segment after the ending human voice segment of the first human voice segment is greater than or equal to a third threshold; In the case where there is an ending human voice segment of the first human voice segment, determine the first human voice segment as the target human voice segment.
9. The determination device according to claim 6, wherein The determining unit is specifically configured to: In the case where the first human voice segment does not exist in the multiple human voice segments, determine the first human voice segment among the multiple human voice segments as the target human voice segment.
10. The determination device according to any one of claims 6-9, characterized in that The determining unit is specifically configured to: Determine whether there is an ending human voice segment of the target human voice segment in the human voice segments after the target human voice segment; the time interval between the ending human voice segment of the target human voice segment and the next adjacent human voice segment after the ending human voice segment of the target human voice segment is greater than or equal to a third threshold; In the case where there is an ending human voice segment of the target human voice segment, determine the music starting from the target human voice segment and ending at the ending human voice segment of the target human voice segment as the audition music; In the case where there is no ending human voice segment of the target human voice segment, determine the music with a preset duration after the start of the target human voice segment as the audition music.
11. An electronic device, characterized in that, Includes: Processor; A memory for storing the processor-executable instructions; Wherein, the processor is configured to execute the instructions to implement the method according to any one of claims 1 to 5.
12. A computer-readable storage medium, characterized in that, When the computer-executable instructions stored in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is capable of executing the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Music playing method and mobile terminal
CN108090140A
Audio stem identification systems and methods
EP3796305A1