Data Processing Method, Apparatus, Device, Storage Medium and Computer Program Product
By analyzing the vocal probability sequence and the audio energy value sequence for the to be processed, the vocal start time is automatically positioned, which solves the problems of low efficiency and poor accuracy of the vocal start position position in the prior art, and realizes efficient and accurate determination of the vocal start time.
Patent Information
- Application Number
- CN202111022361.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-01
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-09-01
AI Technical Summary
In the prior art, the positioning of the starting position information of the vocals usually relies on manual annotation, which is inefficient and poorly accurate, resulting in poor application effect based on the starting point of the vocals.
By performing data processing on the audio to be processed, the target person's voice audio is extracted and the vocal probability sequence is calculated, and the vocal start time is initially located. Then, precise positioning is performed based on the audio energy value sequence to adjust the accuracy of the vocal start time.
Automatic positioning of the starting position of the vocals is achieved, improving the efficiency and accuracy of the starting time of the vocals is determined, and avoiding the limitations of manual labeling.
Smart Images

Figure CN114329042B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a data processing method, apparatus, device, storage medium, and computer program product. Background Art
[0002] With the development of modern digital media technologies, people's demands for various audio and video are becoming increasingly rich and diverse. In a music player, it not only has the function of simply playing music or other audio, but also integrates various functions to enhance the user experience. The information of the starting position of the human voice has always been a hot research topic, and based on the information of the starting position of the human voice, various audio can be automatically processed in modern media management, such as quickly locating song content, lyric alignment, lyric recognition, etc.
[0003] The existing positioning of the starting position information of the human voice is usually achieved by manual annotation. Objectively speaking, this method not only consumes human resources but also has low efficiency; subjectively speaking, due to different annotation standards for different people, there may be inconsistent starting points of the human voice, which may lead to poor application effects based on the starting point of the human voice. Based on this, it is necessary to design an efficient and accurate way to determine the starting time of the human voice. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, apparatus, device, storage medium, and computer program product, which can effectively improve the efficiency and accuracy of determining the starting time of the human voice in audio.
[0005] On the one hand, an embodiment of this application provides a data processing method, including:
[0006] Determine a target human voice audio according to the audio to be processed, and determine a human voice probability sequence of the target human voice audio. The human voice probability sequence includes the human voice probabilities of the human voice audio corresponding to each first unit time, and the human voice probabilities of the human voice audio corresponding to each first unit time are sorted in chronological order;
[0007] If a first human voice starting time is determined according to the human voice probability sequence, determine a reference human voice audio from the target human voice audio according to the first human voice starting time;
[0008] Determine an audio energy value sequence of the reference human voice audio. The audio energy value sequence includes the audio energy values of the human voice audio corresponding to each second unit time, and the audio energy values of the human voice audio corresponding to each second unit time are sorted in chronological order;
[0009] If a second human voice starting time is determined according to the audio energy value sequence, determine the second human voice starting time as the starting time of the human voice in the audio to be processed.
[0010] An embodiment of the present application provides a data processing device on the one hand, including:
[0011] A determination module, configured to determine a target human voice audio according to the audio to be processed, and determine a human voice probability sequence of the target human voice audio, where the human voice probability sequence includes the human voice probabilities of the human voice audios corresponding to each first unit time, and the human voice probabilities of the human voice audios corresponding to each first unit time are sorted in chronological order;
[0012] The determination module is further configured to, if a first human voice start time is determined according to the human voice probability sequence, determine a reference human voice audio from the target human voice audio according to the first human voice start time;
[0013] The determination module is further configured to determine an audio energy value sequence of the reference human voice audio, where the audio energy value sequence includes the audio energy values of the human voice audios corresponding to each second unit time, and the audio energy values of the human voice audios corresponding to each second unit time are sorted in chronological order;
[0014] The determination module is further configured to, if a second human voice start time is determined according to the audio energy value sequence, determine the second human voice start time as the human voice start time of the audio to be processed.
[0015] An embodiment of the present application provides a computer device on the one hand, including: a processor, a memory, and a network interface; the processor is connected to the memory and the network interface, where the network interface is used to provide network communication functions, the memory is used to store program codes, and the processor is used to call the program codes to execute the data processing method in the embodiment of the present application.
[0016] An embodiment of the present application provides a computer-readable storage medium on the one hand, where the computer-readable storage medium stores a computer program, and the computer program includes program instructions, and when the program instructions are executed by a processor, the data processing method in the embodiment of the present application is executed.
[0017] Correspondingly, an embodiment of the present application provides a computer program product or a computer program, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method provided on the one hand in the embodiment of the present application.
[0018] In the embodiment of the present application, by extracting the target human voice audio from the audio to be processed, estimating the human voice probability of the target human voice audio, obtaining a human voice probability sequence for indicating the human voice audio corresponding to each first unit time, then roughly locating the starting time of the human voice according to the human voice probability sequence to obtain the starting time of the human voice, determining the reference human voice audio according to the starting time of the human voice obtained by the rough positioning, narrowing the determination of the starting time of the human voice within a smaller time range, and then re-locating the starting time of the human voice based on the audio energy values in the audio energy value sequence of the reference human voice audio, the starting time of the first human voice can be adjusted to a more accurate position. The whole process can automatically locate the starting position of the human voice, improve the efficiency of determining the starting time of the human voice, and through the combination of rough positioning and accurate positioning, the accuracy of the starting time of the human voice can be higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 is an architecture diagram of a data processing system provided by an embodiment of the present application;
[0021] Figure 2 is a flowchart of a data processing method provided by an embodiment of the present application;
[0022] Figure 3 is a flowchart of a singing recognition algorithm provided by an embodiment of the present application;
[0023] Figure 4 is a flowchart of another data processing method provided by an embodiment of the present application;
[0024] Figure 5 is a schematic diagram of the effect of source separation of a music segment provided by an embodiment of the present application;
[0025] Figure 6 is a schematic diagram of a human voice probability sequence of a target human voice audio provided by an embodiment of the present application;
[0026] Figure 7 is a distribution schematic diagram of an audio energy value sequence provided by an embodiment of the present application;
[0027] Figure 8 is a flowchart of another data processing method provided by an embodiment of the present application;
[0028] Figure 9 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0029] Figure 10 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific embodiments
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the protection scope of the present application.
[0031] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. applied based on the cloud computing business model. It can form a resource pool, be used as needed, and is flexible and convenient. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system. This can only be achieved through cloud computing.
[0032] Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. Alternatively, the SaaS can be directly deployed on the IaaS. PaaS is the platform for software operation, such as databases, web containers, etc. SaaS is various business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS. The data processing solution provided by the present application can be the function provided by the PaaS service, which can support third-party software to call relevant interfaces to obtain the voice start time through executing the data processing solution, and apply the obtained voice start time to specific functions, such as skipping the song prelude, positioning the main song and chorus, etc.
[0033] Please refer to Figure 1 , Figure 1It is an architecture diagram of the data processing system provided by the embodiments of the present application. As Figure 1 shown, it includes a terminal device 101 and a server 100. The terminal device 101 and the server 100 can be communicatively connected by wired or wireless means.
[0034] The terminal device 101 can collect voice data through a sound pickup device to generate audio, or combine with a camera device to collect image data to generate video. It can also obtain audio or video through other means (such as downloading or copying). The terminal device 101 can upload these video or audio data to a third-party application or a client function platform (such as a website of a web client). The server 100 processes the video or audio uploaded by the terminal device 101, determines the starting time of the human voice in the audio, and uses this starting time of the human voice to implement corresponding functions in different application scenarios. Exemplarily, for detecting the starting time of the human voice in a song, the prelude of the song can be skipped according to the obtained first starting time of the human voice, and directly locate to the place where the human voice starts to appear, because the prelude of the song is usually various musical instruments or other accompaniment sounds. Skipping the prelude can help users quickly enter the singing part of the song. Optionally, the accompaniment connecting the verse and the chorus of the song can also be skipped according to other starting times of the human voice, improving the user experience. Of course, different starting times of the human voice can have different application scenarios. How the starting time of the human voice determined by the server 100 is specifically applied and to what kind of scenarios are not limited in the embodiments of the present application. It should be noted that the above terminal device 101 can be a smart phone, a tablet computer, a vehicle-mounted terminal, a smart voice interaction device, a smart home appliance, a smart wearable device, a personal computer and other devices.
[0035] Server 100 can process the video or audio data uploaded by the terminal device 101 or obtain audio or video data from other databases for processing. Since the main processing object of the data processing algorithm installed on Server 100 is audio data, when Server 100 receives video data, it is necessary to preprocess the video data to extract the included audio as the audio to be processed. The data processing algorithm includes different functional modules. The processing content of Server 100 for the audio to be processed may include performing sound source separation on the audio segments obtained by segmenting the audio to be processed or directly performing sound source separation processing on the entire audio to be processed, so as to separate the human voice and other sounds, and obtain the target human voice audio and other audio. For example, if the audio to be processed is a song, it is mainly to separate the accompaniment and the human voice, which can reduce the interference of other sounds and make the subsequent processing more efficient and accurate. According to the human voice probability sequence of the target human voice audio, it is possible to predict whether there is a human voice in each unit of time. If the first human voice start time can be determined using the human voice probability sequence, it is necessary to continue to use the first human voice start time to locate the final human voice start time, which can further ensure the accuracy of the human voice start time. Specifically, first, a section of audio is selected as the reference human voice audio from the target human voice probability using the first human voice start time. After determining the second human voice start time according to the audio energy value sequence of the reference human voice audio, the second human voice start time is the human voice start time of the final audio to be processed.
[0036] It can be found that after Server 100 processes the audio to be processed to obtain the target human voice audio, the determination of the human voice probability sequence and the determination of the audio energy value sequence are essentially based on this target human voice audio. Therefore, more precisely, the key data object of the data processing algorithm of Server 100 is the human voice audio. If the initial first human voice start time is determined through the human voice probability sequence of the human voice audio, then a more refined second human voice start time is determined according to the audio energy value sequence of the human voice audio. If it can be successfully determined, the second human voice start time can be used as the final human voice start time of the audio to be processed, which can make the accuracy of the human voice start time higher. The entire process can automatically determine the human voice start time, thereby effectively improving the efficiency of determining the human voice start time.
[0037] It can be understood that the method provided by the embodiments of the present application can be executed by a computer device (such as Server 100). Server 100 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0038] Further, for ease of understanding, the following embodiments mentioned in this application are all described by taking a server (such as the server 100 in the corresponding embodiment above) as an example. Please refer to Figure 1 for the server 100 in the corresponding embodiment above) as an example for illustration. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a data processing method provided by an embodiment of this application. The data processing method may at least include the following steps S101 - S104, where:
[0039] S101, determine a target human voice audio according to the audio to be processed, and determine a human voice probability sequence of the target human voice audio.
[0040] In one embodiment, the audio to be processed may be audio data obtained by the server from a terminal device or other databases, or audio data extracted from video data. In different application scenarios, the type of the audio to be processed may be different. For example, the audio to be processed may be a song or audio extracted from a video, or it may be a conversation, a dubbing, etc. The specific content of the audio to be processed is not limited herein. The target human voice audio is the human voice audio separated from the audio to be processed, which may be the entire human voice audio determined according to the audio to be processed, that is, the duration of the target human voice audio is the same as that of the audio to be processed, or it may be a part of the human voice audio, that is, the duration of the target human voice audio is less than that of the audio to be processed. For obtaining the target human voice audio, some existing sound source separation tools or sound source separation algorithms can be used to extract the human voice. The extracted audio can be considered as an audio including only the human voice, that is, without any other background sounds (such as accompaniment sounds, noise, etc.). Of course, if the audio to be processed is an audio without human voice, such as pure music, the target human voice audio determined according to the audio to be processed may be empty or the target human voice audio obtained by misidentifying some instrumental sounds similar to human voice as human voice. For such a situation, the result of whether to execute the subsequent steps can be given according to the determined human voice probability. Strictly speaking, for the entire solution, being able to separate the real human voice audio from the audio to be processed is the primary condition for determining the starting time of the human voice. However, for the situation where the separated human voice audio is meaningless, it can also be further determined according to the subsequent processing. Therefore, the specific content of the target human voice audio in this step is not limited.
[0041] In one embodiment, the human voice probability sequence includes the human voice probabilities of the human voice audio corresponding to each first unit time in the target human voice audio, and the human voice probabilities of the human voice audio corresponding to each first unit time are sorted in chronological order. Here, the first unit time may be 30 milliseconds (ms), 1 second (s), or it may also be 2 seconds or other unit times, which is not limited herein.
[0042] Since there is not always human voice in the target human voice audio. For example, in a 30-second target human voice audio, there is human voice only from the 16th second to the 30th second. In the embodiments of the present application, whether there is human voice at each moment in the target human voice audio is measured by the human voice probability. By estimating the possibility of the existence of human voice in the corresponding human voice audio within each first unit time, the human voice probabilities sorted in chronological order are obtained, and a human voice probability sequence is obtained. For example, taking the first unit time as 1 second and a 30-second target human voice audio as an example to illustrate the human voice probability sequence, the human voice probability sequence is denoted as P, and the human voice probabilities included therein are denoted as p i , then P = {p0, p1, p2, …, p 29}, which respectively represent the human voice probability of the human voice audio corresponding to the 1st second, the human voice probability of the human voice audio corresponding to the 2nd second, …, the human voice probability of the human voice audio corresponding to the 30th second. That is, the human voice probability sequence can be regarded as a discrete sequence with one probability value corresponding to 1 second. Correspondingly, the processing of the target human voice audio is performed sequentially according to the time order so as to obtain the human voice probabilities arranged in the time order. The determination method of the human voice probability sequence is not limited herein.
[0043] Optionally, the method for determining the human voice probability sequence of the target human voice audio may be: performing Fourier transform processing on the target human voice audio to obtain the spectrogram of the target human voice audio; using an audio processing network to process the spectrogram to obtain the human voice probability sequence of the target human voice audio. The Fourier transform processing of the target human voice audio refers to performing fast Fourier transform (FFT) or short-time Fourier transform processing on the target human voice audio to obtain the corresponding spectrogram, and then inputting the spectrogram into an audio processing network, such as a deep convolutional network, to obtain the final human voice probability. In this embodiment, MobileNetV2 (a depthwise separable convolutional network) is adopted in terms of the network (i.e., the audio processing network); in terms of accuracy, this solution performs processing every preset time period (for example, 1 second), and a rough localization at the second level can be obtained; in terms of effect, the accuracy rate of this solution on the test set reaches 83.8%. When the target human voice audio is the human voice in a song, the above steps correspond to a singing recognition algorithm, and the schematic diagram of the algorithm can be referred to Figure 3 , Figure 3 . The audio signal in
[0044] S102, if the first human voice start time is determined according to the human voice probability sequence, then the reference human voice audio is determined from the target human voice audio according to the first human voice start time.
[0045] In one embodiment, the voice probability sequence includes voice probabilities arranged in chronological order. Usually, the probability value ranges from 0 to 1, indicating the magnitude of the possibility of voice. The larger the probability value (i.e., the closer it is to 1), the greater the possibility of voice in that unit of time. However, a non-zero voice probability in the sequence does not necessarily mean that there is voice at that moment. Therefore, the starting time of voice in the target voice audio can be determined based on whether the probability value meets a pre-set condition, and this is taken as the starting time of the first voice. It should be noted that the starting time of voice can be a time point or a time period. In this embodiment, a time point is used as an example for illustration. If the starting (time) point of voice appears hereinafter, it refers to the starting time of voice. Only when the starting time of the first voice in the target voice audio is determined can the subsequent steps of determining the reference voice audio be executed. Because if the probability values in the voice probability sequence do not meet the pre-set conditions, it means that there is no voice in the target voice audio, and the starting time of the first voice cannot be determined, and the subsequent steps can stop being executed.
[0046] In one embodiment, determining the starting time of the first voice is the result of preliminary positioning. It is also necessary to determine the reference voice audio based on the starting time of the first voice. Since the reference voice audio is a part of the target voice audio, its time granularity is smaller than that of the target voice audio. Therefore, according to the corresponding rules, a more accurate starting time of voice can be further determined from the reference voice audio.
[0047] S103, determine the audio energy value sequence of the reference voice audio.
[0048] In one embodiment, the audio energy value sequence includes the audio energy values of the voice audio corresponding to each second unit of time of the reference voice audio, and the audio energy values of the voice audio corresponding to each second unit of time are sorted in chronological order. Similar to the voice probability sequence of the target voice audio, the audio energy value sequence of the reference voice audio is also sorted in chronological order. The difference is that the second unit of time here is a time with a finer granularity or a smaller magnitude than the first unit of time. For example, if the first unit of time can be accurate to seconds (s), then the second unit of time can be accurate to milliseconds (ms), or if the first unit of time is accurate to 1 second, then the second unit of time can be accurate to 0.1 second. Another example is that if the first unit of time is 30 ms, then the second unit of time can be 5 ms. Exemplarily, the starting time of the first voice is the 27th second, and the determined reference voice audio is the voice audio from the 26.5th second to the 27.5th second in the target voice audio. The corresponding audio energy value sequence is denoted as E, and the audio energy value is denoted as e i, the audio energy value sequence of the reference human voice audio can be expressed as the audio energy value corresponding to each 0.1 second, that is, E = {e0, e1, e2, …, e9}. Among them, the calculation of the audio energy value can be to first calculate the power spectrum of the reference human voice audio, and then map the power spectrum to the decibel value by taking the logarithm, and use this decibel value as the final audio energy value, which can reduce the calculation magnitude of the audio energy and improve the processing efficiency of subsequent steps.
[0049] S104. If the second human voice start time is determined according to the audio energy value sequence, then the second human voice start time is determined as the human voice start time of the audio to be processed.
[0050] In an embodiment, similar to the first human voice start time, by making a judgment on the audio energy values included in the audio energy value sequence under certain conditions, the second human voice start time can be determined. It should be noted that the second human voice start time is a more accurate description relative to the first human voice start time, that is, in terms of intuitive numbers, the decimal points to which the second human voice start time and the first human voice start time are accurate are different, and the measurement of precision is different. By further adjusting the first human voice start time through the audio energy value sequence, the expression of the human voice start time can be more accurate. Optionally, the prerequisite for determining the human voice start time of the audio to be processed is to determine the second human voice start time according to the audio energy value sequence of the reference human voice audio. If it can be determined, it can be determined as the human voice start time, which is a key factor for the human voice start time of the audio to be processed. Of course, if it cannot be determined, the first human voice start time located for the first time can also be used as the human voice start time of the audio to be processed or the first human voice start time can be determined as an invalid human voice start time, and the human voice start time is determined again.
[0051] It should be noted that the above steps are the processing flow for a target human voice audio. If multiple target human voice audios are determined according to the audio to be processed, the same processing steps can also be adopted for each target human voice audio. According to specific application scenarios, the number of multiple target human voice audios processed can be adjusted accordingly, and whether the human voice start time of the final audio to be processed is one or more can also be determined according to specific requirements. For example, when this solution is applied to an application scenario such as skipping the prelude of a song, only the first human voice start time needs to be determined, and in the case of multiple target human voice audios, the processing of other target human voice audios can be stopped after the first human voice start time is determined, saving computing resources, and dividing the processing into multiple target human voice audios can also efficiently determine the human voice start time.
[0052] The embodiments of this application can be applied to the solution of automatically and quickly calculating the starting point of the human voice in music in various forms. Taking the web interface as an example, the specific operation steps and product presentation forms can be as follows: First, the user uploads a video or audio URL (Uniform Resource Locator), and the algorithm in the background server calculates the starting time point of the human voice in the music. Then, the starting time point of the human voice in the music is returned through the web interface. If there is no human voice, -1 is returned. The function of determining the starting point of the human voice can be deployed on the PAAS (Platform as a Service) service platform. If a third-party application wants to apply the starting point of the human voice to implement certain functions, it can process by calling the relevant interfaces provided by the PAAS service and obtain the processing result. At this time, when the third-party application is specifically facing users, the user can directly upload a video or audio file on the web side, and the background automatically extracts the indicated address of the video or audio file (such as the above URL), so that the PAAS service can obtain the video or audio data according to the indicated address and process it, and return the processing result to the background server of the third-party application to implement the corresponding function.
[0053] In summary, the embodiments of this application have at least the following advantages:
[0054] Determining the target human voice audio according to the audio to be processed and sending the target human voice audio to the subsequent processing link can reduce the interference of non-human voice audio and improve the accuracy of locating the starting time of the human voice. By calculating the human voice probability of the target human voice audio, accurately obtaining the probability of having a human voice at each moment, locking the accuracy of the starting time of the first human voice according to the human voice probability sequence, and then determining the reference human voice audio and calculating its audio energy value in the case of obtaining the starting time of the first human voice to accurately locate the starting moment of the human voice. Among them, different rules are used to screen the starting position of the human voice for different data information, and the accuracy of the starting time of the human voice is ensured by both the human voice probability and the audio energy value. Compared with using only the moment of energy mutation as the starting time of the human voice, this method can further improve the reliability and accuracy of the starting time of the human voice, and the whole process is automatically completed by the computer device according to the corresponding algorithm instructions, which can effectively improve the efficiency of determining the starting time of the human voice.
[0055] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a data processing method provided by the embodiments of this application. The data processing method can at least include the following steps S201-S205, where:
[0056] S201, Obtain the audio to be processed and perform segmentation processing on the audio to be processed to obtain at least two audio segments.
[0057] In one embodiment, the server can obtain the audio to be processed from a terminal device or a database. For example, it extracts the audio track from a video or audio file input by the terminal device and uses it as the input in subsequent processing steps. The obtained audio to be processed can be a piece of music or other audio data, which is not limited here. The segmentation processing of the audio to be processed can be evenly divided into G equal parts according to time to obtain G audio segments. Here, G is generally an integer greater than or equal to 2, and the value range can be 3-5, that is, evenly divided into 3 to 5 parts. Of course, sometimes it may not be possible to ensure complete equal division. Usually, it is allowed that the duration of the last audio segment is different from that of the previous audio segments. The G equal parts of audio segments obtained by the final segmentation processing can be successively sent into the subsequent calculation process. When applied to the scenario of skipping the prelude of a song, the segmentation strategy adopted here divides the original music into G segments on average, which can greatly reduce the calculation time and save computing resources. This is because in most music files with vocals, the start time of the vocals will start immediately after the end of the music prelude. The different segments are sent into the subsequent calculation process in chronological order. When the start position of the vocals is calculated in a certain segment, it stops. That is, when the start time of the vocals is obtained for a certain equal part of the audio segment, the processing of the remaining equal parts of the audio segment can be stopped. Utilizing the prior condition that vocals often start at a relatively early position in the music, acceleration can be achieved in most cases. When the start time position of the vocals has been calculated in the previous segment of the music, it is not necessary to calculate the subsequent music segments, thus saving calculation time. When G = 1, it corresponds to the case of no segmentation. When G takes a value within the optimal range of segmentation, the subsequent calculation of the start time of the vocals can achieve the effect of obtaining the result in one calculation.
[0058] S202. Perform sound source separation processing on at least two audio segments to obtain the vocal audio of each audio segment. Sequentially determine the vocal audio of each audio segment as the target vocal audio according to time order, and determine the vocal probability sequence of the target vocal audio.
[0059] In one embodiment, the main function of performing sound source separation on the audio segments obtained by segmenting the audio to be processed is to strip out the vocal track and obtain the vocal audio of each audio segment for subsequent operations. Taking sound source separation as a preprocessing operation for the audio, the corresponding preprocessing module can perform sound source separation on the input audio. Here, the input audio is the input audio segment, and it is finally separated into two or more audio tracks. Among them, the vocal audio of the vocal track will be sent into the subsequent calculation process. When the audio segment is a music segment, please refer to Figure 5 , Figure 5 which is a schematic diagram of the effect of sound source separation of a music segment provided by an embodiment of the present application. Sound source separation can divide the input music segment into two tracks: vocals and accompaniment.
[0060] Optionally, the algorithm used for sound source separation can be implemented using the open-source Spleeter algorithm. This algorithm models the original audio based on the encoding-decoding structure of the U-net network and can achieve efficient and accurate sound source separation. The U-net network is obtained through supervised training using original audio with vocals and original audio without vocals. In the specific application process, the input audio is transformed to obtain a spectrogram, and then the image processing network U-net is used to process the spectrogram. Finally, the vocal spectrogram is obtained and inverse-transformed into vocal audio, and this vocal audio is output. This step strips the vocal track through the sound source separation algorithm, which can exclude the interference of the accompaniment track on the algorithm in the processing of songs and avoid the algorithm spending more resources to determine whether the current intense signal is vocals or accompaniment. In addition, in some cases with good separation effects, simply using the audio volume of the vocal track can also complete the positioning of the vocal starting point.
[0061] In one embodiment, the vocal audio obtained by performing sound source separation on the audio segment is used as the target vocal audio in chronological order, which can correspond to the vocal audio of all or part of the audio segment. Exemplarily, when performing sound source separation on 3 audio segments, 3 corresponding vocal audios can be obtained, namely V1 (time range: 0 - 20 seconds), V2 (time range: 21 - 40 seconds), and V3 (time range: 41 - 55 seconds). Then, these 3 vocal audios can all be used as the target vocal audio first, and the vocal starting time can be calculated in chronological order. Or, in chronological order, the first vocal audio V1 can be used as the target vocal audio first. When the vocal starting point cannot be determined in this target vocal audio, the second vocal audio V2 can be used as the target vocal audio for further processing. When the vocal starting point is obtained, the third vocal audio V3 will not be used as the target vocal audio. In this way, only part of the vocal audio is used as the target vocal audio. Conversely, if the vocal starting point cannot be obtained from the second vocal audio, the third vocal audio V3 needs to be used as the target vocal audio and input into the corresponding module for calculating the vocal starting time.
[0062] In addition, the method for determining the vocal probability sequence can refer to the content of the foregoing embodiment. It should be noted that since in the sound source separation step, after separating the vocal audio from other types of audio, a vocal spectrogram is first obtained, and then it is inverse-transformed through the inverse fast Fourier transform to obtain the target vocal audio. And in the vocal recognition module (calculating the vocal probability), the target vocal audio will be transformed back into a spectrogram. Therefore, if the modular design is not considered, the vocal spectrogram obtained in the sound source separation can be directly used as the input, such as Figure 3The spectrogram of the convolutional neural network shown can omit the inverse transformation process of the human voice audio and the transformation process of the target human voice audio signal, saving computing resources and having stronger coupling between steps. The modular design separates the two functions of sound source separation and human voice recognition, and also has excellent processing effects. In addition, the order of the audio segmentation process and the sound source separation process in the above steps can also be interchanged, that is, first perform the sound source separation operation on the audio to be processed to obtain the entire human voice audio, and then segment the entire human voice audio to obtain human voice audio segments. These human voice audio segments can be determined as the target human voice audio in chronological order, and the subsequent processing flow can also be executed. The order of these two steps is not limited in this application.
[0063] S203. If the first human voice start time is determined according to the human voice probability sequence, the reference human voice audio is determined from the target human voice audio according to the first human voice start time.
[0064] In one embodiment, the implementation of determining the starting time of the first human voice according to the human voice probability sequence may include: determining the first human voice probability that is greater than or equal to the probability threshold from the human voice probability sequence, and determining the first candidate time corresponding to the first human voice probability that is greater than or equal to the probability threshold; determining a reference time interval according to the first candidate time, and determining the average human voice probability according to the human voice probabilities corresponding to the times within the reference time interval; if the average human voice probability is greater than or equal to the probability threshold, then determining the first candidate time as the starting time of the first human voice. The human voice probability sequence determined according to the target human voice audio includes human voice probabilities arranged in chronological order. After obtaining the human voice probabilities corresponding to each first unit of time (such as per second), the rough positioning of the starting time of the human voice can be completed through a simple rule. This method corresponds to an optional implementation of a rough positioning rule. By traversing the human voice probability sequence, the first human voice probability that is greater than or equal to the probability threshold in the human voice probability sequence can be determined, which corresponds to the first possible starting time of the human voice. The time corresponding to this human voice probability is the first candidate time, which means that the first candidate time is the position where the human voice starts to appear. Among them, the probability threshold can be set according to manual experience or calculated based on the results of multiple tests. To ensure the accuracy of the rough positioning, further determination is required. Taking the first candidate time as the starting time, a period of time after the first candidate time is used as the reference time interval. For example, if the first candidate time is 26 seconds, the reference time interval can be taken as the time interval from 27 to 37 seconds according to the specified rule. Then, some or all of the human voice probabilities corresponding to the times within the reference time interval need to be selected for averaging to obtain the average human voice probability. The average human voice probability is compared with the probability threshold to determine the starting time of the first human voice. Optionally, the probability threshold for comparing the average human voice probability here is the same as the probability threshold for comparing the human voice probability when determining the first candidate time.
[0065] Exemplarily, taking the target human voice audio corresponding to the human voice audio in a song as an example for illustration, Figure 6 A schematic diagram of the human voice probability sequence of a target human voice audio is shown. Among them, the time range of the target human voice audio is 0 - 50 seconds. In the first 25 seconds, the singing probability (i.e., the human voice probability) is relatively low, which is the prelude area here. After 25 seconds, the singing probability is very high, and the human voice part starts. Specifically, a threshold (i.e., the probability threshold) can be set in advance. Figure 6 The shown probability threshold is approximately 0.96. When the probability at a certain moment first exceeds this threshold, and the average value of the singing probabilities of the first K (K can be equal to the following M) in the next 2K seconds is also greater than this threshold, this moment is regarded as the rough positioning moment of the starting time of the human voice. As Figure 6After the value shown is first greater than the threshold at approximately the 26th second, and the average value of the singing probabilities for a period of time after the 26th second is greater than the threshold, it can finally be determined that the 26th second is the starting time of the vocal sound for rough positioning.
[0066] Optionally, the method for determining the average value of the vocal sound probabilities based on the vocal sound probabilities corresponding to times within the reference time interval may include: sorting the vocal sound probabilities corresponding to times within the reference time interval in descending order; determining the average value of the vocal sound probabilities based on the top M vocal sound probabilities after sorting, where M is a positive integer. The top M vocal sound probabilities arranged in descending order refer to the M vocal sound probabilities with the largest values, and taking the average of these M vocal sound probabilities can obtain the average value of the vocal sound probabilities. Optionally, the reference time interval may be the next 2M seconds, that is, an even time interval after the first candidate time, such as 10 seconds or 20 seconds, and the vocal sound probability is taken as half of this time interval, that is, the vocal sound probabilities arranged in the first half are selected for averaging. For example, for a reference time interval from 27 seconds to 37 seconds, the average value of the vocal sound probabilities can be determined by selecting the top 5 vocal sound probabilities arranged in descending order.
[0067] As an alternative method, the determination of the average value of the vocal sound probabilities can also be to take the average of all the vocal sound probabilities corresponding to times within the reference time interval. For example, for the vocal sound audio corresponding to a reference time interval from 27 seconds to 37 seconds, there are 10 vocal sound probabilities, and the final average value of the vocal sound probabilities is the average value of these 10 vocal sound probabilities. It can also be the vocal sound probability ranked at the Mth position among the vocal sound probabilities arranged in ascending order, that is, taking the average of the smaller several vocal sound probabilities to obtain the average value of the vocal sound probabilities. It can also be to take the average of M consecutive vocal sound probabilities. For example, for the vocal sound probabilities corresponding to times from 27 seconds to 32 seconds within the above reference time interval, the embodiments of the present application do not limit the determination of the average value of the vocal sound probabilities here. If the average value of the vocal sound probabilities determined by the above steps meets the conditions, here it means being greater than or equal to the probability threshold, then the first candidate time can be determined as the starting time of the first vocal sound.
[0068] It can be found that the method for determining the starting time of the first human voice is that after selecting the first candidate time, it is also necessary to combine whether the preset conditions are met within a period of time to refer to the rationality of the first candidate time. This is a rule for combining multiple conditions. As a way of rough positioning, it can ensure more accurate results. This screening rule is designed based on the principle that there will still be human voices after the start of the normal human voice. Because in the embodiments of the present application, the rough positioning result itself is regarded as a moment with a higher probability of human voice, and at the same time, it is also limited that there is a higher probability of human voice in a future period of time. Especially when applied to the human voice in a song, this is because in most cases, the starting time of the human voice will start singing a whole sentence or a paragraph continuously, which can avoid extreme situations. For example, there is no human voice in a period before and after the appearance of a tone word. The starting time of the human voice determined only by the condition that the first one is greater than or equal to the probability threshold may not only lack accuracy, but also has certain limitations in its application scenarios.
[0069] In one embodiment, the method for determining the reference human voice audio according to the starting time of the first human voice may be: determining an intercepted time interval according to the starting time of the first human voice, the starting time of the intercepted time interval is before the starting time of the first human voice, and the end time of the intercepted time interval is after the starting time of the first human voice; determining the human voice audio segment corresponding to the intercepted time interval in the target human voice audio as the reference human voice audio. That is, taking the starting time of the first human voice as the reference time point, and taking the range of a period of time before and a period of time after as the final intercepted time interval. Since the starting time of the first human voice is a value within the time range corresponding to the target human voice audio, the intercepted time interval is also a period of time within the time range corresponding to the target human voice audio. The human voice audio corresponding to the intercepted time interval is the reference human voice audio. For example, the starting time of the first human voice is the 26th second. Taking the 26th second as the reference time point, taking the interval of 0.4 seconds before and 0.6 seconds after, that is, the time range from 25.6 seconds to 26.6 seconds is the intercepted time interval. It can be found that the time granularity of the intercepted time interval is smaller than that of the target human voice audio, which can make the determination of the starting time of the human voice more accurate. For the specific range of the intercepted time interval, through experiments, it is found that a better effect is obtained by taking an interval of 1.6 seconds in the fine positioning of the starting point of the human voice in a song. Of course, other interval ranges can also be obtained through experiments according to specific scenarios. Here, the time range of the intercepted time interval is not limited.
[0070] S204. Determine the audio energy value sequence of the reference human voice audio.
[0071] In one embodiment, the final fine positioning can be completed by calculating the audio energy in an interval near the rough positioning moment. The starting time of the first human voice is used as the rough positioning moment, and the reference human voice audio corresponds to the human voice audio in an interval on the left and right of this rough positioning moment. For this purpose, in the first step, the power spectrum of each moment of the reference human voice audio can be calculated and mapped to a decibel value to equivalent the energy information, and the decibel value is used to represent the level of the current audio energy. The finally determined sequence of audio energy values is also the decibel values of each moment arranged in chronological order.
[0072] S205, if the starting time of the second human voice is determined according to the sequence of audio energy values, then the starting time of the second human voice is determined as the starting time of the human voice of the audio to be processed.
[0073] In one embodiment, the method for determining the starting time of the second human voice according to the sequence of audio energy values is similar to the method for determining the starting time of the first human voice according to the sequence of human voice probabilities, that is, in the sequence of audio energy values, the first second candidate time greater than or equal to the energy threshold can be determined first, and then, taking the second candidate time as the reference, if the energy thresholds corresponding to the times after it meet the set conditions, the second candidate time can be determined as the starting time of the second human voice. The difference is that the set conditions here can refer to that the continuously sorted audio energy values in the sequence of audio energy values exceed the pre-specified energy threshold, or that more than 90% of the audio energy values among the consecutive first Y (Y is a positive integer) after the first audio energy value greater than or equal to the energy threshold are greater than or equal to the energy threshold.
[0074] Optionally, the steps of the method for determining the starting time of the second human voice may include: determining the first audio energy value greater than or equal to the energy threshold from the sequence of audio energy values, and determining the second candidate time corresponding to the first audio energy value greater than or equal to the energy threshold; if the audio energy values corresponding to the first N times after the corresponding time are all greater than or equal to the energy threshold, then the second candidate time is determined as the starting time of the second human voice, where N is a positive integer. Exemplarily, please refer to Figure 7 , Figure 7 is a schematic diagram of a sequence of audio energy values provided by an embodiment of the present application. The energy threshold is 30 decibels. Since the time corresponding to the audio energy value is in milliseconds, the discrete sequence of audio energy values is Figure 7 shown in the distribution of the sequence of audio energy values looks like it is continuous. As Figure 7As shown, the audio energy value at approximately the 26.4th second first exceeded the energy threshold, and the sequence of audio energy values corresponding to the time period from 26.4 seconds to 27.6 seconds was above the energy threshold for a period of time afterwards, meeting the requirement of continuously exceeding the energy threshold N times. Therefore, the 26.4th second is the starting time of the second human voice and can be used as the final starting point of the human voice. In this embodiment, an energy range can be arbitrarily specified for calculation based on the starting time of the first human voice, and the final positioning accuracy can reach the millisecond level.
[0075] It should be noted that if the starting time of the second human voice cannot be determined from the audio energy value sequence, that is, a more accurate starting time of the human voice cannot be further determined, the starting time of the first human voice can be selected as the starting time point of the audio to be processed, or determined in combination with other conditions.
[0076] Optionally, if there are multiple target human voice audio files, and the starting time of the first human voice is determined in the i-th target human voice audio file, when determining the starting time of the second human voice in the (i + 1)-th target human voice audio file, it is necessary to determine whether the human voice audio corresponding to the starting time of the second human voice is continuous with the human voice part in the i-th target human voice audio file. If so, it means that the human voice corresponding to the starting time of the second human voice is actually continuous with the human voice in the previous audio segment and is not a real starting point of the human voice. Therefore, the starting time of the second human voice needs to be excluded. Otherwise, it can be used as the starting time of the human voice of the audio to be processed. At this time, multiple starting times of the human voice can be determined for a single audio to be processed, and these multiple starting times of the human voice can be applied to scenarios such as singing alignment or skipping the intermediate accompaniment of song transitions. In addition, after determining a starting point of the human voice, if the adjacent human voice audio is a segment without human voice, then the starting point of the human voice can be detected for the audio segment after this target human voice audio without human voice, which can save detection time.
[0077] Based on the above-described embodiments, this solution can achieve rapid calculation of the starting time point of music human voice, which can also be used for the function of skipping the prelude in a music player, improving the experience of music player users. Combining the detailed description of the above steps, the solution of this embodiment can be summarized as the following content, including: five steps of inputting music, segmenting, preprocessing, singing recognition, and energy calculation. The specific flowchart can be seen in Figure 8 the schematic diagram given. A simple explanation of each step is as follows:
[0078] a) Input music: Input a video or audio file, and extract the audio track as the input music for the algorithm;
[0079] b) Segmenting: Divide the music evenly into G equal parts (G is a positive integer) according to time, and send them into the subsequent calculation process in sequence until the starting time of the human voice is obtained in a certain equal part, then stop;
[0080] c) Pretreatment: Separate the sound source of the music input in step b), strip its vocal track for subsequent operations;
[0081] d) Singing recognition: Calculate the probability of someone singing at each moment for the vocal track input in step c), so as to complete the rough positioning of the starting point, accurate to seconds;
[0082] e) Energy calculation: Calculate the audio energy value at each moment around the rough positioning of the starting point obtained in step d), and take the moment when the energy first breaks through a certain threshold as the starting point of the human voice, accurate to milliseconds.
[0083] Among them, steps b) and c) can also be exchanged. The ultimate goal is to obtain equally divided human voice audio. By using the time prior of the prelude and performing segmented processing, the calculation time can be greatly shortened. In addition, introducing the sound source separation technology in the pretreatment part, stripping the vocal track into the calculation link, can improve the accuracy of starting point positioning. Using the singing recognition algorithm can accurately give the probability of singing human voice at each moment, with an accuracy reaching the second level. Using the audio energy information to accurately locate the starting moment of the human voice can make the accuracy of the final result reach the millisecond level.
[0084] In summary, the embodiments of the present application have at least the following advantages:
[0085] The audio to be processed is segmented to obtain audio segments. When applying this solution to the scenario of calculating the starting time of the human voice in music, using the time prior condition of the prelude can greatly shorten the calculation time and improve the calculation efficiency. The human voice audio of the audio segments is sequentially used as the target human voice audio for processing. In the rough positioning process, the first vocal probability exceeding the threshold and the vocal probability within a period of time after it are used for joint evaluation, which can ensure the accuracy of the rough positioning of the human voice. In the fine positioning process, it is determined by using the audio energy values within a period of time interval around the rough positioning moment, and in the specific screening, the preset condition of the first continuously exceeding energy threshold audio energy value is used to reposition the starting time of the human voice, which can make the positioning accuracy reach the millisecond level and further improve the accuracy of the starting time of the human voice.
[0086] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a data processing device provided by the embodiments of the present application. The above data processing device can be a computer program (including program code) running in a computer device. For example, the data processing device is an application software; the device can be used to execute the corresponding steps in the method provided by the embodiments of the present application. As Figure 9 shown, the data processing device 90 may include: a determination module 901.
[0087] A determination module 901, configured to determine a target human voice audio according to the audio to be processed, and determine a human voice probability sequence of the target human voice audio, where the human voice probability sequence includes the human voice probabilities of the human voice audios corresponding to each first unit time, and the human voice probabilities of the human voice audios corresponding to each first unit time are sorted in chronological order;
[0088] The determination module 901 is further configured to, if a first human voice start time is determined according to the human voice probability sequence, determine a reference human voice audio from the target human voice audio according to the first human voice start time;
[0089] The determination module 901 is further configured to determine an audio energy value sequence of the reference human voice audio, where the audio energy value sequence includes the audio energy values of the human voice audios corresponding to each second unit time, and the audio energy values of the human voice audios corresponding to each second unit time are sorted in chronological order;
[0090] The determination module 901 is further configured to, if a second human voice start time is determined according to the audio energy value sequence, determine the second human voice start time as the human voice start time of the audio to be processed.
[0091] In one embodiment, the determination module 901 is further configured to: determine the first human voice probability that is first greater than or equal to a probability threshold from the human voice probability sequence, and determine a first candidate time corresponding to the first human voice probability that is first greater than or equal to the probability threshold; determine a reference time interval according to the first candidate time, and determine a human voice probability mean value according to the human voice probabilities whose corresponding times are within the reference time interval; if the human voice probability mean value is greater than or equal to the probability threshold, determine the first candidate time as the first human voice start time.
[0092] In one embodiment, the determination module 901 is specifically configured to: sort the human voice probabilities whose corresponding times are within the reference time interval in descending order; determine a human voice probability mean value according to the first M human voice probabilities arranged in the front after sorting, where M is a positive integer.
[0093] In one embodiment, the determination module 901 is further configured to determine the first audio energy value that is first greater than or equal to an energy threshold from the audio energy value sequence, and determine a second candidate time corresponding to the first audio energy value that is first greater than or equal to the energy threshold; if the first N audio energy values after the corresponding time are all greater than or equal to the energy threshold, determine the second candidate time as the second human voice start time, where N is a positive integer.
[0094] In one embodiment, the determining module 901 is specifically configured to: determine an interception time interval according to the starting time of the first human voice, where the starting time of the interception time interval is before the starting time of the first human voice, and the ending time of the interception time interval is after the starting time of the first human voice; determine the human voice audio segment corresponding to the interception time interval in the target human voice audio as the reference human voice audio.
[0095] In one embodiment, the determining module 901 includes an obtaining unit 9011 and a processing unit 9012, where:
[0096] The obtaining unit 9011 is configured to obtain the audio to be processed and perform segmentation processing on the audio to be processed to obtain at least two audio segments;
[0097] The processing unit 9012 is configured to perform sound source separation processing on at least two audio segments to obtain the human voice audio of each audio segment, and sequentially determine the human voice audio of each audio segment as the target human voice audio according to the time sequence.
[0098] In one embodiment, the processing unit 9012 included in the determining module 901 is further configured to: perform Fourier transform processing on the target human voice audio to obtain the spectrogram of the target human voice audio; use the audio processing network to process the spectrogram to obtain the human voice probability sequence of the target human voice audio.
[0099] It can be understood that the functions of the functional modules of the data processing device described in the embodiments of the present application can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions of the above method embodiments, which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either.
[0100] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a computer device 1000 provided in the embodiments of the present application. The computer device 1000 may include independent devices (such as one or more of a server, a node, a terminal, etc.), or may include components inside an independent device (such as a chip, a software module, or a hardware module, etc.). The computer device 1000 may include at least one processor 1001 and a communication interface 1002. Further optionally, the computer device 1000 may further include at least one memory 1003 and a bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 are connected through the bus 1004.
[0101] Among them, the processor 1001 is a module for performing arithmetic operations and / or logical operations, and can specifically be one or a combination of multiple processing modules such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor unit (MPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a coprocessor (assisting the central processing unit to complete corresponding processing and applications), and a microcontroller unit (MCU).
[0102] The communication interface 1002 can be used to provide information input or output for the at least one processor. And / or, the communication interface 1002 can be used to receive data sent externally and / or send data to the outside, and can be a wired link interface including, for example, an Ethernet cable, or a wireless link (Wi-Fi, Bluetooth, universal wireless transmission, and other short-range wireless communication technologies, etc.) interface.
[0103] The memory 1003 is used to provide storage space, and data such as an operating system and computer programs can be stored in the storage space. The memory 1003 can be one or a combination of multiple types such as a random access memory (RAM), a read-only memory (ROM), an erasable programmable read only memory (EPROM), or a compact disc read-only memory (CD-ROM), etc.
[0104] At least one processor 1001 in the computer device 1000 is used to call a computer program stored in at least one memory 1003 for executing the foregoing data processing method, such as the data processing method described in the foregoing Figure 2 、 Figure 4 embodiments shown.
[0105] In a possible implementation manner, the processor 1001 in the computer device 1000 is used to call a computer program stored in at least one memory 1003 for performing the following operations:
[0106] Determine the target human voice audio according to the audio to be processed, and determine the human voice probability sequence of the target human voice audio. The human voice probability sequence includes the human voice probabilities of the human voice audios corresponding to each first unit time, and the human voice probabilities of the human voice audios corresponding to each first unit time are sorted in chronological order. If the first human voice start time is determined according to the human voice probability sequence, then determine the reference human voice audio from the target human voice audio according to the first human voice start time. Determine the audio energy value sequence of the reference human voice audio. The audio energy value sequence includes the audio energy values of the human voice audios corresponding to each second unit time, and the audio energy values of the human voice audios corresponding to each second unit time are sorted in chronological order. If the second human voice start time is determined according to the audio energy value sequence, then determine the second human voice start time as the human voice start time of the audio to be processed.
[0107] In one embodiment, the processor 1001 is further configured to: determine the first human voice probability that is greater than or equal to the probability threshold from the human voice probability sequence, and determine the first candidate time corresponding to the first human voice probability that is greater than or equal to the probability threshold. Determine the reference time interval according to the first candidate time, and determine the human voice probability mean according to the human voice probabilities corresponding to the times within the reference time interval. If the human voice probability mean is greater than or equal to the probability threshold, then determine the first candidate time as the first human voice start time.
[0108] In one embodiment, when the processor 1001 determines the human voice probability mean according to the human voice probabilities corresponding to the times within the reference time interval, it is specifically configured to: sort the human voice probabilities corresponding to the times within the reference time interval in descending order. Determine the human voice probability mean according to the first M human voice probabilities after sorting, where M is a positive integer.
[0109] In one embodiment, the processor 1001 is further configured to: determine the first audio energy value that is greater than or equal to the energy threshold from the audio energy value sequence, and determine the second candidate time corresponding to the first audio energy value that is greater than or equal to the energy threshold. If the first N audio energy values corresponding to the times after the second candidate time are all greater than or equal to the energy threshold, then determine the second candidate time as the second human voice start time, where N is a positive integer.
[0110] In one embodiment, when the processor 1001 determines the reference human voice audio from the target human voice audio according to the first human voice start time, it is specifically configured to: determine the cut time interval according to the first human voice start time. The start time of the cut time interval is before the first human voice start time, and the end time of the cut time interval is after the first human voice start time. Determine the human voice audio segment corresponding to the cut time interval in the target human voice audio as the reference human voice audio.
[0111] In one embodiment, when the processor 1001 determines the target human voice audio based on the audio to be processed, it is specifically used to: obtain the audio to be processed and segment the audio to be processed to obtain at least two audio segments; perform sound source separation processing on the at least two audio segments to obtain the human voice audio of each audio segment, and determine the human voice audio of each audio segment as the target human voice audio in chronological order.
[0112] In one embodiment, when the processor 1001 determines the vocal probability sequence of the target vocal audio, it is specifically used to: perform Fourier transform processing on the target vocal audio to obtain a spectrogram of the target vocal audio; and use an audio processing network to process the spectrogram to obtain a vocal probability sequence of the target vocal audio.
[0113] It should be understood that the computer device 1000 described in the embodiment of the present application can execute the above Figure 2 as well as Figure 4 The description of the data processing method in the corresponding embodiment can also be performed as described above. Figure 9 The description of the data processing device 90 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated here either.
[0114] In addition, it should be pointed out here that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the computer device 1000 of the data processing method mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, the computer program can execute the data processing method mentioned above. Figure 2 and Figure 4 The description of the above data processing method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0115] The above computer-readable storage medium may be the data processing device provided in any of the foregoing embodiments or the internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.
[0116] In one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method provided in one aspect of the embodiments of the present application.
[0117] In one aspect of the present application, there is provided another computer program product. The computer program product includes a computer program or computer instructions, and when the computer program or the computer instructions are executed by a processor, the steps of the data processing method provided in the embodiments of the present application are implemented.
[0118] Finally, it should also be noted that the relational terms such as "first" and "second" in the description and claims of the present application and the above drawings are used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal device including the said element.
[0119] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A data processing method, characterized in that, The method includes: Determining a target human voice audio according to the audio to be processed, and determining a human voice probability sequence of the target human voice audio, where the human voice probability sequence includes the human voice probabilities of the human voice audio corresponding to each first unit time, and the human voice probabilities of the human voice audio corresponding to each first unit time are sorted in chronological order; If a first human voice start time is determined according to the human voice probability sequence, determining a reference human voice audio from the target human voice audio according to the first human voice start time; Determining an audio energy value sequence of the reference human voice audio, where the audio energy value sequence includes the audio energy values of the human voice audio corresponding to each second unit time, and the audio energy values of the human voice audio corresponding to each second unit time are sorted in chronological order; If a second human voice start time is determined according to the audio energy value sequence, determining the second human voice start time as the human voice start time of the audio to be processed.
2. The method according to claim 1, wherein The method further includes: Determining the first human voice probability that is greater than or equal to a probability threshold from the human voice probability sequence, and determining a first candidate time corresponding to the first human voice probability that is greater than or equal to the probability threshold; Determining a reference time interval according to the first candidate time, and determining a human voice probability mean according to the human voice probabilities corresponding to the times within the reference time interval; If the human voice probability mean is greater than or equal to the probability threshold, determining the first candidate time as the first human voice start time.
3. The method according to claim 2, wherein The determining the human voice probability mean according to the human voice probabilities corresponding to the times within the reference time interval includes: Sorting the human voice probabilities corresponding to the times within the reference time interval in descending order; Determining the human voice probability mean according to the first M human voice probabilities in the sorted order, where M is a positive integer.
4. The method according to claim 1, characterized in that, The method further includes: Determining the first audio energy value that is greater than or equal to an energy threshold from the audio energy value sequence, and determining a second candidate time corresponding to the first audio energy value that is greater than or equal to the energy threshold; If the audio energy values of the first N audio energy values after the corresponding time are all greater than or equal to the energy threshold, determining the second candidate time as the second human voice start time, where N is a positive integer.
5. The method according to claim 1, wherein The determining the reference human voice audio from the target human voice audio according to the first human voice start time includes: Determining an interception time interval according to the first human voice start time, where the start time of the interception time interval is before the first human voice start time, and the end time of the interception time interval is after the first human voice start time; Determining the human voice audio segment corresponding to the interception time interval in the target human voice audio as the reference human voice audio.
6. The method according to claim 1, characterized in that The determining the target human voice audio according to the audio to be processed includes: Obtaining the audio to be processed and performing segmentation processing on the audio to be processed to obtain at least two audio segments; Performing sound source separation processing on the at least two audio segments to obtain the human voice audio of each audio segment, and sequentially determining the human voice audio of each audio segment as the target human voice audio according to the time order.
7. The method according to claim 1, characterized in that, The determination of the human voice probability sequence of the target human voice audio includes: Performing Fourier transform processing on the target human voice audio to obtain a spectrogram of the target human voice audio; Processing the spectrogram by using an audio processing network to obtain the human voice probability sequence of the target human voice audio.
8. A data processing device, characterized in that, Including: A determination module, configured to determine a target human voice audio according to the audio to be processed, and determine the human voice probability sequence of the target human voice audio, where the human voice probability sequence includes the human voice probabilities of the human voice audio corresponding to each first unit time, and the human voice probabilities of the human voice audio corresponding to each first unit time are sorted in chronological order; The determination module is further configured to, if a first human voice start time is determined according to the human voice probability sequence, determine a reference human voice audio from the target human voice audio according to the first human voice start time; The determination module is further configured to determine an audio energy value sequence of the reference human voice audio, where the audio energy value sequence includes the audio energy values of the human voice audio corresponding to each second unit time, and the audio energy values of the human voice audio corresponding to each second unit time are sorted in chronological order; The determination module is further configured to, if a second human voice start time is determined according to the audio energy value sequence, determine the second human voice start time as the human voice start time of the audio to be processed.
9. A computer device, characterized in that, Including: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface, where the network interface is used to provide network communication functions, the memory is used to store program codes, and the processor is used to call the program codes to execute the data processing method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, where the computer program includes program instructions, and when the program instructions are executed by a processor, the data processing method according to any one of claims 1-7 is executed.
11. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, the steps of the data processing method according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Audio data processing method and device and computer storage medium
CN109920446A
Human voice detection method and device, electronic equipment and computer readable storage medium
CN112967738A