Voice positioning method, device, computer-readable medium and electronic device
Through the speech recognition model, the spectrum information of the speech information is processed, and the subject speech in the speech information is identified and positioned, which solves the problems of low accuracy of speech positioning and poor model performance in the prior art, and achieves efficient and accurate speech positioning.
Patent Information
- Application Number
- CN202210080156.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-01-24
AI Technical Summary
In the prior art, there are problems such as low accuracy of speech positioning, long use, difficulty in obtaining labeled data, and poor model performance.
By obtaining voice information, processing is used to obtain spectrum information, input it into the speech recognition model, identifying the main speech through the model, obtaining the main speech information and the probability curve, and determining the starting and ending time point of the speech based on the local extreme points of the probability curve.
It improves the accuracy and timeliness of speech positioning, avoids the high cost and low model accuracy caused by manual annotation of data, and enhances the user viscosity and user experience of the product.
Smart Images

Figure CN114420097B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of artificial intelligence, and particularly relates to a voice localization method, a voice localization device, a computer-readable medium, and an electronic device. Background Art
[0002] With the development of multimedia technology, people often use electronic devices to record audio or video. In order to extract the human voice and the corresponding time in the audio or video, it is usually necessary to separate the human voice from the background sound and then localize the human voice.
[0003] Currently, there are mainly two methods for voice localization. One is the localization method based on sound source separation. However, this method depends on the accuracy of sound source separation. Since sound source separation itself is not perfect, it will bring some misjudgments. In addition, other human voices in the audio or video will also be judged as the target human voice, resulting in misjudgments. Moreover, sound source separation is time-consuming and will increase the resource occupancy of voice localization. The other is the scheme based on convolutional neural network prediction. However, this scheme depends on the annotation of data. It is difficult to obtain the annotation data itself, and manual annotation will occupy a large amount of manpower. If a model trained with weakly annotated data is used to identify and localize the voice, there will be a problem of low accuracy. Summary of the Invention
[0004] The purpose of this application is to provide a voice localization method, a voice localization device, a computer-readable medium, and an electronic device, which can overcome the problems of low voice localization accuracy, long time consumption, difficult acquisition of annotation data, and poor model performance in the related art.
[0005] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of this application.
[0006] According to one aspect of the embodiments of this application, a voice localization method is provided. The method includes: obtaining voice information, processing the voice information to obtain spectrum information corresponding to the voice information, where the voice information includes background sound and main voice; inputting the spectrum information into a voice recognition model, and recognizing the main voice in the spectrum information through the voice recognition model to obtain main voice information, where the main voice information includes a main voice probability curve; determining the start and end time points corresponding to the main voice in the voice information according to local extreme points in the main voice probability curve.
[0007] According to one aspect of the embodiments of the present application, a voice positioning device is provided. The device includes: an information processing module configured to obtain voice information and process the voice information to obtain spectrum information corresponding to the voice information, where the voice information includes background sound and main voice; a voice recognition module configured to input the spectrum information into a voice recognition model and recognize the main voice in the spectrum information through the voice recognition model to obtain main voice information, where the main voice information includes a main voice probability curve; a voice positioning module configured to determine start and end time points corresponding to the main voice in the voice information according to local extreme points in the main voice probability curve.
[0008] In some embodiments of the present application, the spectrum information is a Mel spectrogram; based on the above technical solution, the information processing module is configured to: frame and window the voice information, and perform Fourier transform on the windowed voice information to obtain a spectrogram corresponding to the voice information; filter the spectrogram through a Mel scale filter to obtain the Mel spectrogram.
[0009] In some embodiments of the present application, the voice recognition model includes a convolutional network module, a feature enhancement network module, a long short-term memory network module, and a classification prediction module; based on the above technical solution, the voice recognition module includes: a convolutional unit configured to perform segmented feature extraction on the spectrum information through the convolutional network module to obtain a plurality of spectrum feature maps; an enhancement unit configured to perform downsampling, upsampling, and backpropagation on each of the spectrum feature maps through the feature enhancement network module to obtain spectrum enhancement feature maps corresponding to each of the spectrum feature maps; a fusion unit configured to fuse deep semantic and shallow time information in each of the spectrum enhancement feature maps through the long short-term memory network module to obtain fusion feature information; a prediction unit configured to predict the main voice in the fusion feature information through the classification prediction module to obtain the main voice information.
[0010] In some embodiments of the present application, based on the above technical solution, the convolutional network module includes a plurality of convolutional network units with the same structure, and each convolutional network unit includes a first convolutional unit, a second convolutional unit, a pooling layer, and a dropout layer. At the same time, both the first convolutional unit and the second convolutional unit include a two-dimensional convolutional layer, a batch normalization layer, and an activation function layer.
[0011] In some embodiments of the present application, the feature enhancement network module includes a first convolutional network unit and a second convolutional network unit, and the structures of the first convolutional network unit and the second convolutional network unit are the same as that of the convolutional network unit; Based on the above technical solutions, the enhancement unit is configured to: perform downsampling on the spectral feature map through the first convolutional network unit to obtain a first feature map, and perform downsampling on the first feature map through the second convolutional network unit to obtain a second feature map; perform upsampling on the second feature map to obtain a third feature map, and at the same time, perform a convolution operation on the first feature map using a 1×1 convolutional kernel, and splice the third feature map and the first feature map after convolution processing to obtain a fourth feature map; perform upsampling on the fourth feature map to obtain a fifth feature map, and at the same time, perform a convolution operation on the spectral feature map using a 1×1 convolutional kernel, and splice the fifth feature map and the spectral feature map after convolution processing to obtain the spectral enhancement feature map; wherein, the step size corresponding to the upsampling is the same as the step size corresponding to the downsampling.
[0012] In some embodiments of the present application, based on the above technical solutions, the voice localization module is configured to: divide the main voice probability curve into multiple main voice intervals according to any two adjacent wave valleys in the main voice probability curve; obtain local extreme points in each of the main voice intervals, mark the time point corresponding to the maximum value point as the start time point of the main voice, and mark the time point corresponding to the minimum value point as the end time point of the main voice.
[0013] In some embodiments of the present application, based on the above technical solutions, the voice localization device further includes: a sample acquisition module, configured to acquire a voice sample and automatically generated main voice annotation information corresponding to the voice sample; a model training module, configured to train a voice recognition model to be trained according to the voice sample and the main voice annotation information to obtain the voice recognition model.
[0014] In some embodiments of the present application, based on the above technical solutions, the sample acquisition module is configured to: separate the sound source of the voice sample to obtain a background sound waveform diagram and a main voice waveform diagram; slice the background sound waveform diagram and the main voice waveform diagram according to a preset time interval, and determine the energy ratio between the main voice energy and the background sound energy corresponding to each time slice; divide the voice sample into multiple voice intervals according to the start time points of the main voices of each sentence in the voice sample; respectively use each of the voice intervals as a target voice interval, obtain the target energy ratio corresponding to the start time point of the target voice interval, and determine the maximum energy ratio according to the target energy ratio and the lower bound of the energy ratio; compare the energy ratio corresponding to each time slice in the target voice interval with the maximum energy ratio, determine the main voice interval according to the consecutive time slices in the target voice interval whose energy ratio is greater than or equal to the maximum energy ratio, and label the main voice interval to form the voice annotation information.
[0015] In some embodiments of the present application, the speech recognition model to be trained includes a convolutional network module to be trained, a feature enhancement network module to be trained, a long short-term memory network module to be trained, and a classification prediction module to be trained; based on the above technical solutions, the model training module includes: a first training unit configured to fix the parameters of the long short-term memory network module to be trained and the classification prediction module to be trained, and train the convolutional network module to be trained and the feature enhancement network module to be trained according to the voice sample and the main voice annotation information to obtain a converged convolutional network module and a feature enhancement network module; a second training unit configured to fix the parameters of the convolutional network module and the feature enhancement network module, and train the long short-term memory network module to be trained and the classification prediction module to be trained according to the voice sample and the main voice annotation information to obtain a converged long short-term memory network module and a classification prediction module.
[0016] In some embodiments of the present application, based on the above technical solutions, the first training unit is configured to: divide the voice sample into multiple groups according to a preset quantity, randomly intercept a voice segment with a preset length from each group of the voice samples; input the Mel spectrogram corresponding to the voice segment into the speech recognition model to be trained, and recognize the main voice in the Mel spectrogram corresponding to the voice segment through the speech recognition model to be trained to obtain main voice prediction information; determine the main voice prediction error according to the main voice prediction information and the main voice annotation information, and optimize the parameters of the convolutional network module to be trained and the feature enhancement network module to be trained according to the main voice prediction error until the convolutional network module and the feature enhancement network module are obtained.
[0017] In some embodiments of the present application, based on the above technical solutions, the second training unit is configured to: obtain the maximum duration in the voice samples, align the durations of other voice samples with the maximum duration by padding with zeros, and divide the voice samples into multiple groups according to a preset number; input the Mel spectrograms corresponding to each group of the voice samples into a voice recognition model to be trained including a trained convolutional network module and a feature enhancement module, and recognize the main voice in the Mel spectrogram corresponding to the voice samples through the voice recognition model to be trained, so as to obtain main voice prediction information; determine a main voice prediction error according to the main voice prediction information and the main voice annotation information, and optimize the parameters of the long short-term memory network module and the classification prediction network module according to the main voice prediction error until the long short-term memory network module and the classification prediction network module are obtained.
[0018] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the voice localization method in the above technical solutions.
[0019] According to one aspect of the embodiments of the present application, there is provided an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the voice localization method in the above technical solutions by executing the executable instructions.
[0020] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable medium. The processor of the electronic device reads the computer instructions from the computer-readable medium, and the processor executes the computer instructions, so that the electronic device executes the voice localization method in the above technical solutions.
[0021] In the technical solutions provided in the embodiments of the present application, by using a voice recognition model to process the spectral information corresponding to the voice information, the main voice information in the voice information is obtained, and the main voice information includes a main voice probability curve, and then the start and end time points corresponding to the main voice in the voice information are determined according to the main voice probability curve. On the one hand, the present application can accurately locate the main voice in the voice information, improve the accuracy and timeliness of voice localization; on the other hand, it can avoid the high cost and low model accuracy caused by manual annotation of data; on the other hand, it can improve the user stickiness and user experience of the product.
[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other accompanying drawings based on these drawings without creative efforts.
[0024] Figure 1 Schematically shows an exemplary system architecture block diagram applying the technical solution of the present application.
[0025] Figure 2 Schematically shows a schematic flow diagram of the voice localization method in the present application.
[0026] Figure 3 Schematically shows a schematic architecture diagram of the voice recognition model in the present application.
[0027] Figure 4 Schematically shows a schematic structural diagram of the convolutional network unit in the present application.
[0028] Figure 5 Schematically shows a schematic flow diagram of obtaining the main voice information through the voice recognition model in the present application.
[0029] Figure 6 Schematically shows a schematic flow diagram of obtaining the spectrum enhancement feature map in the present application.
[0030] Figure 7 Schematically shows a schematic flow diagram of obtaining the main voice annotation information in the present application.
[0031] Figure 8 Schematically shows the lyrics lrc file marked with the starting time point of the lyrics in the present application.
[0032] Figure 9 Schematically shows a schematic flow diagram of local training in the present application.
[0033] Figure 10 Schematically shows a schematic flow diagram of global training in the present application.
[0034] Figure 11 Schematically shows a block diagram of the structure of the voice localization device in the present application.
[0035] Figure 12 Schematically shows a block diagram of the computer system structure of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0037] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.
[0038] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0039] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0040] Figure 1 An exemplary system architecture block diagram applying the technical solution of this application is schematically shown.
[0041] As Figure 1 shown, the system architecture 100 may include a terminal device 110, a network 120, and a server 130. The terminal device 110 may include various electronic devices such as a smart phone, a tablet computer, a laptop computer, etc. Further, the terminal device 110 may also be a device including a voice recording unit, or a voice recording device, such as a voice recorder, a desktop computer connected to an external microphone, etc. The server 130 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The network 120 may be a communication medium of various connection types capable of providing a communication link between the terminal device 110 and the server 130, for example, it may be a wired communication link or a wireless communication link.
[0042] According to implementation requirements, the system architecture in the embodiments of the present application may have any number of terminal devices, networks, and servers. For example, the server 130 may be a server group composed of multiple server devices. In addition, the technical solutions provided in the embodiments of the present application may be applied to the terminal device 110, or may be applied to the server 130, or may be jointly implemented by the terminal device 110 and the server 130. The present application does not make any special limitations on this.
[0043] In some embodiments of the present application, the user obtains voice information through the terminal device 110 and transmits the voice information to the server 130 through the network 120. The voice information may be, for example, songs, TV / movie clips, and other audio-visual materials, such as live audio-visual materials, and the voice information includes background sound and main voice. Specifically, the accompaniment in a song is the background sound, and the human voice is the main voice. The background music in a TV / movie clip is the background sound, and the dialogue and monologue are the main voice. The music and noisy human voices in live audio-visual materials are the background sound, and the voice of the speaker is the main voice. After receiving the voice information, the server 130 can process it to obtain the corresponding spectrum information, which is information of frequencies recognizable by the human ear. Then, a speech recognition model is called, and the spectrum information is input into the speech recognition model. The main voice in the spectrum information is recognized through the speech recognition model to obtain the main voice information. Further, the start and end time points corresponding to the main voice in the voice information can be determined according to the main voice information, so as to realize the positioning of the main voice in the voice information.
[0044] In some embodiments of the present application, the voice positioning device may also be configured in the terminal device 110. After the user determines the voice information to be positioned in the terminal device 110, the terminal device 110 can process it to obtain the corresponding spectrum information. Then, a speech recognition model is called, and the spectrum information is input into the speech recognition model. The main voice in the spectrum information is recognized through the speech recognition model to obtain the main voice information. Further, the start and end time points corresponding to the main voice in the voice information can be determined according to the main voice information, so as to realize the positioning of the main voice in the voice information. Specifically, the main voice information includes a main voice probability curve, and the start and end time points corresponding to the main voice in the voice information can be determined according to the local extreme points in the main voice probability curve.
[0045] In some embodiments of the present application, the speech recognition model set in the terminal device 110 or the server 130 is a machine learning model for voice positioning based on artificial intelligence technology.
[0046] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0047] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0048] Computer Vision Technology (CV) Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes for object recognition, measurement, and other machine vision, and further performing graphic processing to make the computer process images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image information annotation, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping.
[0049] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0050] In the related art of the present application, taking the positioning of singing in a song as an example, there are mainly two positioning schemes: positioning singing based on sound source separation and positioning singing based on convolutional neural network prediction.
[0051] For the singing positioning scheme based on sound source separation, first, sound source separation is performed on the song to obtain the vocal track and the accompaniment music track. Then, the energy of each small time interval on the vocal track and the accompaniment music track is calculated. Finally, singing positioning is performed according to the relationship between the ratio of the vocal energy to the accompaniment music energy and a preset threshold. Specifically, when the ratio of the vocal energy to the accompaniment music energy is greater than or equal to the preset threshold, it is determined that there is singing in this time interval; when the ratio of the vocal energy to the accompaniment music energy is less than the preset threshold, it is determined that there is no singing in this time interval.
[0052] For the singing positioning scheme based on convolutional neural network prediction, it depends on labeled data. However, due to the large workload of complete annotation, a weak annotation method is usually adopted, that is, it is annotated whether there is singing in each song. If there is singing, it is determined that the whole song is singing. After obtaining a basic model by training the convolutional neural network with weakly labeled samples, the prediction results of the model are used to correct the labels, and then the convolutional neural network model is retrained according to the corrected labels. After repeating many times, a stable convolutional neural network is determined and can be used for positioning the singing in the song.
[0053] However, the above two schemes have corresponding drawbacks. Regarding the positioning of song singing based on sound source separation, first, the scheme based on sound source separation depends on the accuracy of sound source separation. Since the sound source separation method itself is not perfect, it will bring some misjudgments. Second, for some live songs, there will be cheers from the audience, and there will be promotional remarks, emotional prelude dialogues of the singer in some songs, all of which will be judged as vocals, resulting in misjudgments of singing. Finally, the sound source separation algorithm itself is time-consuming and will increase the resource occupancy of singing positioning. Regarding the positioning of song singing based on convolutional neural network, the scheme based on convolutional neural network depends on the standard of data. It is difficult to obtain data annotation itself. Manual annotation will consume a large amount of manpower, and the scheme based on weak annotation is relatively poor in performance, thus resulting in a low accuracy rate of the singing positioning prediction results.
[0054] In view of the problems existing in the related art, the following will make a detailed description of the technical solutions such as the speech positioning method, speech positioning device, computer-readable medium, and electronic device provided by the present application in combination with specific embodiments.
[0055] Figure 2 Schematically shows a schematic flowchart of the steps of the speech positioning method in an embodiment of the present application. This speech positioning method can be executed by a terminal device or a server, or jointly executed by a terminal device and a server. AsFigure 2 As shown, the voice localization method in the embodiment of the present application may mainly include the following steps S210 to S230.
[0056] Step S210: Acquire voice information, and process the voice information to acquire frequency spectrum information corresponding to the voice information, wherein the voice information includes background sound and main voice;
[0057] Step S220: inputting the spectrum information into a speech recognition model, and recognizing the main speech in the spectrum information through the speech recognition model to obtain main speech information, wherein the main speech information includes a main speech probability curve;
[0058] Step S230: determining the start and end time points corresponding to the main speech in the speech information according to the local extreme value points in the main speech probability curve.
[0059] In the voice localization method provided in the embodiment of the present application, the spectrum information corresponding to the voice information is processed by using a voice recognition model to obtain the main voice information in the voice information, and then the start and end time points corresponding to the main voice in the voice information are determined according to the main voice information. On the one hand, the present application can accurately locate the main voice in the voice information and improve the accuracy and timeliness of voice localization; on the other hand, it can avoid the high cost and low model accuracy caused by manual annotation of data; on the other hand, it can improve the user stickiness and user experience of the product.
[0060] The specific implementation of each method step of the voice localization method is described in detail below.
[0061] In step S210, voice information is acquired, and the voice information is processed to acquire frequency spectrum information corresponding to the voice information, wherein the voice information includes background sound and main voice.
[0062] In one embodiment of the present application, the voice information includes background sound and main voice. For example, the voice information can be a song. A song consists of a tune and lyrics. The human voice sings the lyrics according to the beat of the tune. Then the tune used for accompaniment is the background sound, and the human voice is the main voice. It can also be a TV series / movie clip. The character dialogue contained in the clip is the main voice, and the interlude is the background sound. Of course, it can also be other types of voice information, such as a video shot containing human voice or a recorded voice containing human voice, etc. It is worth noting that the human voice can be the sound of a real person or the sound of a virtual character. For example, the sound in a song sung by a virtual singer can also be regarded as the human voice.
[0063] In an embodiment of the present application, after obtaining the voice information, it is necessary to process it to obtain a data structure recognizable by the speech recognition model. In the embodiment of the present application, the voice information can be processed to obtain the spectrum information corresponding to the voice information, and the spectrum information is the Mel spectrogram corresponding to the voice information. Specifically, first, the voice information can be preprocessed, and the preprocessed voice information can be subjected to short-time Fourier transform to obtain the spectrogram corresponding to the voice information; then, the spectrogram is filtered by a Mel filter to obtain the Mel spectrogram.
[0064] Among them, the preprocessing performed on the voice information can specifically be to frame the sound signal in the voice information, then window the sound frames obtained by framing, then perform Fourier transform on each frame of the sound signal, and finally stack the results of each frame along a preset dimension to obtain the spectrogram. Since the obtained spectrogram is large and the unit of frequency is Hz, the audible frequency range of the human ear is 20 - 20000 Hz, but the human ear is not linearly sensitive to the Hz unit, but is sensitive to low Hz and insensitive to high Hz. Therefore, in order to obtain appropriate-sized sound features, the spectrogram is usually transformed into Mel spectrum through a Mel-scale filter bank, and the Hz frequency is converted into Mel frequency, so that the human ear's perception of frequency becomes linear. The transformation formula is shown in formula (1):
[0065]
[0066] Among them, f is the frequency corresponding to each time point in the spectrogram, and m is the Mel frequency.
[0067] In step S220, the spectrum information is input into the speech recognition model, and the main speech in the spectrum information is recognized by the speech recognition model to obtain the main speech information, and the main speech information includes the main speech probability curve.
[0068] In an embodiment of the present application, after obtaining the Mel spectrogram corresponding to the voice information, the speech recognition model can be called to process it to obtain the main speech in the voice information. The speech recognition model in the present application is a composite model. Figure 3 shows a schematic diagram of the architecture of the speech recognition model, as Figure 3As shown in the figure, the speech recognition model 300 includes a convolutional network module 301, a feature enhancement network module 302, a long short-term memory network (LSTM) module 303, and a classification prediction module 304. Among them, the convolutional network module 301 includes multiple convolutional network units with the same structure. For example, it can be 4, or 5, etc.; the feature enhancement network module 302 includes a first convolutional network unit 302-1 and a second convolutional network unit 302-2, and the structures of the first convolutional network unit 302-1 and the second convolutional network unit 302-2 are the same as the structure of the convolutional network unit; the classification prediction module 304 is composed of a fully connected layer FC and a softmax layer.
[0069] In an embodiment of the present application, Figure 4 shows a schematic structural diagram of a convolutional network unit, as Figure 4 shown, the convolutional network unit 400 includes a first convolutional unit 401, a second convolutional unit 402, a pooling layer 403, and a dropout layer 404 connected in sequence. Among them, the first convolutional unit 401 and the second convolutional unit 402 have the same composition, and both include a two-dimensional convolutional layer (conv 2d), a batch normalization layer (BN), and an activation function layer connected in sequence. In the embodiment of the present application, the two-dimensional convolutional layer is a convolutional layer that performs convolution in both the time and frequency dimensions; the activation function used in the activation function layer is the ReLu function to increase the non-linear segmentation ability of the network and avoid gradient explosion during backpropagation; the dropout layer 404 can perform random dropout after obtaining the information output by the pooling layer 403 to prevent overfitting.
[0070] Next, based on Figure 3 the structure of the speech recognition model shown in the figure and Figure 4 the structure of the convolutional network unit shown in the figure, an explanation is given on how to obtain the main speech information through the speech recognition model.
[0071] Figure 5 shows a schematic flowchart of obtaining the main speech information through the speech recognition model, as Figure 5 shown, in step S501, the convolutional network module is used to perform segmented feature extraction on the spectrum information to obtain multiple spectrum feature maps; in step S502, the feature enhancement network module is used to perform downsampling and then upsampling on each of the spectrum feature maps and backpropagate in the reverse direction to obtain spectrum enhancement feature maps corresponding to each of the spectrum feature maps; in step S503, the long short-term memory network module is used to fuse the deep semantics and shallow time information in each of the spectrum enhancement feature maps to obtain fused feature information; in step S504, the classification prediction module is used to predict the main speech in the fused feature information to obtain the main speech information.
[0072] It should be noted that during the training process of the speech recognition model, it is divided into two parts: local training and global training. The convolutional network module 301 and the feature enhancement network module 302 are trained simultaneously to obtain optimized parameters, and the long short-term memory network module 303 and the classification prediction module 304 are trained simultaneously to obtain optimized parameters. When training the convolutional network module 301 and the feature enhancement network module 302, partial speech samples intercepted from the speech samples are used for training. For example, 60s of speech samples are used as training data. When training the long short-term memory network module 303 and the classification prediction module 304, the durations of all speech samples are aligned with the longest speech sample by padding with zeros, and then the speech samples are used for training. Therefore, when using the speech recognition model to process speech information, the convolutional network module 301 can only perform segmented feature extraction on the speech information. The feature enhancement network module 302 performs downsampling and then upsampling on the spectral feature maps obtained after segmented feature extraction and backpropagates them in reverse to achieve feature enhancement. The long short-term memory network module 303 fuses the deep semantics and shallow time information in all the spectrally enhanced feature maps to obtain fusion feature information corresponding to the speech information, and the classification prediction module 304 predicts the main speech based on the fusion feature information to obtain the main speech information.
[0073] Furthermore, in step S502, the processing of the spectral feature map by the feature enhancement network module 302 is divided into two parts. The first part is to perform downsampling on the spectral feature map through the first convolutional network unit 302-1 and the second convolutional network unit 302-2. The second part is to perform upsampling on the features obtained after downsampling and backpropagate them in reverse. During the backpropagation process, it is also necessary to splice the feature map generated by upsampling with the feature map of the same size during downsampling, so that the finally obtained spectrally enhanced feature map contains both the deep semantics obtained by downsampling and the shallow information obtained by upsampling.
[0074] Figure 6 shows a schematic flow diagram for obtaining the spectrally enhanced feature map, such as Figure 6As shown, in step S601, the first convolutional network unit downsamples the spectral feature map to obtain a first feature map, and the second convolutional network unit downsamples the first feature map to obtain a second feature map; in step S602, the second feature map is upsampled to obtain a third feature map, and at the same time, a 1×1 convolutional kernel is used to perform a convolutional operation on the first feature map, and the third feature map and the first feature map after convolutional processing are concatenated to obtain a fourth feature map; in step S603, the fourth feature map is upsampled to obtain a fifth feature map, and at the same time, a 1×1 convolutional kernel is used to perform a convolutional operation on the spectral feature map, and the fifth feature map and the spectral feature map after convolutional processing are concatenated to obtain the spectral enhancement feature map; where the step size corresponding to the upsampling is the same as the step size corresponding to the downsampling.
[0075] In an embodiment of the present application, after the main voice in the fusion feature information is predicted by the classification prediction module 304, main voice information can be obtained. The main voice information includes a main voice probability curve, and each point on the main voice probability curve is the probability of the existence of the main voice at the corresponding time point. Based on the main voice probability curve, the start and end time points corresponding to the main voice in the voice information can be obtained, realizing the positioning of the main voice in the voice information.
[0076] In step S230, according to the local extreme points in the main voice probability curve, the start and end time points corresponding to the main voice in the voice information are determined.
[0077] In an embodiment of the present application, after obtaining the main voice information, a main voice probability curve can be formed according to the probability of the existence of the main voice corresponding to each time point, and based on the main voice probability curve, the start and end time points corresponding to the main voice in the voice information are determined. When determining the start and end time points corresponding to the main voice according to the main voice probability curve, first, it can be clearly understood that the curve between any two adjacent wave valleys on the main voice probability curve corresponds to a main voice interval. For example, when the voice information is a song, the curve between two adjacent wave valleys corresponds to a singing interval. Therefore, the voice probability curve can be divided into multiple main voice intervals according to any two adjacent wave valleys in the main voice probability curve; then, the local extreme points in the main voice interval can be obtained, and the time point corresponding to the maximum value point is marked as the start time point of the main voice, and the time point corresponding to the minimum value point is marked as the end time point of the main voice. Specifically, there must be a time point when the main voice starts in the rising probability curve from the first wave valley to the wave peak, and there must be a time point when the main voice ends in the descending curve from the wave peak to the second wave valley. Therefore, the start time point and the end time point of the main voice can be obtained by calculating the local extreme points of the discrete derivative.
[0078] After obtaining the start and end time points of the main voice, the time interval corresponding to the main voice can be returned in text form through an interface. For example, in a song, there are three lines of lyrics, and each line of lyrics is sung by a human voice. After determining the start time point and end time point of the human voice corresponding to each line of lyrics, the singing intervals "[00:00:30, 00:01:00], [00:01:20, 00:02:30], [00:03:30, 00:05:00]" can be returned.
[0079] In an embodiment of the present application, in order to improve the accuracy of voice localization, before processing the Mel spectrogram using a voice recognition model, a large number of voice samples are also required to train the voice recognition model to be trained, so as to obtain a stable voice recognition model.
[0080] Before training the voice recognition model to be trained, a large number of voice samples and the corresponding main voice annotation information are required, so as to train the voice recognition model to be trained according to the voice samples and the main voice annotation information, so as to obtain a voice recognition model.
[0081] In an embodiment of the present application, a batch of voice samples can be collected, and then the corresponding main voice annotation information can be automatically generated.
[0082] Figure 7 The flowchart showing the acquisition of the main voice annotation information is as Figure 7 shown. The process of obtaining the main voice annotation information includes at least steps S701-S705, which are specifically as follows:
[0083] In step S701, the voice source of the voice sample is separated to obtain a background sound waveform diagram and a main voice waveform diagram.
[0084] In an embodiment of the present application, when annotating the main voice in the voice sample, it can be annotated according to the magnitude relationship between the main voice energy and the background sound energy. In order to obtain the main voice energy and the background sound energy, it is necessary to separate the voice source of the voice sample to extract the main voice waveform diagram and the background sound waveform diagram from the voice sample, and then calculate the main voice energy and the background sound energy according to the main voice waveform diagram and the background sound waveform diagram.
[0085] In step S702, the background sound waveform diagram and the main voice waveform diagram are sliced according to a preset time interval, and the energy ratio between the main voice energy and the background sound energy corresponding to each time slice is determined.
[0086] In one embodiment of the present application, after obtaining the main speech waveform diagram and the background sound waveform diagram, the main speech waveform diagram and the background sound waveform diagram can be sliced according to a preset time interval, which can be set according to actual needs. For example, it can be 0.5 s. After slicing is completed, parameters such as the amplitude and frequency of the main speech waveform diagram and the background sound waveform diagram corresponding to each time slice can be extracted, and then the main speech energy and the background sound energy corresponding to each time slice can be calculated. Finally, the main speech energy and the background sound energy corresponding to the same time slice can be compared to obtain the energy ratio between the two.
[0087] In step S703, the speech sample is divided into a plurality of speech intervals according to the start time points of the main speeches in each sentence of the speech sample.
[0088] In one embodiment of the present application, when obtaining a speech sample, the start time points of each main speech manually marked in the speech sample are available. The speech sample can be divided into a plurality of speech intervals according to the start time points of each main speech. For example, when the speech sample is a song, an lrc file corresponding to the song needs to be collected at the same time, and the start time is marked for each sentence of lyrics, as Figure 8 shown. Further, the song can be divided into a plurality of lyric intervals according to the start time points of each sentence of lyrics.
[0089] In step S704, each speech interval in the plurality of speech intervals is respectively used as a target speech interval, the target energy ratio corresponding to the start time point of the target speech interval is obtained, and the maximum energy ratio is determined according to the target energy ratio and the lower bound of the energy ratio.
[0090] In one embodiment of the present application, for each main speech, when speaking starts, an energy ratio between the main speech energy and the background sound energy is generated. Therefore, the energy ratio corresponding to the start time point of each main speech can be used as the target energy ratio, and based on the target energy ratio, which time point in the speech interval is the end time point of the main speech can be determined. Then, the main speech interval can be determined according to the start time point and the end time point. When determining the main speech interval, a coefficient can be applied to the target energy ratio of each target speech interval, and then the maximum energy ratio is determined according to the processed target energy ratio and the lower bound of the energy ratio. Finally, the energy ratio corresponding to each time slice of the target speech interval is compared with the maximum energy ratio to determine the time interval corresponding to the main speech.
[0091] Among them, the maximum energy ratio can be described as max(ax, b), where a is a coefficient manually set and satisfies 0 < a < 1, x is the target energy ratio, and b is the lower bound of the energy ratio, which can be set according to empirical values.
[0092] In step S705, compare the energy ratios corresponding to each time slice in the target voice interval with the maximum energy ratio, determine the main voice interval according to the continuous time slices in the target voice interval whose energy ratios are greater than or equal to the maximum energy ratio, and label the main voice interval to obtain the main voice annotation information.
[0093] In an embodiment of the present application, after determining the maximum energy ratio, compare the energy ratios corresponding to each time slice in the target voice interval with the maximum energy ratio. If there are multiple consecutive time slices whose corresponding energy ratios are greater than or equal to the maximum energy ratio, the main voice interval can be determined according to these consecutive time slices, indicating that there is a main voice in this main voice interval. Furthermore, the main voice interval can be labeled to obtain the main voice annotation information. By comparing the energy ratios corresponding to each time slice in the target voice interval with the maximum energy ratio, invalid main voices in the voice sample can be filtered out, such as the situation where there is a sentence of text without any human voice reading.
[0094] Through the process as Figure 7 shown, the main voice annotation information corresponding to the obtained voice sample can be automatically labeled, and the voice sample and the main voice annotation information can be used for the training of the voice recognition model to be trained. Compared with the weakly labeled samples, the sample annotation method in the present application improves the accuracy and precision of the model.
[0095] In an embodiment of the present application, after determining the voice sample and the corresponding main voice annotation information, the voice recognition model to be trained can be trained. Similar to the structure of the Figure 3 shown voice recognition model, the voice recognition model to be trained includes a convolutional network module to be trained, a feature enhancement network module to be trained, a long short-term memory network module to be trained, and a classification prediction module to be trained. When training the model, it is divided into two parts, namely local training and global training. Among them, in local training, the parameters of the long short-term memory network module to be trained and the classification prediction module to be trained are fixed, and the convolutional network module to be trained and the feature enhancement network module to be trained are trained according to the voice sample and the main voice annotation information to obtain a converged convolutional network module and a feature enhancement network module; in global training, the parameters of the trained convolutional network module and feature enhancement network module are fixed, and the long short-term memory network module to be trained and the classification prediction module to be trained are trained according to the voice sample and the main voice annotation information to obtain a converged long short-term memory network module and a classification prediction module, and then a converged voice recognition model is obtained.
[0096] Next, the processes of local training and global training will be described in detail.
[0097] Figure 9The schematic diagram of the local training process is shown. As Figure 9 shown, in step S901, according to a preset quantity, the voice samples are divided into multiple groups, and voice segments of a preset length are randomly intercepted from each group of the voice samples; in step S902, the Mel spectrogram corresponding to the voice segment is input into the voice recognition model to be trained, and the main voice in the Mel spectrogram corresponding to the voice segment is recognized through the voice recognition model to be trained, so as to obtain main voice prediction information; in step S903, according to the main voice prediction information and the main voice annotation information, the main voice prediction error is determined, and according to the main voice prediction error, the parameters of the convolutional network module to be trained and the feature enhancement network module to be trained are optimized until the convolutional network module and the feature enhancement network module are obtained.
[0098] Among them, in step S901, when using voice samples to train the voice recognition model to be trained, all the voice samples are traversed, b voice samples are selected each time, and voice segments of a preset length are randomly intercepted from each voice sample, and then the voice recognition model to be trained is trained according to the Mel spectrogram corresponding to the intercepted voice segment. The preset length can be set according to actual needs. For example, it can be 60s. The preset length cannot exceed the duration of the shortest voice sample and cannot exceed the actual limit of video memory occupancy. Under the condition of meeting these two conditions, the longer the preset length is, the better, so that it can learn longer time dependence relationships more. The voice segments of the preset length intercepted from b voice samples can form a group of training samples, and multiple groups of training samples can be obtained by traversing all the voice samples. In step S902, the Mel spectrogram corresponding to the intercepted voice segment of the preset length is input into the voice recognition model to be trained, and each module in the voice recognition model to be trained processes the Mel spectrogram in turn to obtain main voice prediction information. After each group of training samples pass through the convolutional network module to be trained and the feature enhancement network module to be trained, a three-dimensional tensor is formed, with the dimension of (b, n, d), where b is the total number of each group of voice samples, n is the time dimension corresponding to the sample length, and d is the feature dimension of each time point. After passing through the long short-term memory network module to be trained and the classification prediction module to be trained, a binary classification prediction of whether there is a main voice can be made for each time point, and then a two-dimensional matrix X with the dimension of (b, n) is obtained. In step S903, combined with the main voice annotation information, it can be judged whether there is a main voice at each time point. If there is, it is marked with 1, and if not, it is marked with 0. In this way, a labeled two-dimensional matrix Y with the dimension of (b, n) can also be obtained. According to the two-dimensional matrices X, Y and the loss function, the main voice prediction error can be determined, and the parameters can be adjusted backward according to the main voice prediction error until the optimal parameters of the convolutional network module to be trained and the feature enhancement network module to be trained are obtained.
[0099] The loss function adopted in the embodiments of this application may be the cross-entropy loss function, and its calculation formula is shown in Formula (2):
[0100]
[0101] where i is the number of speech samples, j is the time dimension, b is the total number of each group of speech samples, n is the time dimension corresponding to the sample length, X i,j is the main speech prediction information, and Y i,j is the main speech annotation information.
[0102] Figure 10 shows a schematic flowchart of global training. As Figure 10 shown, in step S1001, obtain the maximum duration in the speech samples, align the durations of other speech samples with the maximum duration by padding with zeros, and divide the speech samples into multiple groups according to a preset quantity; in step S1002, input the Mel spectrogram corresponding to each group of speech samples into the speech recognition model to be trained that includes a trained convolutional network module and a feature enhancement module, and identify the main speech in the Mel spectrogram corresponding to the speech samples through the speech recognition model to be trained, so as to obtain the main speech prediction information; in step S1003, determine the main speech prediction error according to the main speech prediction information and the main speech annotation information, and optimize the parameters of the long short-term memory network module and the classification prediction network module according to the main speech prediction error until the long short-term memory network module and the classification prediction network module are obtained.
[0103] When performing global training on the speech recognition model to be trained, b speech samples can also be selected from the speech samples to form a set of training samples. By polling all the speech samples, multiple sets of training samples can be obtained. Different from local training, in step S1001, all segments of each speech sample are selected, with the longest one as the standard. If the length is not enough, it is aligned by padding with zeros. This is to learn the temporal dependence relationship of the entire speech sample and there is no need to intercept the speech sample. Correspondingly, to overcome the limitation of video memory, the parameters of the entire convolutional network module and feature enhancement network module need to be fixed without backpropagation and parameter update. In step S1002, each module in the speech recognition model to be trained processes the Mel spectrogram corresponding to each set of speech samples in turn to obtain the main speech prediction information. After each set of training samples passes through the convolutional network module and the feature enhancement network module, a three-dimensional tensor is formed, with dimensions (b, n’, d), where b is the total number of speech samples in each group, n’ is the time dimension corresponding to the sample length. Since the longest speech sample is selected as the standard, different from the intercepted speech samples in local training, n’ is different from n, and d is the feature dimension at each time point. After passing through the long short-term memory network module to be trained, a three-dimensional feature tensor is still obtained, with dimensions (b, n’, d). Further, after being processed by the classification prediction module to be trained, a binary classification prediction of whether there is a main speech can be made for each time point, and then a two-dimensional matrix X’ with dimensions (b, n’) is obtained. In step S1003, combined with the main speech annotation information, it can be determined whether there is a main speech at each time point. If it exists, it is marked with 1, and if it does not exist, it is marked with 0. In this way, an annotated two-dimensional matrix Y’ with dimensions (b, n) can also be obtained. According to the two-dimensional matrices X’, Y’ and the loss function, the main speech prediction error can be determined, and the parameters can be adjusted backward according to the main speech prediction error until the optimal parameters of the long short-term memory network module to be trained and the classification prediction module to be trained are obtained, and then a converged speech recognition model is obtained. The loss function used to determine the main speech prediction error can be the same as the loss function used in local training, both being the cross-entropy loss function, or different.
[0104] The technical solution of this application can be applied to the main speech localization scenarios in speech information such as song singing localization and movie dialogue localization. To make the technical solution of this application clearer, the following takes song singing localization as an example for illustration.
[0105] If the user wants to skip the prelude of a song and start listening directly from the first line of lyrics, after determining the song to be listened to, the user can trigger a singing positioning control on the display interface of the terminal device. In response to the user's touch operation, the background can process the song selected by the user to obtain the start time point and end time point of each line of lyrics in the song. Specifically, the background can obtain the corresponding spectrogram by performing a short-time Fourier transform on the song; then perform Mel filtering to obtain the corresponding Mel spectrogram; then call the trained speech recognition model to recognize the singing in the Mel spectrogram through the speech recognition model to output the probability of the presence of human singing at each time point; finally, form a singing probability curve based on the probabilities corresponding to each time point. The curve between any two adjacent wave valleys in the singing probability curve corresponds to a singing interval. Furthermore, by calculating the local extreme points of the discrete reciprocal for each singing interval, the specific start time point and end time point of each singing interval can be obtained, realizing singing positioning. After completing the singing positioning, the prelude can be skipped and playback can start directly from the first line of lyrics. Further, when the user wants to skip the interlude, it can also be realized according to the singing positioning information.
[0106] The voice positioning method in this application processes the spectrum information corresponding to the voice information by using a speech recognition model to obtain the main voice information, which includes the main voice probability curve, and determines the start and end time points of the main voice in the voice information according to the local extreme points in the main voice probability curve. On the one hand, this application can improve the accuracy and timeliness of voice positioning; on the other hand, the speech recognition model used is trained with automatically labeled voice samples, which has better performance than the model trained with weakly labeled voice samples, and avoids manual labeling, improving the labeling efficiency and accuracy; on the other hand, it can improve the user stickiness and user experience of products using voice positioning.
[0107] It should be noted that although the steps of the method in this application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution, etc.
[0108] The following introduces the device embodiments of this application, which can be used to execute the voice positioning method in the above embodiments of this application. Figure 11 The structural block diagram of the voice positioning device provided by the embodiments of this application is schematically shown. As Figure 11 shown, the voice positioning device 1100 includes: an information processing module 1110, a speech recognition module 1120, and a voice positioning module 1130. Specifically:
[0109] An information processing module 1110, configured to obtain voice information, process the voice information to obtain spectrum information corresponding to the voice information, where the voice information includes background sound and main voice; a voice recognition module 1120, configured to input the spectrum information into a voice recognition model, and recognize the main voice in the spectrum information through the voice recognition model to obtain main voice information, where the main voice information includes a main voice probability curve; a voice localization module 1130, configured to determine start and end time points corresponding to the main voice in the voice information according to local extreme points in the main voice probability curve.
[0110] In some embodiments of the present application, the spectrum information is a Mel spectrogram; based on the above technical solution, the information processing module 1110 is configured to: frame and window the voice information, and perform Fourier transform on the windowed voice information to obtain a spectrogram corresponding to the voice information; filter the spectrogram through a Mel scale filter to obtain the Mel spectrogram.
[0111] In some embodiments of the present application, the voice recognition model includes a convolutional network module, a feature enhancement network module, a long short-term memory network module, and a classification prediction module; based on the above technical solution, the voice recognition module 1120 includes: a convolutional unit, configured to perform segmented feature extraction on the spectrum information through the convolutional network module to obtain a plurality of spectrum feature maps; an enhancement unit, configured to downsample, upsample, and backpropagate each of the spectrum feature maps through the feature enhancement network module to obtain spectrum enhancement feature maps corresponding to each of the spectrum feature maps; a fusion unit, configured to fuse deep semantic and shallow time information in each of the spectrum enhancement feature maps through the long short-term memory network module to obtain fusion feature information; a prediction unit, configured to predict the main voice in the fusion feature information through the classification prediction module to obtain the main voice information.
[0112] In some embodiments of the present application, based on the above technical solution, the convolutional network module includes a plurality of convolutional network units with the same structure, and each convolutional network unit includes a first convolutional unit, a second convolutional unit, a pooling layer, and a dropout layer, and both the first convolutional unit and the second convolutional unit include a two-dimensional convolutional layer, a batch normalization layer, and an activation function layer.
[0113] In some embodiments of the present application, the feature enhancement network module includes a first convolutional network unit and a second convolutional network unit, and the structures of the first convolutional network unit and the second convolutional network unit are the same as that of the convolutional network unit; Based on the above technical solution, the enhancement unit is configured to: downsample the spectral feature map through the first convolutional network unit to obtain a first feature map, and downsample the first feature map through the second convolutional network unit to obtain a second feature map; upsample the second feature map to obtain a third feature map, and at the same time, perform a convolution operation on the first feature map using a 1×1 convolutional kernel, and splice the third feature map and the convolved first feature map to obtain a fourth feature map; upsample the fourth feature map to obtain a fifth feature map, and at the same time, perform a convolution operation on the spectral feature map using a 1×1 convolutional kernel, and splice the fifth feature map and the convolved spectral feature map to obtain the spectral enhancement feature map; wherein, the step size corresponding to the upsampling is the same as the step size corresponding to the downsampling.
[0114] In some embodiments of the present application, based on the above technical solution, the voice localization module 1130 is configured to: form a main voice probability curve according to the main voice information; divide the main voice probability curve into multiple main voice intervals according to any two adjacent wave troughs in the voice probability curve; obtain the local extreme points in each of the main voice intervals, mark the time point corresponding to the maximum value point as the starting time point of the main voice, and mark the time point corresponding to the minimum value point as the ending time point of the main voice.
[0115] In some embodiments of the present application, based on the above technical solution, the voice localization device 1100 further includes: a sample acquisition module, configured to acquire a voice sample and automatically generated main voice annotation information corresponding to the voice sample; a model training module, configured to train a voice recognition model to be trained according to the voice sample and the main voice annotation information to obtain the voice recognition model.
[0116] In some embodiments of the present application, based on the above technical solution, the sample acquisition module is configured to: separate the sound source of the voice sample to obtain a background sound waveform diagram and a main voice waveform diagram; slice the background sound waveform diagram and the main voice waveform diagram according to a preset time interval, and determine the energy ratio between the main voice energy and the background sound energy corresponding to each time slice; divide the voice sample into multiple voice intervals according to the start time points of the main voices in each sentence of the voice sample; respectively use each of the voice intervals as a target voice interval, obtain the target energy ratio corresponding to the start time point of the target voice interval, and determine the maximum energy ratio according to the target energy ratio and the lower bound of the energy ratio; compare the energy ratio corresponding to each time slice in the target voice interval with the maximum energy ratio, determine the main voice interval according to the continuous time slices in the target voice interval whose energy ratio is greater than or equal to the maximum energy ratio, and label the main voice interval to form the voice annotation information.
[0117] In some embodiments of the present application, the speech recognition model to be trained includes a convolutional network module to be trained, a feature enhancement network module to be trained, a long short-term memory network module to be trained, and a classification prediction module to be trained; based on the above technical solution, the model training module includes: a first training unit configured to fix the parameters of the long short-term memory network module to be trained and the classification prediction module to be trained, and train the convolutional network module to be trained and the feature enhancement network module to be trained according to the voice sample and the main voice annotation information to obtain a converged convolutional network module and a feature enhancement network module; a second training unit configured to fix the parameters of the convolutional network module and the feature enhancement network module, and train the long short-term memory network module to be trained and the classification prediction module to be trained according to the voice sample and the main voice annotation information to obtain a converged long short-term memory network module and a classification prediction module.
[0118] In some embodiments of the present application, based on the above technical solution, the first training unit is configured to: divide the voice sample into multiple groups according to a preset number, and randomly intercept voice segments of a preset length from each group of the voice samples; input the Mel spectrogram corresponding to the voice segment into the speech recognition model to be trained, and identify the main voice in the Mel spectrogram corresponding to the voice segment through the speech recognition model to be trained to obtain main voice prediction information; determine the main voice prediction error according to the main voice prediction information and the main voice annotation information, and optimize the parameters of the convolutional network module to be trained and the feature enhancement network module to be trained according to the main voice prediction error until the convolutional network module and the feature enhancement network module are obtained.
[0119] In some embodiments of the present application, based on the above technical solutions, the second training unit is configured to: obtain the maximum duration in the voice samples, align the durations of other voice samples with the maximum duration by padding with zeros, and divide the voice samples into multiple groups according to a preset number; input the Mel spectrograms corresponding to each group of the voice samples into a voice recognition model to be trained including a trained convolutional network module and a feature enhancement module, and recognize the main voice in the Mel spectrograms corresponding to the voice samples through the voice recognition model to be trained, so as to obtain main voice prediction information; determine a main voice prediction error according to the main voice prediction information and the main voice annotation information, and optimize the parameters of the long short-term memory network module and the classification prediction network module according to the main voice prediction error until the long short-term memory network module and the classification prediction network module are obtained.
[0120] The specific details of the voice localization device provided in each embodiment of the present application have been described in detail in the corresponding method embodiments, and will not be repeated here.
[0121] Figure 12 Schematically shows a block diagram of a computer system of an electronic device for implementing the embodiments of the present application. The electronic device can be a terminal device 110 or a server 130 as shown in Figure 1 the figure.
[0122] It should be noted that Figure 12 the computer system 1200 of the electronic device shown is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present application.
[0123] As Figure 12 shown, the computer system 1200 includes a central processing unit 1201 (Central Processing Unit, CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1202 (Read-Only Memory, ROM) or the program loaded from the storage part 1208 into the random access memory 1203 (Random Access Memory, RAM). In the random access memory 1203, various programs and data required for system operation are also stored. The central processing unit 1201, the read-only memory 1202, and the random access memory 1203 are connected to each other through a bus 1204. The input / output interface 1205 (Input / Output interface, that is, I / O interface) is also connected to the bus 1204.
[0124] In some embodiments, the following components are connected to the input / output interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a local area network card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. The drive 1210 is also connected to the input / output interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read from it can be installed into the storage section 1208 as needed.
[0125] Specifically, according to the embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit 1201, various functions defined in the system of the present application are executed.
[0126] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium can be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0128] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0129] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable an electronic device to execute the method according to the embodiments of the present application.
[0130] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.
[0131] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A voice positioning method, characterized in that, Including: Obtain voice information, process the voice information to obtain spectral information corresponding to the voice information, where the voice information includes background sound and main voice; Input the spectral information into a voice recognition model, and recognize the main voice in the spectral information through the voice recognition model to obtain main voice information, where the main voice information includes a main voice probability curve; Each point on the main voice probability curve is the probability that the main voice exists at the corresponding time point in the voice information; Determine the start and end time points corresponding to the main voice in the voice information according to the local extreme points in the main voice probability curve; Among them, the determining the start and end time points corresponding to the main voice in the voice information according to the local extreme points in the main voice probability curve includes: Divide the main voice probability curve into multiple main voice intervals according to any two adjacent wave troughs in the main voice probability curve; Obtain the local extreme points in each main voice interval, mark the time point corresponding to the maximum value point as the start time point of the main voice, and mark the time point corresponding to the minimum value point as the end time point of the main voice.
2. The method according to claim 1, characterized in that, The spectral information is a Mel spectrogram; The processing the voice information to obtain spectral information corresponding to the voice information includes: Perform frame division and windowing on the voice information, and perform Fourier transform on the windowed voice information to obtain a spectrogram corresponding to the voice information; Perform filtering processing on the spectrogram through a Mel scale filter to obtain the Mel spectrogram.
3. The method according to claim 1, wherein The voice recognition model includes a convolutional network module, a feature enhancement network module, a long short-term memory network module, and a classification prediction module; The recognizing the main voice in the spectral information through the voice recognition model to obtain main voice information includes: Perform segmented feature extraction on the spectral information through the convolutional network module to obtain multiple spectral feature maps; Perform downsampling, upsampling, and backpropagation on each spectral feature map through the feature enhancement network module to obtain a spectral enhancement feature map corresponding to each spectral feature map; Fuse the deep semantics and shallow time information in each spectral enhancement feature map through the long short-term memory network module to obtain fused feature information; Predict the main voice in the fused feature information through the classification prediction module to obtain the main voice information.
4. The method according to claim 3, characterized in that, The convolutional network module includes multiple convolutional network units with the same structure, and each convolutional network unit includes a first convolutional unit, a second convolutional unit, a pooling layer, and a dropout layer. At the same time, both the first convolutional unit and the second convolutional unit include a two-dimensional convolutional layer, a batch normalization layer, and an activation function layer.
5. The method according to claim 4, wherein The feature enhancement network module includes a first convolutional network unit and a second convolutional network unit, and the structures of the first convolutional network unit and the second convolutional network unit are the same as the structure of the convolutional network unit; Performing downsampling, then upsampling, and then backpropagating each of the spectrum feature maps through the feature enhancement network module to obtain spectrum enhancement feature maps corresponding to the spectrum feature maps, including: Performing downsampling on the spectrum feature map through the first convolutional network unit to obtain a first feature map, and performing downsampling on the first feature map through the second convolutional network unit to obtain a second feature map; Performing upsampling on the second feature map to obtain a third feature map, simultaneously performing a convolution operation on the first feature map using a 1×1 convolutional kernel, and splicing the third feature map and the first feature map after convolution processing to obtain a fourth feature map; Performing upsampling on the fourth feature map to obtain a fifth feature map, simultaneously performing a convolution operation on the spectrum feature map using a 1×1 convolutional kernel, and splicing the fifth feature map and the spectrum feature map after convolution processing to obtain the spectrum enhancement feature map; Wherein, the step size corresponding to the upsampling is the same as the step size corresponding to the downsampling.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Obtaining a speech sample and automatically generated main speech annotation information corresponding to the speech sample; Training a speech recognition model to be trained according to the speech sample and the main speech annotation information to obtain the speech recognition model.
7. The method according to claim 6, wherein The obtaining of the automatically generated main speech annotation information corresponding to the speech sample includes: Performing sound source separation on the speech sample to obtain a background sound waveform map and a main speech waveform map; Slicing the background sound waveform map and the main speech waveform map according to a preset time interval, and determining the energy ratio between the main speech energy and the background sound energy corresponding to each time slice; Dividing the speech sample into a plurality of speech intervals according to the start time points of the main speech in each sentence of the speech sample; Taking each of the speech intervals as a target speech interval, obtaining a target energy ratio corresponding to the start time point of the target speech interval, and determining a maximum energy ratio according to the target energy ratio and a lower bound of the energy ratio; Comparing the energy ratio corresponding to each time slice in the target speech interval with the maximum energy ratio, determining a main speech interval according to the continuous time slices in the target speech interval whose energy ratio is greater than or equal to the maximum energy ratio, and annotating the main speech interval to obtain the main speech annotation information.
8. The method according to claim 6, wherein The speech recognition model to be trained includes a convolutional network module to be trained, a feature enhancement network module to be trained, a long short-term memory network module to be trained, and a classification prediction module to be trained; The training of the speech recognition model to be trained according to the speech sample and the main speech annotation information to obtain the speech recognition model includes: Fixing the parameters of the long short-term memory network module to be trained and the classification prediction module to be trained, and training the convolutional network module to be trained and the feature enhancement network module to be trained according to the speech sample and the main speech annotation information to obtain a converged convolutional network module and feature enhancement network module; Fix the parameters of the convolutional network module and the feature enhancement network module, and train the to-be-trained long short-term memory network module and the to-be-trained classification prediction module according to the speech sample and the main speech annotation information, so as to obtain a converged long short-term memory network module and classification prediction module.
9. The method according to claim 8, wherein The training of the to-be-trained convolutional network module and the to-be-trained feature enhancement network module according to the speech sample and the speech annotation information to obtain a converged convolutional network module and feature enhancement network module includes: Divide the speech samples into multiple groups according to a preset quantity, and randomly intercept speech segments of a preset length from each group of the speech samples; Input the Mel spectrogram corresponding to the speech segment into the to-be-trained speech recognition model, and recognize the main speech in the Mel spectrogram corresponding to the speech segment through the to-be-trained speech recognition model to obtain main speech prediction information; Determine the main speech prediction error according to the main speech prediction information and the main speech annotation information, and optimize the parameters of the to-be-trained convolutional network module and the to-be-trained feature enhancement network module according to the main speech prediction error until the convolutional network module and the feature enhancement network module are obtained.
10. The method according to claim 8, characterized in that, The training of the to-be-trained long short-term memory network module and the to-be-trained classification prediction module according to the speech sample and the main speech annotation information to obtain a converged long short-term memory network module and classification prediction module includes: Obtain the maximum duration in the speech samples, align the durations of other speech samples with the maximum duration by padding with zeros, and divide the speech samples into multiple groups according to a preset quantity; Input the Mel spectrogram corresponding to each group of the speech samples into the to-be-trained speech recognition model including the trained convolutional network module and feature enhancement module, and recognize the main speech in the Mel spectrogram corresponding to the speech sample through the to-be-trained speech recognition model to obtain main speech prediction information; Determine the main speech prediction error according to the main speech prediction information and the main speech annotation information, and optimize the parameters of the long short-term memory network module and the classification prediction module according to the main speech prediction error until the long short-term memory network module and the classification prediction module are obtained.
11. A voice positioning device, characterized in that, Including: An information processing module configured to obtain speech information and process the speech information to obtain spectrum information corresponding to the speech information, where the speech information includes background sound and main speech; A speech recognition module configured to input the spectrum information into a speech recognition model and recognize the main speech in the spectrum information through the speech recognition model to obtain main speech information, where the main speech information includes a main speech probability curve; Each point on the main speech probability curve is the probability that the main speech exists at the corresponding time point in the speech information; A speech localization module configured to determine the start and end time points corresponding to the main speech in the speech information according to the local extreme points in the main speech probability curve; Among them, determining the start and end time points corresponding to the main voice in the voice information according to the local extreme points in the main voice probability curve includes: Dividing the main voice probability curve into a plurality of main voice intervals according to any two adjacent wave troughs in the main voice probability curve; Obtaining the local extreme points in each of the main voice intervals, marking the time point corresponding to the maximum value point as the start time point of the main voice, and marking the time point corresponding to the minimum value point as the end time point of the main voice.
12. A computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the voice localization method according to any one of claims 1 to 10 is implemented.
13. An electronic device, characterized in that, Including: A processor; And A memory for storing executable instructions of the processor; Among them, the processor is configured to execute the voice localization method according to any one of claims 1 to 10 by executing the executable instructions.
14. A computer program product, characterized in that, Including a computer program carried on a computer-readable storage medium, and when the computer program is executed by a processor, the voice localization method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Audio signal processing method and device
CN110782908A
Voice signal segmentation model training method, device and computer equipment
CN111243619A