Method, apparatus and computer device for generating time stamps
By extracting the acoustic features and phoneme sequences of the target audio, combined with the first language alignment model, the timestamps of dialect songs are automatically generated, which solves the problem of difficulty in generating timestamps of dialect songs, and achieves more accurate and efficient timestamp generation.
Patent Information
- Application Number
- CN202210310807.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-28
AI Technical Summary
It is difficult for prior art to automatically generate timestamps for dialect songs because dialects lack a unified dictionary and rules for pronunciation.
The phonetic text and its timestamp are determined by extracting acoustic features from the target audio and obtaining a phoneme sequence based on the phoneme set of the second language, and inputting a first language alignment model to determine the state of each frame of audio.
In the case where the first language does not have a phoneme set, the voice timestamp of the target audio can be automatically generated, reducing the cost of manpower labeling and improving the accuracy of the timestamp.
Smart Images

Figure CN115050351B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, computer device, and computer-readable storage medium for generating a timestamp. Background Art
[0002] Automatic timestamping of lyrics takes the song audio and its corresponding text and uses an automatic alignment algorithm to determine the start and end times of each syllable in the audio. This automatic timestamping significantly reduces the manual annotation costs of timestamping and simplifies the production process for musicians.
[0003] Common song languages include Chinese, English, and Cantonese. These languages have strictly standardized and unified phonetic symbols or phoneme dictionaries, which can be used to train an automatic alignment model and automatically generate timestamps for lyrics. However, dialect songs also constitute a significant portion of Chinese songs, such as Minnan, Sichuanese, and Henanese. These dialects differ from Mandarin and cannot be directly translated into Chinese phoneme dictionaries. There are also no pronunciation dictionaries for these dialects. Furthermore, the pronunciation of the same characters varies across dialects, with no uniform pattern. Therefore, automatically generating timestamps for lyrics in dialects is an urgent problem that needs to be solved. Summary of the Invention
[0004] The embodiments of the present application provide a method, apparatus, computer device, and computer-readable storage medium for generating a timestamp, which can automatically generate a timestamp for dialect lyrics.
[0005] In one aspect, an embodiment of the present application provides a method for generating a timestamp, the method comprising:
[0006] extracting acoustic features from a target audio, the target audio corresponding to the first language;
[0007] Obtaining a phoneme sequence corresponding to the target audio, where the phoneme sequence comes from a phoneme set of a second language;
[0008] Input the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language alignment model to obtain the state corresponding to each frame of the target audio;
[0009] The speech text corresponding to each frame in the target audio is determined based on the state corresponding to each frame of audio, and the timestamp of the speech text is determined according to the speech text.
[0010] In one aspect, an embodiment of the present application provides a timestamp generating device, the device comprising a processing unit and a determining unit:
[0011] a processing unit for extracting acoustic features from a target audio, the target audio corresponding to a first language;
[0012] The processing unit is further configured to obtain a phoneme sequence corresponding to the target audio, the phoneme sequence being from a phoneme set of a second language;
[0013] The processing unit is further configured to input the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language alignment model to obtain a state corresponding to each frame of the target audio;
[0014] The determination unit is used to determine the voice text corresponding to each frame in the target audio based on the state corresponding to each frame of audio, and determine the timestamp of the voice text according to the voice text.
[0015] On the one hand, an embodiment of the present application provides a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the above-mentioned generation of a timestamp.
[0016] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is read and executed by a processor of a computer device, the computer device performs the above-mentioned generation of a timestamp.
[0017] In one aspect, embodiments of the present application provide a computer program product, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described generation of a timestamp.
[0018] In the method proposed in the present application, acoustic features are first extracted from the target audio, and the target audio corresponds to the first language; the phoneme sequence corresponding to the target audio is obtained, and the phoneme sequence comes from the phoneme set of the second language; then, the phoneme sequence corresponding to the target audio and the acoustic features of the target audio are input into the first language alignment model to obtain the state corresponding to each frame of the target audio; finally, the speech text corresponding to each frame in the target audio is determined based on the state corresponding to each frame of audio, and the timestamp of the speech text is determined based on the speech text. In the scenario where there is no corresponding phoneme set in the first language, the present application represents the phoneme sequence corresponding to the target audio based on the phoneme set of the second language, thereby converting the phoneme sequence of the target audio into a state sequence and obtaining the speech timestamp of the target audio. Based on the method described in the present application, even when there is no phoneme set in the first language corresponding to the target audio, the speech timestamp of the target audio can be automatically generated, which not only helps the staff reduce the workload, but also makes the generated speech timestamp more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 This is a schematic diagram of the structure of a timestamp generation system provided in an embodiment of the present application;
[0021] Figure 2 This is a flowchart of a method for generating a timestamp provided in an embodiment of the present application;
[0022] Figure 3 is a schematic diagram of a phoneme sequence and state sequence provided in an embodiment of the present application;
[0023] Figure 4 This is a schematic diagram of a time frame corresponding state provided by an embodiment of the present application;
[0024] Figure 5 This is a flowchart of another method for generating a timestamp provided in an embodiment of the present application;
[0025] Figure 6 This is a flowchart of another method for generating a timestamp provided in an embodiment of the present application;
[0026] Figure 7 This is a schematic diagram of the structure of a timestamp generating device provided in an embodiment of the present application;
[0027] Figure 8 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0029] It should be noted that the terms "first" and "second" in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature designated as "first" or "second" may explicitly or implicitly include at least one such feature.
[0030] The embodiments of this application involve artificial intelligence (AI) technology. Artificial intelligence refers to theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, enabling them to have the capabilities of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary discipline covering a wide range of fields, encompassing both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.
[0031] The method proposed in this application belongs to the field of speech processing within artificial intelligence. Automatic lyric timestamping involves inputting song audio and the corresponding text content, and then using an automatic alignment algorithm to determine the start and end times of each word in the audio. Automatically generating lyric timestamps through this automatic lyric alignment algorithm can reduce manual annotation costs, thereby simplifying the audio production process for musicians.
[0032] Common song languages include Chinese, English, and Cantonese. These languages have strictly standardized and unified pronunciation phonetic symbols or phoneme dictionaries, which can be used to train an automatic alignment model to automatically generate lyrics timestamps. However, dialect songs, such as Minnan, Sichuanese, and Henanese, also constitute a significant portion of the Chinese repertoire. These dialects differ from Mandarin and cannot be directly translated into Chinese phoneme dictionaries. Furthermore, the pronunciation of the same characters varies across dialects, with no consistent pattern. Therefore, training an automatic alignment model for these dialect songs is a significant challenge.
[0033] In a specific implementation, the above-mentioned method for generating a timestamp can be executed by a computer device, which can be a terminal device or a server. The terminal device can be, for example, a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart car, etc., but is not limited thereto; the server can be, for example, an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery servers (CDNs), and big data and artificial intelligence platforms.
[0034] Alternatively, the above-mentioned method for generating a timestamp may be performed jointly by the terminal device and the server. For example, see Figure 1 As shown: the terminal device 101 can first obtain the target audio and send the target audio to the server 102. Accordingly, after receiving the target audio, the server 102 can extract the acoustic features and then obtain the phoneme sequence corresponding to the target audio; then, the server 102 inputs the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language alignment model to obtain the state corresponding to each frame of the target audio. Finally, the server 102 determines the corresponding speech text of each frame in the target audio based on the state corresponding to each frame of audio, determines the timestamp of the speech text according to the speech text, and sends the speech timestamp to the terminal device.
[0035] This application addresses the scenario where the first speech corresponding to the target audio does not have a phoneme set. Based on the phoneme set of the second language, the speech in the target audio is mapped to a phoneme sequence, thereby generating a speech timestamp of the corresponding target audio based on the phoneme sequence of the target audio. Based on the method of this application, even if the first language corresponding to the target audio does not have a phoneme set, the speech timestamp of the target audio can be automatically generated. This not only helps reduce the workload of staff, but also makes the generated speech timestamp more accurate.
[0036] It can be understood that the system architecture diagram described in the embodiment of the present application is intended to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.
[0037] Based on the above explanation, the following Figure 2The flowchart shown in FIG. 1 further illustrates the method for generating a timestamp proposed in the embodiment of the present application. In the embodiment of the present application, the method for generating a timestamp is mainly described by taking the above-mentioned computer device as an example. Figure 2 The method for generating a timestamp may specifically include steps S201 to S204:
[0038] S201. A computer device extracts acoustic features from a target audio, where the target audio corresponds to a first language.
[0039] In the embodiment of the present application, the target audio refers to audio with speech, such as songs, recordings, etc. The target audio corresponding to the first language means that the speech in the target audio is the first language. The first language may refer to a foreign language, a dialect, etc. For example, the target audio is a recording of classmate A using his hometown dialect to report his safety to his parents. The first language corresponding to the target audio is classmate A’s hometown dialect. Acoustic features refer to physical quantities that represent the acoustic characteristics of speech, and are also a general term for the acoustic performance of various sound elements. Acoustic features can be used to analyze the characteristics of the target audio. Optionally, the acoustic features can be Mel-Frequency Ceptral Coefficients (MFCC) features. The target audio can be local audio data of the computer device, or it can be audio data sent to the computer device by other devices. The embodiment of the present application does not limit this.
[0040] S202: The computer device obtains a phoneme sequence corresponding to the target audio, where the phoneme sequence comes from a phoneme set of a second language.
[0041] In the embodiments of this application, a phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzing based on the pronunciation actions in a syllable, one action constitutes one phoneme. For example, the Chinese syllable "ā" has only one phoneme, "ài" has two phonemes, "dài" has three phonemes, etc. A second language is a language different from the first language. The pronunciation of the second language is similar to that of the first language. Usually, it is a commonly used language and has a general phoneme dictionary. For example, the second language can be Mandarin, etc. The phoneme set of the second language refers to multiple phonemes used to represent the pronunciation of the second language. For example, if the second language is Mandarin, the phoneme set of the second language can be pinyin; if the second language is English, the phoneme set of the second language can be phonetic symbols. It should be noted that the phoneme set can have other forms in addition to pinyin. The embodiments of this application do not limit this. Using the phoneme set of the second language to represent the phoneme sequence corresponding to the target audio can effectively solve the problem that there is no phoneme set for the first language corresponding to the target audio. Exemplarily, the target audio is the audio of Little A saying "shoes" in Sichuan dialect. The pronunciation of "鞋" in Sichuan dialect is similar to that of "孩" in Mandarin, and the pronunciation of "子" in Sichuan dialect is the same as that of "子" in Mandarin. According to the phoneme set of Mandarin, the phoneme sequence corresponding to this target audio can be represented as "háizi". Another exemplarily, the target audio is the audio of Little A saying "吃饭" in Shanghai dialect. The pronunciation of "吃" in Shanghai dialect is similar to that of "恰" in Mandarin, and the pronunciation of "饭" in Shanghai dialect is the same as that of "饭" in Mandarin. According to the phoneme set of Mandarin, the phoneme sequence corresponding to this target audio can be represented as "qiāfàn".
[0042] In a possible implementation manner, the computer device obtains the phoneme sequence corresponding to the target audio. The specific implementation manner can be: the computer device inputs the acoustic features of the target audio into the second language recognition model to obtain the phoneme sequence corresponding to the target audio. The second language recognition model is used to convert the speech in the audio into a phoneme sequence. Exemplarily, the target audio is the audio of Little A saying "shoes" in Sichuan dialect. The first language corresponding to this target audio is Sichuan dialect. The pronunciation of "鞋子" in Sichuan dialect is similar to that of "孩子" in Mandarin. Therefore, the recognition model of the second language recognizes the phoneme sequence corresponding to the target audio as "háizi". Based on this implementation method, the staff can automatically generate the phoneme sequence of the target audio through the second language recognition model, simplifying the operation and also improving the work efficiency of the staff.
[0043] In a possible implementation, the computer device obtains the phoneme sequence corresponding to the target audio. The specific implementation can be as follows: The computer device obtains the speech text of the target audio; The computer device determines the phoneme sequence corresponding to the target audio based on the speech text of the target audio and the pronunciation dictionary of the first language. The speech text is the text corresponding to the speech in the target audio, and can specifically be the text stored locally on the computer device, or the text sent to the computer device by other devices. The pronunciation dictionary of the first language is used to represent the correspondence between the characters and phonemes of the first language, and the pronunciation dictionary of the first language can be constructed based on multiple sample audios and the speech texts corresponding to the sample audios. For the specific construction process, please refer to Figure 5 the description corresponding to the embodiment. This application does not elaborate here. Exemplarily, the text data corresponding to the target audio is "shoes", and the first language corresponding to the target audio is Sichuan dialect. According to the correspondence between the characters and phonemes in the pronunciation dictionary of the first language, the phoneme sequence corresponding to the target audio can be expressed as "háizi". Based on this implementation, the computer device only needs to obtain the phoneme sequence corresponding to the target audio according to the speech text of the target audio, which is beneficial to obtaining a more accurate phoneme sequence.
[0044] S203. The computer device inputs the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language alignment model to obtain the state corresponding to each frame of the target audio.
[0045] In the embodiments of this application, the state is the audio representation form with the smallest granularity. One phoneme can correspond to one or more states. In the audio data, one frame of audio feature corresponds to one state. Exemplarily, as Figure 3 shown, the text corresponding to the speech in the target audio is "I and you", and the corresponding phoneme sequence can be expressed as "wǒhénǐ". The total length of the target audio corresponds to 10 frames. "w" corresponds to two frames, and "ǒ" also corresponds to two frames. Among the two frames corresponding to "w", the states are different. The state corresponding to the first frame is state 1, and the state corresponding to the second frame is state 2.
[0046] In a possible implementation, the computer device inputs the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language acoustic model to obtain the state corresponding to each frame of the target audio. The specific implementation can be as follows: The computer device calls the language model in the first language alignment model to process the phoneme sequence to obtain the state sequence corresponding to the target audio; The computer device inputs the state sequence corresponding to the target audio and the acoustic features of the target audio into the acoustic model of the first language alignment model to determine the probabilities of multiple states corresponding to each frame of the audio data; The computer device obtains the state corresponding to each frame of the target audio based on the preset state transition probability and the probabilities of multiple states corresponding to each frame of the target audio.
[0047] Among them, the first language alignment model includes two models, namely a language model and an acoustic model. The language model in the first language alignment model is used to convert a phoneme sequence into a state sequence. For example, the language model can be a Hidden Markov Model (HMM), or the model can also be other models, such as an end-to-end singing recognition model, which is not limited in this embodiment of the present application. Among them, the acoustic model in the first language alignment model is used to determine the probability of each frame of audio corresponding to the state of the target audio according to the acoustic features and state sequence of the target audio. The acoustic model can be a Gaussian Mixture Model (GMM) or a Deep Neural Networks (DNN) model. Or the acoustic model can also be a model obtained by training a sample text based on the first language and suitable for a second language acoustic model. The second language acoustic model is used to determine the probability of each frame of audio corresponding to the state in the second language audio data according to the acoustic features and state sequence of the second language audio data. Its specific implementation method can be found in the subsequent Figure 6 Description of the corresponding embodiment.
[0048] It should also be noted that the preset state transition probability refers to the probability that each state in the target audio will transition to another state or remain in the original state in the next frame. For example, when it is known that the first frame is state 1, the probability that the next frame state is still state 1 is 20%, and the probability of being state 2 is 80%.
[0049] For example, Figure 4 As shown, the target audio is a recording of "You and Me," with a total length of 10 frames. Based on the previous steps, the computer device can determine the state sequence corresponding to the target audio. Based on the state sequence, the 10 frames of the target audio can correspond to five states: State 1, State 2, State 3, State 4, and State 5. Using the acoustic model, it can be determined that in the first frame, the probability of State 1 is 50%, the probability of State 2 is 20%, and so on. In the second frame, the probability of State 1 is 10%, the probability of State 2 is 0%, and so on. Based on preset state transition probabilities, it can be determined that if the first frame is known to be in State 1, the probability of the second frame being in State 1 is 60%; and if the first frame is known to be in State 1, the probability of the second frame being in State 2 is 20%. A decoding graph can be constructed based on the probabilities of each frame of the target audio corresponding to multiple states and the state transition probabilities corresponding to each state. Using the Viterbi decoding algorithm, the decoding graph is searched for the globally optimal decoding path, thereby determining the state corresponding to each frame of the target audio. This implementation enables more accurate speech timestamp generation.
[0050] Optionally, the method further includes: the computer device obtains a first language association model based on a plurality of sample texts trained, the sample texts corresponding to the first language; the computer device determines a preset state transition probability based on the pronunciation dictionary of the first language and the first language association model. The first language association model may be an n-gram model. Since the language habits of the first language are different from those of the second language, the probability of each phrase appearing in the sample text corresponding to the first language and the sample text corresponding to the second language is also different. The n-gram model can count the probability of each n phrase appearing. For example, when n is 3, the frequency of occurrence of a triple phrase such as "you and me" in the lyrics text is calculated to obtain the probability of the triple appearing. The larger N is, the richer the context information of the word is. The first language association model is used to determine the preset state transition probability corresponding to the first language, so that the state corresponding to each frame of the target audio is more accurate and the generated speech timestamp is more accurate.
[0051] S204. The computer device determines the speech text corresponding to each frame in the target audio based on the state corresponding to each frame of audio, and determines the timestamp of the speech text according to the speech text.
[0052] In an embodiment of the present application, after determining the state corresponding to each frame of audio, the computer device can determine the phoneme corresponding to each frame of audio based on the correspondence between the state and the phoneme, and then determine the phonetic text corresponding to each frame based on the correspondence between the phoneme and the phonetic text. In this way, the time period corresponding to each word in the target audio can be determined based on the phonetic text corresponding to each frame, and the voice timestamp of the target audio can be obtained.
[0053] Based on the method described in the embodiments of this application, in scenarios where there is no corresponding phoneme set in the first language, the phoneme sequence corresponding to the target audio can be expressed based on the phoneme set of the second language, thereby converting the phoneme sequence of the target audio into a state sequence to obtain the speech timestamp of the target audio. Based on the method of this application, even if there is no phoneme set in the first language corresponding to the target audio, the speech timestamp of the target audio can be automatically generated, which not only helps reduce the workload of staff, but also makes the generated speech timestamp more accurate.
[0054] See Figure 5 , is a flow chart of another method for generating a timestamp disclosed in an embodiment of the present invention. This method for generating a timestamp can be executed by a computer device, specifically server 102 in a timestamp generation system. This embodiment primarily illustrates the process of constructing a pronunciation dictionary for a first language. The text processing method specifically includes steps S501 to S506. Among them:
[0055] S501: A computer device extracts acoustic features from a first sample audio, where the first sample audio corresponds to a first language.
[0056] In the embodiments of the present application, the number of the first sample audios can be one or more, which is not limited in the embodiments of the present application. The first sample audios correspond to the first language and are used to train the pronunciation dictionary of the first language.
[0057] S502. The computer device inputs the acoustic features of the first sample audio into the second language recognition model to obtain the phoneme sequence corresponding to the first sample audio.
[0058] Among them, the second language recognition model is the same as the second language recognition model described in step S202, and is used to convert the speech in the audio into a phoneme sequence, which is not elaborated in the embodiments of the present application.
[0059] In a possible implementation manner, the second language recognition model can be a hybrid model composed of a DNN model and an HMM model. The specific implementation manner is as follows: The computer device inputs the acoustic features of the first sample audio into the DNN model in the second language recognition model to obtain the probabilities of multiple states corresponding to each frame of the first sample audio; the computer device inputs the probabilities of multiple states corresponding to each frame of the first sample audio and the preset state transition probability into the HMM model in the second language recognition model to obtain the state sequence corresponding to the first sample audio. The computer device can determine the phoneme sequence corresponding to the first sample audio according to the correspondence between the states of the second language and the phonemes of the second language. Based on this implementation manner, when there is no corresponding phoneme set and pronunciation dictionary for the first language, the phoneme sequence of the first sample audio can be represented by the second language recognition model based on the phoneme set of the second language, which is beneficial to constructing a more accurate pronunciation dictionary of the first language subsequently.
[0060] S503. The computer device constructs the pronunciation dictionary of the first language based on the phoneme sequence corresponding to the first sample audio and the speech text of the first sample audio.
[0061] In the embodiments of the present application, the computer device can construct the pronunciation dictionary of the first language according to the correspondence between the phoneme sequence corresponding to the first audio and the speech text of the first sample audio. Exemplarily, the phoneme sequence corresponding to the first sample audio can be expressed as "wǒhénǐ", and the speech text of the first sample is "我和你", then it can be recorded in the pronunciation dictionary of the first language that the phoneme sequence corresponding to "我和你" is "wǒhénǐ".
[0062] S504. The computer device extracts the acoustic features from the target audio and obtains the phoneme sequence corresponding to the target audio. The target audio corresponds to the first language, and the phoneme sequence comes from the phoneme set of the second language. The target audio corresponds to the first language.
[0063] S505: The computer device inputs the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language alignment model to obtain the state corresponding to each frame of the target audio.
[0064] S506. The computer device determines the speech text corresponding to each frame in the target audio based on the state corresponding to each frame of audio, and determines the timestamp of the speech text according to the speech text.
[0065] Among them, the implementation method corresponding to steps S504 to S506 is the same as that of steps S201 to S205, and is not described in detail in this embodiment of the present application.
[0066] See Figure 6 , is a flowchart of another method for generating a timestamp disclosed in an embodiment of the present invention. The timestamp generating method can be executed by a computer device, which can specifically be server 102 in the timestamp generating system. This embodiment is mainly used to illustrate the process of training an acoustic model. The text processing method can specifically include steps S601 to S607. Among them:
[0067] S601: A computer device extracts acoustic features from a second sample audio, where the second sample audio corresponds to a first language.
[0068] In the embodiment of the present application, the number of the second sample audios may be one or more, which is not limited in the embodiment of the present application.
[0069] S602: The computer device determines a phoneme sequence corresponding to the second sample audio according to a pronunciation dictionary of the first language and the speech text of the second sample audio.
[0070] In the embodiment of the present application, the pronunciation dictionary of the first language is constructed as follows Figure 5 As shown, the pronunciation dictionary of the first language is used to represent the correspondence between the characters and phonemes of the first language. Therefore, the computer device can determine the phonemes corresponding to each character based on the speech text of the second sample audio, thereby determining the phoneme sequence corresponding to the second sample audio.
[0071] S603: The computer device calls the language model in the first language alignment model to convert the phoneme sequence into a first state sequence, and inputs the acoustic features of the second sample audio into the acoustic model of the first language alignment model to obtain a second state sequence.
[0072] In the embodiment of the present application, the language model in the first language alignment model is used to convert the phoneme sequence into a state sequence. For details, please refer to the description in step S201, which is not described in detail in this embodiment of the present application. The acoustic model of the first language alignment model is used to convert the acoustic features into a state sequence. For details, please refer to the description in step S204, which is not described in detail in this embodiment of the present application. Optionally, before training, since there is no acoustic model suitable for the first language, the acoustic model can be a trained acoustic model suitable for the second language.
[0073] S604: The computer device trains an acoustic model of the first language alignment model based on the first state sequence and the second state sequence.
[0074] In this embodiment of the present application, since the first state sequence is a speech text derived from the second sample audio, the first state sequence can be considered as an initialization label. The parameters of the acoustic model of the first language alignment model are then continuously trained and adjusted so that the second state sequence approximates the first state sequence, thereby obtaining a final acoustic model suitable for the first language. Based on this implementation, the resulting acoustic model can be made more suitable for the first language, thereby making the subsequently generated timestamps more accurate.
[0075] S605: The computer device extracts acoustic features from the target audio and obtains a phoneme sequence corresponding to the target audio, where the target audio corresponds to the first language and the phoneme sequence comes from a phoneme set of the second language. The target audio corresponds to the first language.
[0076] S606: The computer device inputs the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language alignment model to obtain the state corresponding to each frame of the target audio.
[0077] S607: The computer device determines the speech text corresponding to each frame in the target audio based on the state corresponding to each frame of audio, and determines the timestamp of the speech text according to the speech text.
[0078] Among them, the implementation method corresponding to steps S605 to S607 is the same as that of steps S201 to S205, and is not described in detail in this embodiment of the present application.
[0079] Based on the above-mentioned method for generating a timestamp, the present invention provides a device for generating a timestamp. Figure 7 , is a schematic diagram of the structure of a timestamp generating device provided in an embodiment of the present application. The timestamp generating device 700 can run the following units:
[0080] A processing unit 701 is configured to extract acoustic features from a target audio, where the target audio corresponds to a first language;
[0081] The processing unit 701 is configured to obtain a phoneme sequence corresponding to the target audio, where the phoneme sequence is from a phoneme set of a second language;
[0082] The processing unit 701 is further configured to input the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language alignment model to obtain a state corresponding to each frame of the target audio;
[0083] The determining unit 702 is configured to determine the speech text corresponding to each frame in the target audio based on the state corresponding to each frame of audio, and determine the timestamp of the speech text according to the speech text.
[0084] In one embodiment, when the processing unit 701 inputs the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language acoustic model to obtain the state corresponding to each frame of audio of the target audio, it can be specifically used to: call the language model in the first language alignment model to process the phoneme sequence to obtain the state sequence corresponding to the target audio; input the state sequence corresponding to the target audio and the acoustic features of the target audio into the acoustic model of the first language alignment model to determine the probabilities of multiple states corresponding to each frame of audio in the audio data; and obtain the state corresponding to each frame of audio of the target audio based on the preset state transition probability and the probabilities of multiple states corresponding to each frame of audio in the target audio.
[0085] In one embodiment, the processing unit 701 is further configured to: obtain a first language association model based on a plurality of sample texts trained, where the sample texts correspond to the first language; and determine a preset state transition probability based on a pronunciation dictionary of the first language and the first language association model.
[0086] In one embodiment, when the processing unit 701 determines the speech text corresponding to each frame in the target audio based on the state corresponding to each frame of audio, it can be specifically used to determine the speech text corresponding to each frame in the target audio based on the pronunciation dictionary of the first language and the state corresponding to each frame of audio.
[0087] In one embodiment, the processing unit 701 is further used to: extract acoustic features from a first sample audio, where the first sample audio corresponds to a first language; input the acoustic features of the first sample audio into a second language recognition model to obtain a phoneme sequence corresponding to the first sample audio; and construct a pronunciation dictionary for the first language based on the phoneme sequence corresponding to the first sample audio and the speech text of the first sample audio.
[0088] In one embodiment, the processing unit 701 is further used to: extract acoustic features from a second sample audio, where the second sample audio corresponds to a first language; determine a phoneme sequence corresponding to the second sample audio based on a pronunciation dictionary of the first language and a phonetic text of the second sample audio; call a language model in the first language alignment model to convert the phoneme sequence into a first state sequence; input the acoustic features of the second sample audio into the acoustic model of the first language alignment model to obtain a second state sequence; and train the acoustic model of the first language alignment model based on the first state sequence and the second state sequence.
[0089] In one embodiment, when obtaining the phoneme sequence corresponding to the target audio, the processing unit 701 is specifically configured to: input the acoustic features of the target audio into the second language recognition model to obtain the phoneme sequence corresponding to the target audio.
[0090] In one embodiment, when obtaining the phoneme sequence corresponding to the target audio, the processing unit 701 is specifically configured to: obtain the speech text of the target audio; and determine, based on the speech text of the target audio, the phoneme sequence corresponding to the target audio in the pronunciation dictionary of the first language.
[0091] According to another embodiment of the present application, Figure 7 The various units in the timestamp generation device shown can be individually or all combined into one or several other units to constitute, or one (or some) of the units can be further split into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the timestamp generation device may also include other units. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0092] According to another embodiment of the present application, the program can be executed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements. Figure 2 A computer program (including program code) for each step involved in the corresponding method shown in Figure 7 The device for generating a timestamp as shown in the embodiment of the present application is used to implement the method for generating a timestamp. The computer program can be recorded on a computer-readable recording medium, for example, and loaded into the computing device through the computer-readable recording medium and run therein.
[0093] In an embodiment of the present application, acoustic features are extracted from target audio, and the target audio corresponds to a first language; a phoneme sequence corresponding to the target audio is obtained, and the phoneme sequence comes from a phoneme set of a second language; the phoneme sequence corresponding to the target audio and the acoustic features of the target audio are input into a first language alignment model to obtain a state corresponding to each frame of the target audio; based on the state corresponding to each frame of the audio, the corresponding speech text of each frame in the target audio is determined, and the timestamp of the speech text is determined based on the speech text. In a scenario where there is no corresponding phoneme set in the first language, the present application represents the phoneme sequence corresponding to the target audio based on the phoneme set of the second language, thereby obtaining a speech timestamp of the target audio. Based on the method described in the present application, even when there is no phoneme set in the first language corresponding to the target audio, a speech timestamp of the target audio can be automatically generated, which not only helps reduce the workload of the staff, but also makes the generated speech timestamp more accurate.
[0094] Based on the description of the above method embodiment and apparatus embodiment, the present application embodiment also provides a computer device. Figure 8 , the computer device 800 includes at least a processor 801, a communication interface 802 and a computer storage medium 803. The processor 801, the communication interface 802 and the computer storage medium 803 can be connected via a bus or other means. The computer storage medium 803 can be stored in the memory 804 of the computer device 800, and the computer storage medium 803 is used to store a computer program, and the computer program includes program instructions. The processor 801 is used to execute the program instructions stored in the computer storage medium 803. The processor 801 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.
[0095] In one embodiment, the processor 801 described in the embodiment of the present application can be used to implement a method for generating a timestamp, specifically including: extracting acoustic features from the target audio, the target audio corresponding to the first language; obtaining a phoneme sequence corresponding to the target audio, the phoneme sequence coming from a phoneme set of the second language; inputting the phoneme sequence corresponding to the target audio into a language model to obtain a state sequence corresponding to the target audio; inputting the state sequence corresponding to the target audio and the acoustic features of the target audio into the acoustic model to obtain a state corresponding to each frame of audio of the target audio; determining a speech timestamp corresponding to the spoken text of each frame in the target audio based on the state corresponding to each frame of audio, and so on.
[0096] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer storage medium provides a storage space that stores the operating system of the computer device. In addition, one or more instructions suitable for being loaded and executed by the processor 801 are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage; optionally, it can also be at least one computer storage medium located away from the aforementioned processor.
[0097] In one embodiment, the processor may load and execute one or more instructions stored in the computer storage medium to implement the above-mentioned Figure 2 、 Figure 5 or Figure 6 The corresponding steps of the method in the embodiment of the method for generating a timestamp are shown; in a specific implementation, one or more instructions in the computer storage medium are loaded by the processor 801 and execute the following steps:
[0098] extracting acoustic features from a target audio, the target audio corresponding to the first language;
[0099] Obtaining a phoneme sequence corresponding to the target audio, where the phoneme sequence comes from a phoneme set of a second language;
[0100] Input the phoneme sequence corresponding to the target audio and the acoustic features of the target audio into the first language alignment model to obtain the state corresponding to each frame of the target audio;
[0101] The speech text corresponding to each frame in the target audio is determined based on the state corresponding to each frame of audio, and the timestamp of the speech text is determined according to the speech text.
[0102] In one embodiment, when the phoneme sequence corresponding to the target audio and the acoustic features of the target audio are input into the first language acoustic model to obtain the state corresponding to each frame of audio of the target audio, the one or more instructions can be loaded by the processor and specifically executed: calling the language model in the first language alignment model to process the phoneme sequence to obtain the state sequence corresponding to the target audio; inputting the state sequence corresponding to the target audio and the acoustic features of the target audio into the acoustic model of the first language alignment model to determine the probabilities of multiple states corresponding to each frame of audio in the audio data; obtaining the state corresponding to each frame of audio of the target audio based on the preset state transition probability and the probabilities of multiple states corresponding to each frame of audio in the target audio.
[0103] In one embodiment, the one or more instructions may also be loaded by the processor and execute the following steps: obtaining a first language association model based on training of multiple sample texts, where the sample texts correspond to the first language; and determining a preset state transition probability based on a pronunciation dictionary of the first language and the first language association model.
[0104] In one embodiment, when determining the speech text corresponding to each frame in the target audio based on the state corresponding to each frame of audio, the one or more instructions can be loaded and specifically executed by the processor: determining the speech text corresponding to each frame in the target audio based on the pronunciation dictionary of the first language and the state corresponding to each frame of audio.
[0105] In one embodiment, the one or more instructions may also be loaded by the processor and execute the following steps: extracting acoustic features from a first sample audio, where the first sample audio corresponds to a first language; inputting the acoustic features of the first sample audio into a second language recognition model to obtain a phoneme sequence corresponding to the first sample audio; and constructing a pronunciation dictionary for the first language based on the phoneme sequence corresponding to the first sample audio and the speech text of the first sample audio.
[0106] In one embodiment, the one or more instructions can also be loaded by the processor and execute the following steps: extracting acoustic features from a second sample audio, the second sample audio corresponding to the first language; determining a phoneme sequence corresponding to the second sample audio based on a pronunciation dictionary of the first language and a phonetic text of the second sample audio; calling a language model in the first language alignment model to convert the phoneme sequence into a first state sequence; inputting the acoustic features of the second sample audio into the acoustic model of the first language alignment model to obtain a second state sequence; and training the acoustic model of the first language alignment model based on the first state sequence and the second state sequence.
[0107] In one embodiment, when obtaining the phoneme sequence corresponding to the target audio, the one or more instructions can be loaded and specifically executed by the processor: inputting the acoustic features of the target audio into the second language recognition model to obtain the phoneme sequence corresponding to the target audio.
[0108] In one embodiment, when obtaining the phoneme sequence corresponding to the target audio, the one or more instructions can be loaded by the processor and specifically executed: obtaining the speech text of the target audio; determining the phoneme sequence corresponding to the target audio based on the speech text of the target audio and the pronunciation dictionary of the first language.
[0109] In an embodiment of the present application, acoustic features are extracted from target audio, and the target audio corresponds to a first language; a phoneme sequence corresponding to the target audio is obtained, and the phoneme sequence comes from a phoneme set of a second language; the phoneme sequence corresponding to the target audio and the acoustic features of the target audio are input into a first language alignment model to obtain a state corresponding to each frame of the target audio; based on the state corresponding to each frame of the audio, the corresponding speech text of each frame in the target audio is determined, and the timestamp of the speech text is determined based on the speech text. In a scenario where there is no corresponding phoneme set in the first language, the present application represents the phoneme sequence corresponding to the target audio based on the phoneme set of the second language, thereby obtaining a speech timestamp of the target audio. Based on the method described in the present application, even when there is no phoneme set in the first language corresponding to the target audio, a speech timestamp of the target audio can be automatically generated, which not only helps reduce the workload of the staff, but also makes the generated speech timestamp more accurate.
[0110] It can be understood that the system architecture diagram described in the embodiment of the present application is intended to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.
Claims
1. A method for generating a timestamp, characterized in that, the method includes: extracting acoustic features from a target audio, where the target audio corresponds to a first language; obtaining a phoneme sequence corresponding to the target audio, where the phoneme sequence comes from a phoneme set of a second language, and the second language is different from the first language; invoking a language model in a first language alignment model to process the phoneme sequence to obtain a state sequence corresponding to the target audio; inputting the state sequence corresponding to the target audio and the acoustic features of the target audio into an acoustic model of the first language alignment model to determine the probabilities of multiple states corresponding to each frame of audio in the target audio; obtaining the state corresponding to each frame of audio in the target audio based on a preset state transition probability and the probabilities of multiple states corresponding to each frame of audio; determining the speech text corresponding to each frame in the target audio based on the state corresponding to each frame of audio, and determining the timestamp of the speech text according to the speech text.
2. The method according to claim 1, characterized in that, the method further includes: extracting acoustic features from a first sample audio, where the first sample audio corresponds to the first language; inputting the acoustic features of the first sample audio into a second language recognition model to obtain a phoneme sequence corresponding to the first sample audio; constructing a pronunciation dictionary of the first language based on the phoneme sequence corresponding to the first sample audio and the speech text of the first sample audio, where the pronunciation dictionary of the first language is used to determine the phoneme sequence corresponding to the target audio.
3. The method according to claim 2, characterized in that, the method further includes: training a first language association model based on multiple sample texts, where the sample texts correspond to the first language; determining the preset state transition probability based on the pronunciation dictionary of the first language and the first language association model.
4. The method according to claim 2, characterized in that, the determining the speech text corresponding to each frame in the target audio based on the state corresponding to each frame of audio includes: determining the speech text corresponding to each frame in the target audio based on the pronunciation dictionary of the first language and the state corresponding to each frame of audio.
5. The method according to claim 1, characterized in that, the method further includes: extracting acoustic features from a second sample audio, where the second sample audio corresponds to the first language; determining the phoneme sequence corresponding to the second sample audio according to the pronunciation dictionary of the first language and the speech text of the second sample audio; invoking the language model in the first language alignment model to convert the phoneme sequence into a first state sequence; inputting the acoustic features of the second sample audio into the acoustic model of the first language alignment model to obtain a second state sequence; training the acoustic model of the first language alignment model based on the first state sequence and the second state sequence.
6. The method according to claim 2, characterized in that, the obtaining the phoneme sequence corresponding to the target audio includes: inputting the acoustic features of the target audio into the second language recognition model to obtain the phoneme sequence corresponding to the target audio.
7. The method according to claim 2, wherein, the obtaining of the phoneme sequence corresponding to the target audio includes: obtaining the speech text of the target audio; determining the phoneme sequence corresponding to the target audio based on the speech text of the target audio and the pronunciation dictionary of the first language.
8. A computer device, wherein, the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method for generating a timestamp according to any one of claims 1 to 7.
9. A computer-readable storage medium, wherein, the computer-readable storage medium stores one or more computer programs, and the one or more computer programs are adapted to be loaded and executed by a processor to execute the method for generating a timestamp according to any one of claims 1 to 7.
10. A computer program product, comprising computer instructions, wherein, the computer instructions are stored in a computer-readable storage medium and are adapted to be read and executed by a processor of a computer device, so that the computer device executes the method for generating a timestamp according to any one of claims 1 - 7.
Citation Information
Patent Citations
Method for determining lyric timestamp information and training method of acoustic model
CN112735429A
Aligning a transcript to audio data
US8131545B1