A hybrid statement alignment method, electronic device, and storage medium

By using neural network models to identify and process normative and non-canonical word blocks in mixed texts, combined with forced alignment and speech recognition technology, the problem of poor alignment of text in multilingual Chinese is solved, and efficient and accurate speech alignment is achieved.

CN112800220BActive Publication Date: 2025-06-10SHANGHAI GAUDIAN INTELLIGENT TECHNOLOGY GROUP CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110076849.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-20
Publication Date
2025-06-10
Estimated Expiration
2041-01-20

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently process Chinese text containing multiple languages ​​or symbols, resulting in unsatisfactory voice alignment and affecting user interaction experience.

Method used

The pre-trained neural network model is used to identify normative and non-normalistic blocks in mixed texts, and combine forced alignment algorithms and speech recognition technology to obtain audio text time links.

Benefits of technology

Accurate phoneme sequence alignment for multilingual Chinese texts is achieved, reducing the computational volume, improving the speed and performance of speech alignment, and improving the alignment effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112800220B_ABST
    Figure CN112800220B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for aligning hybrid sentences, an electronic device, and a storage medium, including: obtaining hybrid text and target audio; the hybrid text includes canonical chunks and non-canonical chunks; there is a temporal correspondence between the hybrid text and the target audio; using a pre-trained neural network model to identify the non-canonical chunks and canonical chunks in the hybrid text to generate a classification result; according to the classification result, align the hybrid text and the target audio to obtain an audio-text time link. The technical effect of the present invention: Targetedly identify foreign languages or other non-canonical chunks in the text, only need to perform local speech recognition, greatly reducing the computational amount and improving the overall speed and performance of speech alignment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to natural language processing, and particularly to a method for aligning hybrid sentences, an electronic device, and a storage medium. Background Art

[0002] In order to make the lip movements of a virtual human when speaking realistic, vivid, and accurate, it is necessary to align the speech with the content of the speech. However, the Chinese text answered by the virtual human sometimes involves a large amount of content mixed with multiple languages or symbols, such as foreign language characters, special symbols, etc. The current speech alignment technology has an unsatisfactory processing effect on these unknown non-standard language chunks, which greatly affects the user's interaction experience. Currently, there are mainly two methods for aligning Chinese texts with mixed multiple languages: dictionary-based alignment and sentence-level speech recognition text alignment. Both of these methods have obvious deficiencies and cannot meet the alignment requirements of high accuracy, simplicity, and speed. Specifically as follows:

[0003] Dictionary-based alignment requires building dictionaries for all languages in the world and mapping foreign language characters to phonemes. When aligning, after language discrimination, the models of each language are used to uniformly map all characters into phonemes for forced alignment. Although this method has high accuracy, the workload is huge - it is necessary to build pronunciation dictionaries for all words in all languages in the world. In addition, the pronunciations of special symbols, abbreviations, and numbers cannot be enumerated.

[0004] Sentence-level speech recognition text alignment, according to the audio and the reference text pair, uses a speech recognition engine to decode the entire audio data to obtain the speech recognition text. The dynamic programming algorithm is used for maximum feature matching to achieve sentence-level alignment. However, this method requires large acoustic models and language models, and the computational complexity is very high, reducing the alignment speed. At the same time, the speech recognition result is affected by the recognition ability of the engine and the complexity of the Chinese mixed text, and the alignment effect is poor, and accurate time information cannot be obtained. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method for aligning hybrid sentences, an electronic device, and a storage medium. The specific technical solutions are as follows:

[0006] A method for aligning hybrid sentences includes:

[0007] Obtaining a hybrid text and a target audio; the hybrid text includes a standard language chunk and a non-standard language chunk; the hybrid text and the target audio have a temporal correspondence relationship;

[0008] Using a pre-trained neural network model to identify the non-standard language chunks and standard language chunks in the hybrid text to generate a classification result;

[0009] According to the classification result, align the mixed text and the target audio to obtain the audio-text time link.

[0010] Preferably, it further includes: training the neural network model, specifically including:

[0011] Obtain a training text set; the training text set includes pure Chinese texts and mixed texts;

[0012] Mark the non-standard language chunks in the training set as unknown phonemes;

[0013] Train the neural network model according to the training text set.

[0014] Preferably, the aligning the mixed text and the target audio includes:

[0015] Use a standard language chunk pronunciation dictionary to convert the standard language chunks in the mixed text into a standard phoneme sequence.

[0016] More preferably, the aligning the mixed text and the target audio includes:

[0017] According to the standard phoneme sequence, use the forced alignment algorithm to mark the start and end times of all the standard language chunks and all the non-standard language chunks on the target audio;

[0018] According to the start and end times of each non-standard language chunk on the target audio, perform standard sentence speech recognition on the audio segment corresponding to each non-standard sentence chunk to obtain a non-standard phoneme sequence;

[0019] Merge the standard phoneme sequence and the non-standard phoneme sequence to obtain the audio-text time link.

[0020] More preferably, it further includes: taking the recognition accuracy as the ultimate goal, performing discrimination training on the neural network model.

[0021] More preferably, it further includes: training the neural network model according to the training text set based on the criterion of minimizing the cross-entropy loss.

[0022] More preferably, the neural network model is based on a time-delay deep neural network.

[0023] More preferably, the training text set further includes: text-symbol mixed texts and letter abbreviation texts.

[0024] On the other hand, an electronic device is provided, including a processor, a memory, and a computer program stored in the memory and executable on the processor, where the processor is configured to execute the computer program stored on the memory to implement the method for aligning mixed sentences.

[0025] On the other hand, a storage medium is provided, in which at least one instruction is stored, and the instruction is loaded and executed by a processor to implement the hybrid statement alignment method.

[0026] The present invention has at least the following technical effects:

[0027] (1) By inputting any piece of audio and the corresponding Chinese text containing multiple languages, the start and end times of each phoneme in the phoneme sequence can be accurately output. Compared with the traditional alignment method, this method can targetedly identify foreign languages or other non-standard language chunks in the text, and only need to perform local speech recognition, greatly reducing the computational amount and improving the overall speed and performance of speech alignment;

[0028] (2) Pure Chinese language chunks are not affected by the capabilities of the speech recognition engine, greatly improving the alignment effect;

[0029] (3) It makes full use of the existing Chinese information, maximally avoids misrecognition, and has the advantages of simplicity, rapidity and high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0031] Figure 1 It is a flowchart of Embodiment 1 of the present invention;

[0032] Figure 2 It is a flowchart of Embodiment 2 of the present invention;

[0033] Figure 3 It is a flowchart of Embodiment 3 of the present invention;

[0034] Figure 4 It is a flowchart of the training part of the present invention;

[0035] Figure 5 It is a flowchart of the alignment part of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.

[0037] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0038] For the sake of simplicity of the drawings, only the parts related to the present invention are schematically shown in each figure, and they do not represent the actual structure of the product. In addition, for the sake of simplicity and easy understanding of the drawings, in some figures, only one of the components with the same structure or function is schematically drawn, or only one of them is labeled. In this document, "one" not only means "only this one", but also can mean the case of "more than one".

[0039] It should be further understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0040] In addition, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will describe the specific embodiments of the present invention with reference to the accompanying drawings. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, and other embodiments can be obtained.

[0042] Embodiment 1:

[0043] As Figure 1 、 4 、shown in 5, this embodiment provides a method for aligning hybrid sentences, including:

[0044] S1: Obtain hybrid text and target audio; the hybrid text includes canonical chunks and non-canonical chunks; there is a temporal correspondence between the hybrid text and the target audio;

[0045] S2: Use a pre-trained neural network model to identify non-standard chunks and standard chunks in the mixed text, and generate a classification result;

[0046] S3: According to the classification result, align the mixed text and the target audio to obtain the audio-text time link.

[0047] To make the lip movements of the virtual human realistic, vivid, and accurate when speaking, it is necessary to align the speech with the content of the speech. However, the Chinese text answered by the virtual human sometimes involves a large amount of content mixed with multiple languages or symbols, such as foreign language characters, special symbols, etc. The current speech alignment technology has an unsatisfactory processing effect on these unknown non-standard chunks, which greatly affects the user's interaction experience. Currently, there are mainly two alignment methods for Chinese text mixed with multiple languages: dictionary-based alignment and sentence-level speech recognition text alignment. Both of these methods have obvious deficiencies and cannot meet the alignment requirements of high accuracy, simplicity, and speed. Specifically as follows:

[0048] Dictionary-based alignment requires creating dictionaries for all languages in the world and mapping foreign language characters to phonemes. When aligning, after language discrimination, use the models of each language to uniformly map all characters into phonemes for forced alignment. Although this method has high accuracy, the workload is huge - it is necessary to create pronunciation dictionaries for all words in all languages in the world. In addition, the pronunciations of special symbols, abbreviations, and numbers cannot be exhausted.

[0049] For sentence-level speech recognition text alignment, according to the audio and reference text pair, use a speech recognition engine to decode the entire audio data to obtain the speech recognition text. Use a dynamic programming algorithm for maximum feature matching to achieve sentence-level alignment. However, this method requires large acoustic models and language models, and the computational complexity is very high, which reduces the alignment speed. At the same time, the speech recognition result is affected by the recognition ability of the engine and the complexity of the Chinese mixed text, and the alignment effect is poor, and accurate time information cannot be obtained.

[0050] Therefore, in this embodiment, aiming at the problems existing in the dictionary and the speech recognition engine, a divide-and-conquer method is adopted. Through a neural network model, the text is divided into standard chunks and non-standard chunks. For standard sentence chunks and non-standard sentence chunks, different methods are adopted, which can targetedly identify foreign languages or other non-standard chunks in the text, only perform local speech recognition, greatly reduce the computational complexity, and improve the overall speed and performance of speech alignment. At the same time, pure Chinese chunks are not affected by the ability of the speech recognition engine, which greatly improves the alignment effect.

[0051] Specifically, the canonical statement block is generally a statement entirely in Chinese, while the non-canonical statement block is manifested as more than one foreign language, special symbol, foreign language letter, letter abbreviation, and other non-canonical language blocks. It trains a corresponding neural network model through deep learning, combines the method of forced alignment to perform phoneme-level alignment on the statement, and at the same time marks the non-canonical language blocks in the statement, marking them as an unknown phoneme, finds the corresponding audio according to the time range, performs speech recognition, and at the same time completes the phoneme-level text alignment. Finally, the phoneme time mark sequence of the entire text is output.

[0052] At the same time, for non-Chinese languages, similar means can be adopted. For example, in French, the French text can be regarded as a canonical language block, the non-French part can be regarded as a non-canonical language block, and corresponding adjustments can be made to other content to obtain the corresponding results.

[0053] Embodiment 2:

[0054] Such as Figure 2 、 4 、as shown in 5, this embodiment provides a method for aligning hybrid statements, including:

[0055] S0-1: Obtain a training text set; the training text set includes pure Chinese texts and hybrid texts; further preferably, the training text set also includes: text-symbol hybrid texts, letter abbreviation texts; so that the model can adapt to various different application scenarios.

[0056] S0-2: Mark the non-canonical language blocks in the training set as unknown phonemes;

[0057] S0-3: Train the neural network model according to the training text set; preferably, training the neural network model according to the training text set is based on the criterion of minimizing the cross-entropy loss; that is, using the cross-entropy as its loss function to measure the similarity between the real situation and the predicted situation, thereby accelerating the learning rate and eliminating ambiguity.

[0058] S0-4: Taking the recognition accuracy as the ultimate guide, perform discrimination training on the neural network model;

[0059] S1: Obtain hybrid texts and target audio; the hybrid texts include canonical language blocks and non-canonical language blocks; there is a time correspondence between the hybrid texts and the target audio;

[0060] S2: Use the pre-trained neural network model to identify the non-canonical language blocks and canonical language blocks in the hybrid texts, and generate a classification result;

[0061] S3: According to the classification result, align the hybrid texts and the target audio to obtain the audio-text time link.

[0062] In this embodiment, by constructing Chinese data containing foreign languages in a controllable manner, a TDNN (Time-Delay Neural Network) model for training all Chinese phonemes and a non-canonical chunk is trained. The non-canonical chunk is a multi-language average phoneme (MLAP). The neural network model based on MLAP contains statistical information of any non-Chinese language and can represent various speech units of different languages and different lengths. During alignment, the neural network model that can represent various unknown languages is combined with known Chinese phonemes, and then the chunks of each unknown language can be accurately located. After the chunk location, Chinese speech recognition is performed on it to achieve the maximum approximation representation of foreign languages using Chinese phonemes. The known Chinese phonemes are fused with the phonemes and their start times after speech recognition, and the obtained phoneme label sequence is the final alignment result.

[0063] Embodiment 3:

[0064] As Figure 3 、 4 、shown in 5, this embodiment provides a method for aligning hybrid sentences, including:

[0065] S1: Obtain hybrid text and target audio; the hybrid text contains canonical chunks and non-canonical chunks; there is a temporal correspondence between the hybrid text and the target audio;

[0066] S2: Use a pre-trained neural network model to identify the non-canonical chunks and canonical chunks in the hybrid text and generate a classification result;

[0067] S3-1: Use a canonical chunk pronunciation dictionary to convert the canonical chunks in the mixed text into a canonical phoneme sequence;

[0068] S3-2: According to the canonical phoneme sequence, use a forced alignment algorithm to mark the start and end times of all the canonical chunks and all the non-canonical chunks on the target audio;

[0069] S3-3: According to the start and end times of each non-canonical chunk on the target audio, perform canonical sentence speech recognition on the audio segment corresponding to each non-canonical sentence chunk to obtain a non-canonical phoneme sequence;

[0070] S3-4: Combine the canonical phoneme sequence and the non-canonical phoneme sequence to obtain an audio text time link.

[0071] In the specific usage process, first, the target text is analyzed by a neural network model, and then the phoneme-level forced alignment and the chunk-level speech recognition text alignment based on non-canonical chunks are combined, and finally the time stamps of each phoneme are output.

[0072] More specifically: First, the target audio and the corresponding text are obtained. The text should contain foreign languages, special symbols, and other statements.

[0073] Then, the speech and the text are aligned, the target audio and the corresponding text are forced to be aligned, a millisecond-level time link is established between the text and the speech, and the start time of each phoneme is marked. In this process, the specific alignment strategy of the neural network model can be analyzed as follows:

[0074] Using a Chinese pronunciation dictionary, the Chinese text is converted into phonemes. Combining the Chinese phoneme model and the non-canonical chunk model, and using the forced alignment algorithm, the start time and end time of each Chinese phoneme and non-canonical chunk on the audio are marked. The final output result is: a sequence of time stamps of all phonemes including non-canonical chunks. Speech recognition of non-canonical chunks corresponding to the audio: According to the time range of the non-canonical chunks that have been marked, the corresponding audio is subjected to Chinese speech recognition, the start time of each phoneme is marked, and a sequence of phonemes with time stamps is output, thus realizing the alignment of the entire text.

[0075] Embodiment 5:

[0076] This embodiment provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. The processor is used to execute the computer program stored on the memory to implement the hybrid statement alignment method.

[0077] In this embodiment, the device can be a desktop computer, a notebook, a palm computer, a tablet computer, a mobile phone, a human-computer interaction screen, and other devices. The device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that this is only an example of the device and does not constitute a limitation on the device. It may include more or fewer components than shown in the figure, or combine some components, or different components. Exemplarily: The device may further include an input / output interface, a display device, a network access device, a communication bus, a communication interface, etc. The communication interface and the communication bus may further include an input / output interface. Among them, the processor, the memory, the input / output interface, and the communication interface complete mutual communication through the communication bus. The memory stores a computer program, and the processor is used to execute the computer program stored on the memory to implement the method in the above embodiment.

[0078] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0079] The memory may be an internal storage unit of the device, for example: the hard disk or memory of the device. The memory may also be an external storage device of the device, for example: the plug-in hard disk equipped on the device, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. Further, the memory may also include both the internal storage unit and the external storage device of the device. The memory is used to store the computer program and other programs and data required by the device. The memory may also be used to temporarily store the data that has been output or is to be output.

[0080] A communication bus is a circuit that connects the described elements and enables transmission between these elements. Exemplarily, a processor receives commands from other elements via the communication bus, decrypts the received commands, and performs calculations or data processing based on the decrypted commands. The memory may include program modules, such as, for example, a kernel, middleware, an Application Programming Interface (API), and applications. These program modules may be composed of software, firmware, hardware, or at least two of them. The input / output interface forwards commands or data input by a user through the input / output interface (e.g., sensors, keyboards, touchscreens). The communication interface connects the device to other network devices, user devices, and networks. Exemplarily, the communication interface can be connected to a network through a wired or wireless connection to connect to other external network devices or user devices. Wireless communication may include at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth (BT), Near Field Communication (NFC), Global Positioning System (GPS), and cellular communication, etc. Wired communication may include at least one of the following: Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Asynchronous Transfer Standard Interface (RS-232), etc. The network may be a telecommunications network and a communication network. The communication network may be a computer network, the Internet, the Internet of Things, or a telephone network. The device can be connected to the network through the communication interface, and the protocol used for communication between the device and other network devices can be supported by at least one of the application, the Application Programming Interface (API), middleware, the kernel, and the communication interface.

[0081] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. Exemplarily, the apparatus / device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0082] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0083] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0084] If the above-mentioned integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it may also be completed by sending instructions to relevant hardware through a computer program. The computer program may be stored in a medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented. Among them, the computer program may be in the form of source code, object code, executable file or some intermediate form, etc. The medium may include: any entity or device capable of carrying the computer program, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. Exemplarily, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals. Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each program module is used for illustration. In actual applications, the above-mentioned functions may be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program units or modules to complete all or part of the functions described above. Each program module in the embodiment may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one processing unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software program unit. In addition, the specific names of each program module are only for the convenience of mutual distinction and do not limit the protection scope of the present application.

[0085] Embodiment 6:

[0086] This embodiment provides a storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by a processor to implement the above-mentioned hybrid statement alignment method.

[0087] Technical effects of the present invention:

[0088] (1) By inputting any piece of audio and the corresponding Chinese text containing multiple languages, the start and end times of each phoneme in the phoneme sequence can be accurately output. Compared with traditional alignment methods, it can targetedly identify foreign languages or other non-standard language chunks in the text, only requiring local speech recognition, greatly reducing the computational amount and improving the overall speed and performance of speech alignment;

[0089] (2) Pure Chinese language chunks are not affected by the capabilities of the speech recognition engine, greatly improving the alignment effect;

[0090] (3) It makes full use of the existing Chinese information, maximally avoiding misrecognition, and has the advantages of being simple, fast, and having a high accuracy rate.

[0091] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0092] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. A method for aligning hybrid sentences, characterized in that, it includes: Obtain hybrid text and target audio; the hybrid text contains canonical chunks and non-canonical chunks; there is a temporal correspondence between the hybrid text and the target audio; Use a pre-trained neural network model to identify non-canonical chunks and canonical chunks in the hybrid text, generate a classification result, and perform Chinese speech recognition on the non-canonical chunk pairs, where the non-canonical chunks are multi-language average phonemes, and the non-canonical phonemes of the non-canonical chunks are determined based on Chinese phonemes; According to the classification result, align the hybrid text and the target audio, including: Use a canonical chunk pronunciation dictionary to convert the canonical chunks in the hybrid text into canonical phoneme sequences; According to the canonical phoneme sequence, use the forced alignment algorithm to mark the start and end times of all the canonical chunks and all the non-canonical chunks on the target audio; According to the start and end times of each non-canonical chunk on the target audio, perform canonical sentence speech recognition on the audio segment corresponding to each non-canonical chunk to obtain a non-canonical phoneme sequence; Merge the canonical phoneme sequence and the non-canonical phoneme sequence to obtain an audio text time link.

2. A method for aligning hybrid sentences according to claim 1, characterized in that, it further includes: Train the neural network model, specifically including: Obtain a training text set; the training text set contains pure Chinese text and hybrid text; Mark the non-canonical chunks in the training text set as unknown phonemes; Train the neural network model according to the training text set.

3. A method for aligning hybrid sentences according to claim 2, characterized in that, it further includes: Taking the recognition accuracy as the ultimate goal, perform discrimination training on the neural network model.

4. A method for aligning hybrid sentences according to claim 2, characterized in that, it further includes: Training the neural network model according to the training text set is based on the criterion of minimizing the cross-entropy loss.

5. A method for aligning hybrid sentences according to claim 2, characterized in that, the neural network model is based on a time-delay deep neural network.

6. A method for aligning hybrid sentences according to claim 2, characterized in that, the training text set further includes: text-symbol hybrid text and letter abbreviation text.

7. An electronic device, characterized in that, it includes a processor, a memory, and a computer program stored in the memory and executable on the processor. The processor is used to execute the computer program stored on the memory to implement a method for aligning hybrid sentences according to any one of claims 1-6.

8. A storage medium, characterized in that, at least one instruction is stored in the storage medium, and the instruction is loaded and executed by the processor to implement a method for aligning hybrid sentences according to any one of claims 1-6.

Citation Information

Patent Citations

  • Cross-language speech recognition method and device

    CN110349564A

  • Chinese lip gloss synchronization method based on 3D rendering engine

    CN111161755A

  • Digital virtual human mouth shape driving method based on pinyin or English phonetic symbol reading method

    CN112001323A