A method, device and medium for converting multi-language mixed text into speech
By preprocessing and feature extraction of multilingual mixed texts, and using deep learning or machine learning models to generate language recognition models, the problem of inaccurate text recognition in the existing technology in Chinese and foreign languages is solved, and high-accurate speech conversion is achieved.
Patent Information
- Application Number
- CN202411531864.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-10-30
AI Technical Summary
In the prior art, when converting multilingual mixed text into speech, it is difficult to accurately identify foreign text, resulting in inaccurate speech conversion.
By preprocessing Chinese and foreign text data, extracting feature vectors, and using deep learning or machine learning models for training, a language recognition model is generated, and then the converted text is recognized.
It realizes accurate recognition of foreign texts, improves the accuracy and success rate of speech conversion, saves the cost of post-production speech production correction, and improves work efficiency.
Smart Images

Figure CN119673143B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice conversion, and more specifically, to a method, apparatus, and medium for converting a multi-language mixed text into voice. Background Art
[0002] During the use of a voice platform, there will be a situation where some articles are a mix of Chinese and foreign languages. When converting the article into voice, since some foreign language characters are the same as Chinese characters, for example, some Japanese characters are the same as or similar to Chinese characters, the foreign language will be read as Chinese during the recognition process, resulting in the converted voice not accurately expressing the text of the article. Summary of the Invention
[0003] In view of the deficiencies of the prior art, the present invention provides a method, apparatus, and medium for converting a multi-language mixed text into voice.
[0004] According to one aspect of the present invention, there is provided a method for converting a multi-language mixed text into voice, including:
[0005] Performing data preprocessing on the collected text data of Chinese and foreign languages to obtain a training data set;
[0006] Extracting feature vectors from the training data set to obtain a feature vector data set of the training data set;
[0007] Inputting the feature vector data set into a deep learning model or a machine learning model for training to generate a language recognition model;
[0008] Using the language recognition model to recognize the text to be converted to obtain the language data of the text to be converted, where
[0009] Extracting feature vectors from the training data set to obtain a feature vector data set of the training data set, including:
[0010] Performing n-gram feature vector extraction on the training data set to obtain an n-gram feature vector set of the training data set;
[0011] Performing TF-IDF feature vector extraction based on the n-gram feature vector set to obtain a TF-IDF feature vector set of the training data set;
[0012] Determining a feature vector data set of the training data set according to the n-gram feature vector set and the TF-IDF feature vector set;
[0013] Performing TF-IDF feature vector extraction based on the n-gram feature vector set to obtain a TF-IDF feature vector set of the training data set, including:
[0014] Count the number of occurrences of each n-gram feature vector in the n-gram feature vector set;
[0015] Calculate the total number of n-grams in the training data set;
[0016] Calculate the number of text samples in the training data set that contain each n-gram feature vector;
[0017] Determine the TF feature vector based on the number of occurrences and the total number;
[0018] Determine the IDF feature vector based on the number of text samples;
[0019] Obtain the TF-IDF feature vector set of the training data set based on the TF feature vector and the IDF feature vector.
[0020] Optionally, perform data preprocessing on the collected Chinese and foreign language text data to obtain a training data set, including:
[0021] Clean the text data to obtain a valid text data set;
[0022] Use a pre-trained sentence segmentation model and a BERT language representation model to segment the valid text data set to obtain a training sentence set;
[0023] Use the tokenizer built into the BERT language model to tokenize the training sentence set to obtain the token sequence of each sentence in the training sentence set;
[0024] Generate a training data set based on the token sequence of each sentence;
[0025] Optionally, clean the text data to obtain a valid text data set, including:
[0026] Remove special characters and punctuation marks for different languages in the text data to obtain a cleaned text data set;
[0027] Unify the case of special language text in the cleaned text data set to obtain a unified text data set;
[0028] Remove stop words in the unified text data set to obtain a valid text data set.
[0029] According to another aspect of the present invention, there is provided an apparatus for converting a multi-language mixed text into speech, including:
[0030] A preprocessing module for performing data preprocessing on the collected Chinese and foreign language text data to obtain a training data set;
[0031] A feature extraction module for extracting feature vectors from the training data set to obtain a feature vector data set of the training data set;
[0032] A model training module, configured to input a feature vector data set into a deep learning model or a machine learning model for training to generate a language recognition model;
[0033] An identification module, configured to use the language recognition model to identify the text to be converted, and obtain the language data of the text to be converted, where
[0034] A feature extraction module, including:
[0035] A first extraction sub-module, configured to perform n-gram feature vector extraction on the training data set to obtain an n-gram feature vector set of the training data set;
[0036] A second extraction sub-module, configured to perform TF-IDF feature vector extraction according to the n-gram feature vector set to obtain a TF-IDF feature vector set of the training data set;
[0037] A determination sub-module, configured to determine a feature vector data set of the training data set according to the n-gram feature vector set and the TF-IDF feature vector set;
[0038] The second extraction sub-module includes:
[0039] A statistics unit, configured to count the number of occurrences of each n-gram feature vector in the n-gram feature vector set;
[0040] A first calculation unit, configured to calculate the total number of n-grams in the training data set;
[0041] A second calculation unit, configured to calculate the number of text samples in the training data set that contain each n-gram feature vector;
[0042] A first determination unit, configured to determine a TF feature vector according to the number of occurrences and the total number;
[0043] A second determination unit, configured to determine an IDF feature vector according to the number of text samples;
[0044] A first obtaining unit, configured to obtain a TF-IDF feature vector set of the training data set according to the TF feature vector and the IDF feature vector.
[0045] According to another aspect of the present invention, there is provided a computer-readable storage medium, where the storage medium stores a computer program, and the computer program is used to execute the method described in any one of the above aspects of the present invention.
[0046] According to another aspect of the present invention, there is provided an electronic device, which includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of the above aspects of the present invention.
[0047] Thus, through machine learning, the present invention preprocesses the collected Chinese and foreign language text data to obtain a training data set; extracts feature vectors from the training data set to obtain a feature vector data set of the training data set; inputs the feature vector data set into a deep learning model or a machine learning model for training to generate a language recognition model; and uses the language recognition model to recognize the text to be converted to obtain the language data of the text to be converted. The accurate recognition of foreign languages is realized, the accuracy and success rate of speech conversion are improved, the cost of post-production speech correction is saved, and the work efficiency is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] By referring to the following drawings, the exemplary embodiments of the present invention can be more fully understood:
[0049] Figure 1 is a flowchart of a method for converting a multi-language mixed text into speech provided by an exemplary embodiment of the present invention;
[0050] Figure 2 is a schematic structural diagram of a device for converting a multi-language mixed text into speech provided by an exemplary embodiment of the present invention;
[0051] Figure 3 is the structure of an electronic device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.
[0053] It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0054] Those skilled in the art can understand that the terms "first", "second", etc. in the embodiments of the present invention are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0055] It should also be understood that in the embodiments of the present invention, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0056] It should also be understood that for any component, data or structure mentioned in the embodiments of the present invention, in the case of no clear limitation or contrary revelation in the context, it can generally be understood as one or more.
[0057] In addition, the term "and / or" in the present invention is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.
[0058] It should also be understood that the description of each embodiment of the present invention emphasizes the differences between the embodiments, and the same or similar parts can be referred to each other. For the sake of brevity, they will not be described one by one.
[0059] At the same time, it should be understood that for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0060] The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present invention and its application or use.
[0061] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, the technologies, methods and devices should be regarded as part of the specification.
[0062] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0063] The embodiments of the present invention can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0064] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, target programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0065] Exemplary method
[0066] Figure 1 is a schematic flowchart of a method for converting a multilingual mixed text into speech provided by an exemplary embodiment of the present invention. This embodiment can be applied to an electronic device, such as Figure 1 As shown, the method 100 for converting a multilingual mixed text into speech includes the following steps:
[0067] Step 101, perform data preprocessing on the collected text data in Chinese and foreign languages to obtain a training data set;
[0068] Step 102, extract feature vectors from the training data set to obtain a feature vector data set of the training data set;
[0069] Step 103, input the feature vector data set into a deep learning model or a machine learning model for training to generate a language recognition model;
[0070] Step 104, use the language recognition model to recognize the text to be converted to obtain the language data of the text to be converted, where
[0071] extracting feature vectors from the training data set to obtain a feature vector data set of the training data set includes:
[0072] Perform n-gram feature vector extraction on the training data set to obtain an n-gram feature vector set of the training data set;
[0073] Perform TF-IDF feature vector extraction based on the n-gram feature vector set to obtain a TF-IDF feature vector set of the training data set;
[0074] Determine the feature vector data set of the training data set according to the n-gram feature vector set and the TF-IDF feature vector set;
[0075] Performing TF-IDF feature vector extraction based on the n-gram feature vector set to obtain a TF-IDF feature vector set of the training data set includes:
[0076] Count the number of occurrences of each n-gram feature vector in the n-gram feature vector set;
[0077] Calculate the total number of n-grams in the training dataset;
[0078] Calculate the number of text samples in the training dataset that contain each n-gram feature vector;
[0079] Determine the TF feature vector based on the number of occurrences and the total number;
[0080] Determine the IDF feature vector based on the number of text samples;
[0081] Obtain the TF-IDF feature vector set of the training dataset based on the TF feature vector and the IDF feature vector.
[0082] Optionally, perform data preprocessing on the collected Chinese and foreign language text data to obtain a training dataset, including:
[0083] Clean the text data to obtain a valid text dataset;
[0084] Use a pre-trained sentence segmentation model and a BERT language representation model to segment the valid text dataset to obtain a training sentence set;
[0085] Use the tokenizer built into the BERT language model to tokenize the training sentence set to obtain the token sequence of each sentence in the training sentence set;
[0086] Generate a training dataset based on the token sequence of each sentence;
[0087] Optionally, clean the text data to obtain a valid text dataset, including:
[0088] Remove special characters and punctuation marks for different languages in the text data to obtain a cleaned text dataset;
[0089] Unify the case of special language texts in the cleaned text dataset to obtain a unified text dataset;
[0090] Remove stop words from the unified text dataset to obtain a valid text dataset.
[0091] Specifically, the present invention uses machine learning methods to perform Chinese and foreign language recognition, with a large amount of labeled training data, and performs some steps of data preprocessing and model training. The implementation steps are as follows:
[0092] 1. Data collection: Collect a large amount of Chinese and foreign language text data. For example, it can be found on the Internet or crawled from websites using a crawler program.
[0093] 2. Data preprocessing: preprocess the collected text data, including cleaning data, cutting sentences, and word segmentation. The purpose of this step is to convert the original text data into data that can be input into the model.
[0094] (1) Loading a pre-trained language model: First, load and initialize the selected pre-trained language model, such as BERT.
[0095] (2) Text cleaning and preprocessing: Perform basic cleaning and preprocessing on the collected text data, such as removing special characters, HTML tags, URLs, etc., to ensure the purity and consistency of the text.
[0096] Among them, the cleaning scheme is as follows:
[0097] a. Remove special characters and punctuation marks: Different languages may use different punctuation marks, so they need to be cleaned up according to the characteristics of the specific language.
[0098] For example: Chinese: you need to remove Chinese punctuation marks, such as ",",".","!";
[0099] Before cleaning: "Hello! How is the weather today?";
[0100] After cleaning: "Hello, how is the weather today";
[0101] English: Remove English punctuation marks such as “,”, “.”, “!”, and remove extra spaces.
[0102] Before cleaning: "Hello! How are you today?";
[0103] After cleaning: "Hello How are you today".
[0104] b. Unify capitalization:
[0105] In languages like English, capitalization can reduce the number of synonyms that are repeated.
[0106] For example: English: Convert all text to lowercase.
[0107] Before cleaning: "Hello World";
[0108] After cleaning: "hello world";
[0109] German: Nouns in German usually start with a capital letter, and you can decide whether to use the same capital letter as needed.
[0110] Before cleaning: "Die Katze sitzt auf dem Tisch";
[0111] After cleaning: "die katze sitzt auf dem tisch" (you can choose to make it lowercase, for example).
[0112] c. Remove stop words:
[0113] Different languages have different stop words (such as "的", "是", "和", "the", etc.), which need to be removed during cleaning.
[0114] For example: Chinese: remove common stop words.
[0115] Before cleaning: “I love natural language processing”;
[0116] After cleaning: "Love natural language processing";
[0117] French: Remove stop words in French.
[0118] Before cleaning: "Je suis très heureux de vous rencontrer";
[0119] After cleaning: "suis heureux rencontrer".
[0120] (3) Sentence segmentation: Using the most innovative language model BERT, you can use a pre-trained sentence segmentation model to segment text data into sentences. This method can handle complex language structures, including ellipsis, quotation marks, and brackets.
[0121] Sentence segmentation using the BERT model mainly relies on pre-trained sentence segmentation tools, and BERT itself does not directly provide sentence segmentation functions. This function can be achieved by combining the NLP toolkit and the BERT model's word segmenter. The following is a detailed description and example:
[0122] a. Choose a sentence segmentation tool: Use a natural language processing library like NLTK or spaCy, which provide ready-made sentence segmentation functions.
[0123] b. Load the BERT model and tokenizer: Use Hugging Face’s transformers library to load the pre-trained BERT model and the corresponding tokenizer.
[0124] c. Cut text into sentences: Use the selected tool to split the text into sentences.
[0125] d. Tokenize the sentences: Input the segmented sentences into BERT's tokenizer to convert them into a format that the model can process.
[0126] (4) Tokenization: Using the pre-trained language model BERT, the built-in tokenizer can be used to split each sentence into words or sub-words. This tokenization method can handle complex lexical structures such as compound words, abbreviations, and proper nouns.
[0127] (5) Generate token sequences: Convert the tokenization results of each sentence into a sequence for subsequent processing and analysis. Pre-defined encoding schemes such as WordPiece or Byte-Pair Encoding (BPE) can be used to convert the tokens into an input form acceptable to the model.
[0128] 3. Feature engineering: Extract feature vectors from the preprocessed text data. This includes extracting N-gram features and TF-IDF features.
[0129] a. n-gram model:
[0130] (1) Divide the text into sequences consisting of consecutive n words.
[0131] (2) Build an n-gram vocabulary and map each n-gram to a unique integer index.
[0132] (3) For each text sample, count the number of times each n-gram appears in the text or weight it using TF-IDF.
[0133] Count the number of times each n-gram appears in the text:
[0134] Input text sample: text;
[0135] Input n-gram: n_gram;
[0136] Count the number of times each n-gram appears in the text: count = count the number of occurrences of n-gram in text;
[0137] b. Weight using TF-IDF:
[0138] Input text sample: text;
[0139] Input n-gram: n_gram;
[0140] Input corpus: corpus;
[0141] Count the number of occurrences of each n-gram in the text: count = count the number of occurrences of the n-gram in text;
[0142] Calculate the total number of all n-grams in the text sample: total_ngrams = calculate the total number of n-grams in all text samples in corpus;
[0143] Calculate the number of text samples in the corpus that contain the current n-gram: num_documents = calculate the number of text samples in corpus that contain the current n-gram;
[0144] Calculate TF (Term Frequency): tf = count / total_ngrams;
[0145] Calculate IDF (Inverse Document Frequency): idf = log(total number of text samples in the corpus / (1 + num_documents));
[0146] Calculate the TF-IDF weighted value: tf_idf = tf idf.
[0147] (4) Represent each text sample as a vector, where each dimension of the vector corresponds to the number of occurrences or TF-IDF value of an n-gram in the text.
[0148] 4. Model training: Select a suitable machine learning model or deep learning model, such as SVM, etc., and input the feature data into the model for training.
[0149] 5. Model evaluation: After training the model, it is necessary to evaluate the model to see how well it performs. Commonly used evaluation metrics include accuracy, precision, recall, F1 value, etc.
[0150] 6. Model application: Apply the trained model to an actual project for language recognition.
[0151] Thus, through machine learning, the present invention accurately identifies foreign languages, improves the accuracy and success rate of speech conversion, saves the cost of post-production speech correction, and greatly improves work efficiency.
[0152] Exemplary device
[0153] Figure 2 It is a schematic structural diagram of a device for converting a multi-language mixed text into speech provided by an exemplary embodiment of the present invention. As Figure 2 shown, the device 200 includes:
[0154] A preprocessing module for preprocessing the collected text data in Chinese and foreign languages to obtain a training data set;
[0155] A feature extraction module for extracting feature vectors from the training data set to obtain a feature vector data set of the training data set;
[0156] A model training module for inputting the feature vector data set into a deep learning model or a machine learning model for training to generate a language recognition model;
[0157] An identification module for using the language recognition model to identify the text to be converted to obtain the language data of the text to be converted, where
[0158] The feature extraction module includes:
[0159] A first extraction sub-module for extracting n-gram feature vectors from the training data set to obtain an n-gram feature vector set of the training data set;
[0160] A second extraction sub-module for performing TF-IDF feature vector extraction according to the n-gram feature vector set to obtain a TF-IDF feature vector set of the training data set;
[0161] A determination sub-module for determining the feature vector data set of the training data set according to the n-gram feature vector set and the TF-IDF feature vector set;
[0162] The second extraction sub-module includes:
[0163] A statistics unit for counting the number of occurrences of each n-gram feature vector in the n-gram feature vector set;
[0164] A first calculation unit for calculating the total number of n-grams in the training data set;
[0165] A second calculation unit for calculating the number of text samples in the training data set that contain each n-gram feature vector;
[0166] A first determination unit for determining the TF feature vector according to the number of occurrences and the total number;
[0167] A second determination unit for determining the IDF feature vector according to the number of text samples;
[0168] A first acquisition unit for obtaining the TF-IDF feature vector set of the training data set according to the TF feature vector and the IDF feature vector.
[0169] Optionally, the preprocessing module includes:
[0170] The first acquisition sub-module is used to clean the text data and obtain a valid text data set;
[0171] The second acquisition sub-module is used to perform sentence segmentation on the valid text data set by using a pre-trained sentence segmentation model and a BERT language representation model to obtain a training sentence set;
[0172] The third acquisition sub-module is used to perform word segmentation on the training sentence set by using the word segmenter built in the BERT language model to obtain the word segmentation sequence of each sentence in the training sentence set;
[0173] The generation sub-module is used to generate a training data set according to the word segmentation sequence of each sentence;
[0174] Optionally, the first acquisition sub-module includes:
[0175] The second acquisition unit is used to remove special characters and punctuation marks for different languages in the text data to obtain a cleaned text data set;
[0176] The third acquisition unit is used to unify the case of special language texts in the cleaned text data set to obtain a unified text data set;
[0177] The fourth acquisition unit is used to remove stop words in the unified text data set to obtain a valid text data set.
[0178] Exemplary electronic device
[0179] Figure 3 is the structure of the electronic device provided by an exemplary embodiment of the present invention. As Figure 3 shown, the electronic device 30 includes one or more processors 31 and a memory 32.
[0180] The processor 31 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0181] The memory 32 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 31 may run the program instructions to implement the methods of the software programs of the various embodiments of the present invention described above and / or other desired functions. In one example, the electronic device may further include: an input device 33 and an output device 34, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0182] In addition, the input device 33 may further include, for example, a keyboard, a mouse, and so on.
[0183] The output device 34 may output various information to the outside. The output device 34 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.
[0184] Of course, for simplicity, Figure 3 only some of the components related to the present invention in the electronic device are shown, and components such as a bus, an input / output interface, and so on are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.
[0185] Exemplary computer program product and computer-readable storage medium
[0186] In addition to the above methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present invention described in the "Exemplary Method" section above of this specification.
[0187] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0188] In addition, an embodiment of the present invention may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps in the method of information mining on historical change records according to various embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0189] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0190] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present invention. In addition, the above-disclosed specific details are only for the purposes of illustration and facilitating understanding, rather than limitations. The above details do not limit the present invention to necessarily adopt the above specific details for implementation.
[0191] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments may be referred to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple, and the relevant parts may refer to the partial description of the method embodiments.
[0192] The block diagrams of the devices, systems, equipment, and systems involved in the present invention are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, systems, equipment, and systems may be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and may be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and may be used interchangeably with each other unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and may be used interchangeably with each other.
[0193] The methods and systems of the present invention can be implemented in many ways. For example, the methods and systems of the present invention can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the methods is for illustrative purposes only, and the steps of the methods of the present invention are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present invention can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present invention. Thus, the present invention also covers a recording medium storing a program for executing the methods according to the present invention.
[0194] It should also be noted that in the systems, devices, and methods of the present invention, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present invention. Therefore, the present invention is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0195] The above description has been presented for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present invention to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A method for converting multi-language mixed text into speech, wherein the feature vector is: include: Preprocess the collected Chinese and foreign text data to obtain a training data set; Extracting feature vectors from the training data set to obtain a feature vector data set of the training data set; Inputting the feature vector data set into a deep learning model or a machine learning model for training to generate a language recognition model; The language recognition model is used to recognize the text to be converted, and the language data of the text to be converted is obtained, wherein: Extracting feature vectors from the training data set to obtain a feature vector data set of the training data set includes: Extracting n-gram feature vectors from the training data set to obtain an n-gram feature vector set of the training data set; Extract TF-IDF feature vectors according to the n-gram feature vector set to obtain the TF-IDF feature vector set of the training data set; Determine the feature vector data set of the training data set according to the n-gram feature vector set and the TF-IDF feature vector set; Extracting TF-IDF feature vectors according to the n-gram feature vector set to obtain the TF-IDF feature vector set of the training data set includes: Counting the number of occurrences of each n-gram feature vector in the n-gram feature vector set; Calculate the total number of n-grams in the training data set; Calculate the number of text samples that the training data set contains each n-gram feature vector; Determine a TF feature vector according to the number of times and the total number; Determine an IDF feature vector according to the number of text samples; The TF-IDF feature vector set of the training data set is obtained according to the TF feature vector and the IDF feature vector.
2. The method according to claim 1, characterized in that The collected Chinese and foreign text data are preprocessed to obtain training data sets, including: Cleaning the text data to obtain a valid text data set; Using a pre-trained sentence segmentation model and a BERT language representation model to perform sentence segmentation on the valid text dataset to obtain a training sentence set; The training sentence set is segmented using the built-in word segmenter of the BERT language model to obtain a word segmentation sequence of each sentence in the training sentence set; The training data set is generated according to the word segmentation sequence of each sentence.
3. The method according to claim 2, characterized in that Cleaning the text data to obtain a valid text data set includes: Removing special characters and punctuation marks for different languages in the text data to obtain a cleaned text data set; Unifying the case of special language texts in the cleaned text data set to obtain a unified text data set; Stop words in the unified text data set are removed to obtain the valid text data set.
4. A device for converting multi-language mixed text into speech, Its eigenvectors include: A preprocessing module is used to preprocess the collected Chinese and foreign text data to obtain a training data set; A feature extraction module is used to extract feature vectors from the training data set to obtain a feature vector data set of the training data set; A model training module, used to input the feature vector data set into a deep learning model or a machine learning model for training to generate a language recognition model; A recognition module is used to recognize the text to be converted by using the language recognition model to obtain the language data of the text to be converted, wherein: Feature extraction module, including: A first extraction submodule is used to extract n-gram feature vectors from the training data set to obtain an n-gram feature vector set of the training data set; A second extraction submodule is used to extract TF-IDF feature vectors according to the n-gram feature vector set to obtain the TF-IDF feature vector set of the training data set; A determination submodule, configured to determine the feature vector data set of the training data set according to the n-gram feature vector set and the TF-IDF feature vector set; The second extraction submodule includes: A statistical unit, used for counting the number of occurrences of each n-gram feature vector in the n-gram feature vector set; A first computing unit, configured to compute the total number of n-grams in the training data set; A second calculation unit is used to calculate the number of text samples contained in each n-gram feature vector in the training data set; A first determining unit, configured to determine a TF feature vector according to the number of times and the total number; A second determining unit is used to determine an IDF feature vector according to the number of text samples; The first acquisition unit is used to acquire the TF-IDF feature vector set of the training data set according to the TF feature vector and the IDF feature vector.
5. The device according to claim 4, characterized in that Preprocessing modules include: A first acquisition submodule is used to clean the text data to obtain a valid text data set; A second acquisition submodule is used to perform sentence segmentation on the valid text data set using a pre-trained sentence segmentation model and a BERT language representation model to obtain a training sentence set; The third acquisition submodule is used to perform word segmentation processing on the training sentence set using the word segmenter built into the BERT language model to obtain a word segmentation sequence of each sentence in the training sentence set; The generating submodule is used to generate the training data set according to the word segmentation sequence of each sentence.
6. The device according to claim 5, characterized in that The first acquisition submodule includes: A second acquisition unit is used to remove special characters and punctuation marks in different languages in the text data to obtain a cleaned text data set; A third acquisition unit is used to unify the capitalization of the special language text in the cleaned text data set to obtain a unified text data set; The fourth acquisition unit is used to remove stop words in the unified text data set to acquire the valid text data set.
7. A computer-readable storage medium, wherein the characteristic vector is: The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 3.
8. An electronic device, whose characteristic vector is, The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1 to 3.
Citation Information
Patent Citations
Chinese text classification method for computer
CN103020167A
Text intention classification method and device and readable medium
CN112905795A