Semantic feature extraction model training method and device, equipment and storage medium
By training a semantic feature extraction model by fusing word and pronunciation features, the problem of insufficient semantic representation ability of existing models is solved, and a stronger semantic feature capture and representation capability is achieved.
Patent Information
- Application Number
- CN202110393016.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-13
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2041-08-24
AI Technical Summary
Existing semantic feature extraction models have poor semantic representation capabilities when extracting semantic features from text, and fail to make full use of word and pronunciation features.
By acquiring word representation vector sequences and pronunciation representation vector sequences from word and phrase text corpora, and combining various feature fusion methods such as averaging, concatenation, weighted summation, and self-attention mechanisms, the semantic representation capability of the semantic feature extraction model is improved.
It enhances the semantic feature extraction model's ability to capture word and pronunciation features, thereby improving the model's semantic representation capabilities.
Smart Images

Figure CN113723105B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method, apparatus, device, and storage medium for a semantic feature extraction model. Background Technology
[0002] Semantic feature extraction models are neural network models used to extract semantic features from text, thereby achieving a modeling representation of the text.
[0003] In related technologies, when modeling and representing text, the semantic features of the text are usually extracted from the word features contained in the text through a semantic feature extraction model. However, semantic feature extraction models trained in this way have poor semantic representation capabilities. Summary of the Invention
[0004] This application provides a training method, apparatus, device, and storage medium for a semantic feature extraction model, which can improve the semantic representation capability of the semantic feature extraction model. The technical solution is as follows:
[0005] According to one aspect of the embodiments of this application, a method for training a semantic feature extraction model is provided, the method comprising:
[0006] Obtain the training corpus of the semantic feature extraction model, wherein the training corpus includes word and text corpus of the target language and pronunciation annotation information of the word and text corpus;
[0007] Obtain the word representation vector sequence of the word text corpus, and the pronunciation representation vector sequence of the pronunciation annotation information of the word text corpus;
[0008] The semantic feature extraction model extracts the fused semantic features of the word text corpus from the word representation vector sequence and pronunciation representation vector sequence of the word text corpus.
[0009] Based on the fused semantic features of the word and text corpus, the prediction results corresponding to the pre-training task of the semantic feature extraction model are determined.
[0010] The pre-training loss of the semantic feature extraction model is determined based on the prediction results and the actual results corresponding to the pre-training task, and the parameters of the semantic feature extraction model are adjusted according to the pre-training loss to obtain the pre-trained semantic feature extraction model.
[0011] According to one aspect of the embodiments of this application, a training apparatus for a semantic feature extraction model is provided, the apparatus comprising:
[0012] The corpus acquisition module is used to acquire the training corpus of the semantic feature extraction model, wherein the training corpus includes word and text corpus of the target language and the pronunciation annotation information of the word and text corpus;
[0013] The sequence acquisition module is used to acquire the word representation vector sequence of the word text corpus, and the pronunciation representation vector sequence of the pronunciation annotation information of the word text corpus;
[0014] The feature extraction module is used to extract the fused semantic features of the word and phrase text corpus from the word representation vector sequence and pronunciation representation vector sequence of the word and phrase text corpus through the semantic feature extraction model;
[0015] The result determination module is used to determine the prediction result corresponding to the pre-training task of the semantic feature extraction model based on the fused semantic features of the word and text corpus.
[0016] The parameter adjustment module is used to determine the pre-training loss of the semantic feature extraction model based on the prediction results and the actual results corresponding to the pre-training task, and to adjust the parameters of the semantic feature extraction model according to the pre-training loss to obtain the pre-trained semantic feature extraction model.
[0017] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the above-described training method for the semantic feature extraction model.
[0018] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored in the storage medium, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the training method of the above-described semantic feature extraction model.
[0019] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the training method for the semantic feature extraction model described above.
[0020] The technical solutions provided in this application have at least the following beneficial effects:
[0021] When extracting semantic features from word and phrase text corpora, the semantic feature extraction model not only uses the word and phrase features (i.e., word and phrase representation vector sequences) but also the pronunciation features (i.e., pronunciation representation vector sequences). Compared with related technologies that only consider word and phrase features, the technical solution of this application enables the semantic feature extraction model to capture both word and phrase features and pronunciation features, making full use of multiple features to enhance the semantic representation capability of the model. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the implementation environment of a solution provided in one embodiment of this application;
[0024] Figure 2 This is a flowchart of a training method for a semantic feature extraction model provided in one embodiment of this application;
[0025] Figure 3 This is a schematic diagram of word pronunciation feature fusion provided in one embodiment of this application;
[0026] Figure 4 This is a schematic diagram of word pronunciation feature fusion provided in another embodiment of this application;
[0027] Figure 5 This is a schematic diagram of word pronunciation feature fusion provided in another embodiment of this application;
[0028] Figure 6 This is a schematic diagram of staged training of a model provided in one embodiment of this application;
[0029] Figure 7 This is a flowchart of a model fine-tuning process provided in one embodiment of this application;
[0030] Figure 8 This is a flowchart of a training method for a semantic feature extraction model provided in another embodiment of this application;
[0031] Figure 9 This is a block diagram of a training apparatus for a semantic feature extraction model provided in one embodiment of this application;
[0032] Figure 10 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0034] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0035] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0036] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0037] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learn-by-doing.
[0038] The technical solutions provided in this application involve technologies such as machine learning and natural language processing in artificial intelligence, which are specifically described and illustrated through the following embodiments.
[0039] Please refer to Figure 1 The diagram illustrates an implementation environment for a solution provided in one embodiment of this application. This implementation environment may include a model training device 10 and a model usage device 20.
[0040] The model training device 10 can be an electronic device such as a computer, server, or intelligent robot, or other electronic devices with strong computing power. The model training device 10 is used to train the semantic feature extraction model. In this embodiment, the semantic feature extraction model is a neural network model used to extract high-level semantic features of text. The model training device 10 can use machine learning to train the semantic feature extraction model to enable it to have better semantic representation capabilities.
[0041] Optionally, the model training device 10 can also apply the pre-trained semantic feature extraction model to the target task, and fine-tune the parameters of the semantic feature extraction model using the training samples of the target task, so that the fine-tuned semantic feature extraction model, together with the result prediction network of the target task, can have better result prediction performance for the target task.
[0042] The semantic feature extraction model and the result prediction network for the target task, trained as described above, can form a task prediction model for the target task (referred to as the "target task model" in this application). This target task model can be deployed on the model-using device 20 to predict the results of the target task. The model-using device 20 can be a terminal device such as a mobile phone, computer, smart TV, multimedia playback device, wearable device, medical device, or a server; this application does not limit its scope.
[0043] In this application embodiment, the specific task content of the target task is not limited; it can be any task that requires the use of semantic features of text for result prediction. For example, in the field of video comment recognition, the target task can be a video comment sentiment recognition task, a video comment poetry recognition task, etc., and this application does not limit it in this regard.
[0044] The technical solution of this application will be described below through several embodiments.
[0045] Please refer to Figure 2 The diagram illustrates a flowchart of a training method for a semantic feature extraction model provided in one embodiment of this application. The execution entity for each step of this method may be the entity described above. Figure 1 The model training device 10 is used in the implementation environment of the scheme shown. The method may include the following steps (210-250):
[0046] Step 210: Obtain the training corpus for the semantic feature extraction model. The training corpus includes the target language word text corpus and the pronunciation annotation information of the word text corpus.
[0047] Optionally, text corpora of the target language can be obtained from a data source. Data sources can include external data sources and / or internal data sources. External data sources, also known as internet data sources, can be used to scrape text data from the internet as training corpora, such as data from encyclopedia websites, news websites, or large or authoritative portal websites; this application does not limit this. Internal data sources refer to data owned by the business entity itself. For example, if the business entity is a video service provider, it can obtain video-related text data from an internal data source as training corpora, including but not limited to video titles, text recognized by video OCR (Optical Character Recognition) or ASR (Automatic Speech Recognition), comments, posts, etc.
[0048] Optionally, the target language is Chinese. When the target language is Chinese, the text corpus of words and phrases obtained from the data source can be a text corpus of words and phrases in Simplified Chinese, a text corpus of words and phrases in Traditional Chinese, or a text corpus of words and phrases in both Simplified and Traditional Chinese.
[0049] The phonetic annotation information of a word and phrase text corpus refers to the annotation information that represents the pronunciation of the corresponding word or phrase in the text corpus; it is the phonetic explanation of the word and phrase text corpus. Taking Chinese word and phrase text corpus as an example, its phonetic annotation information is Hanyu Pinyin (referred to as "Pinyin"). There are a total of 439 whole-syllable Pinyin characters in Chinese. For example, a word and phrase text corpus is "Everyone loves listening to Shan Tianfang's crosstalk,...", and its corresponding Pinyin is "da'jia'dou'ai'ting'shan'tian'fang'de'xiang'sheng,...".
[0050] Optionally, a phonetic annotation model can be used to generate phonetic annotation information for word and phrase text corpora. A phonetic annotation model is a pre-trained neural network model used to automatically generate phonetic annotation information for word and phrase text corpora; it can be a CRF (Conditional Random Field) model. Taking Chinese Pinyin annotation as an example, by training on a large amount of Chinese-Pinyin corpus, the model can be equipped with the ability to output the corresponding Pinyin when given a Chinese sentence.
[0051] It should be noted that the technical solution provided in the embodiments of this application is applicable not only to Chinese, but also to other languages such as Japanese, Korean, English, Latin, etc. Under different languages, the specific forms of pronunciation annotation information will also be different. For example, the pronunciation annotation information in Chinese is called pinyin, the pronunciation annotation information in Japanese is called romaji, the pronunciation annotation information in English is called phonetic symbols, etc., and no further examples will be given here. In addition, in the embodiments of this application, unless otherwise specified, the technical solution of this application will be mainly introduced and described by taking Chinese as an example.
[0052] Step 220: Obtain the word representation vector sequence of the word text corpus and the pronunciation representation vector sequence of the pronunciation annotation information of the word text corpus.
[0053] The word text corpus can be a phrase or a sentence, and it includes multiple words. Taking Chinese as an example, the word text corpus can be "I love my motherland", which includes 4 characters.
[0054] In one example, the word text corpus is segmented by taking a single character as a unit. The word text corpus can be segmented into multiple characters. For example, "I love my motherland" can be segmented into 4 characters: "I", "love", "motherland", and "country". Each character has a corresponding representation vector, and the representation vectors corresponding to multiple characters are sequentially concatenated to form the word representation vector sequence.
[0055] In another example, the word text corpus is segmented by taking a word as a unit. The word text corpus can be segmented into multiple words. Each word can be a single character or a multi-character word. For example, "I love my motherland" can be segmented into 3 words: "I", "love", and "motherland". Each word has a corresponding representation vector, and the representation vectors corresponding to multiple words are sequentially concatenated to form the word representation vector sequence.
[0056] Similarly, the pronunciation annotation information of the word text corpus can also be segmented by taking a character (or word) as a unit to obtain the pronunciation annotation information of each character (or word). The pronunciation annotation information of each character (or word) has a corresponding representation vector, and the representation vectors corresponding to the pronunciation annotation information of multiple characters (or words) are sequentially concatenated to form the pronunciation representation vector sequence.
[0057] It should be noted that the segmentation granularity of the word text corpus and its pronunciation annotation information should be correspondingly consistent. That is, when the word text corpus is segmented by taking a single character as a unit, the pronunciation annotation information is also segmented by taking a single character as a unit; when the word text corpus is segmented by taking a single word as a unit, the pronunciation annotation information is also segmented by taking a single word as a unit.
[0058] In the embodiments of this application, the representation vectors of words and pronunciations can be called word vectors or word embeddings, which are quantitative representations of word and pronunciation information in the form of vectors.
[0059] Optionally, a word representation vector generation network is used to generate a word representation vector for each word in the word and text corpus, resulting in a sequence of word representation vectors. A pronunciation representation vector generation network is used to generate a pronunciation representation vector for the pronunciation annotation information of each word in the word and text corpus, resulting in a sequence of pronunciation representation vectors. Specifically, the word representation vector generation network generates the word representation vector for each word, and the pronunciation representation vector generation network generates the pronunciation representation vector for the pronunciation annotation information of each word.
[0060] Step 230: Using a semantic feature extraction model, extract the fused semantic features of the word text corpus from the word representation vector sequence and pronunciation representation vector sequence of the word text corpus.
[0061] In the embodiments of this application, when extracting semantic features from word and phrase text corpora, the semantic feature extraction model not only uses the word and phrase features (i.e., word and phrase representation vector sequences) of the word and phrase text corpora, but also uses the pronunciation features (i.e., pronunciation representation vector sequences) of the word and phrase text corpora. Compared with related technologies that only consider word and phrase features, the technical solution of this application enables the semantic feature extraction model to capture both word and phrase features and pronunciation features, making full use of multiple features to enhance the semantic representation capability of the model.
[0062] In one example, such as Figure 3 As shown, the word representation vector sequence and pronunciation representation vector sequence of the word and word text corpus are fused to obtain the fused representation vector sequence of the word and word text corpus; the fused semantic features of the word and word text corpus are obtained by performing feature extraction processing on the fused representation vector sequence of the word and word text corpus through a semantic feature extraction model.
[0063] Optionally, this application provides the following methods for generating fused representation vector sequences:
[0064] Method 1: Average the word representation vectors and pronunciation representation vectors corresponding to the same word position in the word and word representation vector sequences of the word and word text corpus to obtain the fused representation vector sequence of the word and word text corpus.
[0065] Suppose a text corpus contains n characters, where the word representation vector of the i-th character is Ci, and the pronunciation representation vector of the i-th character is PYi, where n is an integer greater than 1, and i is a positive integer less than or equal to n. In Method 1, the fused representation vector of the i-th character is (Ci + PYi) / 2. Concatenating the fused representation vectors of the above n characters sequentially forms the fused representation vector sequence of the text corpus. This method requires that the word representation vector Ci and the pronunciation representation vector PYi have the same vector dimension for ease of computation.
[0066] Method 2: The word representation vectors and pronunciation representation vectors corresponding to the same word position in the word and word text corpus are concatenated to obtain the fused representation vector sequence of the word and word text corpus.
[0067] Assume a text corpus contains n characters, where the word representation vector of the i-th character is Ci, and the pronunciation representation vector of the i-th character is PYi, where n is an integer greater than 1, and i is a positive integer less than or equal to n. In Method 2, the fused representation vector of the i-th character is [Ci, PYi]. Concatenating the fused representation vectors of the above n characters sequentially forms the fused representation vector sequence of the text corpus. This method does not require the word representation vector Ci and the pronunciation representation vector PYi to have the same vector dimension.
[0068] Method 3: Input the word representation vector sequence and pronunciation representation vector sequence of the word and text corpus into the word and pronunciation fusion network; through the word and pronunciation fusion network, perform weighted summation on the word and pronunciation representation vectors corresponding to the same word position in the word and text corpus, to obtain the fused representation vector sequence of the word and text corpus.
[0069] A word-to-pronunciation fusion network is a neural network used to fuse word representation vectors and pronunciation representation vectors; it can be a fully connected network. Assume a word-text corpus contains n words, where the word representation vector of the i-th word is Ci, and the pronunciation representation vector of the i-th word is PYi, where n is an integer greater than 1, and i is a positive integer less than or equal to n. In method 3, the fused representation vector of the i-th word is Ci*Wc + PYi*Wp, where Wc and Wp are the weight parameters of the word-to-pronunciation fusion network, which can be updated and learned during model training. Similarly, concatenating the fused representation vectors of the above n words sequentially constitutes the fused representation vector sequence of the word-text corpus. If method 3 is used to generate the fused representation vector sequence, the word-pronunciation fusion network is set before the semantic feature extraction model. The input of the word-pronunciation fusion network is the word representation vector sequence and the pronunciation representation vector sequence of the word text corpus, and the output is the fused representation vector sequence of the word text corpus. The input of the semantic feature extraction model is the fused representation vector sequence of the word text corpus, and the output is the fused semantic features of the word text corpus. The word-pronunciation fusion network and the semantic feature extraction model can be trained simultaneously, and their parameters can be adjusted.
[0070] In another example, such as Figure 4 As shown, the word representation vector sequence of the word and phrase text corpus is added to the first type of labeled vector sequence to obtain the updated word and phrase representation vector sequence; the pronunciation representation vector sequence of the word and phrase text corpus is added to the second type of labeled vector sequence to obtain the updated pronunciation representation vector sequence; the first type of labeled vector sequence and the second type of labeled vector sequence are used to distinguish between the word and phrase representation vector sequences and the pronunciation representation vector sequences of the word and phrase text corpus; the updated word and phrase representation vector sequences and the updated pronunciation representation vector sequences are concatenated to obtain the concatenated vector sequence; the concatenated vector sequence is processed by a semantic feature extraction model to obtain the fused semantic features of the word and phrase text corpus.
[0071] Suppose a text corpus contains n characters, where the word representation vector of the i-th character is Ci, and the pronunciation representation vector of the i-th character is PYi, where n is an integer greater than 1 and i is a positive integer less than or equal to n. Then, the word representation vector sequence of the text corpus can be represented as {C1, C2, ..., Cn}, and the pronunciation representation vector sequence can be represented as {PY1, PY2, ..., PYn}. Assume the first type of annotation vector is B1, and the second type of annotation vector is B2. B1 and B2 can be two vectors with discriminative power; for example, B1 is a vector with all zeros, and B2 is a vector with all one elements. Adding the first type of annotation vector sequence to the word representation vector sequence {C1, C2, ..., Cn} of the text corpus yields the updated word representation vector sequence {C1+B1, C2+B1, ..., Cn+B1}. The pronunciation representation vector sequence {PY1,PY2,…,PYn} of the word and phrase text corpus is added to the second type of annotation vector sequence to obtain the updated pronunciation representation vector sequence {PY1+B2,PY2+B2,…,PYn+B2}. Finally, the updated word and phrase representation vector sequence {C1+B1,C2+B1,…,Cn+B1} and the updated pronunciation representation vector sequence {PY1+B2,PY2+B2,…,PYn+B2} are concatenated to obtain the concatenated vector sequence [{C1+B1,C2+B1,…,Cn+B1},{PY1+B2,PY2+B2,…,PYn+B2}].
[0072] In another example, such as Figure 5 As shown, the semantic feature extraction model includes a first extraction sub-model and a second extraction sub-model. The first extraction sub-model extracts semantic features of words from the word representation vector sequence of the word text corpus; the second extraction sub-model extracts pronunciation semantic features from the pronunciation representation vector sequence of the word text corpus; the word semantic features and pronunciation semantic features are fused to obtain the fused semantic features of the word text corpus.
[0073] The first and second extraction sub-models can be two neural network models with identical structures. The first extraction sub-model is used to extract features from the word representation vector sequence, and the second extraction sub-model is used to extract features from the pronunciation representation vector sequence. This approach uses a post-fusion method, that is, firstly, the two sub-models extract word semantic features and pronunciation semantic features from the word representation vector sequence and pronunciation representation vector sequence of the word text corpus, respectively, and then the two aspects of features are fused to obtain the fused semantic features of the word text corpus.
[0074] Optionally, a self-attention mechanism can be used to fuse word semantic features and pronunciation semantic features to obtain fused semantic features of the word text corpus. Using a self-attention mechanism for feature fusion can fully extract important features and improve the semantic representation capability of the fused semantic features.
[0075] It should be noted that, in the embodiments of this application, the model structure of the semantic feature extraction model is not limited. It can be any neural network model structure that can process sequence information, such as the Transformer-Encoder structure, or it can be other model structures such as LSTM (Long Short-Term Memory) network, RNN (Recurrent Neural Network).
[0076] Step 240: Based on the fused semantic features of the word and text corpus, determine the prediction results corresponding to the pre-training task of the semantic feature extraction model.
[0077] Pre-training tasks are defined tasks used to pre-train the semantic feature extraction model. The number of pre-training tasks can be one or multiple, depending on a comprehensive consideration of the model training complexity and accuracy requirements, to design an appropriate number of pre-training tasks and their content.
[0078] In one example, the pre-training task includes a masked word prediction task. This task involves masking a portion of the words in the input text corpus (e.g., replacing them with [MASK] tags) and designing a corresponding word prediction network. This network, combined with the contextual information of the unmasked words in the text corpus, predicts and infers the masked words. Optionally, a certain proportion of the words in the text corpus are replaced with [MASK] tags. A semantic feature extraction model extracts the fused semantic features of the text corpus from the word representation vector sequence and pronunciation representation vector sequence containing the [MASK] tags. Then, the word prediction network determines the prediction result of the masked words in the text corpus based on the fused semantic features. Subsequently, the pre-training loss of the model can be calculated based on the difference between the predicted and actual results of the masked words. It should be noted that in this example, when a word in the text corpus is masked and replaced with a [MASK] marker, the corresponding pronunciation annotation information can also be masked and replaced with a [MASK] marker, or it can remain unmasked (i.e., the pronunciation annotation information corresponding to the word is still input into the model). Experiments have shown that masking both words and their corresponding pronunciation annotation information can, to some extent, help improve the model's learning ability.
[0079] In another example, the pre-training task includes sentence order prediction. Sentence order prediction involves taking two consecutive sentences from a data source, inputting them into the model in either normal or reverse order, and designing a corresponding order prediction network to predict whether the two sentences are in normal or reverse order. If the input is in normal order, the model's prediction objective is 1; otherwise, if the input is in reverse order, the prediction objective is 0. This task enables the model to capture contextual semantic relationships based on the input information.
[0080] Optionally, for two consecutive word-text corpora in the data source (denoted as the first word-text corpus and the second word-text corpus), assuming the normal order of these two word-text corpora is the first word-text corpus first and the second word-text corpus last, then the reversed order is the second word-text corpus first and the first word-text corpus last. For example, if the data source includes "The weather is sunny, suitable for travel", the first word-text corpus is "The weather is sunny", and the second word-text corpus is "Suitable for travel", then if these two word-text corpora are input into the model in the form "The weather is sunny, suitable for travel", then they are input into the model in the normal order. If these two word-text corpora are input into the model in the form "Suitable for travel, sunny weather", then they are input into the model in the reversed order. After extracting fused semantic features from the two word-based text corpora using a semantic feature extraction model, a sequence prediction network is used to predict the sentence order of the first and second word-based text corpora based on the fused semantic features of the first and second word-based text corpora. Subsequently, the pre-training loss of the model can be calculated based on the difference between the predicted and actual sentence order results.
[0081] In some embodiments, the pre-training task of the semantic feature extraction model may include both the masked word prediction task and the sentence order prediction task, so that the semantic feature extraction model can capture not only the contextual information in a single sentence, but also the contextual semantic relationship between adjacent sentences, thereby improving the model's ability to extract and represent semantic features.
[0082] Step 250: Determine the pre-training loss of the semantic feature extraction model based on the prediction results and the actual results corresponding to the pre-training task, and adjust the parameters of the semantic feature extraction model according to the pre-training loss to obtain the pre-trained semantic feature extraction model.
[0083] Optionally, based on the pre-training loss of the semantic feature extraction model, the model parameters are adjusted using gradient descent so that the pre-training loss gradually approaches the optimization target until a pre-set stopping condition is reached (such as the pre-training loss being less than a set threshold or reaching a minimum value), thus completing the pre-training process of the semantic feature extraction model and obtaining the pre-trained semantic feature extraction model.
[0084] Optionally, such as Figure 6 As shown, to reduce the difficulty of model training, this application proposes to divide the pre-training process of the semantic feature extraction model into multiple stages, such as a first stage, a second stage, and a third stage. Specifically, the first stage is used to train the word representation vector generation network, the second stage is used to train the pronunciation representation vector generation network, and the third stage is used to train the semantic feature extraction model.
[0085] In the first stage, only the text corpus containing words is input into the model's word representation vector generation network. At this stage, the phonetic annotation information of the text corpus is not included. The main purpose of training in this stage is to optimize the parameters of the word representation vector generation network, enabling it to generate better word representation vectors. In some other examples, if an open-source pre-trained word representation vector generation network is used directly, this training stage can be omitted.
[0086] In the second stage, the pronunciation annotation information of the word and phrase text corpus is input into the pronunciation representation vector generation network of the model. The word and phrase representation vectors of the text corpus are initialized using the results obtained in the first stage. In this training stage, the parameters of the word and phrase representation vector generation network are not updated; only the parameters of the pronunciation representation vector generation network are updated. The main goal is to adjust the parameters of the pronunciation representation vector generation network to a more optimal state so that it can generate better pronunciation representation vectors.
[0087] In the third stage, the word and text corpus and its pronunciation annotation information are input into the model. The word representation vector generation network is initialized with the training results of the first stage, and the pronunciation representation vector generation network is initialized with the training results of the second stage. Then, the entire model (including the word representation vector generation network, the pronunciation representation vector generation network and the semantic feature extraction model) is pre-trained globally to improve the overall training effect of the model.
[0088] After the above three stages, the semantic feature extraction model has the ability to capture word features and pronunciation features simultaneously for representation modeling. In subsequent use, it can be fine-tuned for specific tasks.
[0089] In summary, the technical solution provided in this application, when extracting semantic features from word and phrase text corpora, not only uses the word and phrase features (i.e., word and phrase representation vector sequences) of the word and phrase text corpora, but also uses the pronunciation features (i.e., pronunciation representation vector sequences) of the word and phrase text corpora. Compared with related technologies that only consider word and phrase features, the technical solution of this application enables the semantic feature extraction model to capture both word and phrase features and pronunciation features, making full use of multiple features to enhance the semantic representation capability of the model.
[0090] In addition, the embodiments of this application provide a variety of ways to fuse word features and pronunciation features. Some methods (such as simple feature splicing) are relatively simple to implement and have a small amount of computation, while others (such as post-fusion based on attention mechanism) are more complex to implement but can fully extract important features and further improve the semantic representation capability of fused semantic features.
[0091] In an exemplary embodiment, the pre-trained semantic feature extraction model described above can be applied to the target task, working in conjunction with the result prediction network for that target task to predict the outcome. When applied to the target task, the parameters of the pre-trained semantic feature extraction model can be fine-tuned using training samples from the target task, making the semantic feature extraction model adaptable to the target task and possessing good semantic representation capabilities under that task. Optionally, as... Figure 7 As shown, the model fine-tuning process may include the following steps (710-750):
[0092] Step 710: Obtain training samples for the target task. The training samples include word text samples of the target language and pronunciation annotation information of the word text samples.
[0093] The target task can be any task that requires the use of semantic features of text to predict results. For example, in the field of video comment recognition, the target task could be a video comment sentiment recognition task, a video comment poetry recognition task, etc., and this application does not limit it in this regard.
[0094] The video comment sentiment recognition task refers to automatically identifying the sentiment tendency of user comments on videos (referred to as "video comments") using a model, i.e., what sentiment category they belong to. For example, sentiment categories can be pre-defined as positive and negative (two categories), or pre-defined as positive, neutral, and negative (three categories). This application does not specifically limit the number or types of sentiment categories; this can be set according to actual needs. By performing sentiment recognition on video comments, it is possible to uncover users' emotional tendencies towards video content (such as the characters in the video) and to infer public opinion.
[0095] The task of identifying poetic video comments involves using a model to automatically determine whether user comments on videos (referred to as "video comments") are in poetic form. Poetic form refers to text content that resembles poetry in form and rhythm, such as having the same number of words per sentence and rhyming. Poetic video comments are generally considered highlight comments. By mining poetic video comments from a large pool of comments and prioritizing their display, the task can enhance the community atmosphere.
[0096] Of course, the above are merely illustrative examples of two forms of target tasks. In other application scenarios, such as in the field of input methods (e.g., Chinese input methods), the target task could be a word association task, predicting associated words based on the character and pronunciation features of the user's input words; or, in intelligent question answering scenarios, the target task could be intelligent reply, automatically generating response information corresponding to the user's input question based on the character and pronunciation features of the input question. Naturally, it can also be applied to other application scenarios, which will not be listed here.
[0097] Step 720: Obtain the word representation vector sequence of the word text sample and the pronunciation representation vector sequence of the pronunciation annotation information of the word text sample.
[0098] The specific implementation process of this step is the same as or similar to the specific implementation process of step 220 in the above embodiment. For details, please refer to the description in the above embodiment, which will not be repeated here.
[0099] Step 730: Using the pre-trained semantic feature extraction model, extract the fused semantic features of the word text sample from the word representation vector sequence and pronunciation representation vector sequence of the word text sample.
[0100] The specific implementation process of this step is the same as or similar to the specific implementation process of step 230 in the above embodiment. For details, please refer to the description in the above embodiment, which will not be repeated here.
[0101] Step 740: The result prediction network for the target task determines the task prediction result corresponding to the word text sample based on the fused semantic features of the word text sample.
[0102] The result prediction network for the target task is used to predict the task outcome. Optionally, this result prediction network can be a classification network, and the number of its output categories can be set according to the classification requirements of the target task. For example, when the target task is video comment sentiment recognition, the number of output categories of the result prediction network can be two, corresponding to the two sentiment categories of positive and negative.
[0103] Step 750: Determine the model training loss based on the task prediction results and actual task results corresponding to the word and text samples, and adjust the parameters of the pre-trained semantic feature extraction model and result prediction network according to the model training loss.
[0104] In this embodiment, the combination of the semantic feature extraction model and the result prediction network can be referred to as the target task model. Optionally, based on the difference between the task prediction result and the actual task result corresponding to the word text sample, the model training loss of the target task model is calculated, and the model parameters (including the parameters of the semantic feature extraction model and the result prediction network) are adjusted using gradient descent, so that the model training loss gradually approaches the optimization target until a pre-set stopping condition is reached (such as the model training loss being less than a set threshold or reaching a minimum value), thus completing the training process of the model on the target task and obtaining the trained target task model. The trained target task model can be deployed online to perform prediction tasks for the target task.
[0105] In summary, by fine-tuning the parameters of the pre-trained semantic feature extraction model using training samples from the target task, the semantic feature extraction model is adapted to the target task and has good semantic representation capabilities under the target task.
[0106] In an exemplary embodiment, taking Chinese as the target language and Pinyin as the pronunciation annotation information as an example, the technical solution of this application will be described and explained. Figure 8 As shown, the method may include the following steps:
[0107] Step 802: Obtain the training corpus for the semantic feature extraction model. The training corpus includes Chinese word and phrase text corpus and the pinyin annotation information of the word and phrase text corpus.
[0108] Step 804: Obtain the word representation vector sequence of the word and phrase text corpus, and the pinyin representation vector sequence of the pinyin annotation information of the word and phrase text corpus.
[0109] Step 806: Using a semantic feature extraction model, extract the fused semantic features of the word text corpus from the word representation vector sequence and the pinyin representation vector sequence.
[0110] Step 808: Based on the fusion of semantic features from the word and text corpus, determine the prediction results corresponding to the pre-training task of the semantic feature extraction model.
[0111] Step 810: Determine the pre-training loss of the semantic feature extraction model based on the prediction results and the actual results corresponding to the pre-training task, and adjust the parameters of the semantic feature extraction model according to the pre-training loss to obtain the pre-trained semantic feature extraction model.
[0112] Optionally, the method may further include the following steps: applying the pre-trained semantic feature extraction model to the target task and fine-tuning the model parameters.
[0113] Step 812: Obtain training samples for the target task. The training samples include Chinese word text samples and their pinyin annotation information.
[0114] Step 814: Obtain the word representation vector sequence of the word text sample and the pinyin representation vector sequence of the pinyin annotation information of the word text sample.
[0115] Step 816: Using the pre-trained semantic feature extraction model, extract the fused semantic features of the word text sample from the word representation vector sequence and the pinyin representation vector sequence of the word text sample.
[0116] Step 818: The result prediction network for the target task determines the task prediction result corresponding to the word text sample based on the fused semantic features of the word text sample.
[0117] Step 820: Determine the model training loss based on the task prediction results and actual task results corresponding to the word and text samples, and adjust the parameters of the pre-trained semantic feature extraction model and result prediction network according to the model training loss to obtain the trained target task model.
[0118] In summary, this application provides a method to enhance the semantic representation capability of a semantic feature extraction model for Chinese based on pinyin features. By introducing pinyin features into the semantic feature extraction model for Chinese and combining pinyin features with word features, the model can fully represent and model Chinese, thereby improving its representation capability, especially for Chinese tasks that depend on pronunciation.
[0119] Furthermore, adopting the technical solution of this application can reduce the difficulty of model learning, making model learning more consistent with the objective logical process of human language learning. Taking Chinese as an example, the process of people learning Chinese also starts with learning pinyin, gradually learning from sound to characters, to words, and to sentences. Introducing pinyin features into the model makes the model learning process more consistent with the human learning process.
[0120] Furthermore, in daily communication, normal communication can be achieved solely through pronunciation without writing Chinese characters, indicating that Pinyin can largely reflect the characteristics of Chinese words and phrases. Introducing Pinyin into the model can enhance its ability to represent Chinese. In addition, for data containing misspelled characters but with correct pronunciation, Pinyin features can also compensate for the information loss caused by errors in character form to some extent.
[0121] Furthermore, this application can integrate the representation of simplified and traditional Chinese. Representing Chinese only through characters and words results in different representations of the same characters and words in simplified and traditional Chinese, but the pronunciation of simplified and traditional Chinese can bring the two representations closer together.
[0122] Furthermore, this application can reduce model size and improve inference speed. In Chinese character representation, there are generally 20,000+ characters and 200,000+ words to consider, but Pinyin only has 400+ characters. If only Pinyin is used to represent Chinese, it can also approximate the effect of character and word representation models, but the vector dimension, hidden layer dimension, etc. can be reduced, thereby reducing model size, reducing computational load, and improving inference speed.
[0123] Experiments revealed that training a semantic feature extraction model for Chinese using the technical solution of this application, which integrates word and pinyin features, achieves a significant improvement in prediction accuracy. Compared to training methods that only consider word features without considering pinyin features, the word-pinyin fusion method (Method 1) improves the model's accuracy by 0.4+ percentage points, Method 2 by 0.7+ percentage points, and Method 3 by 0.9+ percentage points. Therefore, the experimental data clearly demonstrates that introducing pinyin features enhances the semantic representation capability of the semantic feature extraction model for Chinese.
[0124] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0125] Please refer to Figure 9 The diagram illustrates a block diagram of a training apparatus for a semantic feature extraction model provided in one embodiment of this application. The apparatus 900 may include: a corpus acquisition module 902, a sequence acquisition module 904, a feature extraction module 906, a result determination module 908, and a parameter adjustment module 910.
[0126] The corpus acquisition module 902 is used to acquire the training corpus of the semantic feature extraction model, wherein the training corpus includes word and text corpus of the target language and pronunciation annotation information of the word and text corpus.
[0127] The sequence acquisition module 904 is used to acquire the word representation vector sequence of the word text corpus, and the pronunciation representation vector sequence of the pronunciation annotation information of the word text corpus.
[0128] The feature extraction module 906 is used to extract the fused semantic features of the word and phrase text corpus from the word representation vector sequence and pronunciation representation vector sequence of the word and phrase text corpus through the semantic feature extraction model.
[0129] The result determination module 908 is used to determine the prediction result corresponding to the pre-training task of the semantic feature extraction model based on the fused semantic features of the word and text corpus.
[0130] The parameter adjustment module 910 is used to determine the pre-training loss of the semantic feature extraction model based on the prediction results and the actual results corresponding to the pre-training task, and adjust the parameters of the semantic feature extraction model according to the pre-training loss to obtain the pre-trained semantic feature extraction model.
[0131] In an exemplary embodiment, the feature extraction module 906 is used to:
[0132] The word representation vector sequence and pronunciation representation vector sequence of the word text corpus are fused to obtain the fused representation vector sequence of the word text corpus;
[0133] The semantic feature extraction model is used to extract features from the fused representation vector sequence of the word and phrase text corpus to obtain the fused semantic features of the word and phrase text corpus.
[0134] Optionally, the feature extraction module 906 is used for:
[0135] The word representation vector sequence and pronunciation representation vector sequence of the word text corpus are averaged to obtain the fused representation vector sequence of the word text corpus.
[0136] or,
[0137] The word representation vector sequence and pronunciation representation vector sequence of the word text corpus are concatenated to obtain the fused representation vector sequence of the word text corpus.
[0138] or,
[0139] The word representation vector sequence and pronunciation representation vector sequence of the word and text corpus are input into the word and pronunciation fusion network; the word and pronunciation representation vectors corresponding to the same word position in the word and text corpus are weighted and summed by the word and pronunciation fusion network to obtain the fused representation vector sequence of the word and text corpus.
[0140] In an exemplary embodiment, the feature extraction module 906 is used to:
[0141] Add the first type of labeled vector sequence to the word representation vector sequence of the word text corpus to obtain the updated word representation vector sequence.
[0142] The pronunciation representation vector sequence of the word and phrase text corpus is added to the second type of annotation vector sequence to obtain the updated pronunciation representation vector sequence; wherein, the first type of annotation vector sequence and the second type of annotation vector sequence are used to distinguish between the word and phrase representation vector sequence and the pronunciation representation vector sequence of the word and phrase text corpus;
[0143] The updated word representation vector sequence and the updated pronunciation representation vector sequence are concatenated to obtain a concatenated vector sequence;
[0144] The semantic feature extraction model is used to extract features from the concatenated vector sequence to obtain the fused semantic features of the word and text corpus.
[0145] In an exemplary embodiment, the semantic feature extraction model includes a first extraction sub-model and a second extraction sub-model; the feature extraction module 906 is used for:
[0146] The first extraction sub-model extracts semantic features of words from the word representation vector sequence of the word text corpus.
[0147] The second extraction sub-model extracts pronunciation semantic features from the pronunciation representation vector sequence of the word and phrase text corpus.
[0148] The semantic features of the words and the semantic features of the pronunciation are fused to obtain the fused semantic features of the word and text corpus.
[0149] Optionally, the feature extraction module 906 is used to perform fusion processing on the word semantic features and the pronunciation semantic features using a self-attention mechanism to obtain the fused semantic features of the word text corpus.
[0150] In an exemplary embodiment, the sequence acquisition module 904 is configured to:
[0151] The word representation vector is generated by the word representation vector generation network to generate the word representation vector of each word in the word text corpus, thus obtaining the word representation vector sequence;
[0152] A pronunciation representation vector generation network is used to generate pronunciation representation vectors for each word in the text corpus, resulting in a sequence of pronunciation representation vectors.
[0153] Optionally, the pre-training process of the semantic feature extraction model includes a first stage, a second stage, and a third stage; wherein the first stage is used to train the word representation vector generation network, the second stage is used to train the pronunciation representation vector generation network, and the third stage is used to train the semantic feature extraction model.
[0154] In an exemplary embodiment, the result determination module 908 is configured to:
[0155] Based on the fused semantic features of the word and text corpus, the word prediction network determines the prediction results of the masked words in the word and text corpus.
[0156] And / or,
[0157] The sequence prediction network determines the prediction result of the sentence order of the first word text corpus and the second word text corpus based on the fused semantic features of the first word text corpus and the fused semantic features of the second word text corpus.
[0158] In an exemplary embodiment, the corpus acquisition module 902 is further configured to acquire training samples for the target task, the training samples including word and text samples of the target language and pronunciation annotation information of the word and text samples.
[0159] The sequence acquisition module 904 is further configured to acquire the word representation vector sequence of the word text sample and the pronunciation representation vector sequence of the pronunciation annotation information of the word text sample.
[0160] The feature extraction module 906 is further configured to extract the fused semantic features of the word text sample from the word representation vector sequence and pronunciation representation vector sequence of the word text sample using the pre-trained semantic feature extraction model.
[0161] The result determination module 908 is further configured to determine the task prediction result corresponding to the word text sample based on the fused semantic features of the word text sample through the result prediction network of the target task.
[0162] The parameter adjustment module 910 is further configured to determine the model training loss based on the task prediction result and the actual task result corresponding to the word text sample, and adjust the parameters of the pre-trained semantic feature extraction model and the result prediction network according to the model training loss.
[0163] In an exemplary embodiment, the corpus acquisition module 902 is configured to:
[0164] Obtain the word and text corpus of the target language from the data source;
[0165] The phonetic annotation information of the text corpus is generated using a phonetic annotation model.
[0166] In an exemplary embodiment, when the target language is Chinese and the pronunciation annotation information is pinyin annotation information:
[0167] The corpus acquisition module 902 is used to acquire the training corpus of the semantic feature extraction model, wherein the training corpus includes Chinese word and phrase text corpus and the pinyin annotation information of the word and phrase text corpus.
[0168] The sequence acquisition module 904 is used to acquire the word representation vector sequence of the word text corpus, and the pinyin representation vector sequence of the pinyin annotation information of the word text corpus.
[0169] The feature extraction module 906 is used to extract the fused semantic features of the word and phrase text corpus from the word and phrase representation vector sequence and the pinyin representation vector sequence of the word and phrase text corpus through the semantic feature extraction model.
[0170] The result determination module 908 is used to determine the prediction result corresponding to the pre-training task of the semantic feature extraction model based on the fused semantic features of the word and text corpus.
[0171] The parameter adjustment module 910 is used to determine the pre-training loss of the semantic feature extraction model based on the prediction results and the actual results corresponding to the pre-training task, and adjust the parameters of the semantic feature extraction model according to the pre-training loss to obtain the pre-trained semantic feature extraction model.
[0172] In summary, the technical solution provided in this application, when extracting semantic features from word and phrase text corpora, not only uses the word and phrase features (i.e., word and phrase representation vector sequences) of the word and phrase text corpora, but also uses the pronunciation features (i.e., pronunciation representation vector sequences) of the word and phrase text corpora. Compared with related technologies that only consider word and phrase features, the technical solution of this application enables the semantic feature extraction model to capture both word and phrase features and pronunciation features, making full use of multiple features to enhance the semantic representation capability of the model.
[0173] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0174] Please refer to Figure 10This illustration shows a schematic diagram of a computer device according to an embodiment of this application. The computer device can be any electronic device with data computing, processing, and storage functions, and can be implemented as... Figure 1 The illustrated scheme implements a model training device 10 and / or a model usage device 20 in the implementation environment. This computer device is implemented as... Figure 1 When the model training device 10 is used in the implementation environment of the scheme shown, the computer device can be used to implement the training method of the semantic feature extraction model provided in the above embodiments. Specifically:
[0175] The computer device 1000 includes a central processing unit (such as a CPU, GPU, or FPGA) 1001, a system memory 1004 including RAM (Random-Access Memory) 1002 and ROM (Read-Only Memory) 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The computer device 1000 also includes a basic input / output system (I / O system) 1006 to facilitate information transfer between various devices within the server, and a large-capacity storage device 1007 for storing the operating system 1013, application programs 1014, and other program modules 1015.
[0176] In some embodiments, the basic input / output system 1006 includes a display 1008 for displaying information and an input device 1009 for user input, such as a mouse or keyboard. Both the display 1008 and the input device 1009 are connected to the central processing unit 1001 via an input / output controller 1010 connected to the system bus 1005. The basic input / output system 1006 may also include the input / output controller 1010 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, printer, or other types of output devices.
[0177] The mass storage device 1007 is connected to the central processing unit 1001 via a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable media provide non-volatile storage for the computer device 1000. That is, the mass storage device 1007 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0178] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage medium is not limited to the above-mentioned types. The system memory 1004 and mass storage device 1007 described above can be collectively referred to as memory.
[0179] According to an embodiment of this application, the computer device 1000 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 1000 can be connected to the network 1012 via the network interface unit 1011 connected to the system bus 1005, or the network interface unit 1011 can be used to connect to other types of networks or remote computer systems (not shown).
[0180] The memory also includes at least one instruction, at least one program, code set, or instruction set, which is stored in the memory and configured to be executed by one or more processors to implement the training method of the semantic feature extraction model described above.
[0181] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set implements the above-described training method for the semantic feature extraction model when executed by a processor of a computer device.
[0182] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0183] In an exemplary embodiment, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the training method for the semantic feature extraction model described above.
[0184] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0185] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A training method for a semantic feature extraction model, characterized in that, The method includes: The training corpus of the semantic feature extraction model is obtained. The training corpus includes word and text corpus of the target language obtained from the data source and pronunciation annotation information of the word and text corpus generated based on the word and text corpus. The semantic feature extraction model is used to model and represent the text to perform the target task. The target task is to predict the result using the semantic features of the text. Obtain the word representation vector sequence of the word text corpus obtained by the word representation vector generation network, and the pronunciation representation vector sequence of the pronunciation annotation information of the word text corpus obtained by the pronunciation representation vector generation network; The semantic feature extraction model includes a first extraction sub-model that extracts semantic features of words from the word representation vector sequence of the word text corpus. The semantic feature extraction model includes a second extraction sub-model that extracts pronunciation semantic features from the pronunciation representation vector sequence of the word and phrase text corpus. A self-attention mechanism is used to fuse the semantic features of the words and the semantic features of the pronunciation to obtain the fused semantic features of the word text corpus; Based on the fused semantic features of the word and text corpus, the prediction result corresponding to the pre-training task of the semantic feature extraction model is determined; The pre-training loss of the semantic feature extraction model is determined based on the prediction results and the actual results corresponding to the pre-training task, and the parameters of the semantic feature extraction model are adjusted according to the pre-training loss to obtain the pre-trained semantic feature extraction model.
2. The method according to claim 1, characterized in that, The method further includes: The word representation vector sequence and pronunciation representation vector sequence of the word text corpus are fused to obtain the fused representation vector sequence of the word text corpus; The fused semantic features of the word and phrase text corpus are obtained by performing feature extraction processing on the fused representation vector sequence of the word and phrase text corpus through the semantic feature extraction model.
3. The method according to claim 2, characterized in that, The process of fusing the word representation vector sequence and pronunciation representation vector sequence of the word and text corpus to obtain the fused representation vector sequence of the word and text corpus includes: The word representation vector sequence and pronunciation representation vector sequence of the word text corpus are averaged to obtain the fused representation vector sequence of the word text corpus. or, The word representation vector sequence and pronunciation representation vector sequence of the word text corpus are concatenated to obtain the fused representation vector sequence of the word text corpus. or, The word representation vector sequence and pronunciation representation vector sequence of the word and text corpus are input into the word and pronunciation fusion network; the word and pronunciation representation vectors corresponding to the same word position in the word and text corpus are weighted and summed by the word and pronunciation fusion network to obtain the fused representation vector sequence of the word and text corpus.
4. The method according to claim 1, characterized in that, The method further includes: Add the first type of labeled vector sequence to the word representation vector sequence of the word text corpus to obtain the updated word representation vector sequence. The pronunciation representation vector sequence of the word and phrase text corpus is added to the second type of annotation vector sequence to obtain the updated pronunciation representation vector sequence; wherein, the first type of annotation vector sequence and the second type of annotation vector sequence are used to distinguish between the word and phrase representation vector sequence and the pronunciation representation vector sequence of the word and phrase text corpus; The updated word representation vector sequence and the updated pronunciation representation vector sequence are concatenated to obtain a concatenated vector sequence; The fused semantic features of the word and text corpus are obtained by performing feature extraction processing on the concatenated vector sequence through the semantic feature extraction model.
5. The method according to claim 1, characterized in that, The step of obtaining the word representation vector sequence of the word text corpus obtained through the word representation vector generation network, and the pronunciation representation vector sequence of the word text corpus obtained through the pronunciation representation vector generation network, includes: The word representation vector generation network generates word representation vectors for each word in the word text corpus, resulting in the word representation vector sequence. The pronunciation representation vector generation network generates pronunciation representation vectors for each word in the text corpus, resulting in the pronunciation representation vector sequence.
6. The method according to claim 1, characterized in that, The pre-training process of the semantic feature extraction model includes a first stage, a second stage, and a third stage; wherein, the first stage is used to train the word representation vector generation network, the second stage is used to train the pronunciation representation vector generation network, and the third stage is used to train the semantic feature extraction model.
7. The method according to claim 1, characterized in that, The step of determining the prediction result corresponding to the pre-training task of the semantic feature extraction model based on the fused semantic features of the word and text corpus includes: Based on the fused semantic features of the word and text corpus, the word prediction network determines the prediction results of the masked words in the word and text corpus. And / or, The sequence prediction network determines the prediction result of the sentence order of the first word text corpus and the second word text corpus based on the fused semantic features of the first word text corpus and the fused semantic features of the second word text corpus.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Obtain training samples for the target task, the training samples including word and text samples of the target language and pronunciation annotation information of the word and text samples; Obtain the word representation vector sequence of the word text sample, and the pronunciation representation vector sequence of the pronunciation annotation information of the word text sample; The semantic feature extraction model, which has been pre-trained, extracts the fused semantic features of the word text sample from the word representation vector sequence and the pronunciation representation vector sequence of the word text sample. The result prediction network for the target task determines the task prediction result corresponding to the word text sample based on the fused semantic features of the word text sample. The model training loss is determined based on the task prediction results and actual task results corresponding to the word and text samples, and the parameters of the pre-trained semantic feature extraction model and the result prediction network are adjusted according to the model training loss.
9. The method according to any one of claims 1 to 7, characterized in that, The step of obtaining the training corpus for the semantic feature extraction model includes: Obtain the word and text corpus of the target language from the data source; The phonetic annotation information of the text corpus is generated using a phonetic annotation model.
10. The method according to any one of claims 1 to 7, characterized in that, When the target language is Chinese and the pronunciation annotation information is pinyin annotation information, the method includes: Obtain the training corpus of the semantic feature extraction model, the training corpus including Chinese word and phrase text corpus and the pinyin annotation information of the word and phrase text corpus; Obtain the word representation vector sequence of the word and phrase text corpus, and the pinyin representation vector sequence of the pinyin annotation information of the word and phrase text corpus; The semantic feature extraction model extracts the fused semantic features of the word and phrase text corpus from the word representation vector sequence and the pinyin representation vector sequence. Based on the fused semantic features of the word and text corpus, the prediction result corresponding to the pre-training task of the semantic feature extraction model is determined; The pre-training loss of the semantic feature extraction model is determined based on the prediction results and the actual results corresponding to the pre-training task, and the parameters of the semantic feature extraction model are adjusted according to the pre-training loss to obtain the pre-trained semantic feature extraction model.
11. A training device for a semantic feature extraction model, characterized in that, The device includes: The corpus acquisition module is used to acquire the training corpus of the semantic feature extraction model. The training corpus includes word and text corpus of the target language obtained from the data source and pronunciation annotation information of the word and text corpus generated based on the word and text corpus. The semantic feature extraction model is used to model and represent the text to perform the target task. The target task is a task of predicting the result using the semantic features of the text. The sequence acquisition module is used to acquire the word representation vector sequence of the word text corpus obtained by the word representation vector generation network, and the pronunciation representation vector sequence of the pronunciation annotation information of the word text corpus obtained by the pronunciation representation vector generation network; The feature extraction module is used to extract semantic features of words from the word representation vector sequence of the word text corpus through the first extraction sub-model included in the semantic feature extraction model; The feature extraction module is further configured to extract pronunciation semantic features from the pronunciation representation vector sequence of the word and text corpus through the second extraction sub-model included in the semantic feature extraction model; The feature extraction module is further configured to use a self-attention mechanism to fuse the word semantic features and the pronunciation semantic features to obtain the fused semantic features of the word text corpus. The result determination module is used to determine the prediction result corresponding to the pre-training task of the semantic feature extraction model based on the fused semantic features of the word and text corpus. The parameter adjustment module is used to determine the pre-training loss of the semantic feature extraction model based on the prediction results and the actual results corresponding to the pre-training task, and to adjust the parameters of the semantic feature extraction model according to the pre-training loss to obtain the pre-trained semantic feature extraction model.
12. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the training method for the semantic feature extraction model as described in any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The storage medium stores at least one program segment, which is loaded and executed by a processor to implement the training method of the semantic feature extraction model as described in any one of claims 1 to 10.
14. A computer program product, characterized in that, The computer program product includes computer instructions that are executed by a processor to implement the training method for the semantic feature extraction model as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Voice dialogue processing method and system
CN111862977A