Language Model Training, Video Subtitle Verification Method, Device, Equipment and Medium
By splitting Chinese characters into radical structures, a Chinese pre-trained language model based on character disassembly is trained, which solves the problems of large model size, slow inference speed and low recognition accuracy in the prior art, and achieves more efficient and faster language model recognition.
Patent Information
- Application Number
- CN202011529805.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-12-22
AI Technical Summary
The existing pre-trained language model based on words is large in size, slow in reasoning when training smaller models, and in application scenarios where words are not rigorous and wrong vocabulary are high in sensitivity but low robustness, resulting in low recognition accuracy.
By splitting Chinese characters into radical structures, a Chinese pre-trained language model based on character disassembly is trained. The model can receive the internal characteristics of the characters, improve the representation ability, and reduce the model volume through smaller vocabulary parameters (500-2500) and improve the recognition speed.
It improves the recognition accuracy of language models, reduces the time for model training and recognition, and is suitable for small model scenarios that require fast training and efficient recognition.
Smart Images

Figure CN112652295B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device and medium for language model training and video subtitle verification. Background Art
[0002] With the development of science and technology, the field of artificial intelligence has also developed faster and faster. In scenarios such as character recognition and text verification, pre-trained language models based on words and characters are often used to preprocess text or text.
[0003] In the prior art, the pre-trained language model based on words and characters used in scenarios such as character recognition and text verification has a relatively large overall vocabulary (usually exceeding twenty thousand). Although this vocabulary contains a large number of words, it will cause the pre-trained language model to be large in size and slow in inference speed, so it is not suitable for training smaller models. For example, in advertising character recognition, only a smaller model needs to be trained so that the model can recognize the words and characters in the advertising language. If the model in the prior art is used for training, it will cause the trained model to have too many parameters, and then cause the model to have a large amount of calculation during the recognition process, resulting in slow recognition speed. Secondly, in some specific application scenarios where the words are not rigorous and there are many incorrect words, the existing pre-trained language models based on words and characters are more sensitive to words and characters but have lower robustness, so the accuracy of preprocessing text or text will be lower. Summary of the Invention
[0004] Embodiments of the present invention provide a method, apparatus, device and medium for language model training and video subtitle verification to improve the accuracy of language model recognition.
[0005] A method for training a language model includes:
[0006] Obtain a text sample set and an initial character-splitting pre-trained model with initial parameters, where the text sample set includes at least one sample sentence, and one sample sentence includes at least one Chinese character; the initial character-splitting pre-trained model includes a character encoding model and a character decoding model;
[0007] When the sample sentence only contains Chinese characters, input the sample sentence into the initial character-splitting pre-trained model, and perform word segmentation processing on the sample sentence through the character encoding model to obtain each Chinese sample word in the sample sentence;
[0008] Perform radical splitting on all Chinese characters in each Chinese sample word through the character encoding model to obtain the radical decomposition result of each Chinese character;
[0009] Perform granularity splitting on all the radical decomposition results through the character encoding model to obtain a splitting result;
[0010] Perform decoding and recognition on the splitting result through the character decoding model to obtain a sample decoded sentence;
[0011] Determine a text loss value according to the sample decoded sentence and the sample sentence containing only Chinese characters;
[0012] When the text loss value does not reach a preset convergence condition, update and iterate the initial parameters of the initial Chinese character splitting pre-training model until the text loss value reaches the preset convergence condition, and record the initial Chinese character splitting pre-training model after convergence as a Chinese pre-training language model based on Chinese character splitting.
[0013] A language model training device, comprising:
[0014] A data acquisition module, configured to acquire a set of text samples and an initial Chinese character splitting pre-training model with initial parameters, where the set of text samples includes at least one sample sentence, and one sample sentence includes at least one Chinese character; the initial Chinese character splitting pre-training model includes a character encoding model and a character decoding model;
[0015] A word segmentation processing module, configured to, when the sample sentence contains only Chinese characters, input the sample sentence into the initial Chinese character splitting pre-training model, and perform word segmentation processing on the sample sentence through the character encoding model to obtain each Chinese sample word in the sample sentence;
[0016] A radical splitting module, configured to perform radical splitting on all Chinese characters in each Chinese sample word through the character encoding model to obtain a radical decomposition result of each Chinese character;
[0017] A granularity splitting module, configured to perform granularity splitting on all the radical decomposition results through the character encoding model to obtain a splitting result;
[0018] A decoding and recognition module, configured to perform decoding and recognition on the splitting result through the character decoding model to obtain a sample decoded sentence;
[0019] A text loss value determination module, configured to determine a text loss value according to the sample decoded sentence and the sample sentence containing only Chinese characters;
[0020] A convergence judgment module, configured to, when the text loss value does not reach a preset convergence condition, update and iterate the initial parameters of the initial Chinese character splitting pre-training model until the text loss value reaches the preset convergence condition, and record the initial Chinese character splitting pre-training model after convergence as a Chinese pre-training language model based on Chinese character splitting.
[0021] A video subtitle verification method, comprising:
[0022] Obtain a video subtitle verification model and a video to be verified; the video subtitle verification model includes a speech recognition model and a subtitle recognition model; the subtitle recognition model is trained based on a Chinese pre-trained language model that splits Chinese characters; the Chinese pre-trained language model that splits Chinese characters is obtained according to the above language model training method;
[0023] Obtain the speech data in the video to be verified, and perform speech recognition on the speech data through the speech recognition model to obtain a speech sentence corresponding to the speech data;
[0024] Obtain the subtitle sentence corresponding to the speech data in the video to be verified, and perform split recognition on the subtitle sentence through the subtitle recognition model to obtain a split sentence;
[0025] Obtain the similarity between the speech sentence and the split sentence to obtain a sentence similarity;
[0026] When the sentence similarity is greater than a preset similarity threshold, confirm that the video to be verified passes the verification.
[0027] A video subtitle verification device, comprising:
[0028] A model acquisition module, configured to obtain a video subtitle verification model and a video to be verified; the video subtitle verification model includes a speech recognition model and a subtitle recognition model; the subtitle recognition model is trained based on a Chinese pre-trained language model that splits Chinese characters; the Chinese pre-trained language model that splits Chinese characters is obtained according to the above language model training method;
[0029] A speech recognition module, configured to obtain the speech data in the video to be verified, and perform speech recognition on the speech data through the speech recognition model to obtain a speech sentence corresponding to the speech data;
[0030] A split recognition module, configured to obtain the subtitle sentence corresponding to the speech data in the video to be verified, and perform split recognition on the subtitle sentence through the subtitle recognition model to obtain a split sentence;
[0031] A similarity acquisition module, configured to obtain the similarity between the speech sentence and the split sentence to obtain a sentence similarity;
[0032] A video verification module, configured to confirm that the video to be verified passes the verification when the sentence similarity is greater than a preset similarity threshold.
[0033] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned similar case detection method is implemented, or when the processor executes the computer program, the above-mentioned video subtitle verification method is implemented.
[0034] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above-mentioned similar case detection method is implemented, or when the computer program is executed by the processor, the above-mentioned video subtitle verification method is implemented.
[0035] For the above-mentioned language model training, video subtitle verification method, device, equipment and medium, by obtaining a text sample set and an initial Chinese character decomposition pre-training model with initial parameters, the text sample set includes at least one sample sentence, and one sample sentence includes at least one Chinese character; the initial Chinese character decomposition pre-training model includes a character encoding model and a character decoding model; when the sample sentence only includes Chinese characters, input the sample sentence into the initial Chinese character decomposition pre-training model, perform word segmentation on the sample sentence through the character encoding model to obtain each Chinese sample word in the sample sentence; perform radical decomposition on all Chinese characters in each Chinese sample word through the character encoding model to obtain the radical decomposition result of each Chinese character; perform granularity decomposition on all the radical decomposition results through the character encoding model to obtain a decomposition result; perform decoding and recognition on the decomposition result through the character decoding model to obtain a sample decoded sentence; determine a text loss value according to the sample decoded sentence and the sample sentence only including Chinese characters; when the text loss value does not reach a preset convergence condition, update and iterate the initial parameters of the initial Chinese character decomposition pre-training model until the text loss value reaches the preset convergence condition, and record the initial Chinese character decomposition pre-training model after convergence as a Chinese pre-training language model based on character decomposition.
[0036] The present invention trains a Chinese pre-training language model based on character decomposition by splitting Chinese characters into radical structures, so that the Chinese pre-training language model based on character decomposition can receive the internal features of characters, improving its representation ability; and the vocabulary used for granularity decomposition in the character encoding model of this model can better restore the structural type of each character, and the parameters of this vocabulary (500 - 2500) are much smaller than the parameters of the vocabulary in the prior art (usually exceeding twenty thousand), making the model recognition speed fast, and thus facilitating the rapid training of other models. Description of the Drawings
[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0038] Figure 1 is a schematic diagram of an application environment of a language model training method and a video subtitle verification method in an embodiment of the present invention;
[0039] Figure 2 is a flowchart of a language model training method in an embodiment of the present invention;
[0040] Figure 3 is a flowchart of step S13 in a language model training method in an embodiment of the present invention;
[0041] Figure 4 is another flowchart of step S13 in a language model training method in an embodiment of the present invention;
[0042] Figure 5 is a flowchart of a video subtitle verification method in an embodiment of the present invention;
[0043] Figure 6 is a schematic block diagram of a language model training device in an embodiment of the present invention;
[0044] Figure 7 is a schematic block diagram of a radical splitting module in a language model training device in an embodiment of the present invention;
[0045] Figure 8 is another schematic block diagram of a radical splitting module in a language model training device in an embodiment of the present invention;
[0046] Figure 9 is a schematic block diagram of a video subtitle verification device in an embodiment of the present invention;
[0047] Figure 10 is a schematic diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0049] The language model training method provided by the embodiments of the present invention can be applied in an application environment as shown in Figure 1 below. Specifically, this language model training is applied in a language model training system, which includes a client and a server as shown in Figure 1 below. The client communicates with the server through a network to improve the accuracy of language model recognition. Among them, the client, also known as the user side, refers to a program that provides local services for clients corresponding to the server. The client can be installed on, but not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0050] In one embodiment, as shown in Figure 2 below, a language model training method is provided. Taking the server in Figure 1 as an example, the method includes the following steps:
[0051] S11: Obtain a text sample set and an initial Chinese character splitting pre-training model with initial parameters. The text sample set includes at least one sample sentence, and a sample sentence includes at least one Chinese character; the initial Chinese character splitting pre-training model includes a character encoding model and a character decoding model.
[0052] Among them, the text sample set contains sentences that appear in any application scenario. This text sample set includes at least one sample sentence, and this sample sentence can be any sentence that contains at least one Chinese character. The character encoding model is used to encode the Chinese characters in the sample sentence, and the character encoding model includes a Jieba word segmentation module, a Chinese character splitting module, and a preset BPE vocabulary table. The character decoding model is used to decode and recognize the result output by the character encoding model.
[0053] S12: When the sample sentence contains only Chinese characters, input the sample sentence into the initial Chinese character splitting pre-training model, and perform word segmentation processing on the sample sentence through the character encoding model to obtain each Chinese sample word in the sample sentence.
[0054] Among them, the Chinese sample word refers to the word obtained after word segmentation processing in the sample sentence in the form of words.
[0055] Specifically, after obtaining a text sample set and an initial character splitting pre-training model containing initial parameters, if a sample sentence in the text sample set contains only Chinese characters (i.e., except for Chinese characters, it does not contain English letters, Arabic numerals, etc.), the sample sentence is input into the initial character splitting pre-training model, and the sample sentence is segmented by the Jieba word segmentation module in the character encoding model to obtain each Chinese sample word in the sample sentence. Among them, the Jieba word segmentation module performs word segmentation based on regular word combinations. For example, for example, "I love China", after word segmentation processing, "I", "love", and "China" are obtained. The word "China" is obtained by the Jieba word segmentation module based on regular word combinations.
[0056] S13: performing radical decomposition on all Chinese characters in each Chinese sample word through a character encoding model to obtain a radical decomposition result of each Chinese character.
[0057] Specifically, after inputting the sample sentence into the initial character splitting pre-training model, the sample sentence is segmented by the character encoding model to obtain each Chinese sample word in the sample sentence, and then all Chinese characters in each Chinese sample word are split by the character splitting module in the character encoding model to obtain the radical decomposition result of each Chinese character. For example, after the radical splitting of "安全" in "安全最重要", the radical decomposition results obtained are "宀", "女", "人", and "王".
[0058] S14: performing granular splitting of all radical decomposition results through a character encoding model to obtain splitting results.
[0059] Specifically, after the character encoding model is used to perform radical splitting on all the Chinese characters in each Chinese sample word and obtaining the radical decomposition result of each Chinese character, although each Chinese character has been split into radicals at this time, achieving the effect of reducing the character dimension, but for the radical decomposition results after the radical splitting, the character decoding model in the subsequent steps has no way to recognize how the radical decomposition results specifically correspond to the original Chinese characters, that is, it is unable to combine and restore the Chinese characters based on the radical decomposition results, but instead randomly combines to generate characters, which will lead to a decrease in recognition accuracy.
[0060] Furthermore, in this embodiment, all the radical decomposition results are split into finer-grained units through a pre-set BPE vocabulary in the character encoding model to obtain the splitting results. The radical decomposition results of each character are divided into finer-grained units, so that in the subsequent step S15, the character decoding model can recognize the splitting results and can clearly know how to combine them to restore to the original Chinese characters, improving the recognition accuracy. Among them, the pre-set BPE vocabulary is generated through an open-source package (sentencepiece). This BPE vocabulary can cover 99.95% of the characters in the existing corpus. Therefore, when splitting into finer-grained units through the BPE vocabulary, there is basically no phenomenon that the radical decomposition results cannot be recognized, thus ensuring the recognition accuracy and efficiency. At the same time, the size of this BPE vocabulary is 500 - 2500. Compared with the Chinese vocabulary in the prior art (usually exceeding twenty thousand), this BPE vocabulary is smaller, resulting in a smaller number of model parameters, so that the recognition efficiency is high and the recognition accuracy is high during the character splitting and recognition process.
[0061] S15: Decode and recognize the splitting results through the character decoding model to obtain the sample decoded sentence.
[0062] Specifically, after all the radical decomposition results are split into finer-grained units through the character encoding model to obtain the splitting results, the character decoding model in the initial character splitting pre-training model is used to decode and recognize the splitting results, that is, to combine the splitting results to restore them to the corresponding Chinese characters; after decoding and recognizing all the splitting results, that is, after restoring them to Chinese characters, the sample decoded sentence is obtained by combining all the Chinese characters.
[0063] S16: Determine the text loss value according to the sample decoded sentence and the sample sentence containing only Chinese characters.
[0064] Among them, the text loss value refers to the difference value between the sample decoded sentence and the sample sentence containing only Chinese characters, that is, the ratio of the number of different characters between the sample decoded sentence and the sample sentence containing only Chinese characters to the total number of characters.
[0065] Specifically, after decoding and recognizing the splitting results through the character decoding model to obtain the sample decoded sentence, it is necessary to verify whether the sample decoded sentence is the same as the corresponding sample sentence containing only Chinese characters. Therefore, according to the sample decoded sentence and the sample sentence containing only Chinese characters, it is determined whether each Chinese character corresponds one by one, and thus the text loss value is determined according to the ratio of the number of different characters between the sample decoded sentence and the sample sentence containing only Chinese characters to the total number of characters.
[0066] S17: When the text loss value does not reach the preset convergence condition, update the initial parameters of the initial character-splitting pre-training model until the text loss value reaches the preset convergence condition, and record the initial character-splitting pre-training model after convergence as the Chinese pre-training language model based on character splitting.
[0067] It can be understood that the convergence condition can be the condition that the text loss value is less than the set threshold, that is, when the text loss value is less than the set threshold, stop training; the convergence condition can also be the condition that the text loss value is very small and will not decrease after 10,000 calculations, that is, when the text loss value is very small and will not decrease after 10,000 calculations, stop training, and record the initial character-splitting pre-training model after convergence as the Chinese pre-training language model based on character splitting.
[0068] Further, after determining the text loss value according to the sample decoded sentence and the sample sentence containing only Chinese characters, when the text loss value does not reach the preset convergence condition, adjust the initial parameters of the initial character-splitting pre-training model according to the text loss value, and re-enter the sample sentence into the initial character-splitting pre-training model after adjusting the initial parameters, so that when the text loss value corresponding to the sample sentence reaches the preset convergence condition, select another sample sentence containing only Chinese characters in the text sample set, and execute steps S12 - S16 to obtain the text loss value corresponding to the sample sentence, and when the text loss value does not reach the preset convergence condition, adjust the initial parameters of the initial character-splitting pre-training model again according to the text loss value, so that the text loss value corresponding to the sample sentence reaches the preset convergence condition.
[0069] In this way, after training the initial character-splitting pre-training model with all the sample sentences containing only Chinese characters in the text sample set, the result output by the initial character-splitting pre-training model can continuously approach the accurate result, making the recognition accuracy higher and higher. Until the text loss values corresponding to all the sample sentences containing only Chinese characters reach the preset convergence condition, record the initial character-splitting pre-training model after convergence as the Chinese pre-training language model based on character splitting.
[0070] In this embodiment, by splitting Chinese characters into radical structures, a Chinese pre-training language model based on character splitting is trained, so that the Chinese pre-training language model based on character splitting can receive the internal features of the characters, improving its representation ability; and the vocabulary used for granularity splitting in the character encoding model of this model can better restore the structure type of each character, and the parameters of this vocabulary (500 - 2500) are much smaller than the parameters of the vocabulary in the prior art (usually exceeding 20,000), making the model recognition speed fast, which is conducive to quickly training other models.
[0071] In another specific embodiment, to ensure the privacy and security of the Chinese pre-trained language model based on character splitting in the above embodiment, the Chinese pre-trained language model based on character splitting can be stored in the blockchain. Among them, the blockchain is an encrypted, chained storage structure of transactions formed by blocks.
[0072] For example, the header of each block can include the hash values of all transactions in the block and also the hash values of all transactions in the previous block, so as to achieve anti-tampering and anti-forgery of the transactions in the block based on the hash values; after the newly generated transactions are filled into the block and consensus is reached by the nodes in the blockchain network, they will be appended to the end of the blockchain to form a chained growth.
[0073] In one embodiment, the language model training method further includes the following steps:
[0074] When the sample sentence contains non-Chinese characters, obtain the position information of all non-Chinese characters in the sample sentence, intercept all non-Chinese characters according to the position information, and input the sample sentence after intercepting the non-Chinese characters into the initial character-splitting pre-training model.
[0075] Among them, the position information refers to the position encoding of non-Chinese characters in the sample sentence. Exemplarily, assuming that starting from the first character in the sample sentence, the position information of each character is encoded. The position information of the first character is V1, the position information of the second character is V2. If the third character is a non-Chinese character, the position information of this non-Chinese character is V3.
[0076] Specifically, after obtaining the text sample set, if the sample sentence contains non-Chinese characters, since the initial character-splitting pre-training model will not split non-Chinese characters, it is necessary to intercept the non-Chinese characters in the sample sentence. Obtain the position information of all non-Chinese characters in the sample sentence, intercept all non-Chinese characters according to the position information, input the sample sentence after intercepting the non-Chinese characters into the initial character-splitting pre-training model, and execute steps S12 - S16 in the above embodiment to obtain the corresponding text loss value.
[0077] In one embodiment, as Figure 3 shown, in step S13, that is, by the character encoding model, all Chinese characters in each Chinese sample word are split into radicals to obtain the radical decomposition result of each Chinese character, including:
[0078] S131: When the Chinese character contains a splittable radical structure, perform an initial radical split on each Chinese character to obtain a first decomposed character.
[0079] Among them, the separable radical structure refers to a Chinese character that contains a known radical (such as "mouth", "sun", etc.), and this radical can be separated (that is, in addition to the radical, the Chinese character also contains other parts; for example, for the character "ru", when the "female" radical is separated, there is still "mouth"). The first decomposed character contains the radical structure in the Chinese character ("female" in the character "ru") and the non-radical structure ("mouth" in the character "ru").
[0080] Specifically, after performing word segmentation on the sample sentence through the character encoding model to obtain each Chinese sample word in the sample sentence, perform radical structure detection on all Chinese characters in each Chinese sample word. When a Chinese character contains a separable radical structure, use the character splitting module in the character encoding model to perform the first radical splitting on all Chinese characters containing the separable radical structure to obtain the first decomposed character corresponding to each Chinese character.
[0081] Exemplarily, if the Chinese sample word is "fresh", and this Chinese sample word contains a separable radical structure, then perform the first radical splitting on this Chinese sample word. Then, the first decomposed characters corresponding to "xin" can be obtained as "qin" and "jin"; the first decomposed characters corresponding to "xian" can be obtained as "fish" and "sheep". It can be understood that when performing the first radical splitting on a Chinese sample word, the radical structure in existing dictionaries such as Xinhua Dictionary can be used for splitting and recognition, that is, all existing radical structures can be imported into the character splitting module in the character encoding model.
[0082] Furthermore, when performing the first radical splitting on a Chinese sample word containing a separable radical structure, the corresponding structure will be recorded simultaneously, such as: upper-lower structure, left-right structure, semi-surrounded structure, etc.; Exemplarily, when obtaining the first decomposed characters corresponding to "xin" as "qin" and "jin", it will also be marked. In the character encoding model, a classification block of the left-right structure can be generated to represent that this splitting is a left-right structure splitting.
[0083] S132: Detect whether the first decomposed character is the smallest character unit.
[0084] Among them, the smallest character unit is the unit that cannot be further decomposed in the first decomposed character (for example, "tu" in the character "zai" is the smallest character unit).
[0085] S133: If the first decomposed character is the smallest character unit, record the first decomposed character corresponding to each smallest character unit as the radical decomposition result of the Chinese character corresponding to it.
[0086] Specifically, when a Chinese character contains a splittable radical structure, after performing an initial radical splitting on each Chinese character to obtain a first decomposed character, it is detected whether the first decomposed character corresponding to each Chinese character is the smallest character unit. If there is a first decomposed character that is the smallest character unit, the first decomposed character corresponding to each smallest character unit is recorded as the radical decomposition result of the Chinese character corresponding to it.
[0087] Exemplarily, if the Chinese sample word is "fresh", when performing an initial radical splitting on this Chinese sample word, the first decomposed characters corresponding to "new" can be obtained as "qin" and "jin"; after the first decomposed characters corresponding to "fresh" are obtained as "fish" and "sheep", it is detected whether each first decomposed character is the smallest character unit. For the character "qin", "li" can still be split out, so "qin" is not the smallest character unit; for the character "jin", it cannot be split further, so "jin" is the smallest character unit. Therefore, "jin" is the radical decomposition result of the first decomposed character "jin" in the Chinese character "new". And the overall radical decomposition result of "new" needs to be combined with "jin" after the character "qin" is completely split to form the radical decomposition result of the Chinese character "new".
[0088] In one embodiment, as Figure 4 shown, after step S132, that is, after detecting whether the first decomposed character is the smallest character unit, it further includes:
[0089] S134: If the first decomposed character is not the smallest character unit, then perform a structural analysis on the first decomposed character to obtain a first character structure of the first decomposed character, and perform a secondary radical splitting on the first decomposed character according to the first character structure to obtain a second decomposed character.
[0090] Wherein, the first character structure represents the structural classification in the first decomposed character, and the first character structure includes but is not limited to an upper-lower structure, a left-right structure, a semi-surrounding structure, etc.
[0091] Specifically, after detecting whether the first decomposed character is the smallest character unit, if the first decomposed character is not the smallest character unit, that is, it indicates that the first decomposed character still contains a splittable radical structure, then perform a structural analysis on the first decomposed character to obtain a first character structure of the first decomposed character; perform a secondary radical splitting on the first decomposed character corresponding to it according to the first character structure to obtain a second decomposed character.
[0092] For example, assuming that one of the Chinese characters is "萌", after the initial radical decomposition, the first decomposed characters obtained include "艹" and "明", and "艹" and "明" are checked to verify whether they are the smallest character units. It can be found that "艹" is the smallest character unit, so "艹" is part of the radical decomposition result of "萌", and "明" is not the smallest character unit. It can be further decomposed into two second decomposed characters "日" and "月", and at the same time, the corresponding structural classification is recorded, which is the left and right structure.
[0093] S135: If the second decomposed characters are all minimum character units, the first decomposed characters and the second decomposed characters of the minimum character unit are recorded as a radical decomposition result.
[0094] Specifically, after performing structural analysis on the first decomposed character to obtain the first character structure of the first decomposed character, and performing secondary radical decomposition on the first decomposed character according to the first character structure to obtain the second decomposed character, it is detected whether all the second decomposed characters are the smallest character units. If all the second decomposed characters are the smallest character units, the first decomposed character and the second decomposed character of the smallest character unit are recorded as the radical decomposition results of the corresponding Chinese characters.
[0095] For example, as mentioned above, assuming that the Chinese character is "萌", after the second radical decomposition, the second decomposed characters obtained are "日" and "月". Both of these second decomposed characters cannot be further decomposed, that is, they are the smallest character units. Therefore, the first decomposed character "艹" of the smallest character unit, and the second decomposed characters "日" and "月" of the smallest character unit are recorded as the radical decomposition results of "萌".
[0096] It should be noted that no matter it is the first radical splitting or the second radical splitting, or when the second decomposed character can still be split, other radical splitting, at the same time, the character structure classification corresponding to the splitting is recorded, that is, the upper and lower structure, the left and right structure, the semi-enclosed structure, etc.
[0097] In one embodiment, after step S131, that is, when the Chinese characters contain a separable radical structure, each of the Chinese characters is initially decomposed into radicals to obtain a first decomposed character, the method further includes:
[0098] Check whether the first decomposed character is an existing character; if the first decomposed character is not an existing character, encode the first decomposed character to obtain an encoded character corresponding to the first decomposed character.
[0099] Among them, existing characters refer to characters that can exist independently in the prior art (characters that can be queried), such as: "口", "女", etc. Characters that are not existing characters, such as: after the initial radical splitting of "在", after the splitting out of "土", the remaining part is not a character that exists independently and can be queried in the existing data or corpus. Coded characters refer to characters obtained by special encoding the first decomposed character that is not an existing character.
[0100] Specifically, when a Chinese character contains a decomposable radical structure, each Chinese character is initially decomposed by radicals to obtain the first decomposed character. It is then checked whether each first decomposed character is an existing character. If there is a first decomposed character that is not an existing character, the first decomposed character is encoded to obtain the encoded character corresponding to the first decomposed character.
[0101] Exemplarily, after the initial radical splitting of "在", after "土" is split out, the remaining part is not an independent and queryable character in the existing data or corpus, so the first decomposed character in the "在" character except the "土" character is encoded. The part that is not an existing character can be encoded into any form, such as CDP8861 to represent the part, and CDP8861 is the encoded character corresponding to the first decomposed character that is not an existing character.
[0102] Among them, the method of encoding the first decomposed character can be set according to the user's preference, but it should be noted during the setting process that each first decomposed character that is not an existing character is associated with a unique encoding character, that is, the same encoding character cannot be used to encode different first decomposed characters that are not existing characters, so as to avoid causing the subsequent character decoding module to be unable to accurately recognize, thereby reducing the recognition accuracy.
[0103] In one embodiment, step S13 further includes:
[0104] When the Chinese characters in the Chinese sample words do not contain a divisible radical structure, the Chinese characters are directly recorded as the radical decomposition results corresponding thereto.
[0105] Specifically, after the sample sentence is segmented by the character encoding model to obtain each Chinese sample word in the sample sentence, all Chinese characters in each Chinese sample word are detected for radical structure, and when the Chinese character does not contain a divisible radical structure, the Chinese character is directly recorded as the radical decomposition result corresponding to it. Exemplarily, if the Chinese sample word is "stuttering", the character "口" in the Chinese sample word does not contain a divisible radical structure, that is, it cannot be further split into smaller character units, then "口" is directly recorded as a part of the radical decomposition result corresponding to the Chinese sample word "stuttering".
[0106] In one embodiment, as Figure 5 shown, a method for verifying video subtitles is provided. Taking the server in Figure 1 as an example, the method includes the following steps:
[0107] S31: Obtain a video subtitle verification model and a video to be verified; the video subtitle verification model includes a speech recognition model and a subtitle recognition model; the subtitle recognition model is trained based on a Chinese pre-trained language model that splits Chinese characters; the Chinese pre-trained language model that splits Chinese characters is obtained according to the language model training method in the above embodiment.
[0108] Among them, the video subtitle verification model refers to a model used to verify whether the speech and subtitles in any video segment match. The video subtitle verification model includes a speech recognition model and a subtitle recognition model; it should be emphasized that the subtitle recognition model is trained based on a Chinese pre-trained language model that splits Chinese characters, and the Chinese pre-trained language model that splits Chinese characters is obtained according to the language model training method in the above embodiment. Since in the video subtitle verification scenario, the subtitles are all conventionally recorded in common corpora and there is no need to use a model with large parameters for verification, in order to use a model with small parameters for verification, the Chinese pre-trained language model that splits Chinese characters obtained by the language model training method in the above embodiment is used to train the subtitle recognition model, which can not only meet the verification requirements, but also has small parameters and fast recognition speed. The video to be verified refers to a video that needs to verify whether the subtitles and speech match.
[0109] S32: Obtain the speech data in the video to be verified, and perform speech recognition on the speech data through the speech recognition model to obtain a speech sentence corresponding to the speech data.
[0110] Among them, the speech data refers to the audio data in the video to be verified. The speech sentence refers to the speech text corresponding to the speech data.
[0111] Specifically, after obtaining the video subtitle verification model and the video to be verified, obtain the speech data in the video to be verified, and perform speech recognition on the speech data through the speech recognition model in the video subtitle verification model to obtain the speech text corresponding to the speech data, that is, the speech sentence.
[0112] S33: Obtain the subtitle sentence corresponding to the speech data in the video to be verified, and perform split recognition on the subtitle sentence through the subtitle recognition model to obtain split sentences.
[0113] Specifically, after obtaining the video subtitle verification model and the video to be verified, obtain the subtitle sentences corresponding to the speech data in the video to be verified, and perform split recognition on the subtitle sentences through the subtitle recognition model in the video subtitle verification model to obtain split sentences. Since the subtitle recognition model is trained based on a Chinese pre-trained language model that splits characters, when performing split recognition on the subtitle sentences through the subtitle recognition model, the characters in the subtitle sentences can be better restored.
[0114] S34: Obtain the similarity between the speech sentence and the split sentence to obtain the sentence similarity.
[0115] Specifically, after obtaining the speech sentence and the split sentence, compare the similarity between the speech sentence and the split sentence to obtain the sentence similarity, that is, determine whether the characters in the speech sentence and the characters in the split sentence correspond one by one to determine whether the speech data and the subtitle sentence match.
[0116] S35: When the sentence similarity is greater than or equal to the preset similarity threshold, confirm that the video to be verified passes the verification.
[0117] Among them, the preset similarity threshold can be set by the user according to the matching requirements. Exemplarily, the preset similarity threshold can be 0.9, 0.95, etc.
[0118] Specifically, after obtaining the similarity between the speech sentence and the split sentence to obtain the sentence similarity, compare the sentence similarity with the preset similarity threshold. When the sentence similarity is greater than or equal to the preset similarity threshold, confirm that the video to be verified passes the verification. When the sentence similarity is less than the preset similarity threshold, it indicates that the similarity between the speech sentence and the split sentence does not meet the standard, that is, the speech data and the subtitle sentence do not match, and the subtitle sentence needs to be adjusted again to make the speech data and the subtitle sentence match, so as to make the video to be verified pass the verification.
[0119] Exemplarily, for the speech data of the video to be verified, if the recognized speech sentence is "Where to have lunch at noon"; assume that the obtained subtitle sentence corresponding to the speech data is "What to eat at noon", at this time, if the split sentence obtained after recognition by the subtitle recognition model trained based on the Chinese pre-trained language model that splits characters is "What to eat at noon", calculate the similarity between the speech sentence and the split sentence. Assume that the obtained sentence similarity is 0.47, and the preset similarity threshold is 0.9, then it is considered that the speech sentence and the split sentence do not match, and the subtitle sentence corresponding to the speech data is adjusted to determine the similarity between the speech sentence and the adjusted subtitle sentence until the sentence similarity between the speech sentence and the subtitle sentence is greater than the preset similarity threshold, then confirm that the video to be verified passes the verification.
[0120] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0121] In one embodiment, a language model training device is provided, which corresponds one-to-one with the language model training method in the above embodiment. As Figure 6 shown, the language model training device includes a data acquisition module 11, a first word segmentation processing module 12, a first radical splitting module 13, a first granularity splitting module 14, a first decoding and recognition module 15, a text loss value determination module 16, and a first convergence judgment module 17. The detailed description of each functional module is as follows:
[0122] The data acquisition module 11 is configured to acquire a set of text samples and an initial character splitting pre-training model with initial parameters. The set of text samples includes at least one sample sentence, and one sample sentence includes at least one Chinese character. The initial character splitting pre-training model includes a character encoding model and a character decoding model.
[0123] The word segmentation processing module 12 is configured to, when the sample sentence only includes Chinese characters, input the sample sentence into the initial character splitting pre-training model, and perform word segmentation processing on the sample sentence through the character encoding model to obtain each Chinese sample word in the sample sentence.
[0124] The radical splitting module 13 is configured to perform radical splitting on all Chinese characters in each Chinese sample word through the character encoding model to obtain the radical decomposition result of each Chinese character.
[0125] The granularity splitting module 14 is configured to perform granularity splitting on all the radical decomposition results through the character encoding model to obtain the splitting result.
[0126] The decoding and recognition module 15 is configured to perform decoding and recognition on the splitting result through the character decoding model to obtain a sample decoded sentence.
[0127] The text loss value determination module 16 is configured to determine the text loss value according to the sample decoded sentence and the sample sentence that only includes Chinese characters.
[0128] The convergence judgment module 17 is configured to, when the text loss value does not reach the preset convergence condition, update and iterate the initial parameters of the initial character splitting pre-training model until the text loss value reaches the preset convergence condition, and record the initial character splitting pre-training model after convergence as a Chinese pre-training language model based on character splitting.
[0129] Preferably, the language model training device further includes the following modules:
[0130] A position information acquisition module 21, configured to, when the sample sentence contains non-Chinese characters, acquire the position information of all the non-Chinese characters in the sample sentence, intercept all the non-Chinese characters according to the position information, and input the intercepted sample sentence into the initial Chinese character splitting pre-trained model.
[0131] Preferably, as Figure 7 shown, the first radical splitting module 13 includes the following units:
[0132] A primary radical splitting unit 131, configured to, when the Chinese character contains a splittable radical structure, perform a primary radical splitting on each Chinese character to obtain a first decomposed character.
[0133] A character detection unit 132, configured to detect whether the first decomposed character is the smallest character unit.
[0134] A first recording unit 133, configured to, when the first decomposed character is the smallest character unit, record the first decomposed character corresponding to each smallest character unit as the radical decomposition result corresponding to the Chinese character.
[0135] Preferably, as Figure 8 shown, the first radical splitting module 13 further includes the following units:
[0136] A secondary radical splitting unit 134, configured to, when the first decomposed character is not the smallest character unit, perform a structural analysis on the first decomposed character to obtain a first character structure of the first decomposed character, and perform a secondary radical splitting on the first decomposed character according to the first character structure to obtain a second decomposed character.
[0137] A second recording unit 135, configured to, when all the second decomposed characters are the smallest character units, record the first decomposed character and the second decomposed character of the smallest character unit as the radical decomposition result.
[0138] Preferably, the first radical splitting module 13 further includes the following units:
[0139] An existing character detection unit, configured to detect whether the first decomposed character is an existing character;
[0140] A character encoding unit, configured to, when the first decomposed character is not an existing character, perform encoding on the first decomposed character to obtain an encoded character corresponding to the first decomposed character.
[0141] For the specific limitations of the language model training device, reference can be made to the limitations of the language model training method in the above text, which will not be elaborated here. Each module in the above language model training device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0142] In one embodiment, a video subtitle verification device is provided, and the video subtitle verification device corresponds one-to-one with the video subtitle verification method in the above embodiment. As Figure 9 shown, the video subtitle verification device includes a model acquisition module 31, a speech recognition module 32, a splitting recognition module 33, a similarity acquisition module 34, and a video verification module 35. The detailed description of each functional module is as follows:
[0143] The model acquisition module 31 is used to acquire a video subtitle verification model and a video to be verified; the video subtitle verification model includes a speech recognition model and a subtitle recognition model; the subtitle recognition model is trained based on a Chinese pre-trained language model based on character splitting; the Chinese pre-trained language model based on character splitting is obtained according to the above language model training method;
[0144] The speech recognition module 32 is used to acquire the speech data in the video to be verified, and perform speech recognition on the speech data through the speech recognition model to obtain a speech sentence corresponding to the speech data;
[0145] The splitting recognition module 33 is used to acquire the subtitle sentence corresponding to the speech data in the video to be verified, and perform splitting recognition on the subtitle sentence through the subtitle recognition model to obtain a split sentence;
[0146] The similarity acquisition module 34 is used to acquire the similarity between the speech sentence and the split sentence to obtain a sentence similarity;
[0147] The video verification module 35 is used to confirm that the video to be verified passes the verification when the sentence similarity is greater than a preset similarity threshold.
[0148] For the specific limitations of the video subtitle verification device, reference can be made to the limitations of the video subtitle verification method in the above text, which will not be elaborated here. Each module in the above video subtitle verification device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0149] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 10 . The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data used for the above-mentioned detection of similar cases. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a language model training method or a video subtitle verification method.
[0150] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the language model training method in the above embodiment or the video subtitle verification method in the above embodiment.
[0151] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the language model training method in the above embodiment or the video subtitle verification method in the above embodiment.
[0152] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0153] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0154] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for training a language model, characterized in that, Including: Obtain a set of text samples and an initial character-splitting pre-training model with initial parameters, where the set of text samples includes at least one sample sentence, and one sample sentence includes at least one Chinese character; The initial character-splitting pre-training model includes a character encoding model and a character decoding model; When the sample sentence only contains Chinese characters, input the sample sentence into the initial character-splitting pre-training model, and perform word segmentation on the sample sentence through the character encoding model to obtain each Chinese sample word in the sample sentence; Perform radical splitting on all the Chinese characters in each Chinese sample word through the character encoding model to obtain the radical decomposition result of each Chinese character; Perform granularity splitting on all the radical decomposition results through the character encoding model to obtain the splitting result; Perform decoding recognition on the splitting result through the character decoding model to obtain a sample decoded sentence; Determine a text loss value according to the sample decoded sentence and the sample sentence that only contains Chinese characters; When the text loss value does not reach the preset convergence condition, update and iterate the initial parameters of the initial character-splitting pre-training model until the text loss value reaches the preset convergence condition, and record the initial character-splitting pre-training model after convergence as a Chinese pre-training language model based on character splitting.
2. The method for training a language model according to claim 1, characterized in that, Before inputting the sample sentence into the initial character-splitting pre-training model, it includes: When the sample sentence contains non-Chinese characters, obtain the position information of all the non-Chinese characters in the sample sentence, intercept all the non-Chinese characters according to the position information, and input the sample sentence after intercepting the non-Chinese characters into the initial character-splitting pre-training model.
3. The method for training a language model according to claim 1, characterized in that, The step of performing radical splitting on all the Chinese characters in each Chinese sample word through the character encoding model to obtain the radical decomposition result of each Chinese character includes: When the Chinese character contains a splittable radical structure, perform an initial radical split on each Chinese character to obtain a first decomposed character; Detect whether the first decomposed character is the smallest character unit; If the first decomposed character is the smallest character unit, record the first decomposed character corresponding to each smallest character unit as the radical decomposition result of the corresponding Chinese character.
4. The method for training a language model according to claim 3, characterized in that, After detecting whether the first decomposed character is the smallest character unit, it further includes: If the first decomposed character is not the smallest character unit, perform structure analysis on the first decomposed character to obtain the first character structure of the first decomposed character, and perform a secondary radical split on the first decomposed character according to the first character structure to obtain a second decomposed character; If all the second decomposed characters are the smallest character units, record the first decomposed character and the second decomposed character of the smallest character unit as the radical decomposition result.
5. The method for training a language model according to claim 3, characterized in that, After performing an initial radical split on each Chinese character to obtain a first decomposed character, it further includes: Detect whether the first decomposed character is an existing character; If the first decomposed character is not an existing character, encode the first decomposed character to obtain an encoded character corresponding to the first decomposed character.
6. A method for verifying video subtitles, characterized in that, Including: Obtain a video subtitle verification model and a video to be verified; The video subtitle verification model includes a speech recognition model and a subtitle recognition model; The subtitle recognition model is trained based on a Chinese pre-trained language model with character decomposition; the Chinese pre-trained language model with character decomposition is obtained according to the language model training method described in any one of claims 1 to 5; Obtain the speech data in the video to be verified, and perform speech recognition on the speech data through the speech recognition model to obtain a speech sentence corresponding to the speech data; Obtain the subtitle sentence corresponding to the speech data in the video to be verified, and perform split recognition on the subtitle sentence through the subtitle recognition model to obtain a split sentence; Obtain the similarity between the speech sentence and the split sentence to obtain a sentence similarity; When the sentence similarity is greater than a preset similarity threshold, confirm that the video to be verified passes the verification.
7. A device for training a language model, characterized in that, Including: A data acquisition module for obtaining a text sample set and an initial character decomposition pre-trained model with initial parameters. The text sample set contains at least one sample sentence, and one sample sentence contains at least one Chinese character; the initial character decomposition pre-trained model includes a character encoding model and a character decoding model; A first word segmentation processing module for, when the sample sentence contains only Chinese characters, input the sample sentence into the initial character decomposition pre-trained model, and perform word segmentation processing on the sample sentence through the character encoding model to obtain each Chinese sample word in the sample sentence; A first radical splitting module for splitting all Chinese characters in each Chinese sample word through the character encoding model to obtain a radical decomposition result of each Chinese character; A first granularity splitting module for performing granularity splitting on all the radical decomposition results through the character encoding model to obtain a splitting result; A first decoding and recognition module for performing decoding and recognition on the splitting result through the character decoding model to obtain a sample decoded sentence; A text loss value determination module for determining a text loss value according to the sample decoded sentence and the sample sentence containing only Chinese characters; A first convergence judgment module for, when the text loss value does not meet the preset convergence condition, updating and iterating the initial parameters of the initial character decomposition pre-trained model until the text loss value meets the preset convergence condition, and recording the initial character decomposition pre-trained model after convergence as a Chinese pre-trained language model with character decomposition.
8. A video subtitle verification device, characterized in that, Including: A model acquisition module for obtaining a video subtitle verification model and a video to be verified; The video subtitle verification model includes a speech recognition model and a subtitle recognition model; the subtitle recognition model is trained based on a Chinese pre-trained language model with character decomposition; the Chinese pre-trained language model with character decomposition is obtained according to the language model training method described in any one of claims 1 to 5; A speech recognition module, configured to obtain speech data in the video to be verified, and perform speech recognition on the speech data through the speech recognition model to obtain a speech sentence corresponding to the speech data; A splitting and recognition module, configured to obtain a subtitle sentence corresponding to the speech data in the video to be verified, and perform splitting and recognition on the subtitle sentence through the subtitle recognition model to obtain a split sentence; A similarity obtaining module, configured to obtain the similarity between the speech sentence and the split sentence to obtain a sentence similarity; A video verification module, configured to confirm that the video to be verified passes the verification when the sentence similarity is greater than a preset similarity threshold.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the language model training method according to any one of claims 1 to 5 is implemented, or when the processor executes the computer program, the video subtitle verification method according to claim 6 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the language model training method according to any one of claims 1 to 5 is implemented, or when the computer program is executed by a processor, the video subtitle verification method according to claim 6 is implemented.
Citation Information
Patent Citations
Language model training method and device
CN110619120A
Pre-training model processing method and device, downstream task processing method and device and storage medium
CN112016300A