Text regularization method, device, electronic device and storage medium

By performing character encoding and feature vector generation on text, and combining language types for regular processing, the existing text regularization system requires manual construction of rules, and efficient and low-cost text regularization is achieved, and multilingual processing is supported.

CN112765937BActive Publication Date: 2025-05-23PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011644545.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-31
Publication Date
2025-05-23
Estimated Expiration
2040-12-31

AI Technical Summary

Technical Problem

Existing text regular systems require manual construction of complex rules, and different languages ​​need to independently construct rules, resulting in high labor costs and low efficiency.

Method used

By encoding the characters of regular text, the feature vector of each character is generated, and regular processing is performed according to the language type, text regularization without manual writing rules is achieved.

Benefits of technology

It improves the efficiency of text regularity, reduces labor costs, and supports regular processing of text in various languages, expanding usage scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112765937B_ABST
    Figure CN112765937B_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence technology, and specifically to a text regularization method, device, electronic device and storage medium. The method comprises: obtaining a text to be regularized; segmenting the text to be regularized into characters to obtain a plurality of characters; encoding each of the plurality of characters to obtain a first feature vector of each of the plurality of characters, wherein the first feature vector of each of the plurality of characters is used to represent the context information of each of the plurality of characters; regularizing the text to be regularized according to the first feature vector of each of the plurality of characters and the language type of the text to be regularized to obtain a regularized text of the text to be regularized. The present application is conducive to improving the efficiency and accuracy of text regularization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a text regularization method, device, electronic device and storage medium. Background Art

[0002] The establishment of traditional text regularization systems requires a strong linguistic background, and often requires experts in specific fields to manually construct a large number of complex and cumbersome text regularization rules based on linguistic characteristics. At the same time, the linguistic knowledge between different languages ​​is obviously different and cannot be effectively transferred. If text regularization is performed on a new language, a set of text regularization rules needs to be rebuilt.

[0003] In recent years, with the rapid development of artificial intelligence, text regularization systems based on neural networks of encoder and decoder models have begun to appear in the public eye. However, due to the soft classification characteristics of simple encoder and decoder models, simple encoder and decoder models cannot achieve satisfactory text regularization accuracy. Therefore, the current mainstream text regularization system still requires manual construction of a set of specific, complex, and cumbersome text regularization rules, and different text regularization rules need to be constructed for different languages, which requires a lot of manpower and physics, and there may be code redundancy between various text rules.

[0004] Therefore, the existing text regularization process requires manual construction of a text regularization system, which has high labor costs and low text regularization efficiency. Summary of the invention

[0005] The embodiments of the present application provide a text regularization method, device, electronic device and storage medium, which perform text regularization based on the language type of the text to be regularized and the feature vector of each character, thereby improving the efficiency of text regularization and reducing labor costs.

[0006] In a first aspect, an embodiment of the present application provides a text regularization method, comprising:

[0007] Get the text to be regularized;

[0008] Performing character segmentation on the text to be regularized to obtain a plurality of characters;

[0009] Encode each of the multiple characters to obtain a first feature vector of each of the multiple characters, wherein the first feature vector of each of the multiple characters is used to represent context information of each of the multiple characters;

[0010] According to the first feature vector of each character in the multiple characters and the language type of the text to be regularized, the text to be regularized is regularized to obtain a regularized text of the text to be regularized.

[0011] In a second aspect, an embodiment of the present application provides a text regularization device, comprising:

[0012] An acquisition unit, used to acquire the text to be regularized;

[0013] A processing unit, used for segmenting the text to be regularized into characters to obtain a plurality of characters;

[0014] Encode each of the multiple characters to obtain a first feature vector of each of the multiple characters, wherein the first feature vector of each of the multiple characters is used to represent context information of each of the multiple characters;

[0015] According to the first feature vector of each character in the multiple characters and the language type of the text to be regularized, the text to be regularized is regularized to obtain a regularized text of the text to be regularized.

[0016] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, the processor is connected to a memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device performs the method described in the first aspect.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program enables a computer to execute the method described in the first aspect.

[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer is operable to cause the computer to execute the method described in the first aspect.

[0019] Implementing the embodiments of the present application has the following beneficial effects:

[0020] It can be seen that in the embodiment of the present application, the regular text is first segmented into characters, and then each character is encoded to obtain the first feature vector of each character; finally, the regular text is regularized according to the first feature vector of each character and the language type of the text to be regularized, which means that the regularization of the regular text can be completed without manually writing regular rules, thereby improving the efficiency of text regularization and saving labor costs. In addition, in the process of text regularization, the language type of the text to be regularized will be combined to achieve regularization of texts in various languages, so that the text regularization method of the present application has more usage scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 A flowchart of a text regularization method provided in an embodiment of the present application;

[0023] Figure 2 A schematic diagram of a flow chart of encoding and decoding processing of non-standard characters provided in an embodiment of the present application;

[0024] Figure 3 A schematic diagram of encoding and decoding of non-standard characters by an encoder and a decoder provided in an embodiment of the present application;

[0025] Figure 4 A block diagram of the functional units of a text regularization device provided in an embodiment of the present application;

[0026] Figure 5 A schematic diagram of the structure of a text regularization device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0028] The terms "first", "second", "third" and "fourth" etc. in the specification and claims of the present application and the drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0029] Reference to "embodiments" herein means that a particular feature, result, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0030] See also Figure 1 , Figure 1 A flowchart of a text regularization method provided in an embodiment of the present application. The method is applied to a text regularization device. The method comprises the following steps:

[0031] 101: The text regularization device obtains the text to be regularized.

[0032] Exemplarily, the text to be regularized can be manually input by a user in the information input field of the text regularization device, or can be automatically read by the text regularization device from a text library. For example, the text to be regularized can be a document to be regularized, and the text regularization device can read the text to be regularized from the document in sequence. Therefore, this application does not limit the acquisition of the text to be regularized.

[0033] 102: The text regularization device performs character segmentation on the regular text to obtain a plurality of characters.

[0034] For example, the regular text can be segmented into characters by a word segmenter to obtain multiple characters, for example, the regular text can be segmented into characters by a word2vec word segmenter, wherein the characters can be English words, Chinese words, French words or special symbols, such as "$", " / ", and so on.

[0035] 103: The text regularization device encodes each of the multiple characters to obtain a first feature vector of each of the multiple characters, wherein the first feature vector of each of the multiple characters is used to represent context information of each of the multiple characters.

[0036] Exemplarily, each of the multiple characters is encoded to obtain a character vector corresponding to each of the multiple characters. Specifically, each character is segmented to obtain a letter string of each character; each letter in the letter string of each character is encoded to obtain a letter vector corresponding to each letter; finally, the letter vector of each letter is encoded to obtain a character vector of each character. For example, the character "Achieve" is processed into a letter string of "A", "c", "h", "i", "e", "v", "e", and the letter vector of each letter in the letter string is used as the input of the encoder to model the character vector of the character "Achieve". Then, a first text corresponding to character A is constructed with character A as the center, wherein character A is any one of the multiple characters, and the first text includes X characters before character A in the regular text, character A, and Y characters after character A in the regular text, wherein X and Y are integers greater than or equal to 1; then, the character vectors corresponding to each character in the first text are spliced ​​(i.e., horizontally spliced) to obtain a first feature vector corresponding to the character A, wherein the first feature vector corresponding to the character A is used to represent contextual information of character A in the first text.

[0037] It should be understood that if there are not X characters before the character A, for example, the character A is the first character or the last character in the text to be regularized, the first text can be constructed for the character A by filling in preset characters (for example, the start character S can be filled in).

[0038] 104: The text regularization device performs regularization processing on the text to be regularized according to the first feature vector of each character in the plurality of characters and the language type of the text to be regularized to obtain a regularized text of the text to be regularized.

[0039] Exemplarily, based on the first feature vector of each character in the multiple characters, the attributes of each character in the multiple characters are determined, wherein the attributes of each character include standard characters or non-standard characters; then, the standard characters in the text to be regularized are used as the regular characters of the marked characters, that is, the standard characters themselves are used as the regular characters of the standard characters, and according to the language type of the text to be regularized and the first feature vector corresponding to the non-standard characters in the text to be regularized, the standard characters are encoded and decoded to obtain the regular characters of the non-standard characters; finally, the regular characters of the standard characters in the text to be regularized are combined with the regular characters of the non-standard characters to obtain the regular text of the text to be regularized.

[0040] For example, a standard character refers to a character whose pronunciation and writing are the same. For example, for the character "year", its pronunciation and writing are the same, that is, both are "year", and the regular character of the standard character is itself. For example, the non-standard characters involved in this application include but are not limited to the following:

[0041] Dates, currencies, addresses, letters, cardinal numbers, ordinal numbers, URLs, units of measure, fractions, decimals, phone numbers, times, digits, punctuation, and foreign words.

[0042] Furthermore, with character B as the center, a second text corresponding to the character B is constructed, the second text including M characters before the character B in the text to be regularized, the character B, and N characters after the character B in the text to be regularized, wherein the character B is any non-standard character in the text to be regularized, and M and N are both integers greater than or equal to 1; then, each character in the second text is encoded by Byte-Pair Encoding (BPE) to obtain a second feature vector for each character in the second text.

[0043] Specifically, each character in the second text is split into letter strings, and the letter strings of each character are combined according to the frequency of occurrence of the letter strings of all characters in the second text to obtain a new letter string for each character; then, the new letter string of each character is input into the encoder for encoding to obtain the second feature vector of each character in the second text. The problem of unregistered words in the second text can be solved by double-byte encoding. Then, the second feature vector of each character in the second text is input into the Transformer-XL network for feature extraction to obtain the third feature vector corresponding to character B, wherein the third feature vector of character B is used to represent the context information of character B in the second text; finally, according to the first feature vector of character B, the third feature vector of character B and the language type of the text to be regularized, the character B is encoded and decoded to obtain the regular character corresponding to character B. The process of encoding and decoding character B will be described in detail later, and no further description will be given here.

[0044] It can be seen that in the embodiment of the present application, the regular text is first segmented into characters, and then each character is encoded to obtain the first feature vector of each character; finally, the regular text is regularized according to the first feature vector of each character and the language type of the text to be regularized, which means that the regularization of the regular text can be completed without manually writing regular rules, thereby improving the efficiency of text regularization and saving labor costs. In addition, in the process of text regularization, the language type of the text to be regularized will be combined to achieve regularization of texts in various languages, so that the text regularization method of the present application has more usage scenarios.

[0045] See also Figure 2 , Figure 2 A schematic flow chart of a coding and decoding method provided in an embodiment of the present application. The method is applied to a text regularization device. The method comprises the following steps:

[0046] 201: Perform word embedding processing on character B to obtain a fourth eigenvector of character B.

[0047] Exemplarily, performing word embedding processing on character B is actually performing mapping processing on character B to obtain the fourth feature vector of character B. For example, the ASCII code of character B can be used as the fourth feature vector of character B.

[0048] 202: Encode the attribute of character B to obtain a word class vector of character B.

[0049] Exemplarily, encoding the attribute of character B means mapping the word class to which character B belongs to obtain the word class vector of character B. For example, if character B is "currency", the GB232 code of "currency" is used as the word class vector of character B.

[0050] It should be understood that although the attributes of character B have been classified by the first eigenvector of each character, the process of classifying the attributes of character B only classifies whether the character is a standard character or a non-standard character, and does not perform detailed classification on non-standard characters. Therefore, the first eigenvector of each character can only be used to distinguish whether each character is a standard character or a non-standard character, and cannot perform further distinction on non-standard characters. Here, after character B is classified more finely, the word class vector of each non-standard character is mapped to obtain a more detailed category of each non-standard character.

[0051] 203: Encode the language type of the text to be regularized to obtain a language vector of the text to be regularized, and use the language vector as an encoding parameter of the encoder and a decoding parameter of the decoder.

[0052] Similarly, the language type of the text to be regularized is mapped to obtain the language vector of the text to be regularized. For example, the GB2312 code of the Chinese representation of the language type (for example, the language types are "English", "Chinese", "French", etc.) can be used as the language vector of the language type.

[0053] 204: Input the fourth feature vector of character B to the encoder for encoding, encode character B, and obtain a fifth feature vector of character B.

[0054] Exemplarily, the encoder may be a neural network based on a long short-term memory network, a bidirectional long short-term memory network, or a recurrent network. This application does not limit the type of encoder.

[0055] Exemplarily, character B is encoded according to the hidden vector output by the encoder last encoding, the fourth eigenvector of character B and the encoding parameters of the encoder (i.e., the language vector) to obtain the fifth eigenvector corresponding to character B and the hidden vector corresponding to character B.

[0056] It should be understood that, when character B is the first non-standard character to be encoded, the hidden vector output by the encoder in the last encoding is a preset hidden vector, such as a zero vector. In addition, if only character B is encoded in this encoding, the hidden vector finally output by the encoder is the hidden vector generated by the encoding process of character B. If other non-standard characters need to be encoded, the hidden vector corresponding to character B is used as the hidden vector of the next non-standard character to be encoded.

[0057] It should be understood that if character B is one of multiple consecutive non-standard characters with the same attributes (i.e., the word classes are exactly the same, for example, they are all dates in non-standard words) in the text to be regularized, in order to speed up the encoding efficiency and encoding accuracy of these multiple non-standard characters, these multiple non-standard characters can be encoded together without encoding any of them individually.

[0058] For example, Figure 3 As shown, there are multiple consecutive non-standard characters with the same attributes in the regular text: [X 1 ,X 2 ,…,X n ]; for multiple non-standard characters [X 1 ,X 2 ,…,X n ] to embed each non-standard character in the word, and obtain the fourth feature vector of each non-standard character; then, based on the preset hidden layer vector e 0 and the encoding parameters of the encoder, for the first non-standard character X of the multiple non-standard characters 1 Perform the first encoding to obtain the first non-standard character X1 The fifth eigenvector Y1 of 1 ; Further, based on the hidden vector e output by the first encoding 1 and the encoding parameters of the encoder, for the second non-standard character X of the plurality of non-standard characters 2 Perform the second encoding to get the second non-standard character X 2 The corresponding fifth eigenvector, and the hidden vector e corresponding to the second encoding 2 ; Repeat the above steps to obtain the last non-standard character X among the multiple non-standard characters n The fifth eigenvector of the last encoding output, and the hidden vector e n Among them, the hidden vector output by the last encoding contains the contextual semantic information of these multiple non-standard characters. In this way, these multiple non-standard characters [X 1 ,X 2 ,…,X n ] Encoding is successful, and the fifth eigenvector corresponding to these multiple non-standard characters is output [Y 1 ,Y 2 ,…,Y n ].

[0059] For example, if the regular text is "Achieve record net income of about $1 billion during the year", the non-standard characters identified are "$", "1", and "billion", and these three non-standard characters have the same attributes and are continuous. Therefore, the three non-standard characters can be encoded continuously, and the fifth feature vectors of the three non-standard characters and the hidden vector obtained by the encoder's last encoding are output together. Specifically, the characters "$", "1" and "billion" are first word-embedded to obtain the fourth eigenvector of each non-standard character; then, the fourth eigenvectors of these three characters are used as the input of the encoder, and the encoder first encodes the character "$" for the first time based on the initial hidden vector (i.e., the zero vector) and the fourth eigenvector of the character "$" to obtain the fifth eigenvector of the character "$" and the hidden vector of the first encoding; then, the encoder encodes the character "1" for the second time based on the hidden vector obtained by the first encoding and the fourth eigenvector of the character "1" to obtain the fifth eigenvector of the character "1" and the hidden vector of the second encoding; then, the encoder encodes the character "billion" for the third time based on the new hidden vector of the second encoding and the fourth eigenvector of the character "billion" to obtain the fifth eigenvector of the character "billion" and the last hidden vector; the last hidden vector contains the full-text semantic information of the three non-standard characters.

[0060] 205: Input the word class vector of character B and the fifth eigenvector of character B into the decoder, decode character B, and obtain the regular text of character B.

[0061] Exemplarily, the decoder may be a neural network based on a long short-term memory network, a bidirectional long short-term memory network or a recurrent network. The present application does not limit the type of the decoder.

[0062] Exemplarily, the hidden vector of the last decoded output of the decoder is subjected to an attention mechanism operation with the fifth eigenvector corresponding to character B to obtain the sixth eigenvector corresponding to character B. The attention mechanism can be a general attention mechanism operation. For example, the fifth eigenvector corresponding to character B can be used as a key-value pair, i.e., a key-value vector-value vector; then, the hidden vector of the last decoded output of the decoder is used as a query vector to perform an attention mechanism operation to obtain the sixth eigenvector corresponding to character B. The subsequent attention mechanism operation is similar to this and will not be described again.

[0063] It should be understood that if character B is the first character to be decoded, the hidden vector output by the decoder for the last decoding is the hidden vector output by the encoder for the last encoding; if character B is not the first character to be decoded, the hidden vector output by the decoder for the last decoding is the hidden vector generated when the decoder decoded the previous character. Since the hidden vector output by the decoder for the last decoding (for example, the hidden vector output by the encoder for the last encoding) contains the contextual semantic information of character B, the key information of this decoding can be retained through the attention mechanism operation, thereby improving the decoding accuracy.

[0064] Furthermore, the word class vector of character B, the third feature vector of character B, the sixth feature vector of character B and the decoding result of the last decoding of the decoder are concatenated to obtain the target feature vector of character B; character B is decoded according to the decoding parameters (language vector) of the encoder and the target feature vector of character B to obtain the regular character corresponding to character B. That is, the decoding parameters of the decoder are used to operate the target feature vector to obtain the probability of falling into each character in the standard dictionary, and the standard character corresponding to the maximum probability is used as the regular character of character B.

[0065] The decoding result of the last decoding of the decoder is the decoding result generated by the decoder in the process of decoding the character last time (i.e., the feature vector of the regular character of the last character). It should be understood that if character B is the first character to be decoded, the decoding result of the last decoding is the feature vector of the preset character. For example, if the preset character is the start character S, the feature vector of the start character S is spliced ​​to indicate the start of this decoding.

[0066] Similarly, if character B is one of multiple consecutive non-standard characters with the same attributes (i.e., the word classes are exactly the same, for example, they are all dates in non-standard words) in the text to be regularized, in order to speed up the decoding efficiency and decoding accuracy of the non-standard characters, these multiple non-standard characters will be decoded in sequence according to their fifth eigenvectors, rather than decoding any non-standard character in isolation.

[0067] For example, Figure 3 As shown, the hidden vector e0 output by the encoder for the last encoding is used to 1 ,X 2 ,…,X n The fifth eigenvector [Y 1 ,Y 2 ,…,Y n ] performs attention mechanism operation to obtain a sixth eigenvector. It should be understood that since the hidden vector output by the encoder for the last encoding will contain the multiple non-standard characters [X 1 ,X 2 ,…,Xn ], the attention mechanism operation will focus the decoding attention on the first character to be decoded, thereby improving the decoding accuracy. Then, the sixth feature vector, the word class vector L of the multiple non-standard characters, the third feature vector H of the multiple non-standard characters, and the feature vector ( Figure 3 ) are concatenated to obtain the target feature vector of the first non-standard character to be decoded, wherein, since the attributes of the multiple non-standard characters are the same, the part-of-speech vectors of the multiple non-standard characters can be the part-of-speech vectors of any one of the multiple non-standard characters, and the third feature vectors of the multiple non-standard characters are the average of the third feature vectors of each non-standard character in the multiple non-standard characters. Finally, based on the target feature vector of the first non-standard character to be decoded, the first non-standard character to be decoded is decoded to obtain the decoding result Z of the first decoding. 1 (i.e., the first regular character of the non-standard character to be decoded), and the hidden vector d of the first decoding 1 ; Then, use the hidden vector d of the first decoding 1 、The decoding result Z of the first decoding 1 , a plurality of non-standard character part-of-speech vectors L and a third feature vector H, and a plurality of non-standard character fifth feature vectors [Y 1 ,Y 2 ,…,Y n ], perform the second decoding, and obtain the decoding result of the second decoding (that is, the regular character of the second non-standard character to be decoded) Z 2 , and the hidden vector of the second decoding; repeat the above steps until the multiple non-standard characters [X 1 ,X 2 ,…,X n ] for each non-standard character in the regular character [Z 1 ,Z 2 ,…,Z n ] to stop decoding.

[0068] For example, let’s take the non-standard character “$1billion” as an example to illustrate the decoding process. During the first decoding process, the hidden vector output by the encoder for the last encoding and the fifth eigenvectors of the three non-standard characters are used to perform an attention mechanism operation to obtain a sixth eigenvector (because the character “1” is to be regularized for the first time, the sixth eigenvector focuses on the character “1”); then, the sixth eigenvector, the word class vector (the word class vectors of the three non-standard characters are the same), the third eigenvector (this third eigenvector is the average of the third eigenvectors of each non-standard character) and the eigenvector of the start symbol S are concatenated to obtain a target feature vector; based on the target feature The vector is decoded for the first time to obtain the vector of the character "1" (after mapping this vector, the regular character of the character "1" is "one") and the hidden vector corresponding to the character "1"; then, the second decoding is performed, and the hidden vector output by the first decoding is used to perform an attention mechanism operation with the fifth eigenvectors of the above three characters to obtain a sixth eigenvector. This sixth eigenvector, the word class vector, the third eigenvector and the vector of the character "1" output by the first decoding are concatenated to obtain a target eigenvector vector, which is input into the decoder for decoding to obtain the character "b The vector of "billion" (after mapping this vector, the regular character of "billion" is "billion") and the hidden vector of the second decoding; then, the third decoding is performed, and the hidden vector of the second decoding is used to perform the attention mechanism operation with the fifth eigenvector of the above three characters to obtain a sixth eigenvector, and the sixth eigenvector, the word class vector, the third eigenvector vector and the character "billion" output by the second decoding (the character corresponding to the regular character of "billion") are concatenated to obtain a target eigenvector, and this target eigenvector is input into Decoding is performed in the decoder to obtain the vector of the character "$" (after mapping, the regular vector of "$" is "dollars") and a hidden vector of the decoder; finally, the hidden vector output by the third decoding is used to perform an attention mechanism operation with the fifth eigenvectors of the above three characters to obtain a sixth eigenvector, and the sixth eigenvector, the word class vector, the third eigenvector vector, and the vector of the character "$" output by the second decoding are concatenated to obtain a target eigenvector, which is input into the decoder for decoding, and the end symbol "end" is decoded to indicate that the decoding stops.

[0069] Therefore, through the above encoding and decoding process, the three consecutive standard characters "$1billion" can be regularized into one billion dollars at one time, and then the above text to be regularized can be regularized into "Achieverecord net income of about one billion dollars during the year".

[0070] It can be seen that in the application embodiment, in the process of encoding and decoding non-standard characters, an attention mechanism is used to improve the accuracy of each encoding and decoding. In addition, multiple consecutive non-standard characters with the same attributes can be encoded and decoded synchronously, and information is mutually borrowed during the encoding and decoding process, which improves the efficiency and accuracy of encoding and decoding.

[0071] See also Figure 4 , Figure 4 The functional unit composition block diagram of a text regularization device provided in an embodiment of the present application. The text regularization device 400 includes: an acquisition unit 401 and a processing unit 402, wherein:

[0072] An acquisition unit 401 is used to acquire the text to be regularized;

[0073] The processing unit 402 is used to segment the text to be regularized into characters to obtain a plurality of characters;

[0074] Encode each of the multiple characters to obtain a first feature vector of each of the multiple characters, wherein the first feature vector of each of the multiple characters is used to represent context information of each of the multiple characters;

[0075] According to the first feature vector of each character in the multiple characters and the language type of the text to be regularized, the text to be regularized is regularized to obtain a regularized text of the text to be regularized.

[0076] In some possible implementations, in encoding each of the multiple characters to obtain a first feature vector of each of the multiple characters, the processing unit 402 is specifically configured to:

[0077] Encode each character in the plurality of characters to obtain a character vector corresponding to each character in the plurality of characters;

[0078] Taking character A as the center, constructing a first text corresponding to the character A, the first text including X characters before the character A in the text to be regularized, the character A, and Y characters after the character A in the text to be regularized, the character A is any one of the multiple characters, wherein X and Y are both integers greater than or equal to 1;

[0079] The character vectors corresponding to each character in the first text are concatenated to obtain a first feature vector of the character A, where the first feature vector of the character A is used to represent context information of the character A in the first text.

[0080] In some possible implementations, in performing regularization on the text to be regularized according to the first feature vector of each character in the plurality of characters and the language type of the text to be regularized to obtain the regularized text of the text to be regularized, the processing unit 402 is specifically configured to:

[0081] Determining an attribute of each of the plurality of characters according to a first feature vector of each of the plurality of characters, wherein the attribute of each of the plurality of characters includes a standard character or a non-standard character;

[0082] Using the standard characters in the text to be regularized as the regular characters of the marked characters;

[0083] According to the language type and the first feature vector corresponding to the non-standard character in the text to be regularized, encoding and decoding the non-standard character to obtain a regular character of the non-standard character;

[0084] The regular characters of the standard characters and the regular characters of the non-standard characters in the text to be regularized are combined to obtain the regular text of the text to be regularized.

[0085] In some possible implementations, in terms of performing encoding and decoding processing on the non-standard characters according to the language type and the first feature vector corresponding to the non-standard characters in the text to be regularized to obtain regular characters of the non-standard characters, the processing unit 402 is specifically configured to:

[0086] Taking character B as the center, construct a second text corresponding to the character B, wherein the second text includes M characters before the character B in the text to be regularized, the character B, and N characters after the character B in the text to be regularized, wherein the character B is any non-standard character in the text to be regularized, wherein M and N are both integers greater than or equal to 1;

[0087] Encode each character in the second text by double-byte encoding to obtain a second feature vector for each character in the second text;

[0088] Inputting the second feature vector of each character in the second text into the Transformer-XL network to obtain a third feature vector corresponding to the character B, wherein the third feature vector of the character B is used to represent context information of the character B in the second text;

[0089] According to the attribute of the character B, the third feature vector of the character B and the language type, the character B is encoded and decoded to obtain a regular character corresponding to the character B.

[0090] In some possible implementations, in performing encoding and decoding processing on the character B according to the attribute of the character B, the third feature vector of the character B, and the language type to obtain the regular character corresponding to the character B, the processing unit 402 is specifically configured to:

[0091] Performing word embedding processing on the character B to obtain a fourth feature vector of the character B;

[0092] Encode the attribute of the character B to obtain a word class vector corresponding to the character B;

[0093] Encode the language type to obtain a language vector, and use the language vector as an encoding parameter of an encoder and a decoding parameter of a decoder;

[0094] Inputting the fourth feature vector of the character B into the encoder, encoding the character B, and obtaining a fifth feature vector of the character B;

[0095] The word class vector of the character B and the fifth feature vector of the character B are input into the decoder, and the character B is decoded to obtain the regular text corresponding to the character B.

[0096] In some possible implementations, in inputting the fourth feature vector of the character B to the encoder for encoding to obtain the fifth feature vector of the character B, the processing unit 402 is specifically configured to:

[0097] The character B is encoded according to the hidden layer vector output by the encoder last time, the fourth eigenvector of the character B and the encoding parameter of the encoder to obtain the fifth eigenvector of the character B.

[0098] In some possible implementations, in inputting the word class vector of the character B and the fifth feature vector of the character B into the decoder, decoding the character B, and obtaining the regular text corresponding to the character B, the processing unit 402 is specifically configured to:

[0099] Perform an attention mechanism operation on the hidden layer vector output by the last decoding of the decoder and the fifth eigenvector corresponding to the character B to obtain a sixth eigenvector corresponding to the character B;

[0100] The word class vector of the character B, the third feature vector of the character B, the sixth feature vector of the character B, and the decoding result of the last decoding by the decoder are concatenated to obtain a target feature vector of the character B;

[0101] According to the encoding parameters of the encoder and the target feature vector of the character B, the character B is decoded to obtain the regular character corresponding to the character B.

[0102] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 As shown, the electronic device 500 includes a transceiver 501, a processor 502 and a memory 503. They are connected via a bus 504. The memory 503 is used to store computer programs and data, and can transmit the data stored in the memory 503 to the processor 502.

[0103] The processor 502 is used to read the computer program in the memory 503 and perform the following operations:

[0104] Control transceiver 501 to obtain the text to be regularized;

[0105] Performing character segmentation on the text to be regularized to obtain a plurality of characters;

[0106] Encode each of the multiple characters to obtain a first feature vector of each of the multiple characters, wherein the first feature vector of each of the multiple characters is used to represent context information of each of the multiple characters;

[0107] According to the first feature vector of each character in the multiple characters and the language type of the text to be regularized, the text to be regularized is regularized to obtain a regularized text of the text to be regularized.

[0108] In some possible implementations, in encoding each of the multiple characters to obtain a first feature vector of each of the multiple characters, the processor 502 is specifically configured to perform the following operations:

[0109] Encode each character in the plurality of characters to obtain a character vector corresponding to each character in the plurality of characters;

[0110] Taking character A as the center, constructing a first text corresponding to the character A, the first text including X characters before the character A in the text to be regularized, the character A, and Y characters after the character A in the text to be regularized, the character A is any one of the multiple characters, wherein X and Y are both integers greater than or equal to 1;

[0111] The character vectors corresponding to each character in the first text are concatenated to obtain a first feature vector of the character A, where the first feature vector of the character A is used to represent context information of the character A in the first text.

[0112] In some possible implementations, in terms of performing regularization on the text to be regularized according to the first feature vector of each character in the multiple characters and the language type of the text to be regularized to obtain the regularized text of the text to be regularized, the processor 502 is specifically configured to perform the following operations:

[0113] Determining an attribute of each of the plurality of characters according to a first feature vector of each of the plurality of characters, wherein the attribute of each of the plurality of characters includes a standard character or a non-standard character;

[0114] Using the standard characters in the text to be regularized as the regular characters of the marked characters;

[0115] According to the language type and the first feature vector corresponding to the non-standard character in the text to be regularized, encoding and decoding the non-standard character to obtain a regular character of the non-standard character;

[0116] The regular characters of the standard characters and the regular characters of the non-standard characters in the text to be regularized are combined to obtain the regular text of the text to be regularized.

[0117] In some possible implementations, in terms of performing encoding and decoding processing on the non-standard characters according to the language type and the first feature vector corresponding to the non-standard characters in the text to be regularized to obtain regular characters of the non-standard characters, the processor 502 is specifically configured to perform the following operations:

[0118] Taking character B as the center, construct a second text corresponding to the character B, wherein the second text includes M characters before the character B in the text to be regularized, the character B, and N characters after the character B in the text to be regularized, wherein the character B is any non-standard character in the text to be regularized, wherein M and N are both integers greater than or equal to 1;

[0119] Encode each character in the second text by double-byte encoding to obtain a second feature vector for each character in the second text;

[0120] Inputting the second feature vector of each character in the second text into the Transformer-XL network to obtain a third feature vector corresponding to the character B, wherein the third feature vector of the character B is used to represent context information of the character B in the second text;

[0121] According to the attribute of the character B, the third feature vector of the character B and the language type, the character B is encoded and decoded to obtain a regular character corresponding to the character B.

[0122] In some possible implementations, in terms of performing encoding and decoding processing on the character B according to the attribute of the character B, the third feature vector of the character B, and the language type to obtain a regular character corresponding to the character B, the processor 502 is specifically configured to perform the following operations:

[0123] Performing word embedding processing on the character B to obtain a fourth feature vector of the character B;

[0124] Encode the attribute of the character B to obtain a word class vector corresponding to the character B;

[0125] Encode the language type to obtain a language vector, and use the language vector as an encoding parameter of an encoder and a decoding parameter of a decoder;

[0126] Inputting the fourth feature vector of the character B into the encoder, encoding the character B, and obtaining a fifth feature vector of the character B;

[0127] The word class vector of the character B and the fifth feature vector of the character B are input into the decoder, and the character B is decoded to obtain the regular text corresponding to the character B.

[0128] In some possible implementations, in terms of inputting the fourth feature vector of the character B to the encoder for encoding to obtain the fifth feature vector of the character B, the processor 502 is specifically configured to perform the following operations:

[0129] The character B is encoded according to the hidden layer vector output by the encoder last time, the fourth eigenvector of the character B and the encoding parameter of the encoder to obtain the fifth eigenvector of the character B.

[0130] In some possible implementations, in terms of inputting the word class vector of the character B and the fifth feature vector of the character B into the decoder, decoding the character B, and obtaining the regular text corresponding to the character B, the processor 502 is specifically configured to perform the following operations:

[0131] Perform an attention mechanism operation on the hidden layer vector output by the last decoding of the decoder and the fifth eigenvector corresponding to the character B to obtain a sixth eigenvector corresponding to the character B;

[0132] The word class vector of the character B, the third feature vector of the character B, the sixth feature vector of the character B, and the decoding result of the last decoding by the decoder are concatenated to obtain a target feature vector of the character B;

[0133] According to the encoding parameters of the encoder and the target feature vector of the character B, the character B is decoded to obtain the regular character corresponding to the character B.

[0134] Specifically, the transceiver 501 may be Figure 4 The acquisition unit 401 and the processor 502 of the text regularization device 400 of the embodiment described above may be Figure 4 The processing unit 402 of the text regularization device 400 of the embodiment described.

[0135] It should be understood that the text regularization device in the present application may include a smart phone (such as an Android phone, an iOS phone, a Windows Phone phone, etc.), a tablet computer, a PDA, a laptop computer, a mobile Internet device MID (Mobile Internet Devices, MID for short) or a wearable device, etc. The above text regularization device is only an example, not an exhaustive list, including but not limited to the above text regularization device. In practical applications, the above text regularization device may also include: a smart vehicle terminal, a computer device, etc.

[0136] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. The computer program is executed by a processor to implement part or all of the steps of any text regularization method recorded in the above method embodiments.

[0137] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps of any text regularization method recorded in the above method embodiments.

[0138] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.

[0139] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0140] In the several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of the units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical or other forms.

[0141] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software program module.

[0143] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or CD-ROM and other media that can store program codes.

[0144] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (English: Read-Only Memory, abbreviated as: ROM), a random access memory (English: Random Access Memory, abbreviated as: RAM), a magnetic disk or an optical disk, etc.

[0145] The embodiments of the present application are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A text regularization method, It is characterized in that include: Get the text to be regularized; Performing character segmentation on the text to be regularized to obtain a plurality of characters; Encode each of the multiple characters to obtain a first feature vector of each of the multiple characters, wherein the first feature vector of each of the multiple characters is used to represent context information of each of the multiple characters; According to the first feature vector of each character in the plurality of characters and the language type of the text to be regularized, regularizing the text to be regularized to obtain a regularized text of the text to be regularized; comprising: Determine the attribute of each character in the plurality of characters according to the first feature vector of each character, wherein the attribute of each character includes a standard character or a non-standard character; and use the standard character in the text to be regularized as a regular character of the standard character; Perform word embedding processing on character B to obtain a fourth feature vector of character B; encode the attribute of character B to obtain a word class vector of character B; encode the language type to obtain a language vector of the text to be regularized, and use the language vector as an encoding parameter of an encoder and a decoding parameter of a decoder, respectively, wherein character B is any non-standard character in the text to be regularized; Input the fourth feature vector of the character B into the encoder, encode the character B, and obtain the fifth feature vector of the character B; input the third feature vector of the character B, the word class vector of the character B, and the fifth feature vector of the character B into the decoder, decode the character B, and obtain the regular text corresponding to the character B, wherein the third feature vector of the character B is used to represent the context information of the character B in the second text, and the second text is a text constructed with the character B as the center in the text to be regularized; The regular characters of the standard characters and the regular characters of the non-standard characters in the text to be regularized are combined to obtain the regular text of the text to be regularized.

2. The method according to claim 1, It is characterized in that The step of encoding each of the plurality of characters to obtain a first feature vector of each of the plurality of characters comprises: Encode each character in the plurality of characters to obtain a character vector corresponding to each character in the plurality of characters; Taking character A as the center, constructing a first text corresponding to the character A, the first text including X characters before the character A in the text to be regularized, the character A, and Y characters after the character A in the text to be regularized, the character A is any one of the multiple characters, wherein X and Y are both integers greater than or equal to 1; The character vectors corresponding to each character in the first text are concatenated to obtain a first feature vector of the character A, where the first feature vector of the character A is used to represent context information of the character A in the first text.

3. The method according to claim 1, It is characterized in that The method further comprises: Taking character B as the center, construct a second text corresponding to character B, wherein the second text includes M characters before character B in the text to be regularized, character B, and N characters after character B in the text to be regularized, wherein M and N are both integers greater than or equal to 1; Encode each character in the second text by double-byte encoding to obtain a second feature vector for each character in the second text; The second feature vector of each character in the second text is input into the Transformer-XL network to obtain a third feature vector corresponding to the character B, where the third feature vector of the character B is used to represent context information of the character B in the second text.

4. The method according to claim 1, It is characterized in that The fourth feature vector of the character B is input into the encoder for encoding to obtain the fifth feature vector of the character B, including The character B is encoded according to the hidden layer vector output by the encoder last time, the fourth eigenvector of the character B and the encoding parameter of the encoder to obtain the fifth eigenvector of the character B.

5. The method according to claim 1, It is characterized in that The step of inputting the third feature vector of the character B, the part-of-speech vector of the character B, and the fifth feature vector of the character B into the decoder, decoding the character B, and obtaining the regular text corresponding to the character B includes: Perform an attention mechanism operation on the hidden layer vector output by the last decoding of the decoder and the fifth eigenvector corresponding to the character B to obtain a sixth eigenvector corresponding to the character B; The word class vector of the character B, the third feature vector of the character B, the sixth feature vector of the character B, and the decoding result of the last decoding by the decoder are concatenated to obtain a target feature vector of the character B; According to the encoding parameters of the encoder and the target feature vector of the character B, the character B is decoded to obtain the regular character corresponding to the character B.

6. A text regularization device, It is characterized in that The text regularization device is used to implement the method according to any one of claims 1 to 5, and the text regularization device comprises: An acquisition unit, used to acquire the text to be regularized; A processing unit, used for segmenting the text to be regularized into characters to obtain a plurality of characters; Encode each of the multiple characters to obtain a first feature vector of each of the multiple characters, wherein the first feature vector of each of the multiple characters is used to represent context information of each of the multiple characters; According to the first feature vector of each character in the multiple characters and the language type of the text to be regularized, the text to be regularized is regularized to obtain a regularized text of the text to be regularized.

7. An electronic device, It is characterized in that include: A processor and a memory, the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Text normalization method, device and equipment and storage medium

    CN110765733A

  • Text normalizing method and device, electronic equipment and storage medium

    CN111832248A