A method for address element recognition based on T-BiLSTM and CRF with specified position forgetting

By introducing T-BiLSTM and CRF methods of a specified location forgetting mechanism in address feature recognition, the difficulty of separating multiple elements in long-sequence address feature recognition is solved, and the accuracy and robustness of the recognition are improved.

CN114880999BActive Publication Date: 2025-05-13ZHEJIANG BANGSUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210578633.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2025-05-13
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

Existing address feature recognition methods are difficult to correctly separate multiple address elements when processing long sequences, and traditional methods are limited in semantic information extraction.

Method used

The address element recognition method based on T-BiLSTM and CRF based on the specified position forgetting, and segment the long sequence into short sequences by forgetting information at the specified word segmentation position, thereby solving the difficulties of deep learning methods on long sequences.

Benefits of technology

It improves the accuracy and robustness of address feature recognition, can capture address feature information more effectively, and solves the problem that multiple address elements cannot be separated by long sequence labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114880999B_ABST
    Figure CN114880999B_ABST
Patent Text Reader

Abstract

The present invention discloses an address element recognition method based on T-BiLSTM and CRF with forgotten specified positions. The method constructs a neural network based on BiLSTM with forgotten specified time steps. The neural network is named T-BiLSTM. The method first converts the address text encoding into a vector matrix based on word information; then the address vectors are respectively input into the address word segmentation network of BiLSTM for word segmentation; after obtaining the word segmentation vector, it is combined with the address vector and input into the T-BiLSTM neural network; finally, the result based on the T-BiLSTM neural network is annotated using the conditional random field CRF to obtain the address elements of each level of the address. Compared with the traditional address element method based on BiLSTM-CRF, the method separates the word segmentation and annotation tasks, introduces the word segmentation forgetting information when identifying elements, and has better accuracy and robustness for new word recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Chinese word segmentation in natural language processing, and in particular to an address element recognition method based on T-BiLS TM and CRF with designated position forgetting. Background Art

[0002] With the development of express delivery and e-commerce industries, there is a huge amount of address information. Real address verification and address matching are becoming more and more important. The address element recognition task is an important part of address verification and address matching, and it plays a decisive role in whether address verification and address matching can be successful. Traditional address element recognition generally adopts a rule-based address element recognition method, which can only achieve certain results on standard sequences, and the rules are numerous. Recently, machine learning-based methods have been used to avoid writing numerous rules, but the semantic information extracted is limited. Based on the deep learning method, since there are a large number of address sequences starting with provinces, cities and districts in the real address sequence, the model will learn this distribution, resulting in the inability to distinguish the provinces, cities and districts that appear in the later positions, and this method will encounter the problem of combining multiple address elements in long sequences.

[0003] In view of the above shortcomings in address element recognition, a T-BiLSTM and CRF address element recognition method based on specified position forgetting is proposed. The model of this method will combine the word segmentation information to forget the information at the specified word segmentation position, split the long sequence into short sequences, and thus solve the problems encountered by the above deep learning methods. Summary of the invention

[0004] The purpose of the present invention is to address the deficiencies of the prior art and propose an address element recognition method based on T-BiLSTM and CRF with specified position forgetting.

[0005] The object of the present invention is to achieve the following technical solution: a method for identifying address elements based on T-BiLSTM and CRF with forgotten specified positions, comprising the following steps:

[0006] Step 1: Use a crawler to crawl the address text on the Internet to obtain an initial address data set, perform data preprocessing on the obtained initial address data set, perform manual address element segmentation and labeling on the preprocessed address data set, obtain address element data after segmentation and labeling, perform statistical deduplication on address characters to obtain a character set, and convert the address element data into an address character id set according to the character set;

[0007] Step 2: randomly initialize the character set obtained in step 1 as a feature vector, and convert the address character ID set obtained in step 1 into an address feature vector matrix according to the feature vector;

[0008] Step 3: Input the address feature vector matrix obtained in step 2 into the BiLSTM model to obtain a semantic feature matrix;

[0009] Step 4: Input the semantic feature matrix obtained in step 3 into the Dense fully connected layer to obtain the sequence segmentation information;

[0010] Step 5: Input the address feature vector matrix obtained in step 2 and the sequence segmentation information obtained in step 4 into T-BiLSTM to obtain the feature matrix, and convert it into a score sequence matrix through a fully connected neural network;

[0011] Step 6: Input the score sequence matrix obtained in step 5 into the conditional random field CRF to obtain the Chinese address element labeling result.

[0012] Furthermore, the step 1 comprises:

[0013] (1) Use a crawler to crawl the address text on the Internet to obtain an initial address data set;

[0014] (2) All characters except Chinese characters, letters, and numbers are removed from the initial data set, and all letters AZ are converted to lowercase az;

[0015] (3) Manually annotate the address elements. The annotation framework uses X, R1, R2, R3, R4, R5, R6, R7, R20, R21, R22, R23, R24, R25, R30, R31, R90, and R99 tags to annotate the address data. The beginning and middle of the word are marked with X, and the character at the end of the word is marked with the corresponding level. R1 represents the province, R2 represents the city, R3 represents the district and county, R4 represents the street and town, R5 represents the road and village, R6 represents the house number and road number, R7 represents the unit, R20 represents the company, R21 represents the scientific and educational institution, R22 represents the medical institution, R23 represents the bank, R24 represents the government agency, R30 represents the residential area, R31 represents the commercial building, R90 represents the direction word, R99 represents the redundant word, and R25 represents other points of interest (POI) except the above.

[0016] (4) Manually segment the address elements. The labeling framework uses 0 and 1 labels to label the address data. The beginning and middle of the word are marked with 0, and the end of the word is marked with 1.

[0017] (5) Perform character count on all initial address texts, then add <pad>Filling characters with <unknow>Unknown characters, let each unique character in the character set as its unique id identifier, map the initial address text with its character id to get the address id sequence. Convert an address into an address character id set:

[0018] address={word1_id,word2_id,…word3_id}

[0019] Among them, wordi_id is the i-th character in the character set.

[0020] Furthermore, the step 2 comprises:

[0021] (1) Each character in the character set obtained in step 1 is randomly initialized as a feature vector to obtain a character feature vector matrix E, where the dimension of E is N*M, where N is the total number of characters in the character set and M is the dimension of each character feature vector.

[0022] (2) The address character id set obtained in step 1 is converted into an address feature vector matrix A according to the character feature vector matrix E. The dimension of A is L*M, where L is the number of characters in the longest sequence of this batch.

[0023] Furthermore, the step 3 comprises:

[0024] (1) The address feature vector matrix A obtained in step 2 is input into the BiLSTM neural network. The bidirectional LSTM uses the concat method to combine vectors. A single unit in the LSTM includes three parts: a forget gate, a memory gate, and an output gate. The calculation formula of the gate control unit at time t is as follows:

[0025] f t =(Wf*[h t-1 ,x t ]+bf)

[0026] i t =(Wi*[h t-1 ,x t ]+bi)

[0027] C_ t =tanh(Wc*[h t-1 ,x t ]+bC)

[0028] C t =ft*C t-1 +it*C_ t

[0029] O t =(Wo*[h t-1 ,x t ]+bo)

[0030] h t =O t *tanh(C t )

[0031] where f t Represents the result of the forget gate at the current moment, i t represents the result of the memory gate at the current moment, C_ t is the temporary memory cell result at the current moment, C t is the memory cell result at the current moment, O t Represents the output gate result at the current moment, h t represents the hidden layer state matrix at the current moment, Wf, Wi, Wo represent the parameter matrices of the forget gate, memory gate, and output gate respectively, Wc represents the memory cell parameter matrix, bf, bi, bo represent the bias of the corresponding gate control unit, bC represents the memory cell result bias, h t-1 Represents the hidden state matrix of the previous moment, x t is the address feature vector at the current moment, C t-1 Memory cell state matrix for the previous moment.

[0032] (2) The address feature vector matrix A is input into the forward LSTM and the backward LSTM respectively to obtain the forward result L_F_O and the backward result L_B_O. The two results are concatenated to obtain the semantic feature matrix L_B with a dimension of L*V as the final BiLSTM neural network result, where V is the number of BiLSTM units.

[0033] Furthermore, the step 4 comprises:

[0034] The semantic feature matrix L_B obtained in step 3 is input into the Dense fully connected layer with two units. After the two units output the results, argmax calculation is performed to obtain the 0, 1 word segmentation results S = {C1, C2...}, Ci∈{0,1}, and the calculation formula is as follows:

[0035] D_O=L_B*D0+bd

[0036] S = argmax(D_O)

[0037] Where D_O is the result of the Dense fully connected layer, S is the argmax word segmentation result, the matrix dimension is 1*L, D0 is the Dense layer parameter, the dimension is V*2, and bd is the bias value of the fully connected layer.

[0038] Furthermore, the step 5 comprises:

[0039] (1) The address feature vector matrix A obtained in step 2 and the word segmentation information S obtained in step 4 are input into T-BiLSTM, and the forward TF-LSTM and the backward TB-LSTM are concatenated by the concat method;

[0040] A single unit in TF-LSTM consists of three parts: forget gate, memory gate, and output gate. The calculation formula of the gate control unit at time t is as follows:

[0041] h_n=h t-1 *(1-S t-1 )+h t0 *(S t-1 )

[0042] f t =(Wf*[h_n,x t ]+bf)

[0043] i t =(Wi*[h_n,x t ]+bi)

[0044] C_ t =tanh(Wc*[h_n,x t ]+bC)

[0045] C t =f t *C t-1 +it*C_ t

[0046] O t =(Wo*[h_n,x t ]+bo)

[0047] h t =O t *tanh(C t )

[0048] Where S t-1 represents the segmentation result of S at the previous moment, h t0 represents the initial hidden state matrix, h_n represents the result of the hidden state matrix at the previous moment after memory calculation using the word segmentation result, and f t Represents the result of the forget gate at the current moment, i t represents the result of the memory gate at the current moment, C_ t is the temporary memory cell result at the current moment, C t is the memory cell result at the current moment, O t Represents the output gate result at the current moment, h t represents the hidden layer state at the current moment, Wf, Wi, Wo represent the parameter matrices of the three gate control units, Wc represents the memory cell parameter matrix, bf, bi, bC, bo represent the bias, h t-1 Represents the hidden state matrix of the previous moment, x t is the address feature vector at the current moment, C t-1 Memory cell state matrix for the previous moment.

[0049] A single unit in TB-LSTM consists of three parts: forget gate, memory gate, and output gate. The calculation formula of the gate control unit at time t is as follows:

[0050] h_n=h t-1 *(1-S t )+h t0 *(S t )

[0051] f t =(Wf*[h_n,x t ]+bf)

[0052] i t =(Wi*[h_n,x t ]+bi)

[0053] C_ t =tanh(Wc*[h_n,x t ]+bC)

[0054] C t =f t *C t-1 +i t *C_ t

[0055] O t =(Wo*[h_n,x t ]+bo)

[0056] h t =O t *tanh(C t )

[0057] Where S t Represents the word segmentation result of S at the current moment. The meanings of other parameters are the same as above.

[0058] Input the address feature vector matrix A into the forward TF-LSTM and the backward TB-LSTM respectively to obtain the forward result TL_F_O and the backward result TL_B_O. Concatenate the two results to obtain the feature matrix TL_B with the dimension of L*V as the final T-BiLSTM neural network result.

[0059] (2) Input TL_B into the fully connected neural network and transform it into the score sequence matrix K, whose calculation formula is as follows:

[0060] K=TL_B*D1

[0061] Where D1 is the fully connected neural network parameter, the dimension is V*M, and M is the number of annotation frame labels.

[0062] Further, the step 6 comprises:

[0063] The score sequence matrix K obtained in step 5 is input into the conditional random field CRF, and the Viterbi algorithm is used to select the maximum probability result sequence as the final prediction result.

[0064] The advantages of the present invention are: the use of a double-layer BiLSTM can capture better address feature information, and the use of T-BiLSTM can solve the data distribution problem and the problem that multiple address elements cannot be separated when annotating a long address sequence. Compared with the traditional address element method based on Bi LSTM-CRF, the present invention separates word segmentation from the labeling task, introduces word segmentation forgetting information when identifying elements, and has better accuracy and robustness for new word recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is a flow chart of the method of the present invention;

[0066] Figure 2 Schematic diagram of the T-BiLSTM-CRF neural network structure.

[0067] Figure 3 Address word segmentation element labeling. DETAILED DESCRIPTION

[0068] The specific implementation modes of the present invention are further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not used to limit the present invention.

[0069] like Figure 1 As shown, the present invention provides a method for identifying address elements based on T-BiLSTM and CRF with specified position forgetting, the method comprising the following steps:

[0070] Step 1: Use a crawler to crawl the address text on the Internet to obtain an initial address data set, perform data preprocessing on the obtained initial address data set, perform manual address element segmentation and labeling on the preprocessed address data set, obtain address element data after segmentation and labeling, perform statistical deduplication on address characters to obtain a character set, and convert the address element data into an address character id set according to the character set; the specific process of step 1 is as follows:

[0071] All characters except Chinese characters, letters, and numbers are removed from the initial text, and all letters AZ are converted to lowercase az.

[0072] like Figure 2 and Figure 3 As shown in the figure, the address elements are manually annotated, and the annotation framework uses X, R1, R2, R3, R4, R5, R6, R7, R20, R21, R22, R23, R24, R25, R30, R31, R90, and R99 tags to annotate the address data. The beginning and middle of the word are marked with X, and the character at the end of the word is marked with the corresponding level. R1 represents the province, R2 represents the city, R3 represents the district and county, R4 represents the street and town, R5 represents the road and village, R6 represents the house number and road number, R7 represents the unit, R20 represents the company, R21 represents the scientific and educational institution, R22 represents the medical institution, R23 represents the bank, R24 represents the government agency, R30 represents the residential area, R31 represents the commercial building, R90 represents the direction word, R99 represents the redundant word, and R25 represents other points of interest POI except the above parts.

[0073] Take "Anhui Province|Xuancheng City|Ningguo City|Ningxin Garden|Chonglixin Supermarket" as an example, where "|" is a word segmentation marker, the corresponding annotation sequence of this address is "XX R1 XX R2 XX R3 XXX R30 XXXX R25"

[0074] The address elements are manually segmented, and the labeling framework uses 0 and 1 labels to annotate the address data. The beginning and middle of the word are marked with 0, and the character at the end of the word is marked with 1.

[0075] Take "Anhui Province|Xuancheng City|Ningguo City|Ningxin Garden|Chonglixin Supermarket" as an example, where "|" is a word segmentation marker, the corresponding annotation sequence of this address is "0 0 1 0 0 1 0 0 1 0 0 0 1 0 0 0 1"

[0076] Perform character count on all initial address texts, then add <pad>Filling characters with <unknow>Unknown characters, let each unique character in the character set as its unique id identifier, map the initial address text with its character id to get the address id sequence. Convert an address into an address character id set:

[0077] address={word1_id,word2_id,…word3_id}

[0078] Among them, wordi_id is the i-th character in the character set.

[0079] Step 2: randomly initialize the character set obtained in step 1 as a feature vector, and convert the address character id set obtained in step 1 into an address feature vector matrix according to the feature vector; the specific process of step 2 is as follows:

[0080] Each character in the character set obtained in step 1 is randomly initialized as a feature vector to obtain a character feature vector matrix E, where the dimension of E is N*M, where N is the total number of characters in the character set and M is the dimension of each character feature vector.

[0081] The address character id set obtained in step 1 is converted into the address feature vector matrix A according to the character feature vector matrix E. The dimension of A is L*M, where L is the number of characters in the longest sequence of this batch.

[0082] Step 3: Input the address feature vector matrix obtained in step 2 into the BiLSTM model to obtain the semantic feature matrix; the specific process of step 3 is as follows:

[0083] The address feature vector matrix A obtained in step 2 is input into the BiLSTM neural network. The bidirectional LSTM uses the concat method to combine vectors. A single unit in the LSTM includes three parts: the forget gate, the memory gate, and the output gate. The calculation formula of the gate control unit at time t is as follows:

[0084] f t =(Wf*[h t-1 ,x t ]+bf)

[0085] i t =(Wi*[h t-1 ,x t ]+bi)

[0086] C_ t =tanh(Wc*[h t-1 ,x t ]+bC)

[0087] C t =f t *C t-1 +i t *C_ t

[0088] O t =(Wo*[h t-1 ,x t ]+bo)

[0089] h t =O t *tanh(C t )

[0090] where f t Represents the result of the forget gate at the current moment, i t represents the result of the memory gate at the current moment, C_ t is the temporary memory cell result at the current moment, C t is the memory cell result at the current moment, O t Represents the output gate result at the current moment, h t represents the hidden layer state matrix at the current moment, Wf, Wi, Wo represent the parameter matrices of the forget gate, memory gate, and output gate respectively, Wc represents the memory cell parameter matrix, bf, bi, bo represent the bias of the corresponding gate control unit, bC represents the memory cell result bias, h t-1 Represents the hidden state matrix of the previous moment, x t is the address feature vector at the current moment, C t-1 Memory cell state matrix for the previous moment.

[0091] The address feature vector matrix A is input into the forward LSTM and the backward LSTM respectively to obtain the forward result L_F_O and the backward result L_B_O. The two results are concatenated to obtain the semantic feature matrix L_B with a dimension of L*V as the final BiLSTM neural network result, where V is the number of BiLSTM units.

[0092] Step 4: Input the semantic feature matrix obtained in step 3 into the Dense fully connected layer to obtain the sequence segmentation information; the specific process of step 4 is as follows:

[0093] The semantic feature matrix L_B obtained in step 3 is input into the Dense fully connected layer with two units. After obtaining the result, argmax calculation is performed to obtain the 0, 1 word segmentation result S = {C1, C2...}, Ci∈{0,1}. The calculation formula is as follows:

[0094] D_O=L_B*D0+bd

[0095] S = argmax(D_O)

[0096] Among them, D_O is the result of the fully connected layer, S is the argmax segmentation result matrix with a dimension of 1*L, D0 is the Dense layer parameter with a dimension of V*2, and bd is the bias value of the fully connected layer.

[0097] Step 5: Input the address feature vector obtained in step 2 and the sequence segmentation information obtained in step 4 into T-BiLSTM to obtain the feature matrix, and convert it into a score sequence matrix through a fully connected neural network; the specific process of step 5 is as follows:

[0098] The address feature vector matrix A obtained in step 2 and the word segmentation information S obtained in step 4 are input into T-BiLSTM, and the forward TF-LSTM and the backward TB-LSTM use the concat method to combine the vectors.

[0099] A single unit in TF-LSTM consists of three parts: forget gate, memory gate, and output gate. The calculation formula of the gate control unit at time t is as follows:

[0100] h_n=h t-1 *(1-S t-1 )+h t0 *(S t-1 )

[0101] f t =(Wf*[h_n,x t ]+bf)

[0102] i t =(Wi*[h_n,x t ]+bi)

[0103] C_ t =tanh(Wc*[h_n,x t ]+bC)

[0104] C t =f t *C t-1 +i t *C_ t

[0105] O t =(Wo*[h_n,x t ]+bo)

[0106] h t =O t *tanh(C t )

[0107] Where S t-1 represents the segmentation result of S at the previous moment, h t0 represents the initial hidden state matrix, h_n represents the result of the hidden state matrix at the previous moment after memory calculation using the word segmentation result, and f t Represents the result of the forget gate at the current moment, i t represents the result of the memory gate at the current moment, C_ t is the temporary memory cell result at the current moment, C t is the memory cell result at the current moment, O t Represents the output gate result at the current moment, h t represents the hidden layer state at the current moment, Wf, Wi, Wo represent the parameter matrices of the three gate control units, Wc represents the memory cell parameter matrix, bf, bi, bC, bo represent the bias, h t-1 Represents the hidden state matrix of the previous moment, x t is the address feature vector at the current moment, C t-1 Memory cell state matrix for the previous moment.

[0108] A single unit in TB-LSTM consists of three parts: forget gate, memory gate, and output gate. The calculation formula of the gate control unit at time t is as follows:

[0109] h_n=h t-1 *(1-S t )+h t0 *(S t )

[0110] f t =(Wf*[h_n,x t ]+bf)

[0111] i t =(Wi*[h_n,x t ]+bi)

[0112] C_ t =tanh(Wc*[h_n,x t ]+bC)

[0113] C t =f t *C t-1 +i t *C_ t

[0114] O t =(Wo*[h_n,x t ]+bo)

[0115] h t =O t *tanh(C t )

[0116] Where S t Represents the word segmentation result of S at the current moment. The meanings of other parameters are the same as above.

[0117] The address feature vector matrix A is input into the forward TF-LSTM and the backward TB-LSTM respectively to obtain the forward result TL_F_O and the backward result TL_B_O. The two results are concatenated to obtain the feature matrix TL_B with the dimension of L*V as the final T-BiLSTM neural network result.

[0118] Input TL_B into the fully connected neural network and transform it into the score sequence matrix K, whose calculation formula is as follows:

[0119] K=TL_B*D1

[0120] Where D1 is the fully connected layer parameter, the dimension is V*M, and M is the number of annotation frame labels.

[0121] Step 6: Input the score sequence matrix obtained in step 5 into the conditional random field CRF to obtain the Chinese address element labeling result; the specific process of step 6 is as follows:

[0122] The score sequence matrix K obtained in step 5 is input into the conditional random field CRF, and the calculation process is as follows:

[0123]

[0124] Where Z(x) is the normalization factor, λ k and μ l is the feature function weight, t k is the kth transfer characteristic function, s l is the characteristic function of the lth state;

[0125] The Viterbi algorithm is used to select the maximum probability result sequence as the final prediction result.

[0126] The model is trained using a forward-backward algorithm, and its loss function calculation formula is as follows:

[0127] Loss=CRF_loss(y_r,p_r)+CE(y_c,p_c)

[0128] Among them, CRF_loss is the CRFloss function, CE is the cross entropy loss function, y_r is the true feature label, p_r is the predicted feature label, y_c is the true word segmentation label, p_c is the predicted word segmentation label, which is the softmax of the fully connected Dense layer result D_O in step 4.

[0129] The above examples are used to explain the present invention rather than to limit the present invention. For those skilled in the art in the technical field of the present invention, replacements or modifications may be made without departing from the scope of protection of the claims. The same uses shall be deemed to be within the scope of protection of the present invention.< / unknow> < / pad> < / unknow> < / pad>

Claims

1. A method for identifying address elements based on T-BiLSTM and CRF with specified position forgetting, characterized in that: The following steps are involved: Step 1: Use a crawler to crawl the address text on the Internet to obtain an initial address data set, perform data preprocessing on the obtained initial address data set, perform manual address element segmentation and labeling on the preprocessed address data set, obtain address element data after segmentation and labeling, perform statistical deduplication on address characters to obtain a character set, and convert the address element data into an address character id set according to the character set; Step 2: randomly initialize the character set obtained in step 1 as a feature vector, and convert the address character ID set obtained in step 1 into an address feature vector matrix according to the feature vector; specifically, the following steps are performed: (2.1) Each character in the character set obtained in step 1 is randomly initialized as a feature vector to obtain a character feature vector matrix E, where the dimension of E is N*M, where N is the total number of characters in the character set and M is the dimension of each character feature vector; (2.2) The address character id set obtained in step 1 is converted into the address feature vector matrix A according to the character feature vector matrix E. The dimension of A is L*M, where L is the number of characters in the longest sequence of this batch; Step 3: Input the address feature vector matrix obtained in step 2 into the BiLSTM model to obtain a semantic feature matrix; Step 4: Input the semantic feature matrix obtained in step 3 into the Dense fully connected layer to obtain the sequence segmentation result; Step 5: Input the address feature vector matrix obtained in step 2 and the sequence segmentation result obtained in step 4 into T-BiLSTM to obtain the feature matrix, and convert it into a score sequence matrix through a fully connected neural network; specifically, the following steps are included: (1) The address feature vector matrix A obtained in step 2 and the word segmentation result S obtained in step 4 are input into T-BiLSTM, and the forward TF-LSTM and the backward TB-LSTM are concatenated by the concat method; A single unit in TF-LSTM consists of three parts: forget gate, memory gate, and output gate. The calculation formula of the gate control unit at time t is as follows: h_n=h t-1 *(1-S t-1 )+h t0 *(S t-1 ) f t =(Wf*[h_n,x t ]+bf) i t =(Wi*[h_n,x t ]+bi) C_ t =tanh(Wc*[h_n,x t ]+bC) C t =f t *C t-1 +i t *C_ t Oh t =(Wo*[h_n,x t ]+bo) h t =O t *tanh(C t ) Where S t-1 represents the segmentation result of S at the previous moment, h t0 represents the initial hidden state matrix, h_n represents the result of the hidden state matrix at the previous moment after memory calculation using the word segmentation result, and f t Represents the result of the forget gate at the current moment, i t represents the result of the memory gate at the current moment, C_ t is the temporary memory cell result at the current moment, C t is the memory cell result at the current moment, O t Represents the output gate result at the current moment, h t represents the current hidden state, Wf, Wi, Wo represent the parameter matrices of the forget gate, memory gate, and output gate, respectively, Wc represents the memory cell parameter matrix, bf, bi, bC, bo represent the bias, h t-1 Represents the hidden state matrix of the previous moment, x t is the address feature vector at the current moment, C t-1 Memory cell state matrix for the previous moment; A single unit in TB-LSTM consists of three parts: forget gate, memory gate, and output gate. The calculation formula of the gate control unit at time t is as follows: h_n=h t-1 *(1-S t )+h t0 *(S t ) f t =(Wf*[h_n,x t ]+bf) i t =(Wi*[h_n,x t ]+bi) C_ t =tanh(Wc*[h_n,x t ]+bC) C t =f t *C t-1 +i t *C_ t Oh t =(Wo*[h_n,x t ]+bo) h t =O t *tanh(C t ) Where S t Represents the word segmentation result of S at the current moment; Input the address feature vector matrix A into the forward TF-LSTM and the backward TB-LSTM respectively to obtain the forward result TL_F_O and the backward result TL_B_O. Concatenate the two results to obtain the feature matrix TL_B with the dimension of L*V as the final T-BiLSTM neural network result. (2) Input TL_B into the fully connected neural network and transform it into the score sequence matrix K, which is calculated as follows: K=TL_B*D1 Where D1 is the fully connected neural network parameter, the dimension is V*M, and M is the number of annotation frame labels; Step 6: Input the score sequence matrix obtained in step 5 into the conditional random field CRF to obtain the Chinese address element labeling result.

2. The address element recognition method based on T-BiLSTM and CRF with specified position forgetting according to claim 1, characterized in that: The step 1 comprises: (1) Use a crawler to crawl the address text on the Internet to obtain an initial address data set; (2) Remove all characters except Chinese characters, letters, and numbers from the initial address data set, and convert all letters AZ to lowercase az; (3) Manually annotate the address elements. The annotation framework uses X, R1, R2, R3, R4, R5, R6, R7, R20, R21, R22, R23, R24, R25, R30, R31, R90, and R99 tags to annotate the address data. The beginning and middle of the word are marked with X, and the character at the end of the word is marked with the corresponding level. R1 represents the province, R2 represents the city, R3 represents the district and county, R4 represents the street and town, R5 represents the road and village, R6 represents the house number and road number, R7 represents the unit, R20 represents the company, R21 represents the scientific and educational institution, R22 represents the medical institution, R23 represents the bank, R24 represents the government agency, R30 represents the residential area, R31 represents the commercial building, R90 represents the direction word, R99 represents the redundant word, and R25 represents other points of interest (POI) except the above. (4) Manually segment the address elements. The annotation framework uses 0 and 1 labels to annotate the address data. The beginning and middle of a word are marked with 0, and the character at the end of the word is marked with 1. (5) Perform character count on all initial address texts, then add <pad>Filling characters with <unknow> Unknown characters, let each unique character in the character set as its unique id identifier, map the initial address text with its character id to get the address id sequence; convert an address into an address character id set:< / unknow> < / pad> address={word1_id,word2_id,…word3_id} Among them, wordi_id is the i-th character in the character set.

3. The address element recognition method based on T-BiLSTM and CRF with specified position forgetting according to claim 1, characterized in that: The step 3 comprises: (1) The address feature vector matrix A obtained in step 2 is input into the BiLSTM neural network. The bidirectional LSTM uses the concat method to combine vectors. A single unit in the LSTM includes three parts: a forget gate, a memory gate, and an output gate. The calculation formula of the gate control unit at time t is as follows: f t =(Wf*[h t-1 ,x t ]+bf) i t =(Wi*[h t-1 ,x t ]+bi) C_ t =tanh(Wc*[h t-1 ,x t ]+bC) C t =f t *C t-1 +i t *C_ t Oh t =(Wo*[h t-1 ,x t ]+bo) h t =O t *tanh(C t ) where f t Represents the result of the forget gate at the current moment, i t represents the result of the memory gate at the current moment, C_ t is the temporary memory cell result at the current moment, C t is the memory cell result at the current moment, O t Represents the output gate result at the current moment, h t represents the hidden layer state matrix at the current moment, Wf, Wi, Wo represent the parameter matrices of the forget gate, memory gate, and output gate respectively, Wc represents the memory cell parameter matrix, bf, bi, bo represent the bias of the corresponding gate control unit, bC represents the memory cell result bias, h t-1 Represents the hidden state matrix of the previous moment, x t is the address feature vector at the current moment, C t-1 Memory cell state matrix for the previous moment; (2) The address feature vector matrix A is input into the forward LSTM and the backward LSTM respectively to obtain the forward result L_F_O and the backward result L_B_O. The two results are concatenated to obtain the semantic feature matrix L_B with a dimension of L*V as the final BiLSTM neural network result, where V is the number of BiLSTM units.

4. The address element recognition method based on T-BiLSTM and CRF with specified position forgetting according to claim 3 is characterized in that: The step 4 comprises: The semantic feature matrix L_B obtained in step 3 is input into the Dense fully connected layer with two units. After the two units output the results, argmax calculation is performed to obtain the 0, 1 word segmentation results S = {C1, C2, ..., CL}, CL∈{0,1}, and the calculation formula is as follows: D_O=L_B*D0+bd S = argmax(D_O) Where D_O is the result of the Dense fully connected layer, S is the argmax word segmentation result, the matrix dimension is 1*L, D0 is the Dense layer parameter, the dimension is V*2, and bd is the bias value of the fully connected layer.

5. The address element recognition method based on T-BiLSTM and CRF with specified position forgetting according to claim 1, characterized in that: The step 6 comprises: The score sequence matrix K obtained in step 5 is input into the conditional random field CRF, and the Viterbi algorithm is used to select the maximum probability result sequence as the final prediction result.

Citation Information

Patent Citations

  • Memory and attention model-based auditory selection method and device

    CN108109619A

  • Address identification method and device and storage medium

    CN108920457A