An improved segmented inference address matching method

By introducing a combination of sectional inference and ESIM model in the address matching method, different geographical elements of addresses are extracted and matched, and the existing methods are solved, and the problem of sensitive address length and low matching accuracy is achieved, achieving higher matching accuracy and semantic understanding.

CN114936627BActive Publication Date: 2025-05-27WUDA GEOINFORMATICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210573269.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2025-05-27
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

The existing address matching method has poor effect when dealing with irregular key addresses and is sensitive to address length, resulting in low matching accuracy.

Method used

The improved sectional inferred address matching method is used to train the address matching model, including the element inference layer model and the element extraction layer matrix, the area, building and road code elements of the address are extracted, and the ESIM inference model is used for matching, and finally, comprehensively determine whether the address matches.

Benefits of technology

Reduces sensitivity to address length, improves sensitivity to digital elements, enhances semantic understanding of addresses, and significantly improves matching accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114936627B_ABST
    Figure CN114936627B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of urban governance systems, and provides an improved segmented inference address matching method. This method first trains three simplified ESIM inference models and three extraction matrices. The text address is extracted through the extraction matrix, and three sub-elements are extracted from the key address and the standard address, including the regional element, the building element, and the road code element. The sub-elements are respectively matched with the simplified ESIM inference model. Finally, according to the matching results of the three sub-elements, it is comprehensively judged whether the address matches. Compared with the existing address matching methods based on deep learning, the method of the present invention reduces the sensitivity to the address length, explicitly distinguishes the differences of different geographical elements, improves the semantic understanding of the address, and has better matching accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of urban governance, and particularly relates to an improved segmented inference address matching method. Background Art

[0002] As one of the basic elements in the urban governance system, an address is a connection hub for multiple elements such as people, events, and things. However, in engineering practice, the addresses collected in front-line operations often do not match the addresses in the system address library. Therefore, how to align the collected address to be matched (hereinafter referred to as the key address) with the unified address (hereinafter referred to as the standard address) saved in the system is a very important link in the urban governance system. The goal of address matching is to determine whether the key address and the standard address point to the same address.

[0003] The existing address matching methods are mainly divided into the following three categories:

[0004] The first is the rule-based address matching method, such as keyword search, calculating edit distance and other methods. This type of method sets certain rules based on the string characteristics of the text address, and judges whether the address pair matches according to the rules. The judgment basis of this type of method is simple, and the judgment effect on the standard address is good, but the effect on irregular key addresses is very poor.

[0005] For example, in the address matching method based on keyword search, the input keyword is "Jingzhou Luzhou Road". Since this type of method only recognizes whether the target address contains these five characters, it may judge "Jingzhou City Luzhou Road" and "Luzhou City Jingzhou Road" as the same address.

[0006] Another example is the address matching method based on edit distance. Virtual address pair 1 ["Building 5, Floor 17, Room 17B of Dafeng Factory", "No. 1, Building 5, Floor 17, Room 17B, Dafeng Community, Guangwei Road, Guangming Sub-district, Guangming District, Jingzhou City, Handong Province"] and virtual address pair 2 ["No. 1, Building 5, Floor 17, Room 17B, Dafeng Community, Guangwei Road, Guangming Sub-district, Guangming District, Jingzhou City, Handong Province", "No. 1, Building 5, Floor 18, Room 18B, Dafeng Community, Guangwei Road, Guangming Sub-district, Guangming District, Jingzhou City, Handong Province"]. Because the number of identical characters between the addresses in virtual address pair 1 is less than that in address pair 2, and the continuous identical text segments between the two addresses in the virtual address pair are shorter, the address matching method based on edit distance will consider that the similarity of virtual address pair 2 is higher than that of virtual address pair 1, but in fact, the similarity of virtual address pair 1 should be higher.

[0007] Second, there are machine learning-based address matching methods, such as support vector machines, bag-of-words statistics, etc. These methods calculate the similarity of address embedding vectors and the similarity values of statistical features through algorithms or models to determine the matching degree of address pairs. Compared with simple rule matching, these address matching methods have better effects on irregular key addresses. However, these methods can only extract shallow text semantics, so the matching effect is not as good as that of deep learning-based address matching methods. Moreover, machine learning-based methods require domain experts to construct sufficient learning samples, consuming a large amount of labor costs.

[0008] For example, address matching based on the bag-of-words model will give a high matching value to two addresses with many common words. For example,

[0009] Virtual Address 1: "Building 5, Unit 2, Room 17B, Dashfeng Community, Guangming District, Jingzhou City, Handong Province"

[0010] Virtual Address 2: "Building 5, Unit 2, Room 71B, Dashfeng Community, Guangming Street, Guangming District, Jingzhou City, Handong Province"

[0011] Because they have many common words and the method can only extract shallow semantics, Virtual Address 1 and Virtual Address 2 will be misjudged as the same address.

[0012] Third, there are deep learning-based address matching methods. Usually, a multi-layer neural network is constructed to embed the address pair into a multi-dimensional vector space, and then the vector similarity of the address pair is compared to determine whether the addresses match. Compared with machine learning-based address matching methods, the neural network can autonomously learn the features of text addresses and learn deeper text semantics, thus reducing the cost of sample annotation and achieving better matching effects.

[0013] However, existing deep learning methods usually directly judge whether an address pair matches. This matching method ignores the differences between the matching characteristics of numerical elements (mostly building numbers, house numbers, and room numbers) and non-numerical elements (mostly regional elements) in text addresses, resulting in the model being sensitive to the address length and less sensitive to the numbers in the address, thereby reducing the matching accuracy of the model.

[0014] For example, the following two virtual addresses:

[0015] Virtual Address 1: "Room 17B, Unit 2, Building 5, Dafeng Community, Guangming District, Jingzhou City"

[0016] Virtual Address 2: "Room 71B, Unit 2, Building 5, Dafeng Community, Guangming Street, Guangming District, Jingzhou City, Handong Province"

[0017] The regional element of Virtual Address 1 is "Dafeng Community, Guangming District, Jingzhou City", and the building element is "Room 17B, Unit 2, Building 5";

[0018] The regional element of virtual address 2 is "Dafeng Community, Guangming Street, Guangming District, Jingzhou City, Handong Province", and the building element is "Room 71B, Unit 2, Building 5".

[0019] For example, "Dafeng Community, Guangming District, Jingzhou City" in virtual address 1 and "Dafeng Community, Guangming District, Jingzhou City" in virtual address 2, although they have different lengths, actually point to the same geographical location. Therefore, the sensitivity of the model to length should be reduced.

[0020] The building element in virtual address 1 is "Room 17B, Unit 2, Building 5", and the building element in virtual address 2 is "Room 71B, Unit 2, Building 5". Although there is only one pair of digital position permutations, they point to completely different locations. Therefore, the sensitivity of the model to numbers should be increased. Summary of the Invention

[0021] In view of the above problems, the purpose of the present invention is to provide an improved segmented inference address matching method, aiming to solve the technical problem of low matching accuracy of existing methods.

[0022] The present invention adopts the following technical solutions:

[0023] The improved segmented inference address matching method includes the following steps:

[0024] Step S1: Train an address matching model. The address matching model includes an element inference layer model and an element extraction layer matrix. The element inference model includes a regional ESIM inference model, a building ESIM inference model, and a road code ESIM inference model. The element extraction layer matrix includes a regional extraction matrix, a building extraction matrix, and a road code extraction matrix;

[0025] Step S2: Input the address pair to be matched, and generate a prediction sample pair through the prediction sample construction module for the address pair to be matched;

[0026] Step S3: Use the address matching model to infer and match the prediction sample pair to obtain the corresponding matching result.

[0027] Further, the specific process of step S3 is as follows:

[0028] S31: Convert the address pair to be matched into a text embedding vector using the bert model, and input it into the three extraction matrices to obtain the key regional element, key building element, and key road code element of the key address, as well as the standard regional element, standard building element, and standard road code element of the standard address;

[0029] S32. Re - input the six elements into the BERT model respectively to obtain the corresponding element word vectors, and then use three ESIM inference models correspondingly to obtain the matching results of three elements, namely the regional element matching result, the building element matching result, and the road code element matching result;

[0030] S33. Finally, based on the matching results of the three elements, comprehensively calculate to obtain the final matching result.

[0031] Furthermore, the training process of the address matching model is as follows:

[0032] S11. Input the samples, and divide the training sample set into training samples and validation samples according to a certain proportion;

[0033] The format of each sample is [regional label sample, building label sample, road code label sample, segmented address label sample], where the formats of the regional label sample, building label sample, and road code label are all [key element, standard element, label], and the format of the segmented address label sample is [text address, label index], and the text address refers to the key address or the standard address;

[0034] S12. Use the BERT model to convert the key elements, standard elements in the regional, building, and road code label samples, and the text address in the segmented address label sample into corresponding word vectors;

[0035] S13. Train the regional ESIM inference model, building ESIM inference model, and road code ESIM inference model through the key element word vectors and standard element word vectors;

[0036] S14. Train the regional extraction matrix, building extraction matrix, and road code extraction matrix through the text address word vectors.

[0037] Furthermore, step S13 and step S14 are trained in parallel.

[0038] Furthermore, in step S12, the key elements and standard elements in the regional, building, and road code label samples are called address elements. The specific process of step S12 is as follows:

[0039] Segment the text address and address elements into characters;

[0040] Use the BERT model to convert the segmented text address and address elements into token encodings, and obtain the corresponding position encodings;

[0041] Input the token encodings and position encodings into the BERT model respectively to obtain their corresponding word vectors.

[0042] Further, in step S13, the three ESIM inference models, namely the area ESIM inference model, the building ESIM inference model, and the road code ESIM inference model, are simplified models, and the training methods are the same. The process is as follows:

[0043] Input the keyword vector of the key element and the standard element vector into the fully connected neural network simultaneously to obtain the hidden layer state vectors of the key element and the standard element;

[0044] Through the alignment operation, obtain the similarity weight matrix of the key element and the standard element;

[0045] Use the similarity weight matrix to perform weighted summation on the hidden layer state vector of the standard element to obtain the key element similarity vector, and use the similarity weight matrix to perform weighted summation on the hidden layer state vector of the key element to obtain the standard element similarity vector;

[0046] Perform subtraction and multiplication on the hidden layer state vector of the key element, the similarity vector, the hidden layer state vector of the standard element, and the similarity vector respectively, and perform soft alignment to obtain the key element information enhancement vector and the standard address information enhancement vector;

[0047] Input the key element information enhancement vector and the standard address information enhancement vector into the bidirectional long short-term memory neural network to obtain the key element matching vector and the standard element matching vector;

[0048] Obtain the key element max pooling vector and the key element average pooling vector by performing pooling operation on the key element matching vector; obtain the standard element max pooling vector and the standard element average pooling vector by performing pooling operation on the standard element matching vector; splice the four obtained pooling vectors to obtain the element matching information vector;

[0049] Input the element matching information vector into the fully connected layer, and obtain the matching value of each category through the softmax function. There are three categories, namely non-matching, matching, and possible matching;

[0050] Use the cross-entropy loss function to calculate the loss value;

[0051] According to the loss value, use the gradient descent method to modify and update the model parameters, and select the parameter version with the highest verification accuracy as the finally trained ESIM inference model.

[0052] Further, in step S14, the training methods of the three extraction matrices, namely the area extraction matrix, the building extraction matrix, and the road code extraction matrix, are the same. The process is as follows:

[0053] Use the extraction matrix to perform dot product on each token encoding vector in the text address vector to obtain the text address vector;

[0054] Extract values for each token in the text address vector, and use the Sigmoid function to obtain the token extraction score;

[0055] Generate an element tag index based on the tag index. According to the token extraction score and the element tag index, use the cross-entropy loss function to calculate the loss value of each predicted token;

[0056] According to the loss value, use the gradient descent method to modify and update the matrix parameters, and select the parameter version with the highest verification accuracy as the finally trained extraction matrix.

[0057] The beneficial effects of the present invention are as follows: The present invention provides an improved segmented inference address matching method for determining whether the to-be-matched address (key address) input by the user and the unified address (standard address) in the address library point to the same destination; specifically, when implemented, three simplified ESIM inference models and three extraction matrices are first trained. The text address is extracted through the extraction matrix, and three sub-elements are extracted from the key address and the standard address, including the regional element, the building element, and the road code element. The sub-elements are respectively matched using the simplified ESIM inference model. Finally, according to the matching results of the three sub-elements, it is comprehensively determined whether the addresses match. Compared with the existing address matching methods based on deep learning, the method of the present invention reduces the sensitivity to the address length, explicitly distinguishes the differences of different geographical elements, improves the semantic understanding of the address, and has a better matching accuracy. Description of the Drawings

[0058] Figure 1 is the flow chart of the improved segmented inference address matching method provided by the embodiment of the present invention;

[0059] Figure 2 is the schematic diagram of the address matching model training provided by the embodiment of the present invention;

[0060] Figure 3 is the schematic diagram of the simplified ESIM inference model training provided by the embodiment of the present invention;

[0061] Figure 4 is the address matching flow chart provided by the embodiment of the present invention;

[0062] Figure 5 is the schematic diagram of the address matching model inference provided by the embodiment of the present invention. Detailed Embodiments

[0063] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0064] Figure 1The flowchart of the address matching method based on sectional inference provided by the embodiment of the present invention is shown. Only the parts related to the embodiment of the present invention are shown for the convenience of description.

[0065] As Figure 1 shown, the improved address matching method based on sectional inference provided in this embodiment includes the following steps:

[0066] Step S1, train the address matching model.

[0067] The address matching model includes an element inference layer model and an element extraction layer matrix. The element inference model includes a region ESIM inference model, a building ESIM inference model, and a road code ESIM inference model. The element extraction layer matrix includes a region extraction matrix, a building extraction matrix, and a road code extraction matrix.

[0068] The training process of the address matching model is combined with Figure 2 shown and includes the following steps:

[0069] S11, input the samples, and divide the training sample set into training samples and validation samples according to a ratio.

[0070] Before training, first divide the labeled training sample set into two parts according to a ratio (9:1 or other ratios): training samples and validation samples. Input the training samples into the address matching model. The model learns all the parameters of the model through the training samples, and then uses the validation samples to test the model with the trained parameters, and saves the parameter version with the highest test accuracy.

[0071] The format of each sample is [region labeled sample, building labeled sample, road code labeled sample, sectional address labeled sample], where the formats of the region labeled sample, building labeled sample, and road code label are all [key element, standard element, label]. In the region labeled sample, the key element refers to the region element in the address to be matched, the standard element refers to the region element of the unified address in the address library, and the label has three types {0, 1, 2}, where 0 represents non-matching, 1 represents matching, and 2 represents possible matching. The building labeled sample and the road code labeled sample are similar to the region labeled sample, except that their elements are the building element and the road code element in the address respectively.

[0072] The format of the sectional address labeled sample is [text address, label index], where the text address refers to the key address or the standard address.

[0073] The label index refers to the text address index of the different elements that have been labeled. For example, 1 represents belonging to the region element, 2 represents belonging to the road code element, and 3 represents belonging to the building element.

[0074] The region element refers to the text segments of the province, city, district, street, community, and residential area in the text address.

[0075] The building element refers to the text segments of the building and room number in the text address.

[0076] The road code element refers to the text segments of the road and house number in the text address.

[0077] For example Figure 2 In the virtual address "Shanshui Group Building, 4th Floor, 10 Beifeng Road, Dafeng Factory, Jingzhou City", "Dafeng Factory, Jingzhou City" is the regional element, "10 Beifeng Road" is the road code element, and "4th Floor, Shanshui Group Building" is the building element.

[0078] Another example is the virtual address "Room 12, 4th Floor, Shanshui Group, Dafeng Factory, South Wind Road, Guangming District". Among them, "Dafeng Factory, Guangming District" is the regional element, "South Wind Road" is the road code element, and "4th Floor, Room 12" is the building element.

[0079] Input sample cases: [("Dafeng Factory, Guangming, Jingzhou City", "Dafeng Factory, Guangming District", 1), ("Building 5, No. 31", "Building 5, No. 13", 0), ("No. 3, South Wind Road", "No. 3, South Wind Road", 1), ("Shanshui Group Building, 4th Floor, 10 Beifeng Road, Dafeng Factory, Jingzhou City", "11111122222233333333")].

[0080] S12. Use the bert model to convert the key elements, standard elements in the regional, building, and road code marked samples, and the text address in the segmented address marked sample into corresponding word vectors.

[0081] The specific process of this step is as follows:

[0082] S121. Split the text address and address elements into characters.

[0083] Here, the key elements and standard elements in the regional, building, and road code marked samples are called address elements. An example of splitting the text address and address elements into characters is as follows:

[0084] For example, the virtual address: "Room 17B, Unit 2, Building 5, Dafeng Community, Guangming Street, Guangming District, Jingzhou City, Handong Province" is split into: [Han, Dong, Sheng, Jing, Zhou, Shi, Guang, Ming, Qu, Guang, Ming, Jie, Dao, Da, Feng, She, Qu, 5, Dong, 2, Dan, Yuan, 1, 7, B, Shi].

[0085] S122. Use the bert model to convert the character-based text address and address elements into token encodings and obtain the corresponding position encodings.

[0086] For example: [Han, Dong, Sheng, Jing, Zhou, Shi, Guang, Ming, Qu, Guang, Ming, Jie, Dao, Da, Feng, She, Qu, 5, Dong, 2, Dan, Yuan, 1, 7, B, Shi]

[0087] The token encoding is: [3727, 691, 4689, 776, 2336, 2356, 1045, 3209, 1277, 1045, 3209, 6125, 6887, 1920, 7599, 4852, 1277, 126, 3406, 123, 1296, 1039, 122, 128, 144, 2147].

[0088] The position encoding is:

[0089] [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25].

[0090] S123. Input the token encoding and the position encoding into the bert model respectively to obtain their corresponding word vectors.

[0091] In this step, after the labeled samples are transformed by the bert model, the corresponding word vector groups are obtained: [(DKe, DSe, labeled), (BKe, BSe, labeled), (CKe, CSe, labeled), (ADDRe, regional element, road code element, building element)]. Among them, DKe and DSe respectively represent the regional key element word vector and the regional standard element word vector; BKe and BSe respectively represent the building key element word vector and the building standard element word vector; CKe and CSe respectively represent the road code key element word vector and the road code standard element word vector; ADDRe represents the text address word vector.

[0092] S13. Train the regional ESIM inference model, the building ESIM inference model, and the road code ESIM inference model through the key element word vectors and the standard element word vectors.

[0093] Use the regional ESIM inference model, the building ESIM inference model, and the road code ESIM inference model to infer whether the key element word vectors and the standard element word vectors in the region, building, and road code match respectively. The training methods of these three models are the same, but the training model samples are different, so the trained parameters will be different.

[0094] In this embodiment, the ESIM inference model is a simplified model, which removes a Bi-LSTM network in the ESIM model. Since the number of parameters is reduced, the resource expenditure of the model can be reduced, and the training and inference speeds of the model can be accelerated. Specifically, when implemented, as Figure 3 shown, the process of this step is specifically as follows:

[0095] S131. Input the key element word vector and the standard element word vector into the fully connected neural network at the same time to obtain the hidden layer state vectors of the key element and the standard element.

[0096] The key element word vector DKe and the standard element word vector DSe are simultaneously input into a fully connected neural network layer to obtain the hidden layer state vector of the key element and the hidden layer state vector of the standard element where

[0097] W 1 is a matrix of dimension d1*d2, which is a learning parameter of the model. d1 is the dimension number of DKe and DSe, and d2 is and the dimension number of is the transpose matrix of W 1 In the present invention, d1 is 768 and d2 is 100 (it can also be other dimension numbers).

[0098] S132. Through the alignment operation, the similarity weight matrix between the key element and the standard element is obtained.

[0099] Through the alignment operation, the similarity weight matrix E between the key element and the standard element is obtained. The alignment operation is as follows: represents the vector of the i-th token in the hidden layer state vector of the key element, represents the vector of the j-th token in the hidden layer state vector of the standard element. i ranges from 0 to the number of key element tokens, and j ranges from 0 to the number of element address tokens.

[0100] S133. Using the similarity weight matrix, the weighted sum of the hidden layer state vectors of the standard elements is calculated to obtain the key element similarity vector, and the weighted sum of the hidden layer state vectors of the key elements is calculated to obtain the standard element similarity vector.

[0101] Using the obtained similarity weight matrix E, the hidden layer state vector of the standard element is weighted and summed to obtain the key element similarity vector Using the obtained similarity weight matrix E, the hidden layer state vector of the key element is weighted and summed to obtain the key element similarity vector

[0102]

[0103]

[0104] where l s represents the number of tokens of the standard element, l K represents the number of tokens of the key element, e ij represents the value of the i-th row and j-th column in the similarity weight matrix E. e in , e mj Similarly.

[0105] S134. Subtract and multiply the hidden layer state vectors and similarity vectors of the key elements, as well as the hidden layer state vectors and similarity vectors of the standard elements, respectively, and perform soft alignment to obtain the key element information enhancement vector and the standard address information enhancement vector.

[0106] Subtract and multiply the vectors related to the key elements, and perform soft alignment to obtain the key element information enhancement vector Similarly, the standard address information enhancement vector can be obtained

[0107] S135. Input both the key element information enhancement vector and the standard address information enhancement vector into a bidirectional long short-term memory neural network to obtain the key element matching vector and the standard element matching vector.

[0108] Input the key element information enhancement vector M k into the bidirectional long short-term memory neural network Bi-LSTM to obtain the key element matching vector V k , and similarly, the standard element matching vector V can be obtained s .

[0109] S136. Obtain the key element max pooling vector and the key element average pooling vector through pooling operations on the key element matching vector; obtain the standard element max pooling vector and the standard element average pooling vector through pooling operations on the standard element matching vector; concatenate the four obtained pooling vectors to obtain the element matching information vector.

[0110] Obtain the key element max pooling vector V through max pooling on the key element matching vector V k , and obtain the key element average pooling vector V through average pooling on the key element matching vector V k,max . Similarly, the standard element max pooling vector V can be obtained k , and the standard element average pooling vector V k,avg . s,max , and the standard element average pooling vector V s,avg .

[0111] The formulas for average pooling and max pooling are as follows:

[0112]

[0113] where V k,i represents the i-th vector in the key element matching vector V k .

[0114] Concatenate the four obtained pooling vectors to obtain the element matching information vector V = [V k,avg , V k,max , V s,avg , V s,max .

[0115] S137. Input the element matching information vector into the fully connected layer, and obtain the matching values for each category through the softmax function. There are three categories in total, namely unmatched, matched, and possibly matched.

[0116] Input the element matching information vector V into the fully connected layer, and obtain the matching values for each category (a total of three categories, 0: unmatched; 1: matched, 2: possibly matched) through the softmax function. The fully connected layer contains two fully connected neural networks, and the activation function between the two networks is the tanh activation function. The matching values output by the softmax function are between 0 and 1.

[0117] S138. Calculate the loss value using the cross-entropy loss function.

[0118] The formula for the loss function is

[0119]

[0120] where y i is the existing labeled category, and p i is the output matching value. If the labeled category is 1, its one-hot label is [0, 1, 0]. If the output matching value is [0.4, 0.2, 0.4], then the loss value is: -(0 * log0.4 + 1 * log0.2 + 0 * log0.4) = -log0.2.

[0121] S139. According to the loss value, use the gradient descent method to modify and update the model parameters, and select the parameter version with the highest validation accuracy as the finally trained ESIM inference model.

[0122] As Figure 2 shown, the loss value of the regional ESIM inference model will be added to the loss values of the other two ESIM inference models and the three extraction matrices to obtain the total loss value, and then the model parameters will be updated using gradient descent. The calculation of the total loss value of the other ESIM inference models is the same.

[0123] The ESIM inference model will traverse the training samples multiple times. After each traversal of the training samples, the accuracy of the model will be tested using the validation samples. The validation process is basically the same as the training process, except that after completing step S137, the category with the largest matching value is selected as the prediction result and compared with the labeled result. If the types are the same, it means the prediction is correct, otherwise it is a prediction error. In the model training stage, the parameter version with the highest validation accuracy will be selected as the finally trained ESIM inference model.

[0124] S14. Train the regional extraction matrix, building extraction matrix, and road code extraction matrix through the text address word vectors.

[0125] In this step, three extraction matrices are used to extract the corresponding three-element parts in the text address vector. The training methods of the three extraction matrices, namely the region extraction matrix, the building extraction matrix, and the road code extraction matrix, are the same, and the process is as follows:

[0126] S141. Use the extraction matrix to perform a dot product on each token encoding vector in the text address token vector to obtain the text address vector.

[0127] Use the extraction matrix to perform a dot product on each token encoding vector in the text address token vector ADDRe to obtain the text address vector res, res = (res 1 , res 2 ,..., res n ), res i is the extraction value of the i-th token.

[0128] Where

[0129] W 2 is a d1 * 1-dimensional matrix, e i is the encoding vector of the i-th token in the text address token vector ADDRe. is the transpose matrix of W 2 .

[0130] S142. For each token extraction value in the text address vector, use the Sigmoid function to obtain the token extraction score.

[0131] For each token extraction value, use the Sigmoid function to obtain the token extraction score Score i :

[0132] Score i = Sigmoid(res i ).

[0133] S143. Generate the element label index based on the label index. According to the token extraction score and the element label index, use the cross-entropy loss function to calculate the loss value of each predicted token.

[0134] During the training phase, generate the element label index based on the label index. For example, when generating the region label index, all indices with a value of 1 in the label index are assigned a value of 1, and the rest are assigned a value of 0. When generating the road code label index, all indices with a value of 2 in the label index are assigned a value of 1, and the rest are assigned a value of 0. When generating the building label index, all indices with a value of 3 in the label index are assigned a value of 1, and the rest are assigned a value of 0. Compare the token extraction score and the element label index, and use the cross-entropy loss function to calculate the loss value of each predicted token. The calculation process of each token loss value is the same as the method in S138.

[0135] During the inference matching stage, for the Score i Extract the lemmatization scores greater than 0.5, extract the indices of the corresponding lemmas, and extract the corresponding lemmas according to the indices.

[0136] For example: Virtual text address:

[0137] "Room 17B, Unit 2, Building 5, Dafeng Community, Guangming Street, Guangming District, Jingzhou City, Handong Province", its index is: [0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25]. After the corresponding lemma encoding vectors are transformed by the region extraction matrix, the corresponding region vectors are obtained. After each lemma extraction value in the region vector passes through the Sigmoid function, the corresponding extraction score is obtained. Suppose the lemma extraction scores score obtained after the virtual text address is processed by the region extraction matrix are [0.8,0.6,0.7,0.7,0.6,0.9,0.7,0.7,0.8,0.1,0.1,0.2,0.2,0.9,0.8,0.9,0.7,0.3,0.4,0.2,0.3,0.2,0.1,0.1,0.1,0.1].

[0138] In the training stage: According to the marked index: [1,1,1,1,1,1,1,1,1,2,2,2,2,1,1,1,1,3,3,3,3,3,3,3,3,3], the region marked index is obtained as [1,1,1,1,1,1,0,0,0,0,1,1,1,1,0,0,0,0,0,0,0,0,0]. The loss value of the first lemma is loss 1 =-log0.8, and the loss value of the last lemma is loss 26 =-log(1 - 0.1)=-log0.9. The comprehensive loss value of the text address is where loss i is the loss value of the i-th lemma.

[0139] When in the inference stage, extract the indices where the lemma scores are greater than 0.5, then the extracted indices are [0,1,2,3,4,5,6,7,8,13,14,15,16], and the extracted lemma fragments are [Han, Dong, Province, Jing, Zhou, City, Guang, Ming, District, Da, Feng, Community].

[0140] S144. According to the loss value, use the gradient descent method to modify and update the matrix parameters, and select the version of the parameter with the highest validation accuracy as the finally trained extraction matrix.

[0141] Such as Figure 2As shown, the loss value of the region extraction matrix will be added to the loss values obtained from the other three ESIM inference models and two extraction matrices to obtain the total loss value, and then the model parameters will be updated using gradient descent. The calculation of the total loss value for other extraction matrices is the same.

[0142] The extraction task of the extraction matrix will traverse the training samples multiple times. After each traversal of the training samples, the accuracy of the extraction by the extraction matrix is tested using the validation samples. The validation process is basically the same as the training process, except that in step S143, the corresponding token index is selected and compared with the marked result. If the indexes are the same, it means the prediction is correct, otherwise it is a prediction error. In the model training stage, the version of the parameter with the highest validation accuracy will be selected as the finally trained extraction matrix.

[0143] In this embodiment, the three models in the feature inference layer and the three extraction matrices in the feature extraction layer can be trained in parallel simultaneously to improve efficiency.

[0144] The above step S1 is the training of the address matching model, which means using the marked training sample set to train the parameters of the address matching model to obtain a trained address matching model. Combining Figure 4 As shown, the following steps S2 - S3 are the inference of the address matching model, which means using the trained address matching model to determine whether the input address pair matches.

[0145] Step S2: Input the address pair to be matched, and generate a prediction sample pair for the address pair to be matched through the prediction sample construction module.

[0146] The format of the address pair to be matched is [key address, standard address 1, standard address 2...... standard address n].

[0147] The address pair to be matched enters the prediction sample construction module to generate a prediction sample pair.

[0148] The prediction sample construction module is to combine each standard address in the address pair to be matched with the key address respectively to generate a prediction sample pair. For example:

[0149] Address pair to be matched: [key address, standard address 1, standard address 2, standard address 3]

[0150] The prediction sample pairs constructed by the prediction sample construction module are:

[0151] Prediction sample pairs: [key address, standard address 1], [key address, standard address 2], [key address, standard address 3].

[0152] Step S3: Infer and match the prediction samples using the address matching model to obtain the corresponding matching results.

[0153] The address pairs to be matched use the BERT model to obtain the embedding vectors KADDRe and SADDRe of the key address and the standard address. The matching result structure of each prediction sample pair is [key address, standard address, matching result, matching value]. Then, the matching results are output in descending order according to the matching value.

[0154] This step is a specific inference process. As Figure 5 shown below:

[0155] S31. Convert the address pairs to be matched into text embedding vectors using the BERT model, and input them into three extraction matrices to obtain the key area elements, key building elements, and key road code elements of the key address, as well as the standard area elements, standard building elements, and standard road code elements of the standard address.

[0156] Input the key address embedding vector KADDRe into the area extraction matrix, building extraction matrix, and road code extraction matrix respectively to obtain the key area element index KDi, key building element index KBi, and key road code element index KCi.

[0157] Then generate the key area elements, key building elements, and key road code elements according to KDi, KBi, and KCi.

[0158] For example, for the virtual key address: [4th Floor, Dashanfeng Group Building, Jingzhou City], the generated key area element is [Dashanfeng, Jingzhou City], the generated key building element is [4th Floor, Dashanfeng Group Building], the generated key road code element is a null value, and the result of the road code ESIM model is forced to be "no information".

[0159] Similar to the key address processing process, obtain the standard area elements, standard building elements, and standard road code elements according to the standard address.

[0160] S32. Re-input the six elements into the BERT model to obtain the corresponding element word vectors, and then use three ESIM inference models respectively to obtain the matching results of the three elements, namely the area element matching result, the building element matching result, and the road code element matching result.

[0161] Re-input the six elements into the BERT model respectively to obtain their respective element word vectors. As Figure 5As shown in the figure, the key area feature word vectors and the standard area feature word vectors are input into the area ESIM inference model to obtain the area feature matching result; the key building feature word vectors and the standard building feature word vectors are input into the building ESIM inference model to obtain the building feature matching result; the key road code feature word vectors and the standard road code feature word vectors are input into the road code ESIM inference model to obtain the road code feature matching result. The ESIM inference model outputs 4 results: ["matched", "not matched", "possibly matched", "no information"]. Among them, only when at least one of a pair of features input into the ESIM model is empty, the "no information" result is output.

[0162] S33. Finally, according to the matching results of the three features, the final matching result is comprehensively calculated.

[0163] The comprehensive calculation method is flexibly designed according to the characteristics of the address to be matched in different regions. For example, it can be set that when the area is matched, the building is not matched, and the road code is matched, the final matching result is "matched"; it can also be set that when the area is matched, the building is not matched, and the road code is matched, the final matching result is "not matched". The execution steps of the ESIM inference model during inference are the same as those during training, except that during the inference process, the corresponding category with the largest matching value output by the model is used as the result output.

[0164] In the embodiment of the present invention, three simplified ESIM inference models and three extraction matrices are trained. The bert model is used to convert the text address pair into a corresponding word vector pair. The three extraction matrices are used to extract three sub-features of the address, including area features, building features, and road code features. The sub-features are respectively matched using the simplified ESIM inference model. Finally, according to the matching results of the three sub-features, it is comprehensively judged whether the address is matched. Compared with the existing address matching method based on deep learning, the method of the present invention reduces the sensitivity to the address length, increases the sensitivity to numbers, and improves the matching accuracy.

[0165] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An improved segmented inference address matching method, characterized in that, the method comprises the following steps: Step S1, train an address matching model, the address matching model includes an element inference layer model and an element extraction layer matrix, wherein the element inference layer model includes a region ESIM inference model, a building ESIM inference model and a road code ESIM inference model, and the element extraction layer matrix includes a region extraction matrix, a building extraction matrix and a road code extraction matrix; Step S2, input the address pair to be matched, and generate a prediction sample pair from the address pair to be matched through a prediction sample construction module; Step S3, perform inference matching on the prediction sample pair using the address matching model to obtain the corresponding matching result; wherein the specific process of step S3 is as follows: S31, convert the address pair to be matched into a text embedding vector using the bert model, and input it into the three extraction matrices to obtain the key region elements, key building elements and key road code elements of the key address, as well as the standard region elements, standard building elements and standard road code elements of the standard address; S32, re-input the six elements into the bert model respectively to obtain the corresponding element word vectors, and then use the three ESIM inference models correspondingly to obtain the matching results of the three elements, namely the region element matching result, the building element matching result and the road code element matching result; S33, finally, comprehensively calculate the final matching result according to the matching results of the three elements; The training process of the address matching model is as follows: S11, input samples, and divide the training sample set into training samples and verification samples according to a ratio; The format of each sample is [region label sample, building label sample, road code label sample, segmented address label sample], wherein the formats of the region label sample, building label sample and road code label are all [key element, standard element, label], and the format of the segmented address label sample is [text address, label index], and the text address refers to the key address or the standard address; S12, use the bert model to convert the key elements, standard elements in the region, building, road code label samples, and the text address in the segmented address label sample into corresponding word vectors; S13, train a region ESIM inference model, a building ESIM inference model and a road code ESIM inference model through the key element word vectors and the standard element word vectors; S14, train a region extraction matrix, a building extraction matrix and a road code extraction matrix through the text address word vectors.

2. The improved segmented inference address matching method according to claim 1, characterized in that, steps S13 and S14 are trained in parallel.

3. The improved segmented inference address matching method according to claim 2, characterized in that, in step S12, the key elements and standard elements in the region, building, road code label samples are called address elements, and the specific process of step S12 is as follows: Split the text address and the address elements into characters; Use the bert model to convert the text address and address elements into wordpiece encodings in characters and obtain the corresponding position encodings; Input the wordpiece encodings and position encodings into the bert model respectively to obtain their corresponding word vectors.

4. The improved segmented inference address matching method according to claim 3, wherein, in step S13, the three ESIM inference models, namely the region ESIM inference model, the building ESIM inference model, and the road code ESIM inference model, are simplified models, and the training methods are the same. The process is as follows: Input the keyword vector of the key element and the standard element word vector into the fully connected neural network simultaneously to obtain the hidden layer state vectors of the key element and the standard element; Through the alignment operation, obtain the similarity weight matrix of the key element and the standard element; Use the similarity weight matrix to perform weighted summation on the hidden layer state vector of the standard element to obtain the key element similarity vector, and use the similarity weight matrix to perform weighted summation on the hidden layer state vector of the key element to obtain the standard element similarity vector; Respectively perform subtraction and multiplication on the hidden layer state vector of the key element, the similarity vector, the hidden layer state vector of the standard element, and the similarity vector, and perform soft alignment to obtain the key element information enhancement vector and the standard address information enhancement vector; Input the key element information enhancement vector and the standard address information enhancement vector into the bidirectional long short-term memory neural network to obtain the key element matching vector and the standard element matching vector; Obtain the key element maximum pooling vector and the key element average pooling vector by performing pooling operation on the key element matching vector; obtain the standard element maximum pooling vector and the standard element average pooling vector by performing pooling operation on the standard element matching vector; splice the four obtained pooling vectors to obtain the element matching information vector; Input the element matching information vector into the fully connected layer, and obtain the matching value of each category through the softmax function. There are three categories in total, namely non-matching, matching, and possible matching; Use the cross-entropy loss function to calculate the loss value; According to the loss value, use the gradient descent method to modify and update the model parameters, and select the parameter version with the highest verification accuracy as the finally trained ESIM inference model.

5. The improved segmented inference address matching method according to claim 4, wherein, in step S14, the training methods of the three extraction matrices, namely the region extraction matrix, the building extraction matrix, and the road code extraction matrix, are the same. The process is as follows: Use the extraction matrix to perform dot product on each token encoding vector in the text address word vector to obtain the text address vector; For each token extraction value in the text address vector, use the Sigmoid function to obtain the token extraction score; Generate the element token index according to the marked index. According to the token extraction score and the element token index, use the cross-entropy loss function to calculate the loss value of each predicted token; According to the loss value, use the gradient descent method to modify and update the matrix parameters, and select the parameter version with the highest verification accuracy as the finally trained extraction matrix.

Citation Information

Patent Citations

  • AI-based house address matching method, storage medium and equipment

    CN113869052A