Chinese address resolution method based on pre-training model

The BiLSTM network is optimized through the decoupling attention mechanism of the DeBERTa model, the CNN structure of the parallel attention mechanism and the GLSMA algorithm, combined with the CRF output, and the shortcomings of the traditional Chinese address resolution method in complex address resolution are solved, and more efficient address resolution effect is achieved.

CN120542413APending Publication Date: 2025-08-26HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510667699.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Traditional Chinese address analysis methods face problems such as place name ambiguity, diversified address formats, difficulty in identifying rare place name, complex hierarchy analysis and insufficient context understanding, resulting in poor performance of the model in practical applications.

Method used

The decoupled attention mechanism of the DeBERTa model is used to capture the structured characteristics and hierarchical dependencies of the address, combine the CNN structure of the parallel attention mechanism to extract local features, and optimize timing information through the BiLSTM network, optimize population distribution using the GLSMA algorithm, and finally use CRF for output analysis.

Benefits of technology

It improves the accuracy and generalization ability of Chinese address resolution, especially when dealing with addresses of multi-level nested structures, it can more accurately identify hierarchical relationships and semantic information, and improves the model's analytical ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542413A_ABST
    Figure CN120542413A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese address resolution method based on a pre-training model, and the method comprises the steps: firstly carrying out the data preprocessing and rule-based BIOES labeling of Chinese address data crawled on a Baidu map open source platform, and obtaining standardized Chinese address data and labels of address elements; capturing the structural features of the address and the dependency relationship between different address hierarchies by using the decoupling attention mechanism of the DeBERTa model; a CNN structure based on a parallel attention mechanism is introduced, local features of an address are better extracted through the parallel attention mechanism, and the short sequence text analysis capability of the model is improved; a BiLSTM network is applied to further model long-sequence text features, time sequence information of address texts is extracted, and GLSMA (improved myxin algorithm) is provided for solving the problems that population distribution is not uniform and local optimum is prone to being caught in the later stage of algorithm iteration due to random population initialization of SMA (myxin algorithm). According to the algorithm, Logistic chaotic mapping is introduced to initialize populations, so that the populations are uniformly distributed, and the convergence speed and optimization efficiency of the algorithm are improved. A genetic learning strategy is introduced, global search is carried out in a solution space through crossover, variation and selection operations, local optimum can be effectively avoided, and the method has higher global search ability. And finally, outputting the analyzed Chinese address by using a CRF (Conditional Random Field).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Chinese address parsing under natural language processing, and specifically relates to Chinese address preprocessing and a deep learning model based on pre-training. Background Art

[0002] Chinese address parsing methods based on pre-trained models are an important research direction in natural language processing (NLP), particularly in geographic information systems (GIS), intelligent customer service, search engines, and map applications. Chinese address parsing involves identifying and extracting valid geographic location information (such as province, city, district, street, and house number) from natural language text and converting this information into a structured data format. This is commonly used for tasks such as geographic location, address auto-completion, and route planning.

[0003] Although traditional Chinese address parsing methods have made progress in accuracy and flexibility, they still face some challenges, including place name ambiguity, diverse address formats, difficulty in identifying rare place names, complex hierarchical parsing problems, insufficient context understanding, and limitations in data annotation. These problems greatly affect the performance of the model in practical applications.

[0004] Chinese address data crawled from the Baidu Map open source platform undergoes data preprocessing and rule-based BIOES annotation to generate standardized Chinese address data and corresponding address element labels. Subsequently, the decoupled attention mechanism in the DeBERTa model is used to capture the structural features of addresses and the dependencies between different address levels. Furthermore, a CNN architecture based on a parallel attention mechanism is introduced to enhance the ability to extract local address features, improving the model's performance in parsing short text sequences. To better model the characteristics of long text sequences, a BiLSTM network is used to further extract temporal information from address text. To address the problem of uneven population distribution and a tendency to fall into local optima in the late iterations of the Slime Mold Algorithm (SMA), the improved Slime Mold Algorithm (GLSMA) is proposed. This improved method uses a logistic chaos map to initialize the population, ensuring a uniform population distribution and improving the algorithm's convergence speed and optimization efficiency. Furthermore, a genetic learning strategy is employed to conduct a global search in the solution space through crossover, mutation, and selection operations, effectively avoiding local optima and enhancing global search capabilities. Finally, a Conditional Random Field (CRF) is used to parse and output the Chinese addresses. Experiments show that this method can perform Chinese address parsing effectively and accurately. Summary of the Invention

[0005] Purpose of the invention: The present invention proposes a Chinese address resolution method based on a pre-trained model to perform Chinese address resolution effectively and accurately.

[0006] Beneficial effects: 1. The present invention uses the decoupled attention mechanism of the DeBERTa model to capture the structural features of the address and the dependencies between different address levels. In particular, when faced with Chinese addresses with multi-level nested structures, it can more accurately identify and understand the hierarchical relationships and semantic information between the various levels. 2. The present invention introduces a CNN structure based on a parallel attention mechanism to better extract the local features of the address with the parallel attention mechanism, thereby improving the model's ability to parse short sequence texts. 3. In terms of optimizing the BiLSTM network, the present invention proposes a GLSMA method, which improves the accuracy and generalization ability of the Chinese address parsing model by introducing Logistic chaotic mapping to initialize the population and adopting a genetic learning strategy.

[0007] Technical solution: The Chinese address parsing method based on the pre-trained model proposed in the present invention includes the following steps:

[0008] (1) The Chinese address data crawled from the Baidu Map open source platform is preprocessed and labeled with BIOES based on rules to obtain standardized Chinese address data and labels of address elements.

[0009] (2) The decoupled attention mechanism of the DeBERTa model is used to capture the structural features of standardized Chinese address data and the dependencies between different address levels.

[0010] (3) A CNN structure based on a parallel attention mechanism is introduced to better extract the local features of Chinese addresses using the parallel attention mechanism, thereby improving the model's ability to parse short sequence texts.

[0011] (4) The BiLSTM network is used to further model the long sequence text features, extract the temporal information of Chinese address text, and propose the GLSMA algorithm to optimize the BiLSTM network structure.

[0012] (5) Apply CRF (Conditional Random Field) to output the parsed Chinese address.

[0013] The step (1) comprises the following steps:

[0014] (11) First, we use crawler technology to obtain a large amount of Chinese address data on the Baidu Map open source platform to ensure the diversity and breadth of the data. By accessing the open API interface, we can obtain the original Chinese address data containing geographic information such as provinces, cities, and counties.

[0015] (12) The obtained original Chinese address data is preprocessed, and the CPCATransformer (Chinese_Province_City_Area_mapper) module is introduced to complete the administrative elements of incomplete Chinese address data.

[0016] (13) After preprocessing, the Chinese address data is annotated with BIOES to identify the different elements in the address. The BIOES annotation method annotates each address element as B (begin), I (inside), O (outside), E (end), and S (single).

[0017] (14) Obtain standardized Chinese address data and labels of address elements.

[0018] The algorithm definition and steps of the decoupled attention mechanism using the DeBERTa model described in step (2) are as follows:

[0019] (21) The DeBERTa model is based on BERT. Each character in the input layer is represented by two vectors, which encode its content and position respectively, and the attention weights between characters are calculated using a disentangled matrix based on their content and relative position.

[0020] The specific steps are: For the token at position i in the sequence, two vectors are used, {Hx} and {P i} j} represents it, and represents its content and relative position to the token at position j respectively. The calculation of the cross attention score between token i and j can be decomposed into four parts:

[0021]

[0022] Among them, {H i} and {P i} j} represents the content of token i and its relative position to token j; {H j} and {P j} i} represents the content of token j and its relative position to token 1.

[0023] That is, the attention weight of a word pair can be calculated as the sum of scores of four attentions (content-to-content, content-to-position, position-to-content, and position-to-position) using the decoupling matrix of its content and position.

[0024] (22) By introducing the decoupled attention mechanism and the position-to-content and position-to-position scores, the model can better understand the Chinese address sequence and capture the structural features of standardized Chinese address data and the dependencies between different address levels.

[0025] The steps for introducing the CNN structure based on the parallel attention mechanism in step (3) are as follows:

[0026] (31) CNN structures based on parallel attention mechanisms, such as Figure 2 shown.

[0027] (32) The specific steps are as follows:

[0028] 1) Input processing: The original input is split into two different scales. Each CNN input is a subsequence of length s (original scale); the corresponding attention mechanism module inputs a subsequence of length sa (cross-scale), with the cross-scale subsequence taking the midpoint of the original scale subsequence as its midpoint. Furthermore, setting sa > s allows the attention module to more comprehensively grasp the context and accurately capture the local features of Chinese addresses.

[0029] 2) The CNN module is composed of multiple layers of stacked one-dimensional networks. Each layer includes a convolutional layer, a batch normalization layer, and a nonlinear layer. Pooling layers are used to aggregate samples. The stacking of convolutional layers creates a hierarchical structure that extracts progressively more abstract features. The module outputs m feature sequences of length n, which can be expressed as (n × m).

[0030] 3) The attention mechanism module consists of two parts: feature aggregation and scale restoration. The feature aggregation part uses a stack of multiple convolution and pooling layers to extract key features from cross-scale subsequences. The last layer uses a 1×1 convolution kernel to mine linear relationships. The scale restoration part restores the key features to (n×m)

[0031] 4) Through the attention module, the influence of important time series features on the model can be increased, and the interference of unimportant features on the model can be suppressed, which effectively solves the problem that the model cannot distinguish the differences in the importance of time series data.

[0032] 5) Parallel feature fusion: The output features of the CNN module are element-wise multiplied by the saliency features output by its corresponding attention mechanism module. The higher the importance of the CNN output feature, the closer the output of the corresponding attention mechanism module is to 1; conversely, the lower the importance of the CNN output feature, the closer the output of the corresponding attention mechanism module is to 0. The value reflects the importance of the local feature, thereby achieving the identification of important local features.

[0033] 6) Use the parallel attention mechanism to better extract the local features of Chinese addresses and improve the model's ability to parse short sequence texts.

[0034] The proposed GLSMA algorithm described in step (4) optimizes the network structure of BiLSTM.

[0035] (41) The specific steps of Logistic Chaotic Mapping are as follows:

[0036] 1) Introduce the Logistic Chaotic Map to initialize the population so that the population is evenly distributed. The formula is as follows:

[0037] x k+1 =μx k (1-x k ) (1.2)

[0038] where x k is the state at time n, and μ is a control parameter that determines the behavior of the system. k+1 is the state at the next time step.

[0039] 2) As the control parameter μ changes, the behavior of the system will change significantly. When μ is small, the system will tend to a fixed value, indicating that the system is stable. At this time, x k It will converge to a constant (fixed point).

[0040] 3) As μ increases, the system's behavior becomes periodic, manifesting as a period doubling phenomenon. That is, the system oscillates periodically within a certain time interval, but the length of each period is twice that of the previous one.

[0041] 4) When μ exceeds a certain critical value, the system will exhibit chaotic behavior, that is, the state of the system is very sensitive to the initial conditions, making it difficult to predict and control.

[0042] 5) Use Logistic chaotic mapping to initialize the population to make the population evenly distributed, thereby improving the convergence speed and optimization efficiency of the algorithm.

[0043] (42) The specific steps of the genetic learning strategy are as follows:

[0044] 1) Introduce a genetic learning strategy to perform a global search in the solution space through crossover, mutation, and selection operations.

[0045] 2) Mutation operation: For individual i, its excellent offspring Obtained by the following formula:

[0046]

[0047] Among them, rand is a random number uniformly distributed between [0, 1], k is a random number in the set {1, 2, ..., N}, N is the population size, X pi and X pk are the historical optimal positions of the i-th and k-th individuals, is the d-dimensional historical optimal position of the i-th individual, is the global optimal position, and f(·) represents the individual fitness. and The convex combination integrates global and local individual information and increases the probability of individuals becoming excellent offspring.

[0048] 3) Mutation operation: The dth dimension of the offspring is mutated according to the probability, and the formula is as follows:

[0049]

[0050] Among them, p m is the mutation probability, UB d With LB d are the upper and lower bounds of the d-dimensional search space, respectively. The mutation operation introduces new gene combinations by changing the chromosomes of individuals, helping the algorithm to escape from the local optimal solution.

[0051] 4) Selection operation: construct elite offspring by selecting outstanding individuals. The formula is as follows:

[0052]

[0053] Among them, E i represents the elite descendants of individual i.

[0054] 5) For each individual, the aforementioned crossover, mutation, and selection operations are iteratively performed, using a convex combination of historical and global optimal positions to generate elite offspring, providing guidance for the evolution of the slime mold individuals. This effectively avoids falling into local optimality and enhances global search capabilities.

[0055] The application of CRF (conditional random field) in step (5) outputs the parsed Chinese address.

[0056] (51) Conditional random fields are a serialization annotation algorithm that is often used for sequence annotation problems. Conditional random fields are applied to the field of NLP to reduce the output error of the sequence modeling layer.

[0057] (52) The definition of conditional random field is as follows:

[0058] 1) Given an input sequence X(X1, X2, ..., X N ), output sequence Y(Y1, Y2, ..., Y N ), i ranges from 1 to n, X is the observed state; Y is the hidden state; X and Y have the same structure, and their definition formulas are as follows:

[0059] P(Y i |X,Y1,Y2,…,Y n )=P(Y i |X,Y i-1 , Y i+1 ) (1.6)

[0060] 2) Reduce the output probability of wrong labels through constraints and use the softmax function for normalization. The discriminant calculation formula is as follows:

[0061]

[0062] Where, Score(X, Y) represents the comprehensive evaluation score; n represents the number of characters; Indicates that the i-th character is marked as y i The probability of; A represents the transfer matrix obtained by CRF learning; represents the true label; W x Represents the predicted label sequence; P(Y|X) represents the corresponding probability of the input sequence and the label sequence.

[0063] 3) When the P(Y|X) value is closer to 1, it means that the predicted result is consistent with the correct labeling result, and the Chinese address data has been effectively trained. Ultimately, the globally optimal Chinese address label sequence is obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 Flowchart of the present invention;

[0065] Figure 2 This is a diagram of the CNN structure based on the parallel attention mechanism of the present invention;

[0066] Figure 3 This is the initialization sequence diagram of the Logistic chaotic map of the present invention;

[0067] Figure 4 This is a parameter optimization diagram for the genetic learning strategy of the present invention;

[0068] Figure 5 A line graph showing the regularization rate and batch size of the BiLSTM network optimized by the GLSMA algorithm of the present invention.

[0069] Figure 6 The Chinese address dataset used in the experiments of this invention;

[0070] Figure 7 Compare the present invention with other Chinese address resolution models; DETAILED DESCRIPTION

[0071] The present invention will be described in further detail below with reference to the accompanying drawings.

[0072] The overall flow chart of the present invention is as follows Figure 1 shown. The present invention performs data preprocessing and rule-based BIOES annotation on the Chinese address data crawled from the Baidu map open source platform to obtain standardized Chinese address data and labels of address elements. The specific number of Chinese address data and the number of label categories are as follows: Figure 6 shown.

[0073] The decoupled attention mechanism of the DeBERTa model is then used to capture the structural features of the address and the dependencies between different address levels. A CNN structure based on the parallel attention mechanism is introduced to better extract the local features of the address and improve the model's ability to parse short sequence texts. Figure 2 As shown;

[0074] The BiLSTM network is used to further model the features of long sequence texts and extract the time sequence information of address texts. In order to solve the problem that the random initialization of the population in SMA (Slime Mold Algorithm) leads to uneven population distribution and the algorithm is prone to falling into local optimality in the later stages of iteration, GLSMA (Improved Slime Mold Algorithm) is proposed. This algorithm introduces Logistic Chaotic Map to initialize the population, making the population uniformly distributed and improving the convergence speed and optimization efficiency of the algorithm. The initialization results are as follows: Figure 3 As shown in Figure 1, when μ = 4, the value of x diverges to 0 and 1, resulting in chaotic phenomena, and the sequence is more random. Introducing a genetic learning strategy, through crossover, mutation, and selection operations, a global search is performed in the solution space, which can effectively avoid falling into local optimality and has a stronger global search capability. The parameter optimization results are shown in Figure 1. Figure 4 As shown, the learning rate is Lr = 0.00001. The GLSMA algorithm optimizes the regularization rate and batch size of the BiLSTM network as shown in Figure 5 shown.

[0075] Finally, CRF (Conditional Random Field) is applied to output the parsed Chinese address.

[0076] The performance of the improved Chinese address resolution model is evaluated using indicators such as Accuracy and F1. The Accuracy calculation formula is:

[0077]

[0078] Where TP (True Positives) is the number of true positives, that is, the number of examples correctly predicted as positive by the model. TN (True Negatives) is the number of true negatives, that is, the number of examples correctly predicted as negative by the model. FP (False Positives) is the number of false positives, that is, the number of examples incorrectly predicted as positive by the model. FN (False Negatives) is the number of false negatives, that is, the number of positive examples incorrectly predicted as negative by the model.

[0079] The F1 score is the harmonic mean of precision and recall. The F1 calculation formula is:

[0080]

[0081] Among them, Precision is the precision rate and Recall is the recall rate.

Claims

1. A Chinese address parsing method based on a pre-trained model, characterized in that (1) The Chinese address data crawled from the Baidu Map open source platform is preprocessed and labeled with BIOES based on rules to obtain standardized Chinese address data and labels of address elements. (2) The decoupled attention mechanism of the DeBERTa model is used to capture the structural features of standardized Chinese address data and the dependencies between different address levels. (3) A CNN structure based on a parallel attention mechanism is introduced to better extract the local features of Chinese addresses using the parallel attention mechanism, thereby improving the model's ability to parse short sequence texts. (4) The BiLSTM network is used to further model the long sequence text features, extract the temporal information of Chinese address text, and propose the GLSMA algorithm to optimize the BiLSTM network structure. (5) Apply CRF (Conditional Random Field) to output the parsed Chinese address.

2. A Chinese address resolution method based on a pre-trained model according to claim 1, characterized in that: The step (1) comprises the following steps: (11) First, we use crawler technology to obtain a large amount of Chinese address data on the Baidu Map open source platform to ensure the diversity and breadth of the data. By accessing the open API interface, we can obtain the original Chinese address data containing geographic information such as provinces, cities, and counties. (12) The obtained original Chinese address data is preprocessed, and the CPCATransformer (Chinese_Province_City_Area_mapper) module is introduced to complete the administrative elements of incomplete Chinese address data. (13) After preprocessing, the Chinese address data is annotated with BIOES to identify the different elements in the address. The BIOES annotation method annotates each address element as B (begin), I (inside), O (outside), E (end), and S (single). (14) Obtain standardized Chinese address data and labels of address elements.

3. The Chinese address resolution method based on a pre-trained model according to claim 1, characterized in that: The algorithm definition and steps of the decoupled attention mechanism using the DeBERTa model described in step (2) are as follows: (21) The DeBERTa model is based on BERT. Each character in the input layer is represented by two vectors, which encode its content and position respectively, and the attention weights between characters are calculated using a disentangled matrix based on their content and relative position. The specific steps are: For the token at position i in the sequence, two vectors are used, {H i } and {P i}j } represents it, indicating its content and its relative position to the token at position j respectively. The calculation of the cross-attention score between tokens i and j can be decomposed into four parts: Among them, {H i } and {P i}j } represents the content of token i and its relative position to token j; {H j } and {P j}i } represents the content of token j and its relative position to token 1. That is, the attention weight of a word pair can be calculated as the sum of scores of four attentions (content-to-content, content-to-position, position-to-content, and position-to-position) using the decoupling matrix of its content and position. (22) By introducing the decoupled attention mechanism and the position-to-content and position-to-position scores, the model can better understand the Chinese address sequence and capture the structural features of standardized Chinese address data and the dependencies between different address levels.

4. The Chinese address resolution method based on a pre-trained model according to claim 1, characterized in that: The steps for introducing the CNN structure based on the parallel attention mechanism in step (3) are as follows: (31) The CNN structure based on the parallel attention mechanism is shown in Figure 2. (32) The specific steps are as follows: 1) Input processing: The original input is split into two different scales. Each CNN input is a subsequence of length s (original scale); the corresponding attention mechanism module inputs a subsequence of length sa (cross-scale), with the cross-scale subsequence taking the midpoint of the original scale subsequence as its midpoint. Furthermore, setting sa > s allows the attention module to more comprehensively grasp the context and accurately capture the local features of Chinese addresses. 2) The CNN module is composed of multiple layers of stacked one-dimensional networks. Each layer includes a convolutional layer, a batch normalization layer, and a nonlinear layer. Pooling layers are used to aggregate samples. The stacking of convolutional layers creates a hierarchical structure that extracts progressively more abstract features. The module outputs m feature sequences of length n, which can be expressed as (n × m). 3) The attention mechanism module consists of two parts: feature aggregation and scale restoration. The feature aggregation part uses a stack of multiple convolution and pooling layers to extract key features from cross-scale subsequences. The last layer uses a 1×1 convolution kernel to mine linear relationships. The scale restoration part restores the key features to (n×m) 4) Through the attention module, the influence of important time series features on the model can be increased, and the interference of unimportant features on the model can be suppressed, which effectively solves the problem that the model cannot distinguish the differences in the importance of time series data. 5) Parallel feature fusion: The output features of the CNN module are element-wise multiplied by the saliency features output by its corresponding attention mechanism module. The higher the importance of the CNN output feature, the closer the output of the corresponding attention mechanism module is to 1; conversely, the lower the importance of the CNN output feature, the closer the output of the corresponding attention mechanism module is to 0. The value reflects the importance of the local feature, thereby achieving the identification of important local features. 6) Use the parallel attention mechanism to better extract the local features of Chinese addresses and improve the model's ability to parse short sequence texts.

5. The Chinese address resolution method based on a pre-trained model according to claim 1, characterized in that: The proposed GLSMA algorithm described in step (4) optimizes the network structure of BiLSTM. (41) The specific steps of Logistic Chaotic Mapping are as follows: 1) Introduce the Logistic Chaotic Map to initialize the population so that the population is evenly distributed. The formula is as follows: x k+1 =μx k (1-x k ) (1.2) where x k is the state at time n, and μ is a control parameter that determines the behavior of the system. k+1 is the state at the next time step. 2) As the control parameter μ changes, the behavior of the system will change significantly. When μ is small, the system will tend to a fixed value, indicating that the system is stable. At this time, x k It will converge to a constant (fixed point). 3) As μ increases, the system's behavior becomes periodic, manifesting as a period doubling phenomenon. That is, the system oscillates periodically within a certain time interval, but the length of each period is twice that of the previous one. 4) When μ exceeds a certain critical value, the system will exhibit chaotic behavior, that is, the state of the system is very sensitive to the initial conditions, making it difficult to predict and control. 5) Use Logistic chaotic mapping to initialize the population to make the population evenly distributed, thereby improving the convergence speed and optimization efficiency of the algorithm. (42) The specific steps of the genetic learning strategy are as follows: 1) Introduce a genetic learning strategy to perform a global search in the solution space through crossover, mutation, and selection operations. 2) Mutation operation: For individual i, its excellent offspring Obtained by the following formula: Among them, rand is a random number uniformly distributed between [0, 1], k is a random number in the set {1, 2, ..., N}, N is the population size, X pi and X pk are the historical optimal positions of the i-th and k-th individuals, is the d-dimensional historical optimal position of the i-th individual, is the global optimal position, and f(·) represents the individual fitness. and The convex combination integrates global and local individual information and increases the probability of individuals becoming excellent offspring. 3) Mutation operation: The dth dimension of the offspring is mutated according to the probability, and the formula is as follows: Among them, p m is the mutation probability, UB d With LB d are the upper and lower bounds of the d-dimensional search space, respectively. The mutation operation introduces new gene combinations by changing the chromosomes of individuals, helping the algorithm to escape from the local optimal solution. 4) Selection operation: construct elite offspring by selecting outstanding individuals. The formula is as follows: Among them, E i represents the elite descendants of individual i. 5) For each individual, the aforementioned crossover, mutation, and selection operations are iteratively performed, using a convex combination of historical and global optimal positions to generate elite offspring, providing guidance for the evolution of the slime mold individuals. This effectively avoids falling into local optimality and enhances global search capabilities.

6. The Chinese address resolution method based on a pre-trained model according to claim 1, characterized in that: The application of CRF (conditional random field) in step (5) outputs the parsed Chinese address. (51) Conditional random fields are a serialization annotation algorithm that is often used for sequence annotation problems. Conditional random fields are applied to the field of NLP to reduce the output error of the sequence modeling layer. (52) The definition of conditional random field is as follows: 1) Given an input sequence X(X1, X2, ..., X N ), output sequence Y(Y1, Y2, ..., Y N ), i ranges from 1 to n, X is the observed state; Y is the hidden state; X and Y have the same structure, and their definition formulas are as follows: P(Y i |X, Y1, Y2,…, Y n )=P(Y i |X, Y i-1 ,Y i+1 ) (1.6) 2) Reduce the output probability of wrong labels through constraints and use the softmax function for normalization. The discriminant calculation formula is as follows: Where, Score(X, Y) represents the comprehensive evaluation score; n represents the number of characters; Indicates that the i-th character is marked as y i The probability of; A represents the transfer matrix obtained by CRF learning; represents the true label; W x Represents the predicted label sequence; P(Y|X) represents the corresponding probability of the input sequence and the label sequence. 3) When the P(Y|X) value is closer to 1, it means that the predicted result is consistent with the correct labeling result, and the Chinese address data has been effectively trained. Ultimately, the globally optimal Chinese address label sequence is obtained.