Text data processing method, apparatus, system, device, and storage medium

By extracting elements, performing hierarchical processing and segmentation on the dialogue text, and utilizing the lookup strategy of the address database, the problem of divergent wording in the dialogue text was solved, and the standardized processing of the dialogue text was achieved.

CN115587124BActive Publication Date: 2026-02-27IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211185415.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2026-02-27
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the issue of divergent wording in dialogue texts, particularly the standardization of address-related elements, failing to meet the standardization requirements of multiple elements.

Method used

By extracting elements, performing hierarchical processing and segmentation on the dialogue text to be processed, using different search strategies to search the address database, obtaining candidate addresses, and sorting the candidate addresses, the dialogue text is standardized.

Benefits of technology

It improves the efficiency of standardized processing of dialogue text data, solves the problem of divergent statements of elements, and realizes standardized processing of multiple elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115587124B_ABST
    Figure CN115587124B_ABST
Patent Text Reader

Abstract

The application discloses a text data processing method, device, system, equipment and storage medium. The method comprises the following steps: obtaining a to-be-processed dialogue text, performing element extraction on the to-be-processed dialogue text, and obtaining a plurality of initial elements; performing hierarchical processing on the plurality of initial elements, and obtaining a plurality of initial addresses; performing segmentation processing on each initial address, and obtaining a plurality of initial sub-addresses; wherein different initial sub-addresses are provided with different search strategies; searching an address database based on each initial sub-address of each initial address and the search strategy matched with each initial sub-address, obtaining a plurality of candidate addresses; sorting the plurality of candidate addresses, obtaining a candidate address queue, and outputting the candidate address queue. The technical scheme of the application solves the problem that the standardization requirement of the to-be-processed dialogue text cannot be met, especially the problem of divergent element statements in the to-be-processed dialogue text data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of natural language, and particularly relates to a text data processing method, device, system, equipment and storage medium. BACKGROUND

[0002] Currently, element standardization is mainly achieved by classifying elements, so as to realize the standardization processing of elements. This is suitable for simple category element standardization, but is not suitable for address elements with divergent statements and large categories. In reality, elements are often not enumerable.

[0003] In addition, there are some technical solutions for address standardization at present, which are realized by a method similar to a World Wide Web (web) service calling a Point of Interest (POI) retrieval, inputting an address and returning address ground supplement and correction. However, this kind of solution can only meet the standardization processing needs of a single element, and cannot meet the standardization needs of input text, especially the problem of divergent statements of elements in dialogue text data. The effect is very limited. SUMMARY

[0004] The present application aims to at least solve one of the technical problems in the related art. To this end, one object of the present application is to provide a text data processing method, device, system, equipment and storage medium.

[0005] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:

[0006] The embodiments of the present application provide a text data processing method, comprising:

[0007] Obtaining a dialogue text to be processed, and performing element extraction on the dialogue text to be processed to obtain a plurality of initial elements;

[0008] Performing hierarchical processing on the plurality of initial elements to obtain a plurality of initial addresses;

[0009] Performing segmentation processing on each initial address to obtain a plurality of initial sub-addresses; wherein different initial sub-addresses are provided with different search strategies;

[0010] Based on each initial sub-address of each initial address and the search strategy matched with each initial sub-address, searching an address database to obtain a plurality of candidate addresses;

[0011] Sorting the plurality of candidate addresses to obtain a candidate address queue, and outputting the candidate address queue.

[0012] Optionally, the hierarchical processing on the plurality of initial elements to obtain a plurality of initial addresses comprises:

[0013] dividing the plurality of initial elements based on a preset division rule to determine a hierarchy matched with each of the initial elements; wherein the preset hierarchy division rule comprises a plurality of hierarchies, and each of the hierarchies is provided with a serial number M;

[0014] performing hierarchical processing on the plurality of initial elements according to the preset hierarchy division rule and the hierarchy matched with each of the initial elements to obtain a plurality of initial addresses; wherein the plurality of initial elements in each of the initial addresses are arranged according to the serial number M; and M is a positive integer.

[0015] Optionally, the segmentation processing on each of the initial addresses to obtain a plurality of initial sub-addresses comprises:

[0016] obtaining an initial format of each of the initial addresses, and comparing the initial format with a reference format; if the initial format is inconsistent with the reference format, adjusting the initial format to the reference format based on the address database; wherein the reference format comprises the initial elements of N hierarchies; wherein N≥M, and N is a positive integer;

[0017] segmenting each of the initial addresses based on the hierarchy corresponding to each of the initial elements and a preset segmentation value, and obtaining a first initial sub-address, a second initial sub-address, a third initial sub-address and a fourth initial sub-address; wherein the preset segmentation value is used to determine the segmentation position of each of the initial addresses;

[0018] the searching of the address database based on each of the initial sub-addresses of each of the initial addresses and the search strategy matched with each of the initial sub-addresses to obtain a plurality of candidate addresses comprises:

[0019] when the initial sub-address is the first initial sub-address, searching the address database based on the first initial sub-address and the first search strategy matched with the first initial sub-address to obtain a first candidate address; or

[0020] when the initial sub-address is the second initial sub-address, searching the address database based on the second initial sub-address and the second search strategy matched with the second initial sub-address to obtain a first search result; searching the address database based on the first search result to obtain a second candidate address; or

[0021] When the initial sub-address is the third initial sub-address, then based on the third initial sub-address and the third search strategy matching the third initial sub-address, the address database is searched to obtain a third candidate address; or

[0022] When the initial sub-address is the fourth initial sub-address, the fourth candidate address is obtained based on the non-empty state of the fourth initial sub-address.

[0023] Optionally, when the initial sub-address is the second initial sub-address, the address database is searched based on the second initial sub-address and a second search strategy matching the second initial sub-address to obtain a first search result, including:

[0024] Based on the second initial sub-address, the initial element at the level of road is obtained;

[0025] The initial elements at the road level are processed to obtain search elements;

[0026] Based on each of the search elements, perform an intersection search on the address database to obtain intersection search results;

[0027] Based on the initial elements at each level that are roads, a road search is performed on the address database to obtain road search results;

[0028] The first search result is obtained based on the intersection search result and the road search result; wherein, obtaining the first search result based on the intersection search result and the road search result includes:

[0029] The address database is searched based on the intersection search results and the road search results to obtain a first index list;

[0030] The first index list is searched based on the intersection search results and the road search results to obtain a first candidate list;

[0031] The first candidate list is searched based on the intersection search results and the road search results to obtain the first search result.

[0032] Optionally, when the initial sub-address is the third initial sub-address, the address database is searched based on the third initial sub-address and a third search strategy matching the third initial sub-address to obtain a third candidate address, including:

[0033] A search is performed between the third initial sub-address and the address database to obtain the address data that matches the third initial sub-address;

[0034] a match degree of the address data and the third initial sub-address is calculated, and the match degree is compared with a match degree threshold;

[0035] If the match degree is greater than the match degree threshold, a third candidate address is obtained based on the address data; or if the match degree is less than the match degree threshold, a second index list is obtained by searching the address database based on the third initial sub-address;

[0036] The second index list is searched based on the third initial sub-address to obtain a second candidate list;

[0037] The second candidate list is searched based on the third initial sub-address to obtain a second search result;

[0038] The address database is searched based on the second search result to obtain the third candidate address.

[0039] Optionally, the plurality of candidate addresses are sorted to obtain a candidate address queue, and the candidate address queue is output, including:

[0040] A priority of each target element in each candidate address is determined based on a preset priority rule;

[0041] A level of the priority of each target element in each candidate address is obtained, and an upper limit of the level of the priority of each candidate address is determined;

[0042] The upper limit of the level of each candidate address is determined as a target level of the candidate address;

[0043] The plurality of candidate addresses are sorted according to the target level of each candidate address to obtain the candidate address queue, and the candidate address queue is output.

[0044] Embodiments of the present application also provide a text data processing apparatus, including:

[0045] An element extraction unit is configured to acquire a to-be-processed dialogue text, and perform element extraction on the to-be-processed dialogue text to obtain a plurality of initial elements;

[0046] A hierarchy unit is configured to perform hierarchy processing on the plurality of initial elements to obtain a plurality of initial addresses;

[0047] A segmentation unit is configured to perform segmentation processing on each initial address to obtain a plurality of initial sub-addresses; wherein different initial sub-addresses are provided with different search strategies;

[0048] The matching unit is configured to search the address database based on each initial sub-address of each initial address and the search strategy matched with each initial sub-address, and obtain a plurality of candidate addresses.

[0049] The output unit is configured to sort the plurality of candidate addresses, obtain a candidate address queue, and output the candidate address queue.

[0050] A text data processing system comprises:

[0051] The elements extraction module, the hierarchical division module, the address search module, and the sorting module are connected in sequence.

[0052] The elements extraction module is configured to acquire a to-be-processed dialogue text, perform element extraction on the to-be-processed dialogue text, and obtain a plurality of initial elements; perform hierarchical processing on the plurality of initial elements, and obtain a plurality of initial addresses.

[0053] The hierarchical division module is configured to perform division processing on each initial address, and obtain a plurality of initial sub-addresses; different initial sub-addresses are provided with different search strategies.

[0054] The address search module is configured to search an address database based on each initial sub-address of each initial address and the search strategy matched with each initial sub-address, and obtain a plurality of candidate addresses.

[0055] The sorting module is configured to sort the plurality of candidate addresses, obtain a candidate address queue, and output the candidate address queue.

[0056] Embodiments of the present application also provide an electronic device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the method described above when executing the computer program.

[0057] Embodiments of the present application also provide a computer-readable storage medium comprising a stored computer program, wherein the computer-readable storage medium controls a device in which the computer-readable storage medium is located to execute the method described above when the computer program runs.

[0058] Embodiments of the present application have the following technical effects:

[0059] The above technical solutions of the present application first perform element extraction on the to-be-processed dialogue text and obtain a plurality of initial elements, perform hierarchical processing on the plurality of initial elements to obtain an initial address, perform segmentation processing on the initial address to respectively obtain a plurality of initial sub-addresses; based on each initial sub-address and a search strategy corresponding to each initial sub-address, search the address database, and then perform standardized processing on the initial address, thereby solving the problem that the standardized demand of the to-be-processed dialogue text cannot be met, especially the problem of divergent expression of elements in the to-be-processed dialogue text data.

[0060] Additional aspects and advantages of the present application will be made apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0061] Figure 1 is a structural schematic diagram of a text data processing system provided by an embodiment of the present application;

[0062] Figure 2 is a flowchart of a text data processing method provided by an embodiment of the present application;

[0063] Figure 3 is a flowchart of a first search strategy provided by an embodiment of the present application;

[0064] Figure 4 is a flowchart of a second search strategy provided by an embodiment of the present application;

[0065] Figure 5 is a flowchart of a third search strategy provided by an embodiment of the present application;

[0066] Figure 6 is a flowchart of a fourth search strategy provided by an embodiment of the present application;

[0067] Figure 7 is a structural schematic diagram of a text data processing apparatus provided by an embodiment of the present application. DETAILED DESCRIPTION

[0068] The embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.

[0069] In order to facilitate the understanding of the embodiments by those skilled in the art, some terms are explained:

[0070] (1) NLP: Natural Language Processing, natural language processing.

[0071] (2) python: a programming language for scripting and rapid application development on most platforms.

[0072] (3) SQL: Structured Query Language, structured query language.

[0073] (4) TFIDF: term frequency-inverse document frequency, a commonly used weighting technique for information retrieval and data mining.

[0074] (5) entityFuzzy: fuzzy entity.

[0075] As shown in Figure 1 The embodiments of the present application also provide a text data processing system, comprising:

[0076] The elements extraction module, the hierarchical division module, the address search module and the sorting module are connected in sequence; wherein the elements extraction module, the hierarchical division module, the address search module and the sorting module can be connected based on the network;

[0077] The elements extraction module is used for obtaining a to-be-processed dialogue text, and performing element extraction on the to-be-processed dialogue text to obtain a plurality of initial elements; and performing hierarchical processing on the plurality of initial elements to obtain a plurality of initial addresses.

[0078] The hierarchical division module is used for performing segmentation processing on each of the initial addresses to obtain a plurality of initial sub-addresses; wherein different initial sub-addresses are provided with different search strategies.

[0079] The address search module is used for searching an address database based on each of the initial sub-addresses of each of the initial addresses and the search strategy matched with each of the initial sub-addresses to obtain a plurality of candidate addresses.

[0080] The sorting module is used for sorting the plurality of candidate addresses to obtain a candidate address queue, and outputting the candidate address queue.

[0081] The embodiments of the present application combine the elements extraction module, the hierarchical division module, the address search module and the sorting module to perform standardized processing on the obtained to-be-processed dialogue text.

[0082] Specifically, first, element extraction is performed on each to-be-processed dialogue text to obtain initial elements of multiple levels such as province, city, district, county, town, street, place, organization, house number, and building;

[0083] The initial elements of each level can also include multiple initial elements. Second, after obtaining the initial elements of multiple levels, the embodiment of the application performs level processing to obtain multiple initial addresses. Each initial address is segmented to obtain multiple initial sub-addresses. The address database is searched based on each initial sub-address to obtain multiple candidate addresses, thereby improving the efficiency of standardizing the to-be-processed dialogue text. Different initial sub-addresses correspond to different search strategies.

[0084] An optional embodiment of the application, as shown in Figure 1 , further includes a data preprocessing module. The data preprocessing module is connected to the level division module based on a network. The data preprocessing module is configured to receive multiple initial sub-addresses corresponding to each initial address, preprocess the multiple initial sub-addresses corresponding to each initial address, and obtain a preprocessing result. The preprocessing includes processing the format and abnormalities of each initial address.

[0085] Further, as shown in Figure 1 , the data preprocessing module is further connected to the address search module based on a network and is configured to send the preprocessing result to the address search module.

[0086] An optional embodiment of the application, as shown in Figure 1 , the address search module is connected to the address database. The address search module is configured to search the address database based on each initial sub-address of each initial address obtained, and then obtain multiple candidate addresses based on the address database.

[0087] An optional embodiment of the application, as shown in Figure 1 , further includes a data post-processing module. The data post-processing module is connected to the address module based on a network. The data post-processing module is configured to receive each candidate address, post-process each candidate address, and obtain a post-processing result. The post-processing includes processing the format and abnormalities of each candidate address.

[0088] Further, as shown in Figure 1 , the data preprocessing module is further connected to the sorting module based on a network and is configured to send the post-processing result to the sorting module.

[0089] As shown in Figure 2 , the embodiment of the application further provides a text data processing method applied to the text data processing system as shown in Figure 1 , which includes the following steps.

[0090] Step S21: obtaining a to-be-processed dialogue text, and performing element extraction on the to-be-processed dialogue text to obtain a plurality of initial elements;

[0091] Embodiments of the present application can first perform element extraction on each to-be-processed dialogue text based on NLP. For specific elements to be extracted, presetting can be performed according to actual needs, for example, 8 levels of elements need to be extracted; wherein the 8 levels of elements can be: province, city, district / county, township / town, road, institution, location, and house number.

[0092] Further, after obtaining a to-be-processed dialogue text, element extraction is performed on the to-be-processed dialogue text based on NLP based on the above-mentioned 8 levels of elements, and a plurality of initial elements are obtained.

[0093] For example, the following initial elements can be obtained:

[0094] 1) A province, B city, C district / county, D township / town, E road, F institution, G location, H house number, A province, b city;

[0095] 2) A province, B city, C district / county, D township / town, E road, F institution, G location, H house number, c district / county, d township / town;

[0096] 3) A province, B city, C district / county, D township / town, E road, F institution, G location, H house number, e road, i road, j road;

[0097] 4) A province, E road, F institution, G location, H house number, e road, i road, j road.

[0098] By analogy, when performing element extraction on the to-be-processed dialogue text, it is possible that due to user semantic repetition, uncertainty, or accent, an initial element of a certain level corresponds to multiple different information; it is also possible that the user has abbreviated the content to be expressed, and the initial element of that level cannot be extracted.

[0099] Step S22: performing hierarchical processing on the plurality of initial elements to obtain a plurality of initial addresses;

[0100] In an optional embodiment of the present application, the hierarchical processing of the plurality of initial elements to obtain a plurality of initial addresses comprises:

[0101] dividing the plurality of initial elements based on a preset division rule to determine a level matched with each initial element; wherein the preset hierarchical division rule comprises a plurality of levels, and each level is provided with a serial number M;

[0102] According to the preset hierarchical division rule and the level matched with each initial element, the plurality of initial elements are processed in levels to obtain a plurality of initial addresses; wherein the plurality of initial elements in each initial address are arranged according to the serial number M; wherein M is a positive integer.

[0103] The embodiments of the present application are explained and described based on the above-mentioned preset eight levels of elements; based on the preset division rule, the level corresponding to the element of each level in the above-mentioned eight levels of elements is determined, wherein in order to facilitate the differentiation of each level, a serial number is set for each level;

[0104] For example, the following preset can be made:

[0105] The serial number of the level of the province is M=1, and the level corresponding to the province is the first level;

[0106] The serial number of the level of the city is M=2, and the level corresponding to the city is the second level;

[0107] The serial number of the level of the district / county is M=3, and the level corresponding to the district / county is the third level;

[0108] The serial number of the level of the township / town is M=4, and the level corresponding to the township / town is the fourth level;

[0109] The serial number of the level of the road is M=5, and the level corresponding to the road is the fifth level;

[0110] The serial number of the level of the agency is M=6, and the level corresponding to the agency is the sixth level;

[0111] The serial number of the level of the location is M=7, and the level corresponding to the location is the seventh level;

[0112] The serial number of the level of the house number is M=8, and the level corresponding to the house number is the eighth level.

[0113] Further, according to the size of M, the initial elements of each level in each initial address are arranged in the embodiments of the present application; specifically, the order of the levels is: first level>second level>third level>fourth level> fifth level> sixth level> seventh level> eighth level;

[0114] Based on the above-mentioned order of levels, a plurality of initial addresses obtained by combining the initial elements of a plurality of levels can be obtained;

[0115] For example: 1) A province, B city, C district / county, D township / town, E road, F agency, G location, H house number;

[0116] 2) A province, B city, F agency, G location, H house number;

[0117] 3) B city, C district / county, D township / town, E road, F agency, G place.

[0118] Step S23: segmenting each of the initial addresses to obtain a plurality of initial sub-addresses; wherein different initial sub-addresses are provided with different search strategies;

[0119] In an optional embodiment of the present application, the segmenting each of the initial addresses to obtain a plurality of initial sub-addresses comprises:

[0120] obtaining an initial format of each of the initial addresses, and comparing the initial format with a reference format, if the initial format is inconsistent with the reference format, adjusting the initial format to the reference format based on the address database; wherein the reference format comprises N levels of the initial elements; wherein N≥M, and N is a positive integer;

[0121] segmenting each of the initial addresses based on the level corresponding to each of the initial elements and a preset segmentation value, and obtaining a first initial sub-address, a second initial sub-address, a third initial sub-address and a fourth initial sub-address; wherein the preset segmentation value is used to determine a segmentation position of each of the initial addresses.

[0122] In an embodiment of the present application, although the above-mentioned 8 levels of initial elements are all address elements, since there is also a strict distinction between address elements in definition, in order to improve the standardization efficiency of the to-be-processed dialogue text, at least one segmentation position that needs to be segmented is determined based on a preset segmentation value, each initial address is segmented based on all the determined segmentation positions to obtain a plurality of initial sub-addresses, and different initial sub-addresses are formulated with different standardization strategies.

[0123] Further, in an embodiment of the present application, each initial address is divided into 4 parts based on the level, including a first initial sub-address, a second initial sub-address, a third initial sub-address and a fourth initial sub-address.

[0124] In an optional embodiment of the present application, when the levels of the plurality of initial elements are M < N = 8 due to the user abbreviating some content, the missing level corresponding initial element needs to be supplemented; specifically, in the supplementing process, the initial element that has been extracted can be supplemented:

[0125] For example: when the initial element obtained includes Hefei City, but the initial element obtained does not include Anhui Province, according to Hefei City, it can be determined that the province corresponds to Anhui Province.

[0126] Further, when the name of the community in the appearing place is repeated, but there is no road, all the roads related to the community with the name need to be obtained and supplemented.

[0127] It should be noted that in the process of supplementing the initial elements, some levels may not exist, such as towns, etc. Therefore, when supplementing the initial elements corresponding to the levels, null values can be supplemented.

[0128] In an optional embodiment of the present application, the initial address is segmented based on the level corresponding to each initial element and a preset segmentation value, and first, second, third, and fourth initial sub-addresses are obtained, including:

[0129] When 1≤m≤the first segmentation value, the first initial sub-address is obtained based on the initial element corresponding to the serial number; wherein m is the serial number of any one level in each initial address; m is a positive integer, and m≤M;

[0130] When m=the second segmentation value, the second initial sub-address is obtained based on the initial element corresponding to the serial number;

[0131] When the third segmentation value≤m≤the fourth segmentation value, the third initial sub-address is obtained based on the initial element corresponding to the serial number;

[0132] When m=M, the fourth initial sub-address is obtained based on the initial element corresponding to the serial number.

[0133] In an optional embodiment of the present application, M, the first segmentation value, the second segmentation value, the third segmentation value, and the fourth segmentation value are preset.

[0134] In an embodiment of the present application, it is assumed that M=8, the first segmentation value is 4, the second segmentation value is 5, the third segmentation value is 6, and the fourth segmentation value is 7.

[0135] When 1≤m≤4, the first initial sub-address is obtained based on the initial element corresponding to the serial number;

[0136] When m=5, the second initial sub-address is obtained based on the initial element corresponding to the serial number;

[0137] When 6≤m≤7, the third initial sub-address is obtained based on the initial element corresponding to the serial number;

[0138] When m=8, the fourth initial sub-address is obtained based on the initial element corresponding to the serial number.

[0139] An optional embodiment of the present application, after obtaining the completion of the segmentation processing of each initial address, the segmentation result also needs to be preprocessed, and then the final first initial sub-address, second initial sub-address, third initial sub-address and fourth initial sub-address can be obtained.

[0140] For example: 1) preprocessing is used to solve the problem of abnormal level;

[0141] The segmentation result is: Yaohai, Changfeng County, Hefei City;

[0142] After preprocessing: Hefei City, Yaohai District, Changfeng County.

[0143] 2) Preprocessing is used to filter invalid information:

[0144] The segmentation result is: 18 buildings 304, 14 buildings 602 in Yue Xidong Village;

[0145] After preprocessing: Yue Xidong Village.

[0146] 3) Preprocessing is used to check the district / county:

[0147] The segmentation result is: Shushan District, Longshan Road;

[0148] After preprocessing: Before searching for the address, preferentially search for the part of the address database corresponding to Shushan District.

[0149] 4) Preprocessing is used for deduplication;

[0150] The segmentation result is: Yaohai, Yaohai District, Hefei City;

[0151] After preprocessing: Hefei City, Yaohai District.

[0152] 5) Preprocessing is used to arrange and combine multiple initial elements with the level of road;

[0153] The segmentation result is: Taihu Road, Ningguo Road;

[0154] After preprocessing: Taihu Road Ningguo Road, Ningguo Road Taihu Road.

[0155] An optional embodiment of the present application, when one of the obtained initial addresses is A province, B city, C district / county, D township / town, E road, F organization, G place, H door number;

[0156] Based on the above method, the following is obtained:

[0157] The first initial sub-address: A province, B city, C district / county, D township / town;

[0158] The second initial sub-address: E road;

[0159] The third initial sub-address: F organization, G place;

[0160] Fourth initial sub-address: H address.

[0161] In one optional embodiment of this application, one of the obtained initial addresses is: Province A, City B, District / County C, Township / Town D, Road E, Location G, and House Number H.

[0162] Based on the above method, we obtain:

[0163] First initial sub-address: Province A, City B, District / County C, Township / Town D;

[0164] Second initial sub-address: E-path;

[0165] Third initial sub-address: Location G;

[0166] Fourth initial sub-address: H address.

[0167] In the embodiments of this application, when no initial element of the organization level is extracted during the element extraction process, if some locations and organizations are duplicated, that is, any initial element of some locations and their corresponding organizations can represent the address information corresponding to the sixth and seventh levels. Therefore, the address information of these two levels can be represented based on the location without supplementing the organization. For example, if the initial element is No. 50 Fengxing New Village (XXX Hotel), then either No. 50 Fengxing New Village or XXX Hotel can be used to represent the address information corresponding to the sixth and seventh levels.

[0168] Step S24: Based on each initial sub-address of each initial address and the search strategy matching each initial sub-address, search the address database to obtain multiple candidate addresses;

[0169] In an embodiment of this application, the address database can be constructed based on a topographic map of a certain area to obtain an initial address database. During the search process of the address database, if abnormal data is found, that is, if the address data that cannot be matched cannot be found in the address database, the address database can add the abnormal data to the address database to enrich the address database.

[0170] In an optional embodiment of this application, the step of searching the address database based on each initial sub-address of each initial address and the search strategy matching each initial sub-address to obtain multiple candidate addresses includes:

[0171] When the initial sub-address is the first initial sub-address, the address database is searched based on the first initial sub-address and a first search strategy matching the first initial sub-address to obtain a first candidate address; or

[0172] When the initial sub-address is the second initial sub-address, then the address database is searched based on the second initial sub-address and a second search strategy matched with the second initial sub-address, and a first search result is obtained; the address database is searched based on the first search result, and a second candidate address is obtained; or

[0173] When the initial sub-address is the third initial sub-address, then the address database is searched based on the third initial sub-address and a third search strategy matched with the third initial sub-address, and a third candidate address is obtained; or

[0174] When the initial sub-address is the fourth initial sub-address, then a fourth candidate address is obtained based on a non-empty state of the fourth initial sub-address.

[0175] Embodiments of the present application do not perform standardization processing on the fourth initial sub-address (i.e. the eighth level).

[0176] An optional embodiment of the present application, as shown in Figure 3 The first search strategy is:

[0177] Step S24a1: obtaining a first initial sub-address;

[0178] Step S24a2: judging whether the first initial sub-address is non-empty;

[0179] Step S24a3: if the first initial sub-address is non-empty, searching the address database based on the first initial sub-address, and obtaining a first candidate address; in the process of searching the address database based on the first initial sub-address, multiple address data may be obtained, and each address data matches one initial element in the first initial sub-address;

[0180] Step S24a4: calculating a first similarity between each address data and the corresponding initial element;

[0181] Step S24a5: comparing the first similarity with a similarity threshold;

[0182] Step S24a6: if the first similarity is greater than the similarity threshold, then the address data is reserved, and then multiple address data with higher matching degrees are obtained, and these address data are spliced according to the sequence numbers of the levels, and then one or more first candidate addresses are obtained.

[0183] Step S24a7: if the first similarity is less than the similarity threshold, then the address data is not reserved.

[0184] Further, the first similarity between each address data and the corresponding initial element can be calculated based on a cosine similarity algorithm.

[0185] Further, after obtaining a plurality of first candidate addresses, post-processing is performed on all the first candidate addresses, including de-duplication and null value filtering (deleting null values) of the plurality of candidate addresses, etc.

[0186] Embodiments of the present application, when the user transmits the initial elements of the province or the level of the province due to the accent or the mispronunciation, based on the above method, in the matching process, based on the size relationship of the first similarity, the mispronunciation or the accent problem can also be corrected, and the final first candidate address is obtained; for example: Anhui Province, Hefei City, Feixi County, xx Town.

[0187] An optional embodiment of the present application, when there is a conflict or repetition of the initial elements of a certain level of a certain first initial sub-address, the first initial sub-address is continued to be segmented, and the address database is continued to be searched;

[0188] For example, the first initial sub-address can be: Anhui Province, Hefei City, Qingyang County, xx Township;

[0189] Then, based on the above segmented results, Anhui Province, Hefei City, Qingyang County, xx Township; Anhui Province, Hefei City, Qingyang County, xx Township, the address database is continued to be searched;

[0190] If Anhui Province, Hefei City, Qingyang County, xx Township can be found, and the first similarity corresponding to the search result is greater than the similarity threshold, then Anhui Province, Hefei City, Qingyang County, xx Township is retained and spliced according to the sequence number of the level to obtain the first candidate address, and the first candidate address is post-processed to obtain the final first candidate address.

[0191] An optional embodiment of the present application, when the initial sub-address is the second initial sub-address, the address database is searched based on the second initial sub-address and the second search strategy matched with the second initial sub-address, and a first search result is obtained, including:

[0192] Based on the second initial sub-address, the initial elements of the level of the road are obtained;

[0193] The initial elements of the level of the road are processed to obtain search elements;

[0194] Based on each search element, the address database is cross-intersection searched to obtain a cross-intersection search result;

[0195] perform a road search on the address database based on the initial element of the road at each level to obtain a road search result;

[0196] obtain the first search result based on the intersection search result and the road search result.

[0197] Embodiments of the present application, as shown in Figure 4 The second search strategy is:

[0198] Step S24b1: obtain a second initial sub-address;

[0199] Step S24b2: determine whether the second initial sub-address is empty

[0200] Step S24b3: if the second initial sub-address is not empty, obtain an initial element of a road at a level based on the second initial sub-address; specifically, address information ending with "road", "intersection", "ring", "block", and "avenue" can be extracted based on a regularization rule to obtain the initial element of the road at the level.

[0201] Step S24b4: since the obtained initial element of the road at the level can include multiple elements, the initial element of the road at the level needs to be processed, for example, permutation and combination of the initial element of the road at the level is performed to realize merging to obtain a search element;

[0202] Step S24b5: perform an intersection search on the address database based on the above-mentioned fuzzy search algorithm (entityFuzzy) and the intersection (i.e., the search element) formed by permutation and combination to obtain an intersection search result;

[0203] For example, the initial element of the road at the level in a certain second initial sub-address includes a first road, a second road, and a third road.

[0204] Based on the above-mentioned algorithm, permutation and combination of the first road, the second road, and the third road can be performed to obtain:

[0205] A first search element: the first road, the second road, and the third road;

[0206] A second search element: the first road, the third road, and the second road;

[0207] A third search element: the second road, the first road, and the third road;

[0208] A fourth search element: the second road, the third road, and the first road;

[0209] A fifth search element: the third road, the first road, and the second road;

[0210] A sixth search element: the third road, the second road, and the first road;

[0211] Step S24b6: search the address database based on the above six search elements respectively to obtain intersection search results;

[0212] Step S24b7: search the address database based on the initial elements of the second initial sub-address (first road, second road and third road) to obtain road search results;

[0213] Step S24b8: obtain first search results based on the road search results and the intersection search results.

[0214] Further, step S24b9: after obtaining the first search results, search the address database based on the first search results to obtain one or more address data matched with the first search results.

[0215] Further, calculate a second similarity between each address data and the first search results based on the cosine similarity algorithm, and compare each second similarity with a similarity threshold value, if the second similarity is greater than the similarity threshold value, the address data is determined as a second candidate address; if the second similarity is less than the similarity threshold value, the address data is not reserved; and further obtain all second candidate addresses matched with the first search results.

[0216] Further, post-process all second candidate addresses and obtain final second candidate addresses, for example: Anhui Province, Hefei City, Feixi County, xx Town, Huazha Road xx.

[0217] In an optional embodiment of the present application, the first search results are obtained based on the intersection search results and the road search results, and the method comprises:

[0218] search the address database based on the intersection search results and the road search results respectively to obtain a first index list;

[0219] search the first index list based on the intersection search results and the road search results respectively to obtain a first candidate list;

[0220] search the first candidate list based on the intersection search results and the road search results respectively to obtain the first search results.

[0221] Embodiments of the present application can search the address database based on the intersection search results and the inverted index algorithm, and then search the address database based on the road search results and the inverted index algorithm to obtain a first index list;

[0222] The address database can be searched based on the TFIDF coarse recall candidate algorithm and the intersection search result, and then the first index list is searched based on the TFIDF coarse recall candidate algorithm and the road search result to obtain a first candidate list.

[0223] The first candidate list can be searched based on the deep learning fine ranking algorithm and the intersection search result, and then the first candidate list is searched based on the deep learning fine ranking algorithm and the road search result to obtain a first search result.

[0224] Further, when the first search result is obtained, each address data in the first search result corresponds to a similarity score, each similarity score is compared with a similarity score threshold, if the similarity score is greater than the similarity score threshold, the address data is retained; if the similarity score is less than the similarity score threshold, the address data is not retained; and a final first search result is obtained.

[0225] In an optional embodiment of the present application, when the initial sub-address is the third initial sub-address, the address database is searched based on the third initial sub-address and a third search strategy matched with the third initial sub-address to obtain a third candidate address, including:

[0226] The third initial sub-address is searched based on the address database to obtain the address data matched with the third initial sub-address;

[0227] The matching degree of the address data and the third initial sub-address is calculated and compared with a matching degree threshold;

[0228] If the matching degree is greater than the matching degree threshold, a third candidate address is obtained based on the address data.

[0229] In the embodiment of the present application, the matching degree of each address data and the corresponding initial element in the third sub-address can be calculated based on the cosine similarity algorithm.

[0230] In an optional embodiment of the present application, when the initial sub-address is the third initial sub-address, the address database is searched based on the third initial sub-address and a third search strategy matched with the third initial sub-address to obtain a third candidate address, further including:

[0231] If the matching degree is less than the matching degree threshold, a second index list is obtained by searching the address database based on the third initial sub-address;

[0232] The second index list is searched based on the third initial sub-address to obtain a second candidate list;

[0233] search the second candidate list based on the third initial sub-address, and obtain a second search result.

[0234] An optional embodiment of the present application is shown in Figure 5 The third search strategy is as follows:

[0235] Step S24c1: obtaining a third initial sub-address;

[0236] Step S24c2: judging whether the third initial sub-address is empty or not;

[0237] Step S24c3: if the third initial sub-address is not empty, searching the address database based on the third initial sub-address, and obtaining the address data matched with the third initial sub-address;

[0238] Step S24c4: obtaining a matching degree of the address data and the third initial sub-address, and comparing the matching degree with a matching degree threshold;

[0239] Step S24c5: if the matching degree is greater than the matching degree threshold;

[0240] Step S24c6: searching the address database based on the third initial sub-address, and obtaining a second index list;

[0241] Step S24c7: searching the second index list based on the third initial sub-address, and obtaining a second candidate list;

[0242] Step S24c8: searching the second candidate list based on the third initial sub-address, and obtaining a second search result.

[0243] Step S24c9: searching the address database based on the second search result, and obtaining a plurality of first candidate addresses.

[0244] An optional embodiment of the present application can search the address database based on an inverted index algorithm and the third initial sub-address, and obtain a second index list;

[0245] The second index list can be searched based on a TFIDF coarse recall candidate algorithm and the third initial sub-address, and a second candidate list is obtained;

[0246] The second candidate list can be searched based on a deep learning fine arrangement algorithm and the third initial sub-address, and the second search result is obtained.

[0247] Further, when the second search result is obtained, each address data in the second search result corresponds to a similarity score, each similarity score is compared with a similarity score threshold, if the similarity score is greater than the similarity score threshold, the address data is retained; if the similarity score is less than the similarity score threshold, the address data is not retained; and then a final second search result is obtained.

[0248] Further, after the final second search result is obtained, the address database is searched based on the final second search result to obtain one or more address data; a third similarity between each address data and each address element (each address element corresponds to an initial element) in the final second search result is calculated based on a cosine similarity algorithm, the third similarity is compared with a similarity threshold, if the third similarity is greater than the similarity threshold, the address data is retained, if the third similarity is less than the similarity threshold, the address data is not retained, and then one or more address data is obtained, and the address data is spliced according to the corresponding hierarchical serial number, and a third candidate address is obtained; for example: Anhui Province, Hefei City, Feixi County, xx Town, xx Hotel.

[0249] Further, all third candidate addresses are post-processed to obtain a final third candidate address.

[0250] For example, the address database can be:

[0251]

[0252] Suppose that based on the third initial sub-address, the address database is directly searched, and no return result is obtained (that is, based on the third initial sub-address, the address data obtained by searching the address database has a matching degree with the corresponding initial element less than a matching degree threshold);

[0253] Then when the third initial sub-address is “Fanhua Yicheng”, the address database is first searched based on the inverted index algorithm and “Fanhua Yicheng” to obtain a second index list; for example, the second index list includes 1000 indexes;

[0254] The second index list is searched based on the TFIDF coarse recall candidate algorithm and “Fanhua Yicheng” to obtain a second candidate list; for example, the second candidate list includes 100 candidates;

[0255] The second candidate list is searched based on the deep learning fine ranking algorithm and “Fanhua Yicheng” to obtain a second search result; and the similarity score corresponding to the second search result (Fanhua Yicheng) is 0.8, 0.8> similarity score threshold (0.68); therefore, the second search result is retained;

[0256] Then the address database is queried based on the second search result (Fanhua Yicheng), and the query method can include but is not limited to the following two methods:

[0257] 1) The pandas dataframe can be used to process data and find the address database under the python development environment: for example:

[0258] df = pd.read_csv(file, encoding="gbk")

[0259] df_city = df.loc[(df["pname"] == "Anhui Province") & (df["adname"] == "Feixi County")]["cityname"].

[0260] 2) SQL statements can also be written to directly find the address database, for example:

[0261] select * from table where location="Fanhua Yicheng".

[0262] Based on the above method, the third candidate address obtained can be: Anhui Province, Hefei City, Feixi County, Fanhua Avenue XX, Fanhua Yicheng; wherein, Anhui Province corresponds to Pname, Hefei City corresponds to City; Feixi County corresponds to Adname, Fanhua Avenue XX corresponds to Road, and Fanhua Yicheng corresponds to Location.

[0263] An optional embodiment of the present application is shown in Figure 6 The fourth search strategy is:

[0264] Step S24d1: obtaining a fourth initial sub-address;

[0265] Step S24d2: judging whether the fourth initial sub-address is empty or not;

[0266] Step S24d3: if not empty, the fourth initial sub-address is de-duplicated, and the fourth candidate address is returned.

[0267] It should be noted that the embodiments of the present application obtain the first initial sub-address, the second initial sub-address, the third initial sub-address and the fourth initial sub-address based on the initial address, and simultaneously perform standardization processing on the first initial sub-address, the second initial sub-address and the third initial sub-address, and obtain the first candidate address, the second candidate address and the third candidate address respectively; the embodiments of the present application do not need to perform standardization processing on the fourth initial sub-address, but directly de-duplicate the fourth initial sub-address to obtain the fourth candidate address.

[0268] Step S25: sorting the plurality of candidate addresses to obtain a candidate address queue, and outputting the candidate address queue.

[0269] In an optional embodiment of the present application, the step of sorting the plurality of candidate addresses to obtain a candidate address queue, and outputting the candidate address queue comprises:

[0270] determining a priority of each target element in each candidate address based on a preset priority rule;

[0271] obtaining a level of the priority of each target element in each candidate address, and determining an upper limit of the level of the priority of each candidate address;

[0272] determining the upper limit of the level of each candidate address as a target level of the candidate address;

[0273] sorting the plurality of candidate addresses according to the target level of each candidate address to obtain the candidate address queue, and outputting the candidate address queue.

[0274] In an embodiment of the present application, the preset priority levels of the seven levels are arranged in descending order as follows: the priority level of the seventh level > the priority level of the sixth level > the priority level of the fifth level > the priority level of the fourth level > the priority level of the third level > the priority level of the second level > the priority level of the first level > the priority level of the eighth level.

[0275] For example, Baohe District, Binhu New District, Hefei;

[0276] The priority level of the Binhu New District is the highest, and thus the priority level of the Binhu New District is the highest.

[0277] Based on the priority levels of the levels, the three candidate addresses are sorted to obtain: Binhu New District, Baohe District, and Hefei.

[0278] Further, for the candidate addresses with the same target level, a target element corresponding to the target level in each candidate address is determined, and an initial element corresponding to the target element is determined according to the target element. The initial element appearing first is arranged in front, and vice versa.

[0279] For example:

[0280] 1) candidate addresses with the same target level;

[0281] Input: Chaohu City, Dongfangjingyuan; No. 11 building; Chaohu City, Dongfangjingyuan community, No. 11 building, unit 106, Jin'an foot bath;

[0282] Before sorting output: Jinxin foot bath, No. 1-13, Tuanjie East Road, Dongfang Scenic Garden, Chaohu City, Hefei City, Anhui Province; Dongfang Scenic Garden, Chaohu City, Hefei City, Anhui Province;

[0283] After sorting output: Dongfang Scenic Garden, Chaohu City, Hefei City, Anhui Province; Jinxin foot bath, No. 1-13, Tuanjie East Road, Dongfang Scenic Garden, Chaohu City, Hefei City, Anhui Province.

[0284] 2) candidate addresses with different target levels;

[0285] Input: Dayang Town, Lu Yang District; Hailiang Blue County; shop house S1-116;

[0286] Before sorting output: Lu Yang District, Hefei City, Anhui Province; Dayang Town, Lu Yang District, Hefei City, Anhui Province; Hailiang Blue County, Qingyuan Road, Lu Yang District, Hefei City, Anhui Province;

[0287] After sorting output: Hailiang Blue County, Qingyuan Road, Lu Yang District, Hefei City, Anhui Province; Dayang Town, Lu Yang District, Hefei City, Anhui Province; Lu Yang District, Hefei City, Anhui Province.

[0288] Embodiments of the present application first perform element extraction on the to-be-processed dialogue text and obtain a plurality of initial elements, perform hierarchical processing on the plurality of initial elements to obtain an initial address, and perform segmentation processing on the initial address to respectively obtain a plurality of initial sub-addresses; based on each initial sub-address and a search strategy corresponding to each initial sub-address, search the address database, and then implement standardized processing on the initial address, thereby solving the problem that the standardized demand of the to-be-processed dialogue text cannot be met, especially the problem of divergent expression of elements in the to-be-processed dialogue text data.

[0289] In an optional embodiment of the present application, after the plurality of candidate addresses are sorted to obtain a candidate address queue and the candidate address queue is output, the following further includes:

[0290] An initial accuracy of each candidate address is calculated;

[0291] The initial accuracy is compared with an accuracy threshold;

[0292] The initial accuracy greater than the accuracy threshold is determined as a target accuracy;

[0293] The candidate address matching the target accuracy is determined as a to-be-feedback address;

[0294] The plurality of to-be-feedback addresses are sorted according to the target level of each to-be-feedback address to obtain a to-be-feedback address queue, and the to-be-feedback address queue is output.

[0295] Further, the initial accuracy can be calculated based on any one of the following ways:

[0296] 1) Initial accuracy = number of correct candidate addresses / total number of to-be-processed dialogue texts;

[0297] In the above formula, based on each to-be-processed dialogue text, a plurality of candidate addresses can be obtained, and the plurality of to-be-processed dialogue texts can be standardized.

[0298] 2) Initial accuracy = 2 * pre * rec / (pre + rec);

[0299] In the above formula,

[0300] pre = number of correct candidate addresses / total number of non-empty outputs;

[0301] rec = number of correct candidate addresses / total number of empty outputs;

[0302] In an embodiment of the present application, the initial accuracy can include top1-initial accuracy and top3-initial accuracy.

[0303] Further, in an embodiment of the present application, based on the initial accuracy, a candidate address queue with low initial accuracy can be deleted, and the candidate address queue can be used for user recommendation or feedback, and can also be used for evaluation of the precision of the system by a service party.

[0304] In an optional embodiment of the present application, when the input to-be-processed dialogue text is:

[0305] A: Hello, what can I do for you? B: Well, hello, I want to ask B: Is my house in this

Changhe Home

junior high school

Changhe Home

Changhe Home

Luyang District

Luyang District

Building Material Factory No. 1

Building Material Factory No. 1

[0306] The candidate address queue output based on the above algorithm of the present application is: Changhe Home, Luyang District, Hefei, Anhui Province; No. 331 Provincial Road North, Chaohu City, Hefei, Anhui Province, 50 meters from the building material factory; Luyang District, Hefei, Anhui Province.

[0307] In an optional embodiment of the present application, when the input to-be-processed dialogue text is:

[0308] A: Which town in Fudong County is the location you are at now? Is it a store or somewhere else? B: Dianbu Town, Fudong County, Dianbu Town, Fudong County, uh, the location of Fudong No. 3 Middle School A: Which road is Fudong No. 3 Middle School on, Dianbu Town? B: Uh, the intersection of Baogong Avenue and Qiaotouji Road A: Is Baogong Avenue or Qiaotouji Road at the intersection? B: Uh, the intersection is on Baogong Avenue. A: Is Baogong Avenue at Fudong No. 3 Middle School, the location of Fudong No. 3 Middle School? B: Uh.

[0309] The candidate address queue output based on the above algorithm of the present application is: the intersection of Baogong Avenue and Qiaotouji Road, Fudong County, Hefei City, Anhui Province; the intersection of Qiaotouji Road and Baogong Avenue, Fudong County, Hefei City, Anhui Province; Baogong Avenue, Fudong County, Hefei City, Anhui Province; Qiaotouji Road, Fudong County, Hefei City, Anhui Province; Dianbu Town, Fudong County, Hefei City, Anhui Province; Fudong County, Hefei City, Anhui Province.

[0310] As shown in Figure 7 Embodiments of the present application also provide a text data processing apparatus 70, comprising:

[0311] An element extraction unit 71 is configured to obtain a to-be-processed dialogue text, and perform element extraction on the to-be-processed dialogue text to obtain a plurality of initial elements;

[0312] A hierarchical unit 72 is configured to perform hierarchical processing on the plurality of initial elements to obtain a plurality of initial addresses;

[0313] A segmentation unit 73 is configured to perform segmentation processing on each of the initial addresses to obtain a plurality of initial sub-addresses; wherein different initial sub-addresses are provided with different search strategies;

[0314] A matching unit 74 is configured to search an address database based on each initial sub-address of each initial address and the search strategy matched with each initial sub-address to obtain a plurality of candidate addresses;

[0315] An output unit 75 is configured to sort the plurality of candidate addresses to obtain a candidate address queue, and output the candidate address queue.

[0316] Optionally, the hierarchical processing on the plurality of initial elements to obtain a plurality of initial addresses comprises:

[0317] dividing the plurality of initial elements based on a preset division rule to determine a hierarchical level matched with each initial element; wherein the preset hierarchical division rule comprises a plurality of hierarchical levels, and each hierarchical level is provided with a serial number M;

[0318] According to a preset hierarchical division rule and the hierarchy matched with each initial element, the plurality of initial elements are processed in hierarchy to obtain a plurality of initial addresses; wherein the plurality of initial elements in each initial address are arranged according to the serial number M; M is a positive integer.

[0319] Optionally, the segmentation processing of each initial address to obtain a plurality of initial sub-addresses comprises:

[0320] An initial format of each initial address is obtained, and the initial format is compared with a reference format; if the initial format is inconsistent with the reference format, the initial format is adjusted to the reference format based on the address database; wherein the reference format comprises the initial elements of N hierarchies; wherein N≥M, and N is a positive integer;

[0321] Based on the hierarchy corresponding to each initial element and a preset segmentation value, each initial address is segmented to obtain a first initial sub-address, a second initial sub-address, a third initial sub-address and a fourth initial sub-address; wherein the preset segmentation value is used to determine the segmentation position of each initial address.

[0322] Wherein, based on each initial sub-address of each initial address and the search strategy matched with each initial sub-address, the address database is searched to obtain a plurality of candidate addresses, comprising:

[0323] When the initial sub-address is the first initial sub-address, the address database is searched based on the first initial sub-address and the first search strategy matched with the first initial sub-address to obtain a first candidate address; or

[0324] When the initial sub-address is the second initial sub-address, the address database is searched based on the second initial sub-address and the second search strategy matched with the second initial sub-address to obtain a first search result; the address database is searched based on the first search result to obtain a second candidate address; or

[0325] When the initial sub-address is the third initial sub-address, the address database is searched based on the third initial sub-address and the third search strategy matched with the third initial sub-address to obtain a third candidate address; or

[0326] When the initial sub-address is the fourth initial sub-address, the fourth candidate address is obtained based on the non-empty state of the fourth initial sub-address.

[0327] Optionally, when the initial sub-address is the second initial sub-address, the address database is searched based on the second initial sub-address and a second search strategy matched with the second initial sub-address, and a first search result is obtained, including:

[0328] The initial element of the hierarchical road is obtained based on the second initial sub-address;

[0329] The initial element of the hierarchical road is processed to obtain a search element;

[0330] The address database is searched based on each search element to obtain a crossroad search result;

[0331] The address database is searched based on each initial element of the hierarchical road to obtain a road search result;

[0332] The first search result is obtained based on the crossroad search result and the road search result, including:

[0333] The address database is searched based on the crossroad search result and the road search result respectively to obtain a first index list;

[0334] The first index list is searched based on the crossroad search result and the road search result respectively to obtain a first candidate list;

[0335] The first search result is obtained by searching the first candidate list based on the crossroad search result and the road search result respectively.

[0336] Optionally, when the initial sub-address is the third initial sub-address, the address database is searched based on the third initial sub-address and a third search strategy matched with the third initial sub-address to obtain a third candidate address, including:

[0337] The address data matched with the third initial sub-address is obtained by searching the address database based on the third initial sub-address;

[0338] The matching degree of the address data and the third initial sub-address is calculated and compared with a matching degree threshold;

[0339] If the matching degree is greater than the matching degree threshold, a third candidate address is obtained based on the address data, or if the matching degree is less than the matching degree threshold, a second index list is obtained by searching the address database based on the third initial sub-address;

[0340] searching the second index list based on the third initial sub-address to obtain a second candidate list;

[0341] searching the second candidate list based on the third initial sub-address to obtain a second search result;

[0342] searching the address database based on the second search result to obtain the third candidate address.

[0343] Optionally, the sorting of the plurality of candidate addresses to obtain a candidate address queue and outputting the candidate address queue comprises:

[0344] determining a priority of each target element in each candidate address based on a preset priority rule;

[0345] obtaining a level of the priority of each target element in each candidate address and determining an upper limit of the level of the priority of each candidate address;

[0346] determining the upper limit of the level of each candidate address as a target level of the candidate address;

[0347] sorting the plurality of candidate addresses according to the target level of each candidate address to obtain the candidate address queue and outputting the candidate address queue.

[0348] Embodiments of the present application also provide an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the method described above.

[0349] Embodiments of the present application also provide a computer readable storage medium, comprising a stored computer program, wherein the computer program controls a device where the computer readable storage medium is located to execute the method described above when the computer program is running.

[0350] In addition, other configurations and functions of the device of the embodiments of the present application are known to those skilled in the art, and to reduce redundancy, they are not described here.

[0351] It is to be appreciated that the above description and the examples that follow are intended to be illustrative only and that changes can be made to the description, either functionally or chronologically, as well as changes being made concerning the order of implementation. The logic and / or steps represented in the flow diagrams and / or described herein can be considered as a sequence of executable instructions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) including one or more wires, discrete or integrated circuits, semiconductor memory modules, or other tangible media. A computer-readable medium can be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) including one or more wires, discrete or integrated circuits, semiconductor memory modules, or other tangible media. Other possibilities can also exist.

[0352] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, the various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, or combinations thereof, can be used: a discrete logic circuit having logic gates for implementing logic functions upon data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0353] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in an appropriate manner.

[0354] In the description of the application, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like is based on the orientation or positional relationship shown in the drawings, and is only for the purpose of facilitating the description of the application and simplifying the description, and therefore cannot be understood as indicating or implying that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the application.

[0355] In addition, the terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise explicitly specified and limited.

[0356] In this application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting", "fixing" and the like should be broadly understood, for example, it can be fixed connection, or detachable connection, or integral; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through intermediate medium, it can be the internal communication of two elements or the interaction relationship between two elements, unless otherwise explicitly limited. For those skilled in the art, the specific meaning of the above terms in this application can be understood according to the specific circumstances.

[0357] In this application, unless otherwise explicitly specified and limited, the first feature is "on" or "under" the second feature, which can be direct contact between the first and second features, or indirect contact between the first and second features through intermediate medium. Moreover, the first feature "above", "above" and "above" the second feature can be directly above or obliquely above the first feature, or only indicate that the horizontal height of the first feature is higher than that of the second feature. The first feature "below", "below" and "below" the second feature can be directly below or obliquely below the first feature, or only indicate that the horizontal height of the first feature is less than that of the second feature.

[0358] Although the embodiments of the application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the application.

Claims

1. A method of processing text data, characterized by, The method comprises the following steps: obtaining a to-be-processed dialogue text, and performing element extraction on the to-be-processed dialogue text to obtain a plurality of initial elements; performing hierarchical processing on the plurality of initial elements to obtain a plurality of initial addresses; performing segmentation processing on each of the initial addresses to obtain a plurality of initial sub-addresses; wherein different initial sub-addresses are provided with different search strategies; based on each of the initial sub-addresses of each of the initial addresses and the search strategy matched with each of the initial sub-addresses, searching the address database to obtain a plurality of candidate addresses; sorting the plurality of candidate addresses to obtain a candidate address queue, and outputting the candidate address queue; the hierarchical processing on the plurality of initial elements to obtain a plurality of initial addresses comprises: dividing the plurality of initial elements based on a preset hierarchical division rule to determine the hierarchy matched with each of the initial elements; wherein the preset hierarchical division rule comprises a plurality of hierarchies, and each of the hierarchies is provided with a serial number M; performing hierarchical processing on the plurality of initial elements according to the preset hierarchical division rule and the hierarchy matched with each of the initial elements to obtain a plurality of initial addresses; wherein the plurality of initial elements in each of the initial addresses are arranged according to the serial number M; M is a positive integer; the segmentation processing on each of the initial addresses to obtain a plurality of initial sub-addresses comprises: obtaining an initial format of each of the initial addresses, and comparing the initial format with a reference format; if the initial format is inconsistent with the reference format, adjusting the initial format to the reference format based on the address database; wherein the reference format comprises the initial elements of N hierarchies; wherein N≥M, and N is a positive integer; segmenting each of the initial addresses based on the hierarchy corresponding to each of the initial elements and a preset segmentation value, and obtaining a first initial sub-address, a second initial sub-address, a third initial sub-address, and a fourth initial sub-address; wherein the preset segmentation value is used to determine the segmentation position of each of the initial addresses; the searching the address database based on each of the initial sub-addresses of each of the initial addresses and the search strategy matched with each of the initial sub-addresses to obtain a plurality of candidate addresses comprises: when the initial sub-address is the first initial sub-address, searching the address database based on the first initial sub-address and a first search strategy matched with the first initial sub-address to obtain a first candidate address; or when the initial sub-address is the second initial sub-address, searching the address database based on the second initial sub-address and a second search strategy matched with the second initial sub-address to obtain a first search result; searching the address database based on the first search result to obtain a second candidate address; or when the initial sub-address is the third initial sub-address, searching the address database based on the third initial sub-address and a third search strategy matched with the third initial sub-address to obtain a third candidate address; or when the initial sub-address is the fourth initial sub-address, searching the address database based on the fourth initial sub-address and a fourth search strategy matched with the fourth initial sub-address to obtain a fourth candidate address. When the initial sub-address is the fourth initial sub-address, a fourth candidate address is obtained based on a non-empty state of the fourth initial sub-address.

2. The method of claim 1, wherein, When the initial sub-address is the second initial sub-address, the address database is searched based on the second initial sub-address and a second search strategy matched with the second initial sub-address, and a first search result is obtained, including: An initial element of a road level is obtained based on the second initial sub-address; The initial element of the road level is processed to obtain a search element; An intersection search is performed on the address database based on each search element to obtain an intersection search result; A road search is performed on the address database based on each initial element of the road level to obtain a road search result; The first search result is obtained based on the intersection search result and the road search result, including: The address database is searched based on the intersection search result and the road search result respectively to obtain a first index list; The first index list is searched based on the intersection search result and the road search result respectively to obtain a first candidate list; The first candidate list is searched based on the intersection search result and the road search result respectively to obtain the first search result.

3. The method of claim 1, wherein, When the initial sub-address is the third initial sub-address, a third candidate address is obtained based on the third initial sub-address and a third search strategy matched with the third initial sub-address, including: The address database is searched based on the third initial sub-address to obtain address data matched with the third initial sub-address; A matching degree between the address data and the third initial sub-address is calculated, and the matching degree is compared with a matching degree threshold; If the matching degree is greater than the matching degree threshold, a third candidate address is obtained based on the address data, or if the matching degree is less than the matching degree threshold, a second index list is obtained by searching the address database based on the third initial sub-address; The second index list is searched based on the third initial sub-address to obtain a second candidate list; The second candidate list is searched based on the third initial sub-address to obtain a second search result; The address database is searched based on the second search result to obtain the third candidate address.

4. The method of claim 1, wherein, The candidate addresses are sorted to obtain a candidate address queue, and the candidate address queue is output, including: A priority of each target element in each candidate address is determined based on a preset priority rule; A level of the priority of each target element in each candidate address is obtained, and an upper limit of the level of the priority of each candidate address is determined; The upper limit of the level of each candidate address is determined as a target level of the candidate address. According to the target level of each of the candidate addresses, the candidate addresses are sorted to obtain a candidate address queue, and the candidate address queue is output.

5. A text data processing apparatus characterized by comprising: The device is used to implement the method in any one of claims 1 to 4, and the device comprises: An element extraction unit is configured to obtain a to-be-processed dialogue text, perform element extraction on the to-be-processed dialogue text, and obtain a plurality of initial elements. A hierarchy unit is configured to perform hierarchy processing on the initial elements to obtain a plurality of initial addresses. A segmentation unit is configured to perform segmentation processing on each of the initial addresses to obtain a plurality of initial sub-addresses, wherein different initial sub-addresses are provided with different search strategies. A matching unit is configured to search an address database based on each of the initial sub-addresses of each of the initial addresses and the search strategy matched with each of the initial sub-addresses to obtain a plurality of candidate addresses. An output unit is configured to sort the candidate addresses to obtain a candidate address queue, and output the candidate address queue.

6. A text data processing system characterized by comprising: The system is used to implement the method in any one of claims 1 to 4, and the system comprises: An element extraction module, a hierarchy division module, an address search module, and a sorting module connected in sequence. The element extraction module is configured to obtain a to-be-processed dialogue text, perform element extraction on the to-be-processed dialogue text, and obtain a plurality of initial elements; and perform hierarchy processing on the initial elements to obtain a plurality of initial addresses. The hierarchy division module is configured to perform segmentation processing on each of the initial addresses to obtain a plurality of initial sub-addresses, wherein different initial sub-addresses are provided with different search strategies. The address search module is configured to search an address database based on each of the initial sub-addresses of each of the initial addresses and the search strategy matched with each of the initial sub-addresses to obtain a plurality of candidate addresses. The sorting module is configured to sort the candidate addresses to obtain a candidate address queue, and output the candidate address queue.

7. An electronic device, comprising: The computer readable storage medium comprises a stored computer program, wherein the computer program, when executed, controls a device in which the computer readable storage medium is located to implement the method in any one of claims 1 to 4.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored computer program, wherein the computer program, when executed, controls a device in which the computer readable storage medium is located to implement the method in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Address standardization method, device, storage medium and computer

    CN107145577A

  • Address standardization method and device

    CN110046352A