Method and related electronic device for performing address matching.

The electronic device and method automate address matching through entity extraction and NLP techniques, addressing format variations to achieve high accuracy and efficiency in shipping instruction processing.

JP7851917B2Active Publication Date: 2026-04-27MAERSK AS
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MAERSK AS
Filing Date
2021-09-10
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing address matching systems face challenges in accurately and efficiently processing shipping instructions due to variations in address formats, leading to inconsistent and costly manual corrections.

Method used

An electronic device and method utilizing entity extraction, natural language processing, and logistic regression to determine similarity scores and weights for automated address matching, improving accuracy and efficiency.

Benefits of technology

Achieves 99% accuracy and 80% recall in address matching, reducing manual intervention and enhancing processing speed and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007851917000001
    Figure 0007851917000001
  • Figure 0007851917000002
    Figure 0007851917000002
  • Figure 0007851917000003
    Figure 0007851917000003
Patent Text Reader

Abstract

A method for address matching, performed by an electronic device, is disclosed. The method includes obtaining a first dataset indicative of a first address. The method includes determining, based on the first dataset, one or more first elements indicative of the first address using an entity extraction technique. The method includes obtaining a second dataset based on at least one of the one or more first elements. The second dataset includes one or more second elements indicative of the second address. The method includes applying natural language processing (NLP) techniques to the first dataset and the second dataset to obtain one or more similarity scores. The method includes obtaining a set of weights associated with corresponding first elements of the first address.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of transporting and shipping goods. The present disclosure relates to a method for performing address matching and related electronic devices.

Background Art

[0002] In the shipping cycle, once a reservation is confirmed, a shipping instruction (SI) is issued by the customer. The SI may be further used to create a bill of lading (B / L). The SI can be considered as a key document containing shipping details. The SI includes a section called party information that describes details regarding the shipper, consignee, first notify party, and their corresponding postal addresses. The aforementioned addresses may already exist in the customer database. The addresses in the SI can be mapped by the addresses in the database.

[0003] However, not all addresses can be successfully mapped. Some addresses have variations (such as changes in text layout, typography, mistakes, typos and / or omitted addresses, etc.) and cannot be identified. The time spent dealing with such incorrect addresses is significant and can lead to inconsistent processing of incorrect addresses. Manual processing is prone to errors and is costly and time-consuming.

Summary of the Invention

[0004] There is a need for a technical process to assist in processing address data. There is a need for a tool that supports address matching and reduces the time spent dealing with incorrect addresses while maintaining and / or improving accuracy and consistency. For example, the disclosed technology enables automatic identification and matching of addresses between two sources.

[0005] Therefore, there is a need for electronic devices and methods for address matching that mitigate, reduce, or address existing shortcomings, and that provide automated, more time-efficient address identification and address matching with improved accuracy, scope, and consistency.

[0006] For example, a method for performing address matching using an electronic device is disclosed. The method includes obtaining a first dataset representing a first address. Based on the first dataset, the method includes determining one or more first elements representing the first address using an entity extraction technique. The method includes obtaining a second dataset based on at least one of the one or more first elements. The method includes the second dataset containing one or more second elements representing a second address. The method includes obtaining one or more similarity scores by applying natural language processing (NLP) techniques to the first and second datasets. The method includes obtaining a set of weights associated with the first elements corresponding to the first address. Based on the set of weights, the method includes determining matching parameters that indicate a match between at least one of the one or more first elements and at least one of the one or more second elements. The method optionally includes outputting the matching parameters.

[0007] An electronic device is disclosed, comprising a memory circuit, a processor circuit, and an interface circuit, wherein the electronic device is configured to perform any of the methods disclosed herein.

[0008] A computer-readable storage medium is disclosed for storing one or more programs, wherein the one or more programs, when executed by an electronic device having a display and a touch sensor surface, include instructions causing the electronic device to perform any of the methods disclosed herein.

[0009] The advantage of this disclosure is that the disclosed electronic devices and methods provide automated, more time-efficient address identification and address matching with improved accuracy, scope, and consistency. As a result, faster address correction can be achieved using the disclosed technology. This results in more robust processing of address correction in SI and SI.

[0010] This disclosure provides, in one or more embodiments, techniques for storing, retrieving, and processing address-related data, which improve the storage space used, the scope and accuracy of address matching, and the speed of processing SI.

[0011] The aforementioned and other features and advantages of this disclosure will be readily apparent to those skilled in the art through the following detailed description of typical embodiments with reference to the accompanying drawings. [Brief explanation of the drawing]

[0012] [Figure 1A] This diagram schematically illustrates the process performed by the electronic device example described herein using the disclosed technology. [Figure 1B] This diagram schematically illustrates the process performed by the electronic device example described herein using the disclosed technology. [Figure 2] This flowchart illustrates a typical method used by electronic devices to perform address matching as disclosed herein. [Figure 3] This is a block diagram illustrating a typical electronic device as described in this disclosure. [Modes for carrying out the invention]

[0013] Various typical embodiments and details are described below, with reference to the figures where applicable. Note that the figures may or may not be drawn to a fixed proportion, and that similar structural or functional elements are represented by similar reference numerals throughout the figures. Also note that the figures are intended solely to facilitate the description of the embodiments. They are not intended to be a limitation on the scope of this disclosure or an exhaustive description of this disclosure. In addition, the illustrated embodiments do not necessarily have all the aspects or advantages illustrated. Aspects or advantages described with a particular embodiment are not necessarily limited to that embodiment and can be performed in any other embodiment (even if not illustrated or explicitly described as such).

[0014] The figures are schematic and simplified for clarity, illustrating only details that aid in understanding this disclosure, while other details are omitted. Throughout, the same reference numerals are used for identical or corresponding parts.

[0015] Figures 1A and 1B schematically illustrate the processes performed by the electronic device example described herein using the disclosed technology.

[0016] Figure 1A shows a first dataset 14 representing the acquired first address 12. For example, the first dataset 14 representing the first address 12 may be obtained from a shipping instruction. The first address 12 may be considered an input address. The source and / or structure of the first dataset may be a shipping instruction. For example, the first dataset may be obtained based on a customer address, which is free text provided by the customer in the shipping instruction as input.

[0017] An example of the first dataset showing the first address may be in the following format: A&B Company (XXXXXX) CO. LTD. Room XXXXX 6th Floor Prime Number NO. 1361 North XXXXXX Street XXXXX Area XXXXX 857578XXXXX Contact: Name XXXXX Phone: XXXXXXXXX Email: XXXXX@XXX.COMXXXXXX Tax Reference: XXXXXXXXXXXX

[0018] One or more first elements 17, 18, 19, and 20 that represent the first address 12 are determined based on the first dataset 14 using an entity extraction technique 16.

[0019] The second dataset 24 is obtained, for example, by performing a search 22 based on at least one of the first elements 17, 18, 19, and 20. The second dataset 24 contains one or more second elements that represent the second address. The second dataset 24 can also be thought of as containing the corresponding database matching (for example, the database matching corresponding to the first address).

[0020] In other words, a first address (as an input address) contains one or more first elements that can be extracted as entities (e.g., customer trade name, broker name, city, country, telephone number, email ID, and tax reference number). Entities are extracted using NLP techniques and used to search for matching addresses in the customer database to generate a second dataset. Different entities are searched in the customer database, and as a first level of the search results, for example, 50 closest matching addresses are taken and constitute part of the second dataset as second addresses.

[0021] Natural language processing (NLP) techniques 26 are applied to the first dataset 14 and the second dataset 24 to obtain one or more similarity scores 28. The second addresses extracted from the second dataset are compared with the first addresses using an NLP technique (e.g., approximate sentence matching technique) to obtain similarity scores.

[0022] The similarity score may be considered, for example, as a parameter indicating the similarity between a first element and a second element for quantifying the similarity between the first element and the second element. The similarity score may be associated with a second element of a second dataset, for example, to indicate how similar the second element is to the corresponding first element. For example, a similarity score indicating a broker name may indicate how similar a second element related to the broker name is to a corresponding first element related to the broker name.

[0023] For example, equal weighting is given to each word at the first address in NLP to reach a second dataset with a similarity score. However, it turns out that some words are given more importance (e.g., weight) than other words when determining the final matching address. For example, words in a customer number have a greater impact on the result than in the case of a city or country name.

[0024] The disclosed technique provides for varying the weights for different words when determining a matching parameter as the final match between a first element of a first address and a corresponding second element of a second address. A set of weights is obtained by predictive analytics techniques to calculate the impact of different words of the first address when determining the result. This may be done, for example, by analyzing many first addresses (e.g., thousands of first addresses) and their corresponding second addresses (e.g., corresponding database matches).

[0025] Obtain a set of weights associated with the corresponding first element of the first address and apply it in NLP techniques in operation 30 to determine a matching parameter 32 indicating a match between at least one of the one or more first elements and at least one of the one or more second elements. In other words, the set of weights may be considered an improved set of weights (e.g., a customized set of weights adapted to the first address).

[0026] FIG. 1B shows an example of an embodiment that illustrates the disclosed process according to an aspect of the present disclosure. In FIG. 1B, a first data set indicating a first address is obtained.

[0027] One or more first elements indicating the first address (e.g., customer number, broker name, email ID, phone number) are determined using entity extraction techniques based on the first data set. For example, the first address (e.g., from a physical document) is fed into the system through PDF extraction or a similar tool. The first address includes multiple entities (e.g., customer number, broker name, email ID, phone number, and tax reference). For example, the first address is free text, the first address does not follow a specific format, and there may be numerous mistakes or anomalies. Entities are extracted by a custom address parsing algorithm.

[0028] A second data set is obtained based on at least one of the one or more first elements, for example, by performing a search 22. For example, to reduce the computational load of the system, a similarity score may be calculated for the most likely existing addresses (part of the second data set). A search service can be used to extract a set of the most likely addresses as the second data set. The search service may be created using all the addresses existing within the system. The entities extracted from the first address parsing are used as criteria in the search.

[0029] Natural language processing (NLP) techniques (e.g., cosine similarity) are applied to the first and second datasets to obtain one or more similarity scores associated with the corresponding second elements in the second dataset. The second dataset and the first address returned from the search service are fed to the NLP technique (e.g., the cosine similarity algorithm). The NLP technique calculates similarity scores (e.g., cosine similarity scores) for each record in the second dataset, associated with the second elements that make up the second dataset.

[0030] A set of weights is obtained that is associated with the corresponding first element of the first address, and applied in NLP techniques in custom machine learning and NLP techniques to determine, for example, the final matched address, along with an optional modified similarity score. For example, the influence factors of different words in the first address on the result are calculated by processing historical data using a logistic regression model to generate a set of weights.

[0031] It should be noted that cosine similarity measures the similarity between two documents. A document can be defined as an address, a sentence, or a paragraph. The default cosine similarity algorithm assigns equal weight to all words in the input document to determine the final match. However, in practice, different words in a document have different influences in determining the result. In this disclosure, regression analysis is used to devise custom weights for each token or word in the input to achieve the desired result. The disclosed technique yields results with 99% accuracy (compared to 95% with conventional cosine similarity) and an increased recall of matched addresses to 80% (compared to 70%).

[0032] In other words, once a standard similarity score is obtained using NLP techniques, the next step is to obtain a refined similarity score using custom weights on the acquired set of output addresses. For example, the top address (e.g., the one with the highest refined similarity score is considered the final output). If the refined similarity score for the final matched address is 0.5 or higher, the final output is considered a match; if the score is less than 0.5, it is explicitly stated that no matching address was found.

[0033] Figure 2 shows a flowchart illustrating an example 100 of a method for performing address matching according to this disclosure, which is performed by an electronic device (for example, the electronic device 300 in Figure 3) disclosed herein. The electronic device includes a memory circuit, a processor circuit, and an interface. The disclosed method may also be considered a method for matching text strings. The disclosed technique may be applied to any type of text string (including text strings that are not addresses).

[0034] Method 100 includes S102 of obtaining a first dataset indicating a first address (for example, via an interface and / or processor circuitry). For example, the first dataset indicating a first address may be obtained from a shipping instruction. The first address may be considered an input address. For example, the first dataset may be extracted from a shipping instruction. In one or more example methods, obtaining the first dataset S102 includes obtaining the first dataset from one or more first systems (for example, via an interface and / or processor circuitry) S102A (for example, retrieving and / or receiving). In one or more example methods, one or more first systems are operational systems. In one or more example methods, one or more operational systems include one or more of a shipping system, an invoice issuing system, and a case management system.

[0035] Method 100 includes determining one or more first elements that indicate a first address (for example, via a processor circuit) using an entity extraction technique based on a first dataset S104. The entity extraction technique can be thought of as an information extraction technique that can automatically extract entities from unstructured and / or semi-structured machine-readable text or documents and other electronically represented sources.

[0036] For example, one or more first elements may be considered as one or more first entities obtained from an entity extraction technique. In one or more example methods, the entity extraction technique includes one or more of the following: parsing techniques and element extraction techniques. For example, the entity extraction technique may use delimiters to, for example, parse the first dataset. For example, the entity extraction technique may use optical character recognition (OCR). For example, the entity extraction technique may use named entity recognition (NER) (for example, training a model using a dictionary of entities). For example, different entities can be extracted using delimiters.

[0037] Method 100 includes obtaining (e.g., retrieving and / or receiving) a second dataset (e.g., via an interface and / or processor circuit) S105 based on at least one of one or more first elements. The second dataset includes one or more second elements that indicate a second address. One or more second elements may indicate one or more second addresses. One or more example methods obtain a second dataset by selecting at least one (e.g., less than all) of one or more first elements. One or more example methods obtain a second dataset based on at least one of one or more first elements, including searching for one or more second systems (e.g., via a processor circuit) S105A using one or more search criteria. One or more example methods find one or more search criteria based on one or more first elements. For example, searching for one or more second systems generates a second dataset. For example, the second dataset may be thought of as a set of search results obtained using a search based on selected entities. In other words, the addresses of one or more second systems are removed to obtain fewer second elements. One or more example methods include obtaining a second dataset based on at least one of one or more first elements S105B, which includes obtaining a second dataset from one or more second systems based on at least one of one or more first elements (e.g., via an interface and / or processor circuitry). One or more example methods include one or more second systems including a customer management system and / or a customer database.

[0038] Method 100 includes applying natural language processing (NLP) techniques (e.g., via a processor circuit) to a first dataset and a second dataset to obtain one or more similarity scores S106. NLP techniques can be thought of as computational techniques that can process and analyze natural language data in discourse (e.g., speech or text). In this disclosure, NLP techniques are applied to text. In one or more example methods, applying NLP techniques to a first dataset and a second dataset to obtain one or more similarity scores S106 includes comparing one or more first elements with one or more second elements of the second dataset (e.g., via a processor circuit) using approximate sentence matching techniques S106A. In other words, for example, one or more ordinary cosine similarity scores are obtained for one or more second elements of the second dataset (e.g., an address list obtained from a search).

[0039] Method 100 includes obtaining a set of weights associated with the corresponding first element of a first address (e.g., via an interface and / or processor circuit) S108. In one or more example methods, obtaining the set of weights S108 includes applying a logistic regression model to one or more similarity scores based on a first and second dataset S108A, and optionally obtaining a set of weights using historical similarity scores. Historical similarity scores are considered to be previous similarity scores. LR may be used to model the probability of events having a particular class (e.g., pass / fail, match / no match). LR may be used to model several classes of events (e.g., determining whether an address contains the name of a particular broker). Each element found in an address may be assigned a probability between 0 and 1 (totaling 1).

[0040] One or more method examples include applying a logistic regression model (e.g., in a processor circuit) to one or more similarity scores to obtain a set of weights (S108A), and determining the set of weights (e.g., via a processor circuit) based on a comparison between at least one of the first elements and the corresponding second element (S108AA).

[0041] One or more similarity scores are obtained by assigning equal weights to all second elements in the second dataset (S106). In this disclosure, the set of weights may be thought of as custom weights for the second elements of the second dataset to obtain refined similarity scores. For example, the set of weights may be generated using logistic regression (LR) rather than execution time (e.g., not in alternate address matching). LR may be performed occasionally when training the set of weights (e.g., for one-time activities). The weights may be derived based on an LR model, in that this contributes to the weights if a match exists.

[0042] For example, the input to LR includes a dependent variable and an independent variable, and LR outputs a set of weights for a second element of a second dataset (e.g., all entities at the first and / or second addresses). The independent variable indicates the presence of different elements (e.g., entities) at the addresses (e.g., by comparing the input with the final matching output). The independent variable may include a similarity score. The dependent variable indicates the status of the final matching (e.g., matched or not matched) comparing the first and second addresses, e.g., the final matching output. The final matching output may be evaluated by manual inspection.

[0043] In some embodiments, LR and / or NLP may use artificial intelligence and / or be trained using supervised or unsupervised machine learning, and the machine learning program may use a neural network (which may be a convolutional neural network, a deep learning neural network, or a combined learning module or program that learns in two or more fields or domains of interest). Machine learning may include identifying and recognizing patterns in existing data to facilitate predictions for subsequent data. The model may be formed based on input examples to make effective and reliable predictions for new inputs.

[0044] In addition to or instead of the above, a machine learning program may be trained by inputting sample datasets or specific data into the program (e.g., image data, text data, report data, and / or numerical analysis). The machine learning program may use deep learning algorithms that may focus primarily on pattern recognition, and may be trained after processing multiple examples. The machine learning program may include object recognition, optical character recognition, and / or natural language processing, individually or in combination. The machine learning program may also include natural language processing, semantic analysis, and / or automated reasoning.

[0045] In supervised machine learning, a processing element may be given input examples and their corresponding outputs, and the processing element may attempt to find a general rule for mapping inputs to outputs, so that when new inputs are given thereafter, the processing element can accurately predict the correct output based on the rule it has found. In unsupervised machine learning, the processing element may need to find its own structure in unmarked input examples. In some embodiments, machine learning techniques may be used to extract data about computer devices, users of computer devices, computer networks hosting computer devices, services running on computer devices, and / or other data.

[0046] Based on such analysis, processing elements may learn how to identify features and patterns, which can then be applied to model training, address and text string analysis, and matching detection.

[0047] Method 100 includes determining a matching parameter (e.g., via a processor circuit) that indicates a matching between at least one of one or more first elements and at least one of one or more second elements, based on a set of weights. For example, the matching parameter may be determined by using a set of weights in an NLP technique. A set of weights that are associated with the corresponding first elements of a first address is obtained and applied in a custom machine learning and NLP technique to determine the matching parameter. The matching parameter may be thought of as a parameter that indicates the degree of matching between a first element(s) and the corresponding second element(s). The matching parameter may indicate whether there is a match or not. For example, the matching parameter may indicate 1 or 0, where 1 is 100% match and 0 is no match at all. For example, the matching parameter may indicate a number between 0 and 1 that needs to be compared to a criterion before determining whether there is a match or not.

[0048] In exemplary embodiments applying the disclosed technology, logistic regression is performed using a dependent variable (e.g., what to predict) and an independent variable (e.g., a feature used for prediction). Based on the feature, address matching may be predicted as successful or unsuccessful. In the example, the variables used may include one or more of the following: Independent variable = ['namepresent', 'addresspresent', 'emailpresent', 'phone_num_present', 'Cosine_Score', 'Azure_Search_Score', 'Count_of_Matching_numbers'], dependent variable = ['Success'].

[0049] In the example, the following equation may also be performed.

[0050] p=1 / [1+exp(-a-B0X0-B1X1-B2X2-B3X3-B4X4-B5X5-B6X6)]

[0051] p is the probability that the matching parameter can be represented (p>=0.5 is a pass [success=1], p<0.5 is a failure [success=0]), a is a constant value, X0 is the value of the variable namepresent [0 / 1] indicating the existence of a name, X1 is the value of the variable addresspresent [0 / 1] indicating the existence of an address, X2 is the value of the variable emailpresent [0 / 1] indicating the existence of an email address, X3 is the value of the variable ph_num_present [0 / 1] indicating the existence of a phone number, X4 is the value of the cosine similarity score Cosine_Score [varies from 0 to 1], X5 is the value of Azure_Search_Score [numerical value], and X6 is the value of Count_of_Matching_numbers [numerical value].

[0052] The coefficients can be thought of as a set of weights obtained by applying LR, and may be as follows: B0 = dependency coefficient for X0, B1 = dependency coefficient for X1, B2 = dependency coefficient for X2, B3 = dependency coefficient for X3, B4 = dependency coefficient for X4, B5 = dependency coefficient for X5, and B6 = dependency coefficient for X6. For example, the following coefficients are calculated from logistic regression: a = -6.19, B0 = 2.3697863, B1 = 0.10784766, B2 = 1.49094974, B3 = 0.65098831, B4 = 5.99906035, B5 = 0.6030922, B6 = 0.97430007.

[0053] In one or more example methods, determining a matching parameter using a set of weights that indicates a matching between at least one of one or more first elements and at least one of one or more second elements (S110) includes calculating an updated similarity score for each second element of a second address (e.g., via a processor circuit) (S110A). The updated similarity score may be considered a refined similarity score. In one or more example methods, determining the matching parameter using a set of weights (S110) includes determining whether the updated similarity score satisfies a criterion (S110B). The criterion may be based on a threshold. The criterion is satisfied when the matching parameter is greater than or equal to the threshold, and not satisfied when the matching parameter is less than the threshold.

[0054] In one or more example methods, determining a matching parameter using a set of weights S110 includes determining (e.g., via a processor circuit) S110D that the matching parameter is a successful match between the second element of the second address and the corresponding first element of the first address when the updated similarity score is determined to meet the criteria. In one or more example methods, determining a matching parameter using a set of weights S110 includes determining that the matching parameter is a failed match between the first address and the second address when the updated similarity score is determined to not meet the criteria. For example, while calculating an updated or refined similarity score(s) using a set of weights (e.g., custom weights), small matching parameters less than 0.5 are said to be no match, and matching parameters greater than or equal to 0.5 are considered to be a match.

[0055] Method 100 includes outputting matching parameters (e.g., via an interface and / or processor circuitry) S111. The output matching parameters may represent a similarity score updated or refined using a set of weights. For example, the matching parameters may be output from one part of the processor circuitry to another part of the processor circuitry (e.g., from a matching parameter generator circuitry to a correction circuitry of an electronic device) so that the first address can be corrected with the matching parameters. The disclosed method can achieve results with 99% accuracy and a wider range of applicability (e.g., how many addresses are covered).

[0056] In one or more method examples, the NLP technique includes the cosine similarity technique. In one or more method examples, one or more similarity scores are cosine similarity scores. Cosine similarity can be thought of as a measure of similarity between two non-zero vectors (the first and second elements in this disclosure) in an inner product space. Cosine similarity can also be defined as equal to the cosine of the angle between the two vectors, which is also equal to the inner product of the same vectors normalized to both have length 1. For example, each element may be conceptually assigned a different dimension, and addresses may be characterized by vectors, with values ​​in each dimension corresponding to the number of times the element appears in the address. Cosine similarity then provides a useful measure of how similar two addresses may be with respect to their elements.

[0057] One or more method examples include applying NLP techniques (e.g., via a processor circuit) to a first dataset and a second dataset to obtain one or more similarity scores S106, and generating one or more cosine similarity scores (e.g., via a processor circuit) for each second element of the second dataset S106B.

[0058] In one or more example methods, method 100 includes correcting the first address (for example, via a processor circuit) by matching parameters S112. In one or more example methods, the second address is the corrected address.

[0059] Figure 3 shows a block diagram of a typical electronic device 300 according to this disclosure. The electronic device 300 includes a memory circuit 301, a processor circuit 302, and an interface 303. The electronic device 300 is configured to perform one of the methods disclosed in Figure 2. In other words, the electronic device 300 is configured to perform address matching or text string matching. For example, the electronic device disclosed herein (e.g., electronic device 300) may be a shipping instruction processing device. For example, the electronic device disclosed herein (e.g., electronic device 300) may be an address correction device. For example, the electronic device disclosed herein (e.g., electronic device 300) may be an invoice processing device.

[0060] The electronic device 300 is configured to acquire a first data set indicating a first address (for example, via the processor circuit 302).

[0061] The electronic device 300 is configured to determine one or more first elements that represent a first address (for example, via the processor circuit 302) using an entity extraction technique based on a first dataset.

[0062] The electronic device 300 is configured to acquire a second dataset (for example, via the processor circuit 302 and / or interface 303) based on at least one of one or more first elements. The second dataset includes one or more second elements that indicate a second address.

[0063] The electronic device 300 is configured to apply natural language processing (NLP) techniques (for example, via the processor circuit 302) to the first dataset and the second dataset to obtain one or more similarity scores.

[0064] The electronic device 300 is configured to obtain (for example, via the processor circuit 302 and / or interface 303 and / or from the memory circuit 301) a set of weights associated with the first element corresponding to the first address.

[0065] The electronic device 300 is configured to determine a matching parameter (for example, via the processor circuit 302) that indicates a matching between at least one of one or more first elements and at least one of one or more second elements, based on a set of weights.

[0066] Optionally, the electronic device 300 is configured to output matching parameters.

[0067] The processor circuit 302 is optionally configured to perform any of the operations disclosed in Figure 2 (for example, one or more of S102A, S104A, S105A, S105B, S106A, S106B, S108A, S108AA, S110A, S110B, S110C, S110D, S112). The operation of the electronic device 300 may be embodied in the form of executable logic routines (for example, lines of code, software programs, etc.) stored on a non-temporary computer-readable medium (for example, a memory circuit 301) and executed by the processor circuit 302.

[0068] Furthermore, the operation of the electronic device 300 can be considered as a method configured for the electronic device 300 to perform. While the described functions and operations may be implemented in software, such functions may also be performed via dedicated hardware or firmware, or any combination of hardware, firmware, and / or software.

[0069] The memory circuit 301 may be one or more of a buffer, flash memory, hard drive, removable media, volatile memory, non-volatile memory, random access memory (RAM), or other suitable devices. In a typical configuration, the memory circuit 301 may include non-volatile memory for long-term data storage and volatile memory that functions as system memory for the processor circuit 302. The memory circuit 301 may exchange data with the processor circuit 302 via a data bus. Control lines and an address bus between the memory circuit 301 and the processor circuit 302 may also be present (not shown in Figure 3). The memory circuit 301 is considered a non-temporary computer-readable medium.

[0070] The memory circuit 301 may be configured to store a set of weights in a portion of the memory.

[0071] In some embodiments, the electronic device 300 may function as a user device.

[0072] In some embodiments, the electronic device 300 may function as a server device.

[0073] Embodiments of the methods and products (electronic devices) described herein are described in the following sections.

[0074] 1. A method for performing address matching using an electronic device, wherein the method is: Obtaining a first dataset that indicates a first address (S102), Based on the first dataset, determine one or more first elements that represent the first address using an entity extraction technique (S104), Obtaining a second dataset based on at least one of the one or more first elements (S105), wherein the second dataset includes one or more second elements indicating a second address, Applying natural language processing (NLP) techniques to the first and second datasets (S106) to obtain one or more similarity scores, Obtaining a set of weights associated with the corresponding first element of the first address (S108), Based on the set of weights, determine a matching parameter that indicates a matching between at least one of the one or more first elements and at least one of the one or more second elements (S110), The method comprising outputting the matching parameters (S111).

[0075] 2. The entity extraction technique is the method described in Clause 1, which includes one or more of the parsing techniques and element extraction techniques.

[0076] 3. Obtaining the second dataset based on at least one of the one or more first elements (S105) includes searching one or more second systems using one or more search criteria (S105A), wherein the one or more search criteria are based on the one or more first elements, as described in any of the preceding clauses.

[0077] 4. The method according to any of the preceding clauses, wherein obtaining the set of weights (S108) includes applying a logistic regression model to the one or more similarity scores based on the first dataset and the second dataset (S108A).

[0078] 5. The method according to any of the preceding clauses, wherein obtaining the one or more similarity scores by applying the NLP technique to the first dataset and the second dataset (S106) includes comparing the one or more first elements with one or more second elements of the second dataset using an approximate sentence matching technique (S106A).

[0079] 6. The NLP technique is the method described in any of the preceding clauses, including the cosine similarity technique.

[0080] 7. The method according to clause 6, wherein the one or more similarity scores are cosine similarity scores.

[0081] 8. The method according to any one of the claims 6 to 7, wherein obtaining one or more similarity scores by applying the NLP technique to the first dataset and the second dataset (S106) includes generating one or more cosine similarity scores for each second element of the second dataset (S106B).

[0082] 9. Applying the logistic regression model to one or more similarity scores (S108A) to obtain the set of weights is: A method of any of the provisions as a dependent provision to provision 4, comprising determining the set of weights based on a comparison of at least one of the first elements with a corresponding second element (S108AA).

[0083] 10. Using the set of weights, determine the matching parameter that indicates the matching between at least one of the one or more first elements and at least one of the one or more second elements (S110), Calculating updated similarity scores for each second element of the second address (S110A), Determining whether the updated similarity score meets the criteria (S110B), The method of any of the preceding clauses, comprising determining that when the updated similarity score is determined to satisfy the criterion, the matching parameter is a successful match between the second element of the second address and the corresponding first element of the first address (S110D).

[0084] 11. The method described above is the method described in any of the preceding clauses, which includes (a) correcting the first address with the matching parameter.

[0085] 12. The method according to any of the preceding clauses, wherein obtaining the first dataset (S102) includes obtaining the first dataset from one or more first systems (S102A), where the one or more first systems are one or more operating systems, and the one or more operating systems include one or more of the following: a shipping system, an invoice issuing system, and a case management system.

[0086] 13. The method according to any of the preceding clauses, wherein obtaining a second dataset based on at least one of the one or more first elements (S105) includes obtaining a second dataset from the one or more second systems based on at least one of the one or more first elements (S105B), the one or more second systems including a customer management system and / or a customer database.

[0087] 14. The method described in any of the preceding clauses, wherein the second address is a corrected address.

[0088] 15. An electronic device including a memory circuit, a processor circuit, and a wireless interface, wherein the electronic device is configured to perform any of the methods described in any of the clauses 1 to 14.

[0089] 16. A computer-readable storage medium storing one or more programs, wherein the one or more programs, when executed by an electronic device, include instructions causing the electronic device to perform any of the methods described in Clauses 1 to 14.

[0090] The use of terms such as "first," "second," "third," and "fourth," "primary," "secondary," and "tertiary" does not imply any order of identification, but is included to identify individual elements. Furthermore, the use of terms such as "first," "second," "third," and "fourth," "primary," "secondary," and "tertiary" does not indicate any order or importance; rather, these terms are used to distinguish one element from another. Note that the terms "first," "second," "third," and "fourth," "primary," "secondary," and "tertiary" are used here and elsewhere simply for marking purposes and are not intended to indicate any specific spatial or temporal order. Moreover, marking the first element does not imply the existence of a second element, and vice versa.

[0091] Figures 1A-13 can be understood to include circuits or operations illustrated by solid lines and circuits or operations illustrated by dashed lines. The circuits or operations included by solid lines are those included in the broadest embodiments. The circuits or operations included by dashed lines are embodiments that may be included in, or are part of, further circuits or operations that may be incorporated in addition to, the circuits or operations of the solid-line embodiments. Naturally, these operations do not have to be performed in the order shown. Furthermore, naturally, not all operations have to be performed. Typical operations may be performed in any order and in any combination.

[0092] Please note that the term "includes" does not necessarily exclude the existence of other elements or steps not listed.

[0093] Note that the terms "a" or "an" preceding an element do not exclude the existence of multiple such elements.

[0094] Furthermore, note that no reference numerals limit the scope of the claims, typical embodiments can be implemented at least partially by both hardware and software, and some “means,” “units,” or “apparatus” can be represented by the same article of hardware.

[0095] The various typical methods, apparatus, nodes, and systems described herein are described in the general context of a method step or process (in one embodiment, which may be implemented by a computer program product embodied in a computer-readable medium (including computer-executable instructions such as program code executed by a computer in a network environment)). Computer-readable media may include removable and non-removable storage devices, such as, but not limited to, read-only memory (ROM), random-access memory (RAM), compact discs (CDs), digital versatile discs (DVDs), and the like. Generally, program circuits may include routines, programs, objects, components, data structures, etc., that perform a specified task or implement an abstract data type for identification. Computer-executable instructions, associated data structures, and program circuits represent examples of program code for performing the steps of the methods disclosed herein. The order of identification of such executable instructions or associated data structures represents examples of corresponding actions for performing the functions described in such steps or processes.

[0096] While the features have been illustrated and described, it will be obvious to those skilled in the art that they are not intended to limit the requested disclosures, and that various modifications and changes may be made without deviating from the scope of the requested disclosures. Therefore, the specifications and drawings should be considered illustrative, not restrictive. The requested disclosures are intended to extend to all substitutes, modifications, and equivalents. [Configuration 1] A method for performing address matching, which is performed by an electronic device, wherein the method is Obtaining a first dataset that indicates a first address (S102), Based on the first dataset, determine one or more first elements that represent the first address using an entity extraction technique (S104), Obtaining a second dataset based on at least one of the one or more first elements (S105), wherein the second dataset includes one or more second elements indicating a second address, Applying natural language processing (NLP) techniques to the first and second datasets (S106) to obtain one or more similarity scores, Obtaining a set of weights associated with the corresponding first element of the first address (S108), Based on the set of weights, determine a matching parameter that indicates a matching between at least one of the one or more first elements and at least one of the one or more second elements (S110), Outputting the matching parameters (S111), Correcting the first address using the matching parameter (S112), Methods that include... [Configuration 2] The entity extraction technique is the method according to configuration 1, wherein the entity extraction technique includes one or more of the syntactic analysis technique and the element extraction technique. [Configuration 3] Obtaining the second dataset based on at least one of the one or more first elements (S105) includes searching one or more second systems using one or more search criteria (S105A), The method according to configuration 1 or 2, wherein the one or more search criteria are based on the one or more first elements. [Structure 4] The method according to any one of configurations 1 to 3, wherein obtaining the set of weights (S108) includes applying a logistic regression model to one or more similarity scores based on the first dataset and the second dataset (S108A). [Composition 5] The method according to any one of configurations 1 to 4, wherein obtaining the one or more similarity scores by applying the NLP technique to the first dataset and the second dataset (S106) includes comparing the one or more first elements with one or more second elements of the second dataset using an approximate sentence matching technique (S106A). [Composition 6] The aforementioned NLP technique is the method according to any one of configurations 1 to 5, including the cosine similarity technique. [Composition 7] The method according to configuration 6, wherein the one or more similarity scores are cosine similarity scores. [Structure 8] The method according to configuration 6 or 7, wherein applying the NLP technique to the first dataset and the second dataset (S106) to obtain one or more similarity scores includes generating one or more cosine similarity scores for each second element of the second dataset (S106B). [Composition 9] Applying the logistic regression model to the one or more similarity scores (S108A) to obtain the set of weights is: A method of any of the above configurations as a dependent term to configuration 4, comprising determining the set of weights based on a comparison of at least one of the first elements with a corresponding second element (S108AA). [Configuration 10] Determining the matching parameter that indicates a matching between at least one of the one or more first elements and at least one of the one or more second elements using the set of weights (S110) is, Calculating updated similarity scores for each second element of the second address (S110A), Determining whether the updated similarity score meets the criteria (S110B), The method according to any one of configurations 1 to 9, comprising determining (S110D) that when the updated similarity score is determined to satisfy the criteria, the matching parameter is a successful match between the second element of the second address and the corresponding first element of the first address. [Composition 11] Obtaining the first dataset (S102) includes obtaining the first dataset from one or more first systems (S102A), The one or more first systems are one or more operating systems, The aforementioned one or more operating systems include one or more of the following: a shipping system, an invoice issuance system, and a case management system. The method described in any of configurations 1 to 10. [Composition 12] Obtaining a second dataset based on at least one of the one or more first elements (S105) includes obtaining a second dataset from one or more second systems based on at least one of the one or more first elements (S105B), The method according to any one of configurations 1 to 11, wherein the one or more second systems include a customer management system and / or a customer database. [Composition 13] The method described above, in which the second address is a corrected address, is one of the configurations 1 to 12. [Composition 14] An electronic device comprising a memory circuit, a processor circuit, and a wireless interface, wherein the electronic device is configured to perform any of the methods described in any of configurations 1 to 13. [Composition 15] A computer-readable storage medium for storing one or more programs, wherein the one or more programs, when executed by an electronic device, include instructions that cause the electronic device to perform any of the methods described in configurations 1 to 13.

Claims

1. A method for performing address matching, which is performed by an electronic device, wherein the method is Obtaining a first dataset that indicates a first address (S102), Based on the first dataset, one or more first elements representing the first address are determined using an entity extraction technique (S104), Obtaining a second dataset based on at least one of the one or more first elements (S105), wherein the second dataset includes one or more second elements indicating a second address, Applying natural language processing (NLP) techniques to the first dataset and the second dataset (S106) to obtain one or more similarity scores, Obtaining a set of weights associated with the corresponding first element of the first address (S108), Based on the set of weights, a matching parameter is determined that indicates a matching between at least one of the one or more first elements and at least one of the one or more second elements (S110), Outputting the matching parameters (S111), Correcting the first address using the matching parameter (S112), Methods that include...

2. The method according to claim 1, wherein the entity extraction technique includes one or more of the syntactic analysis technique and the element extraction technique.

3. Obtaining the second dataset based on at least one of the one or more first elements (S105) includes searching one or more second systems using one or more search criteria (S105A), The method according to claim 1 or 2, wherein the one or more search criteria are based on the one or more first elements.

4. The method according to any one of claims 1 to 3, wherein obtaining the set of weights (S108) includes applying a logistic regression model to one or more similarity scores based on the first dataset and the second dataset (S108A).

5. The method according to any one of claims 1 to 4, wherein obtaining one or more similarity scores by applying the NLP technique to the first dataset and the second dataset (S106) includes comparing one or more first elements with one or more second elements of the second dataset using an approximate sentence matching technique (S106A).

6. The method according to any one of claims 1 to 5, wherein the NLP technique includes a cosine similarity technique.

7. The method according to claim 6, wherein the one or more similarity scores are cosine similarity scores.

8. The method according to claim 6 or 7, wherein applying the NLP technique to the first dataset and the second dataset (S106) to obtain one or more similarity scores includes generating one or more cosine similarity scores for each second element of the second dataset (S106B).

9. Applying the logistic regression model to one or more similarity scores (S108A) to obtain the set of weights is: The method of any one of claims 5 to 8 as a dependent of claim 4, comprising determining the set of weights based on a comparison of at least one of the first elements with a corresponding second element (S108AA).

10. Using the set of weights, determining the matching parameter that indicates the matching between at least one of the one or more first elements and at least one of the one or more second elements (S110) is, Calculating updated similarity scores for each second element of the second address (S110A), Determining whether the updated similarity score meets the criteria (S110B), The method according to any one of claims 1 to 9, comprising determining (S110D) that when the updated similarity score is determined to satisfy the criterion, the matching parameter is a successful match between the second element of the second address and the corresponding first element of the first address.

11. Acquiring the first dataset (S102) includes acquiring the first dataset from one or more first systems (S102A), The one or more first systems are one or more operating systems, The aforementioned one or more operating systems include one or more of the following: a shipping system, an invoice issuance system, and a case management system. The method according to any one of claims 1 to 10.

12. Obtaining a second dataset based on at least one of the one or more first elements (S105) includes obtaining a second dataset from one or more second systems based on at least one of the one or more first elements (S105B), The method according to any one of claims 1 to 11, wherein the one or more second systems include a customer management system and / or a customer database.

13. The method according to any one of claims 1 to 12, wherein the second address is a correction address.

14. An electronic device comprising a memory circuit, a processor circuit, and a wireless interface, wherein the electronic device is configured to perform any of the methods described in any of claims 1 to 13.

15. A computer-readable storage medium for storing one or more programs, wherein the one or more programs, when executed by an electronic device, include instructions that cause the electronic device to execute any of the methods according to claims 1 to 13.

Citation Information

Patent Citations

  • Character string correcting method for address and zip code

    JP2000090192A

  • Knowledge management and retrieval system

    JP2001134590A

  • Map data error correction device

    JP2010231560A

  • Navigation device and navigation system

    JP2010276373A

  • Mobile phone apparatus, confirmation information displaying program, and confirmation information displaying method

    JP2011135284A