An address matching method and matching system based on style conversion

Through the address matching method based on style conversion, the problem of lack of flexibility and adaptability of address matching methods in the prior art is solved, and higher matching accuracy and adaptability are achieved.

CN119357701BActive Publication Date: 2025-05-27WUDA GEOINFORMATICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411907354.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-27
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

The existing address matching methods are based on discriminant methods, lack flexibility and adaptability, and it is difficult to provide accurate matching in complex and changing urban environments and constantly updated address information.

Method used

The address matching method based on style conversion is adopted, and the source address and style address are divided into regional parts and specific address parts through the address segmentation model. The style conversion model is used to convert the specific address part of the source address into a style consistent with the target address, and then the converted address is matched with the target address.

Benefits of technology

It improves the accuracy and flexibility of address matching, can better adapt to address formats and changes in different regions, and enhances the matching ability in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357701B_ABST
    Figure CN119357701B_ABST
Patent Text Reader

Abstract

The present invention provides an address matching method and a matching system based on style conversion. The method includes: separately splitting the source address and the style address to obtain the regional part and the specific address part; inputting the specific address part of the source address and the specific address part of the style address into a style conversion model to obtain the specific address part of the candidate address; splicing the specific address part of the candidate address and the regional part of the style address to obtain the candidate address; comparing the candidate address with the target address, if they are exactly the same, the matching is successful, otherwise, the matching fails. The present invention converts the description style of the source address to be consistent with the description style of the target address, and then matches the source address with the converted style with the target address. Compared with the existing discriminant matching, the accuracy of matching is improved; and the style of the source address is converted through the style conversion model, and there is almost no limitation on the description style of the source address, which improves the flexibility and adaptability of address matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of address matching, and more particularly, to an address matching method and system based on style conversion. Background Art

[0002] Address matching refers to determining whether a pair of input addresses point to the same place. A common application is to input a non-standard source address and compare it with a standard address in an address database to determine if they match. This plays a crucial role in urban governance and service systems. It ensures the accuracy and efficiency of information. For urban managers, being able to accurately identify and locate each specific address is the basis for effective management and services.

[0003] However, despite the increasing importance of address matching, existing methods have some limitations. Currently, most address matching methods are discriminative, which means they rely on pre-set rules or patterns to identify and match addresses. Although this method can provide quick matching results in some cases, it may lack flexibility and adaptability, especially when faced with complex and changing urban environments and continuously updated address information. Summary of the Invention

[0004] The present invention addresses the technical problems existing in the prior art and provides an address matching method and system based on style conversion.

[0005] According to a first aspect of the present invention, there is provided an address matching method based on style conversion, including:

[0006] Respectively segmenting the source address and the style address based on an address segmentation model to obtain the regional part and the specific address part of the source address, and the regional part and the specific address part of the style address, wherein the description style of the specific address part of the style address is consistent with the description style of the specific address part of the target address;

[0007] Inputting the specific address part of the source address and the specific address part of the style address into a style conversion model to perform style conversion on the specific address part of the source address to obtain the specific address part of the candidate address;

[0008] Concatenating the specific address part of the candidate address and the regional part of the style address to obtain a candidate address;

[0009] Comparing the candidate address with the target address. If they are exactly the same, the candidate address matches the target address successfully; otherwise, the match fails.

[0010] Based on the above technical solution, the present invention can also be improved as follows.

[0011] Optionally, the training process of the style conversion model includes:

[0012] Obtain multiple address data, where each piece of address data includes the specific address part of the source address, the specific address part of the style address, and the specific address part of the candidate address;

[0013] Extract source address information, style address information, and candidate address information from the specific address part of the source address, the specific address part of the style address, and the specific address part of the candidate address respectively. The source address information includes the key identifier and description style of the source address, the style address information includes the key identifier and description style of the style address, the candidate address information includes the key identifier and description style of the candidate address, and the description style of the style address is consistent with the description style of the candidate address;

[0014] Construct a training sample data set, where the training sample data set contains multiple training samples, and each training sample includes 3 parts, namely [the specific address part of the source address, source address information], [the specific address part of the style address, style address information], [the specific address part of the candidate address, candidate address information];

[0015] Train the style conversion model based on the training sample data set to obtain a trained style conversion model.

[0016] Optionally, the style conversion model is a large language model. The training of the style conversion model based on the training sample data set to obtain a trained style conversion model includes:

[0017] Construct a prompt for each training sample based on the specific address part of the source address and the specific address part of the style address, input the prompt of each training sample into the style conversion model, obtain the output content of the style conversion model, and extract the specific address part of the candidate address from the output content;

[0018] Calculate the loss function value of the style conversion model according to the specific address part of the candidate address output by the style conversion model and the specific address part of the candidate address in the training sample data set, and specifically optimize the style conversion model with the loss function value to obtain a trained style conversion model.

[0019] Optionally, the prompt for each training sample is constructed according to the following template:

[0020] Please generate the specific address part of the candidate address according to the specific address part of the source address and the specific address part of the style address;

[0021] The specific address part of the source address: {The specific address part of the source address};

[0022] The specific address part of the style address: {The specific address part of the style address};

[0023] Please generate step by step in the following way:

[0024] First step, analyze the key identifiers and description styles of the specific address part of the source address;

[0025] Second step, analyze the key identifiers and description styles of the specific address part of the style address;

[0026] Third step, generate the specific address part of the candidate address according to the source address information and the description style of the style address;

[0027] The style conversion model constructs the output content according to the following template:

[0028] First step, the specific address part of the source address: {The specific address part of the source address};

[0029] The key identifier is: {The key identifier of the source address};

[0030] The description style is: {The description style of the source address};

[0031] Second step, the style address: {The style address}

[0032] The key identifier is: The specific address part of the style address: {The specific address part of the style address};

[0033] The description style is: {The description style of the style address}

[0034] Third step, infer the candidate address information according to the source address information and the description style of the style address:

[0035] The key identifier is: {The key identifier of the candidate address};

[0036] The description style is: {The description style of the candidate address};

[0037] Generate the specific address part of the candidate address according to the candidate address information: {The specific address part of the candidate address}.

[0038] Optionally, use the target address as the style address.

[0039] According to the second aspect of the present invention, there is provided an address matching system based on style conversion, including:

[0040] A splitting module, configured to split a source address and a style address respectively based on an address splitting model, so as to obtain a regional part and a specific address part of the source address, and a regional part and a specific address part of the style address, wherein the description style of the specific address part of the style address is consistent with the description style of the specific address part of the target address;

[0041] A style conversion module, configured to input the specific address part of the source address and the specific address part of the style address into a style conversion model, and perform style conversion on the specific address part of the source address to obtain the specific address part of a candidate address;

[0042] An assembling module, configured to assemble the specific address part of the candidate address and the regional part of the style address to obtain a candidate address;

[0043] A matching module, configured to compare the candidate address with the target address. If they are exactly the same, the candidate address matches the target address successfully; otherwise, the matching fails.

[0044] An address matching method and a matching system based on style conversion provided by the present invention split a source address and a style address respectively to obtain a regional part and a specific address part; input the specific address part of the source address and the specific address part of the style address into a style conversion model to obtain the specific address part of a candidate address; assemble the specific address part of the candidate address and the regional part of the style address to obtain a candidate address; compare the candidate address with the target address. If they are exactly the same, the matching is successful; otherwise, the matching fails. The present invention converts the description style of the source address to be consistent with the description style of the target address, and then matches the source address with the converted style with the target address. Compared with the existing discriminant matching, the matching accuracy is improved; and by converting the style of the source address through a style conversion model, there is almost no limitation on the description style of the source address, and the flexibility and adaptability of address matching are improved. Description of the Drawings

[0045] Figure 1 It is a flowchart of an address matching method based on style conversion provided by the present invention;

[0046] Figure 2 It is an overall flowchart of the address matching method based on style conversion provided by the present invention;

[0047] Figure 3 It is a schematic diagram of a case of the address matching method based on style conversion;

[0048] Figure 4 It is a training flowchart of a style conversion model;

[0049] Figure 5 It is a schematic diagram of a case of training a style conversion model;

[0050] Figure 6 It is a structural block diagram of an address matching system based on style conversion provided by the present invention. Detailed implementation manners

[0051] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or a single embodiment provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. Such combination is not restricted by the order of steps and / or the mode of structural composition, but must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0052] In urban service management, it is usually necessary to input the address of a computer household, and then match the input address with the addresses in the address library to find the accurate address. However, since the description styles of the input address and the addresses in the address library are inconsistent, this increases the difficulty of address matching.

[0053] Almost all existing methods are discriminative address matching methods. This method directly compares the features of the source address and the target address to determine whether the address pair matches. However, all discriminative-based address matching methods are limited by the incompleteness of the training set, and it is difficult for the model or algorithm to learn all discriminative features. For example, in the addresses of City B, there are often addresses such as Room B1, Unit 4, Building 123, while in the addresses of City D, there are often addresses such as No. 4-5, Lane 123. For the address matching method obtained by training in City B, since information such as "lane" and "number" has never appeared in the training samples, it is difficult to capture or identify the features of the addresses in City D. Therefore, it is difficult for such address matching methods to achieve the goal of being applicable in Area D after being trained in Area B.

[0054] Therefore, the present invention provides an address matching method based on style conversion, as Figure 1 shown, the steps of the address matching method based on style conversion include:

[0055] Step 1, respectively segment the source address and the style address based on the address segmentation model to obtain the regional part and the specific address part of the source address and the regional part and the specific address part of the style address, wherein the description style of the specific address part of the style address is consistent with the description style of the specific address part of the target address.

[0056] Among them, Figure 2 shows the overall flowchart of the address matching method based on style conversion, and Figure 3 shows a schematic diagram of a case of the address matching method based on style conversion. First, the source address and the target address are input.

[0057] For example:

[0058] Source address: Room 1-1201, Community C, District B, City A;

[0059] Target address: Room 01, 12th Floor, Building 1, Community C, No. 1, Road E, Street F, District B, City A, Province D.

[0060] Duplicate the target address once as the style address, and use the address segmentation model to separately segment the source address and the style address into a regional part and a specific address part, obtaining the regional part and the specific address part of the source address and the regional part and the specific address part of the style address.

[0061] The present invention uses the open-source address model Mgeo as the address segmentation model. The Where / What model in the Mgeo model can segment an address into two parts, namely the regional part and the specific address part.

[0062] For example:

[0063] Source address: Room 1-1201, Community C, District B, City A;

[0064] The regional part after segmentation is: District B, City A, and the specific address part: Room 1-1201, Community C;

[0065] Style address: Room 01, 12th Floor, Building 1, Community C, No. 1, Road E, Street F, District B, City A, Province D;

[0066] The regional part after segmentation is: No. 1, Road E, Street F, District B, City A, Province D;

[0067] The specific address part: Room 01, 12th Floor, Building 1, Community C.

[0068] Step 2: Input the specific address part of the source address and the specific address part of the style address into the style conversion model to perform style conversion on the specific address part of the source address, obtaining the specific address part of the candidate address.

[0069] It can be understood that in Step 1, both the source address and the style address are segmented, obtaining the regional part and the specific address part of the source address and the regional part and the specific address part of the style address. Input the specific address part of the source address and the specific address part of the style address into the trained style conversion model to perform style conversion on the specific address part of the source address, obtaining the specific address part of the candidate address.

[0070] Among them, the present invention uses a large language model as the style conversion model. The large language model can understand and use language. Such language models are pre-trained with a large amount of text and can generate text, summarize content, and perform translation, rewriting, classification, categorization, and analysis, etc. The large language model described in the present invention refers to an open-source large language model, such as ChatGLM-6b, Qwen-14b, etc. The present invention selects ChatGLM-6b as the style conversion model of the present invention.

[0071] Before using the style conversion model to convert the style of the source address, it is necessary to train the style conversion model.

[0072] Exemplarily, the training process of the style conversion model includes:

[0073] Step 21, obtain multiple address data, and each piece of address data includes the specific address part of the source address, the specific address part of the style address, and the specific address part of the candidate address.

[0074] It can be understood that a training sample data set is collected. The sample data set includes multiple address data, and each piece of address data includes the specific address part of the source address, the specific address part of the style address, and the specific address part of the candidate address. The specific address part of the source address, the specific address part of the style address, and the specific address part of the candidate address can be respectively segmented from the source address, the style address, and the candidate address through a segmentation model.

[0075] Step 22, respectively extract source address information, style address information, and candidate address information from the specific address part of the source address, the specific address part of the style address, and the specific address part of the candidate address. The source address information includes the key identifier and description style of the source address, the style address information includes the key identifier and description style of the style address, the candidate address information includes the key identifier and description style of the candidate address, and the description style of the style address is consistent with the description style of the candidate address.

[0076] Step 23, construct a training sample data set. The training sample data set contains multiple training samples, and each training sample includes 3 parts, namely [the specific address part of the source address, source address information], [the specific address part of the style address, style address information], [the specific address part of the candidate address, candidate address information].

[0077] Specifically, for the training process of the style conversion model, see Figure 4 and Figure 5 , including the following steps:

[0078] First, obtain the training sample dataset. The training sample dataset includes multiple training samples, and the processing method for each training sample is as follows.

[0079] Each piece of data in the training sample contains three parts, namely: [specific address part of the source address, source address information], [specific address part of the style address, style address information], [specific address part of the candidate address, candidate address information]. Among them, the specific address part of the source address is the specific address part segmented from the addresses collected by the business, and the source address information includes the key identifier and description style of the source address; the specific address part of the style address is the address consistent with the description style of the candidate address, and the style address information includes the key identifier and description style of the style address; the specific address part of the candidate address is the unified address corresponding to the source address in the address library, and the candidate address information includes the key identifier and description style of the candidate address.

[0080] Among them, there are only three types of key identifiers in the specific address part, namely: entity, number, and letter. For example, in "Building 1, Unit 1, Room 1B in ABC Community", "ABC Community" is the entity identifier, 1 in "Building 1" is the number identifier, and B in "1B" is the letter identifier.

[0081] In an address, replacing the specific value of the key identifier with the corresponding identifier is the description style of this address. For example, for "Building 1, Unit 1, Room 1B in ABC Community", after replacing with the corresponding identifier, it becomes: number + "Building" + number + "Unit" + number + letter, which is the description style of the address.

[0082] For example:

[0083] Training sample 1:

[0084] Specific address part of the source address: Building 1, Room 1201;

[0085] Source address information:

[0086] Key identifier: {1: number, 1201: number};

[0087] Description style: number + "Building" + number;

[0088] Specific address part of the style address: Building 2, Room 03, 5th floor;

[0089] Style address information:

[0090] Key identifier: {2: number, 5: number, 03: number};

[0091] Description style: number + "Building" + number + "floor" + number + "room";

[0092] Specific address part of the candidate address: Room 01, 12th Floor, Building 1;

[0093] Candidate address information:

[0094] Key identifiers: {1: number, 12: number, 01: number}

[0095] Description style: number + "Building" + number + "Floor" + number + "Room"

[0096] That is, the source address and the candidate address point to the same location, and the description style of the style address should be consistent with that of the candidate address.

[0097] Step 24: Train the style conversion model based on the training sample dataset to obtain the trained style conversion model.

[0098] It can be understood that after obtaining the training sample dataset, the style conversion model is trained based on the training sample dataset. In the present invention, a large language model is used as the style conversion model. In order to enable the large language model to accurately output the desired content, the present invention constructs the prompt words of the large language model and defines the format of the output content of the large language model based on the specific address part of the source address and the specific address part of the style address.

[0099] Exemplarily, the style conversion model is a large language model, and step 24, training the style conversion model based on the training sample dataset to obtain the trained style conversion model, includes:

[0100] Step 241: Construct the prompt word of each training sample based on the specific address part of the source address and the specific address part of the style address, input the prompt word of each training sample into the style conversion model, obtain the output content of the style conversion model, and extract the specific address part of the candidate address from the output content.

[0101] Specifically, according to the specific address part of the source address and the chicken leg address part of the style address, construct the training prompt word and the output content based on the chain of thought:

[0102] According to the training sample, construct the prompt word according to the following template:

[0103] Please generate the specific address part of the candidate address according to the specific address part of the source address and the specific address part of the style address:

[0104] Specific address part of the source address: {Specific address part of the source address};

[0105] Specific address part of the style address: {Specific address part of the style address};

[0106] Please generate step by step as follows:

[0107] First, analyze the key identifiers and description styles of the specific address part of the source address;

[0108] Second, analyze the key identifiers and description styles of the specific address part of the style address;

[0109] Third, generate the specific address part of the candidate address according to the source address information and the description style of the style address.

[0110] For example:

[0111] “

[0112] Please generate the specific address part of the candidate address according to the specific address part of the source address and the specific address part of the style address:

[0113] Specific address part of the source address: Room 1201, Building 1;

[0114] Specific address part of the style address: Room 03, 5th Floor, Building 2;

[0115] Please generate step by step as follows:

[0116] First, analyze the key identifiers and description styles of the source address;

[0117] Second, analyze the key identifiers and description styles of the style address;

[0118] Third, generate the candidate address according to the source address information and the description style of the style address”.

[0119] Correspondingly, the style conversion model constructs the output content according to the following template:

[0120] First, specific address part of the source address: {specific address part of the source address};

[0121] Key identifier: {key identifier of the source address};

[0122] Description style: {description style of the source address};

[0123] Second, style address: {style address}

[0124] Key identifier: Specific address part of the style address: {specific address part of the style address};

[0125] Description style: {description style of the style address}

[0126] Third, infer the candidate address information according to the source address information and the description style of the style address:

[0127] The key identifier is: {the key identifier of the candidate address};

[0128] The description style is: {the description style of the candidate address};

[0129] According to the candidate address information, generate the specific address part of the candidate address: {the specific address part of the candidate address}.

[0130] Step 1, the specific address part of the source address: Room 1201, Building 1;

[0131] The key identifier is: {1: number, 1201: number};

[0132] The description style is: number + "Building" + number;

[0133] Step 2, the specific address part of the style address: Room 03, 5th Floor, Building 2;

[0134] The key identifier is: {2: number, 5: number, 03: number};

[0135] The description style is: number + "Building" + number + "Floor" + number + "Room";

[0136] Step 3, according to the source address information and the description style of the style address, the candidate address information can be inferred:

[0137] The key identifier is: {1: number, 12: number, 01: number};

[0138] The description style is: number + "Building" + number + "Floor" + number + "Room";

[0139] According to the candidate address information, generate the specific address part of the candidate address: Room 01, 12th Floor, Building 1.

[0140] Step 242, calculate the loss function value of the style conversion model according to the specific address part of the candidate address output by the style conversion model and the specific address part of the candidate address in the training sample dataset. Specifically, optimize the style conversion model with the loss function value to obtain the trained style conversion model.

[0141] It can be understood that after constructing the prompt words for each training sample in the training sample dataset, input the prompt words of each training sample into the large language model. According to the output content of the large language model and the constructed output content, calculate the loss through the loss function, and reduce the loss function based on the optimizer built in the large language model, so that the output of the large language model gradually approaches the expected output content, and obtain the trained style conversion model.

[0142] In step 1, the regional part and specific address part of the source address to be matched, as well as the regional part and specific address part of the style address, are obtained respectively through the segmentation model. A prompt is constructed based on the specific address part of the input source address and the specific address part of the style address.

[0143] For example, the input is: [Room 1-1201, Building 1, Block C; Room 01, 12th Floor, Building 1, Block C];

[0144] Construct the prompt:

[0145] "

[0146] Please generate the specific address part of the candidate address according to the specific address part of the source address and the specific address part of the style address:

[0147] Specific address part of the source address: Room 1-1201, Building 1, Block C;

[0148] Specific address part of the style address: Room 01, 12th Floor, Building 1, Block C;

[0149] Please generate step by step in the following way:

[0150] First step, analyze the key identifiers and description styles of the source address;

[0151] Second step, analyze the key identifiers and description styles of the style address;

[0152] Third step, generate the specific address part of the candidate address according to the source address information and the description style of the style address".

[0153] In step 3, the specific address part of the candidate address and the regional part of the style address are concatenated to obtain the candidate address.

[0154] It can be understood that a prompt is constructed according to the specific address part of the source address to be matched and the specific address part of the style address, and the prompt is input into the style conversion model to obtain the output content of the style conversion model.

[0155] For example, when the above prompt is input into the style conversion model, the output content of the style conversion model may be as follows:

[0156] "

[0157] First step, specific address part of the source address: Room 1-1201, Building 1, Block C;

[0158] Key identifiers are: {Block C: entity, 1: number, 1201: number};

[0159] Description style is: entity + number + "-" + number;

[0160] Step 2, the specific address part of the style address: Room 01, 12th Floor, Building 1, Community C;

[0161] The key identifiers are: {Community C: entity, 1: number, 12: number, 01: number};

[0162] The description style is: entity + number + "Building" + number + "Floor" + number + "Room";

[0163] Step 3, according to the source address information and the description style of the style address, the specific address part of the candidate address information can be deduced:

[0164] The key identifiers are: {Community C: entity, 1: number, 12: number, 01: number};

[0165] The description style is: entity + number + "Building" + number + "Level" + number + "Room";

[0166] According to the candidate address information, generate the candidate address: "Room 01, 12th Floor, Building 1, Community C".

[0167] Extract the specific address part of the candidate address from the output content of the style conversion model, and then splice the area part of the style address with the specific address part of the candidate address to obtain the entire candidate address. Then, the description style of the candidate address is consistent with the description style of the target address.

[0168] Step 4, compare the candidate address with the target address. If they are exactly the same, the candidate address matches the target address successfully; otherwise, the match fails.

[0169] It can be understood that compare the spliced candidate address with the target address. If the spliced candidate address is exactly the same as the target address, it is considered that the source address and the target address point to the same location, and output a successful match; otherwise, output a failed match.

[0170] For example, the specific address part of the candidate address extracted from Step 3 is: "Room 01, 12th Floor, Building 1, Community C", and the area part of the style address is "No. 1, Road E, Street F, District B, City A, Province D".

[0171] The spliced candidate address is: "No. 1, Road E, Street F, District B, City A, Province D Room 01, 12th Floor, Building 1, Community C";

[0172] The target address is: "No. 1, Road E, Street F, District B, City A, Province D Room 01, 12th Floor, Building 1, Community C";

[0173] The spliced candidate address and the target address are exactly the same, and output a successful match; in addition, if any character of the spliced candidate address is inconsistent with the target address, output a failed match.

[0174] Figure 6 A system for address matching based on style conversion according to the present invention is provided. The system includes:

[0175] A segmentation module 601, configured to segment the source address and the style address respectively based on an address segmentation model, to obtain a regional part and a specific address part of the source address, and a regional part and a specific address part of the style address, wherein the description style of the specific address part of the style address is consistent with the description style of the specific address part of the target address;

[0176] A style conversion module 602, configured to input the specific address part of the source address and the specific address part of the style address into a style conversion model, perform style conversion on the specific address part of the source address, and obtain the specific address part of the candidate address;

[0177] A splicing module 603, configured to splice the specific address part of the candidate address and the regional part of the style address to obtain a candidate address;

[0178] A matching module 604, configured to compare the candidate address with the target address. If they are exactly the same, the candidate address matches the target address successfully; otherwise, the matching fails.

[0179] It can be understood that a system for address matching based on style conversion provided by the present invention corresponds to the method for address matching based on style conversion provided in the foregoing embodiments. The relevant technical features of the system for address matching based on style conversion can refer to the relevant technical features of the method for address matching based on style conversion, which will not be elaborated herein.

[0180] A method and a system for address matching based on style conversion provided in an embodiment of the present invention regard two addresses with different descriptions but pointing to the same location as two addresses with the same key identifier but different description styles. Therefore, an address matching method is proposed to convert the style of one address into the style of another address, and then determine whether the addresses match according to whether the key identifiers are exactly the same. During the style conversion process, the style conversion model focuses on the extraction of key identifiers and whether the key identifiers of the addresses are the same, while reducing the weight of specific description word features in the description style, so that the style conversion model can correctly distinguish the same geographical element in different regions, improving the flexibility and applicability of the style conversion model, and having higher matching accuracy in scenarios of different regions during training and application.

[0181] It should be noted that in the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0182] Those skilled in the art will understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0183] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0184] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0185] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0186] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0187] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An address matching method based on style conversion, characterized in that: include: Based on the address segmentation model, the source address and the style address are segmented respectively to obtain the region part and the specific address part of the source address and the region part and the specific address part of the style address, wherein the description style of the specific address part of the style address is consistent with the description style of the specific address part of the target address; The specific address part of the source address and the specific address part of the style address are input into the style transfer model, and the style transfer is performed on the specific address part of the source address to obtain the specific address part of the candidate address; Concatenate the specific address portion of the candidate address and the regional portion of the style address to obtain a candidate address; Compare the candidate address with the target address. If they are completely consistent, the candidate address and the target address are matched successfully. Otherwise, the match fails. The training process of the style transfer model includes: Constructing a training sample data set, wherein the training sample data set includes multiple training samples, each of which includes three parts, namely [specific address part of the source address, source address information], [specific address part of the style address, style address information], and [specific address part of the candidate address, candidate address information]; Training the style transfer model based on the training sample data set to obtain a trained style transfer model; The style conversion model is a large language model, and the style conversion model is trained based on the training sample data set to obtain a trained style conversion model, including: constructing a prompt word for each of the training samples based on the specific address part of the source address and the specific address part of the style address, and training the style conversion model based on the prompt word for each of the training samples; The prompt words of each training sample are constructed according to the following template: Please generate the specific address part of the candidate address based on the specific address part of the source address and the specific address part of the style address; The specific address part of the source address: {the specific address part of the source address}; The specific address part of the style address: {the specific address part of the style address}; Please generate it step by step as follows: The first step is to analyze the key identifiers and description styles of the specific address part of the source address; The second step is to analyze the key identifiers and description styles of the specific address part of the style address; The third step is to generate the specific address part of the candidate address according to the source address information and the description style of the style address; The style transfer model constructs output content according to the following template: The first step is the specific address part of the source address: {the specific address part of the source address}; The key identifier is: {key identifier of the source address}; The description style is: {source address description style}; Step 2, style address: {style address}; The key identifier is: the specific address part of the style address: {the specific address part of the style address}; The description style is: {description style of style address} The third step is to infer the candidate address information based on the source address information and the description style of the style address: The key identifier is: {key identifier of the candidate address}; The description style is: {description style of candidate address}; According to the candidate address information, generate the specific address part of the candidate address: {the specific address part of the candidate address}.

2. The address matching method based on style conversion according to claim 1, characterized in that: The step of constructing a training sample data set includes: Acquire multiple pieces of address data, each piece of address data including a specific address portion of a source address, a specific address portion of a style address, and a specific address portion of a candidate address; Extracting source address information, style address information and candidate address information from the specific address part of the source address, the specific address part of the style address and the specific address part of the candidate address respectively, wherein the source address information includes a key identifier and a description style of the source address, the style address information includes a key identifier and a description style of the style address, and the candidate address information includes a key identifier and a description style of the candidate address, and the description style of the style address is consistent with the description style of the candidate address; A training sample data set is constructed based on the specific address part of the source address, the source address information, the specific address part of the style address, the style address information, the specific address part of the candidate address, and the candidate address information.

3. The address matching method based on style conversion according to claim 2, characterized in that: The step of training the style conversion model based on the prompt word of each training sample includes: Inputting the prompt word of each training sample into the style conversion model, obtaining the output content of the style conversion model, and extracting the specific address part of the candidate address from the output content; The loss function value of the style conversion model is calculated according to the specific address part of the candidate address output by the style conversion model and the specific address part of the candidate address in the training sample data set, and the style conversion model is optimized by the loss function value to obtain a trained style conversion model.

4. The address matching method based on style conversion according to claim 1, characterized in that: The target address is used as the style address.

5. An address matching system based on style conversion, characterized in that: include: a segmentation module, configured to segment the source address and the style address respectively based on the address segmentation model to obtain the region part and the specific address part of the source address and the region part and the specific address part of the style address, wherein the description style of the specific address part of the style address is consistent with the description style of the specific address part of the target address; A style conversion module, used for inputting the specific address part of the source address and the specific address part of the style address into the style conversion model, performing style conversion on the specific address part of the source address, and obtaining the specific address part of the candidate address; A splicing module, used for splicing the specific address part of the candidate address and the regional part of the style address to obtain a candidate address; A matching module, used for comparing the candidate address with the target address. If they are completely consistent, the candidate address and the target address are matched successfully; otherwise, the match fails; The training process of the style transfer model includes: Constructing a training sample data set, wherein the training sample data set includes multiple training samples, each of which includes three parts, namely [specific address part of the source address, source address information], [specific address part of the style address, style address information], and [specific address part of the candidate address, candidate address information]; Training the style transfer model based on the training sample data set to obtain a trained style transfer model; The style conversion model is a large language model, and the style conversion model is trained based on the training sample data set to obtain a trained style conversion model, including: constructing a prompt word for each of the training samples based on the specific address part of the source address and the specific address part of the style address, and training the style conversion model based on the prompt word for each of the training samples; The prompt words of each training sample are constructed according to the following template: Please generate the specific address part of the candidate address based on the specific address part of the source address and the specific address part of the style address; The specific address part of the source address: {the specific address part of the source address}; The specific address part of the style address: {the specific address part of the style address}; Please generate it step by step as follows: The first step is to analyze the key identifiers and description styles of the specific address part of the source address; The second step is to analyze the key identifiers and description styles of the specific address part of the style address; The third step is to generate the specific address part of the candidate address according to the source address information and the description style of the style address; The style transfer model constructs output content according to the following template: The first step is the specific address part of the source address: {the specific address part of the source address}; The key identifier is: {key identifier of the source address}; The description style is: {source address description style}; Step 2, style address: {style address}; The key identifier is: the specific address part of the style address: {the specific address part of the style address}; The description style is: {description style of style address} The third step is to infer the candidate address information based on the source address information and the description style of the style address: The key identifier is: {key identifier of the candidate address}; The description style is: {description style of candidate address}; According to the candidate address information, generate the specific address part of the candidate address: {the specific address part of the candidate address}.

Citation Information

Patent Citations

  • Address matching method and device

    CN115470307A