Address rewriting method based on large-scale language model

Through multi-task multi-instruction supervision fine-tuning and reinforcement learning unbiased target alignment methods based on large-scale language model, the problem of inefficient address rewriting in the prior art is solved, and higher address correction accuracy and logistics scheduling accuracy are achieved.

CN120106098APending Publication Date: 2025-06-06HARBIN INST OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510140491.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing address rewriting method is inefficient in dealing with diversified address errors in the logistics industry and cannot meet the diverse address error correction needs in actual logistics distribution.

Method used

A multi-task multi-instruction supervision fine-tuning method based on a large-scale language model is adopted, combining address completion, address rewriting and address word segmentation data of logistics data to form an SFT data set, and trained through reinforcement learning PPO's unbiased target alignment. At the same time, reward functions and retrievals are designed to improve the perception ability and correction accuracy of address text.

Benefits of technology

It improves the scheduling accuracy of the logistics industry, improves the accuracy and efficiency of address rewriting, can effectively reduce package rerouting events caused by abnormal addresses, and reduces the operating costs of logistics companies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106098A_ABST
    Figure CN120106098A_ABST
Patent Text Reader

Abstract

The invention provides an address rewriting method based on a large-scale language model. The method comprises the following steps: carrying out fine tuning on a Chinese large model by utilizing multi-task fine tuning data based on prompt word structures such as address completion, address rewriting, address word segmentation and the like; performing unbiased alignment training on the fine-tuned large model; according to parcel delivery data of the logistics industry, obtaining a logistics text address as a question and answer pair, inputting the question and answer pair to the fine-tuned large model, and finely tuning the large model by using characteristics such as address coding, address decoding and semantic information as a training loss function of a reward model according to a proportion; and performing fine adjustment on the retriever by using the address coding data, acquiring an address with high similarity from the RAG in the address vector database by using the retriever after fine adjustment as a cue word, and assisting the large model after fine adjustment to correct the text address. The method can significantly improve the anti-scheduling rate of the logistics industry, improves the accuracy of the address library, reduces the loss of manpower and material resources, and can be accepted by the logistics industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer natural language processing, and in particular to an address rewriting method based on a large-scale language model. Background Art

[0002] In the logistics industry, accurate address information is essential for efficient delivery. However, in many countries, such as India and China, the prevalence of inaccurate or abnormal addresses poses a major challenge to logistics delivery. These abnormal addresses may be caused by factors such as an imperfect address regulatory framework, complex address structure, and fraudulent behavior, resulting in addresses containing incorrect information, such as missing administrative regions, nested addresses, aliases, irrelevant words, and spelling errors, which cannot be accurately parsed into a standard hierarchical structure.

[0003] In the logistics and delivery process, abnormal addresses can cause packages to be mistakenly sent to the wrong delivery station, which in turn leads to additional transshipment and rerouting operations, increasing logistics costs. Taking the logistics industry as an example, there are about 25,000 rerouting events every day, causing losses of more than $2 million per year. Therefore, developing an effective address rewriting framework to convert user-submitted addresses into a standardized format is of great significance to improving logistics and delivery efficiency.

[0004] Existing address rewriting methods are mainly based on statistical machine translation (SMT) or neural machine translation (NMT), but these methods have many limitations. The SMT method is limited by the representation ability of the statistical model, and the NMT method often needs to be retrained when processing new address data. In addition, most methods can only correct specific types of address errors and cannot meet the diverse address error correction needs in actual logistics distribution. Summary of the invention

[0005] The purpose of the present invention is to solve the problems in the prior art and propose an address rewriting method based on a large-scale language model.

[0006] The present invention is implemented by the following technical solution. The present invention proposes an address rewriting method based on a large-scale language model, and the method comprises the following steps:

[0007] Step 1: Multi-task multi-instruction supervised fine-tuning uses the address completion, address rewriting and address segmentation datasets based on logistics data to form the SFT dataset D; the process of multi-instruction supervised fine-tuning SFT using a large-scale language model to generate text is regarded as autoregressive sampling. In autoregressive language generation, one word is predicted at a time, and each prediction condition is based on the prompt and the previously generated words; given the model input x and the standard output y, the training goal is to find the conditional probability p(y|x)=Π i=1 p(y i |y0:i-1 ,x); thereby fine-tuning all parameters of the large-scale language model;

[0008] Step 2: Use unbiased target alignment based on reinforcement learning PPO, combine the parcel delivery data of logistics, apply the address encoding, address decoding and semantic information of the logistics system, design the reward function, and perform unbiased alignment training on the large-scale language model fine-tuned in step 1;

[0009] Step three, for the trained large-scale language model, design prompt words to obtain highly relevant addresses as prompt words, wherein the retriever that obtains the relevance refines the decoder part according to G2PTL using the logistics address service data to obtain the fine-tuned retriever, and uses the fine-tuned retriever to encode the input rewritten address to obtain embedding, retrieve similar addresses and input them into the trained large-scale language model obtained in step two to obtain the final successfully rewritten address.

[0010] Furthermore, in step one, a multi-task dataset is constructed: a dataset is constructed using address parsing, address entity prediction, and address rewriting, wherein the address parsing task involves the process of decomposing an address into its components, and the address parsing dataset contains 20 million <address, address component> pairs obtained from a logistics address service system; the LLM is fine-tuned on the address parsing task to generate addresses that conform to the standard hierarchical structure; finally, the address parsing, address entity prediction, and address rewriting datasets are mixed together to form the SFT dataset D.

[0011] Furthermore, in step 1, the training objective is expressed as:

[0012]

[0013] Where D is the SFT dataset, π is the large-scale language model, and θ is the parameter.

[0014] Furthermore, in step 2, the reward function r is composed of three scores: semantic score, reverse geocoding score and geocoding score; where the semantic score seman(x,y):

[0015] seman(x,y)=cos(f(x),f(y))

[0016] Where x is the original address, y is the rewritten address, and cos is the cosine similarity function. The formula is: f is a semantic embedding model; the pre-trained model BERT is selected as f; the rewritten address needs to be semantically close to the address obtained by reverse geocoding by passing the coordinates.

[0017] Furthermore, in step 2, the formula for the reverse geocoding score revgeo(y,c) is:

[0018] revgeo(y,c)=cos(f(y),f(reverse(c)))

[0019] Where y is the rewritten address, c is the successfully delivered coordinates, reverse is the reverse geocoding service, and f is the semantic embedding model.

[0020] Furthermore, in step 2, the geocoding score geo(y,c) formula is:

[0021]

[0022] k=dis(geocoding(y),c)

[0023] where dis is the Euclidean distance function, geocoding is the JD geocoding system that maps addresses to coordinates; and weights θ are used to calculate the distance between addresses and coordinates. 1 and θ 2 Normalize geocoding scores to [0,1].

[0024] Further, in step three, first, the retriever identifies and extracts a set of relevant addresses from the database, and then the large-scale language model generates the output retriever. The function of the retriever is to identify all addresses related to the input and determine their priorities; formally, the role of the retriever is encapsulated by the function M, which maps the query q and the database K to a subset K′, such that M:[q,K]→K′, where q is the address to be modified and K′ is a set of addresses related to q; the base model of the retriever is the encoder model E, which maps addresses to representations; then the similarity score between the query address q and each sample in K is calculated:

[0025] s(p,q)=sim(E(q),E(p)),p∈K

[0026] in is the similarity function; finally, the retriever returns the sample with the highest similarity, i.e. K′; the retriever encodes the text information into the embedding space, and performs the search in the embedding space, shifting the relevance criterion to geographical proximity to ensure that the retrieved address is located in a spatial range close to the query address.

[0027] Furthermore, in step three, the BERT encoder architecture is adopted as the base model E, the retriever is initialized using the parameters from the G2PTL text encoder, and then BERT is fine-tuned on the geocoding task.

[0028] The present invention also proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the address rewriting method based on a large-scale language model when executing the computer program.

[0029] The present invention also proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the address rewriting method based on a large-scale language model.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] Aiming at the address re-correction task, the present invention proposes a multi-task multi-instruction supervised fine-tuning, which utilizes the address completion, address rewriting and address segmentation data sets based on logistics data to mix together to form the SFT data set D after full parameter fine-tuning; uses unbiased target alignment based on reinforcement learning PPO, combines the parcel delivery data of logistics, applies the address encoding, address decoding, semantic information and other information of the logistics system, designs the reward function, and performs unbiased alignment training on the fine-tuned large model, thereby improving the perception ability of address text and improving the scheduling accuracy of the logistics industry.

[0032] For the information retrieval part, the decoder part was re-fine-tuned using the logistics system data according to G2PTL to obtain the fine-tuned retrieval method; the corrected address texts have a high accuracy rate and reasonable interpretability, and can be recognized by the logistics industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0034] Figure 1 It is a flow chart of the address rewriting method based on large-scale language model of the present invention;

[0035] Figure 2 It is an illustration of reverse scheduling for the logistics industry;

[0036] Figure 3 This is a schematic diagram of address classification rules for the logistics industry;

[0037] Figure 4 t-NSE graph after fine-tuning the retriever;

[0038] Figure 5 This is a schematic diagram of the anti-dispatch improvement rate in Zhejiang;

[0039] Figure 6 This is a schematic diagram of Yulin’s anti-dispatching improvement rate. DETAILED DESCRIPTION

[0040] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] The present invention proposes an address rewriting method based on a large-scale language model, the method comprising fine-tuning a Chinese large model Baichuan2-7B, Qwen using multi-task fine-tuning data based on prompt word structures such as address completion, address rewriting, and address segmentation; performing unbiased alignment training on the fine-tuned large model; obtaining a logistics text address as a question-answer pair input to the fine-tuned large model based on parcel delivery data in the logistics industry, and fine-tuning the large model using features such as address encoding, address decoding, and semantic information according to a proposed ratio as a training loss function of a reward model; fine-tuning a retriever using address encoding data, and using the fine-tuned retriever to obtain addresses with high similarity from an address vector database RAG as prompt words to assist the fine-tuned large model in correcting text addresses.

[0042] Specifically, combined Figure 1-Figure 6 The present invention proposes an address rewriting method based on a large-scale language model, and applies the address rewriting method of the large language model to the reverse scheduling of the logistics industry. The method comprises the following steps:

[0043] Step 1: Multi-task multi-instruction supervised fine-tuning uses the address completion, address rewriting and address segmentation datasets based on logistics data to form the SFT dataset D; the process of multi-instruction supervised fine-tuning SFT using a large-scale language model to generate text is regarded as autoregressive sampling. In autoregressive language generation, one word is predicted at a time, and each prediction condition is based on the prompt and the previously generated words; given the model input x and the standard output y, the training goal is to find the conditional probability p(y|x)=Π i=1 p(y i |y 0:i-1 ,x); thereby fine-tuning all parameters of the large-scale language model;

[0044] In step one, multi-task dataset construction: Since address semantics differ significantly from the large model corpus, directly using these models for address processing may result in inaccuracies. To address this challenge, the present invention adopts a strategy that involves aggregating a series of address-related tasks to fine-tune the large models, thereby improving their ability to understand standard Chinese addresses. Therefore, a dataset is constructed using address parsing, address entity prediction, and address rewriting, where the address parsing task involves the process of decomposing an address into its components, and the address parsing dataset contains 20 million <address, address component> pairs obtained from the logistics address service system; LLM is fine-tuned on the address parsing task to generate addresses that conform to the standard hierarchy, and finally, the address parsing, address entity prediction, and address rewriting datasets are mixed together to form the SFT dataset D.

[0045] In step 1, the training objective is expressed as:

[0046]

[0047] Where D is the SFT dataset, π is the large-scale language model, and θ is the parameter.

[0048] Step 2: Use unbiased target alignment based on reinforcement learning PPO, combine the parcel delivery data of logistics, apply the address encoding, address decoding and semantic information of the logistics system, design the reward function, and perform unbiased alignment training on the large-scale language model fine-tuned in step 1;

[0049] In step 2, the reward function r is composed of three scores: semantic score, reverse geocoding score and geocoding score; first, the rewritten address should not be too far from the original address in semantics. For example, if the user enters "Harbin Institute of Technology (Campus 1)", the model should not simply rewrite it as "Harbin Institute of Technology", even if the geocoding results of the two addresses are close in space. In order to prevent the rewriting from deviating too far from the user's intention, the present invention designs the semantic score seman(x,y):

[0050] seman(x,y)=cos(f(x),f(y))

[0051] Where x is the original address, y is the rewritten address, and cos is the cosine similarity function. The formula is: f is a semantic embedding model; the pre-trained model BERT is selected as f; secondly, the rewritten address should be semantically close to the address obtained by passing the coordinates for reverse geocoding. The result of reverse geocoding may not be accurate, for example, the room number or building number may not be correct. However, the rewritten address should be semantically close to the address to a certain extent, for example, they are in the same neighborhood or road.

[0052] In step 2, the reverse geocoding score revgeo(y,c) formula is:

[0053] revgeo(y,c)=cos(f(y),f(reverse(c)))

[0054] Where y is the rewritten address, c is the successfully delivered coordinates, reverse is the reverse geocoding service, and f is the semantic embedding model.

[0055] In step 2, the present invention selects a geocoding task to evaluate the rewriting result. The geocoding result of the rewritten address should be close to the coordinates of the successful delivery. Therefore, the present invention designs the geocoding score geo(y,c) formula as:

[0056]

[0057] k=dis(geocoding(y),c)

[0058] where dis is the Euclidean distance function, geocoding is the JD geocoding system that maps addresses to coordinates; and weights θ are used to calculate the distance between addresses and coordinates. 1 and θ 2 The geocoding scores are normalized to [0,1]. 1 Set to 100 meters, θ 2 Set to 1000 meters. When rewritten, the address cannot be recognized as an address by the logistics geocoding service (geocoding failure), and the geocoding score is 0. The three scores are combined with the weight λ 1 ,λ 2 ,λ 3 Addition:

[0059] r(x,y,c)=λ 1 seman(x,y)+λ 2 revgeo(y,c)+λ 3 geo(y,c)

[0060] In the present invention, λ 1 ,λ 2 ,λ 3 They are set to 0.2, 0.2, and 0.6 respectively. RL task formula: At each time step t, the large language model calculates the current state s t Generate the next token as an action, which includes the tokens that have been generated. Then the model passes the reward function Get instant rewards t The training uses proximal policy optimization (PPO) to optimize the large language model. The PPO algorithm can be expressed as:

[0061]

[0062] clip(k t,θ ,1-∈,1+∈)A θ′ (s t ,a t )}]

[0063]

[0064] Where θ′ is the parameter of the fixed strategy, θ is the parameter of the update strategy, and the clipping function Clip(k t,0 ,1-∈,1+∈) will be the ratio k t,θ Restricted to the range [1-∈, 1+∈]. A is the advantage function, which is based on the value network V φ· The value network V φ· By the policy network π 0 Initialization. The formula follows the generalized advantage estimation (GAE):

[0065]

[0066] Where λ is the bias-variance trade-off parameter. To prevent the model from deviating too far from the initialization, KL divergence regularization is added to reward:

[0067] R(s t ,a t )=r(x,y,c)-βKL(π θ ||π 0 )

[0068] The final loss function consists of policy loss and value loss:

[0069]

[0070] clip(k t,θ ,1-∈,1+∈)A θ′ (s t ,a t )}

[0071]

[0072] Where S is the sampling set and T is the number of steps.

[0073] Step three, for the trained large-scale language model, design prompt words to obtain highly relevant addresses as prompt words, wherein the retriever that obtains the relevance refines the decoder part according to G2PTL using the logistics address service data to obtain the fine-tuned retriever, and uses the fine-tuned retriever to encode the input rewritten address to obtain embedding, retrieve similar addresses and input them into the trained large-scale language model obtained in step two to obtain the final successfully rewritten address.

[0074] In step three, first, the retriever identifies and extracts a set of relevant addresses from the database, and then the large-scale language model generates the output retriever. The function of the retriever is to identify all addresses related to the input and determine their priorities; formally, the role of the retriever is encapsulated by the function M, which maps the query q and the database K to a subset K′, such that M:[q,K]→K′, where K′ includes the relevant addresses corresponding to K, where q is the address to be modified and K′ is a set of addresses related to q; the base model of the retriever is the encoder model E, which maps addresses to representations; then the similarity score between the query address q and each sample in K is calculated:

[0075] s(p,q)=sim(E(q),E(p)),p∈K

[0076] in is a similarity function, such as cosine similarity; finally, the retriever returns the sample with the highest similarity, i.e. K′; for efficiency and scalability, the retriever encodes the text information into an embedding space, in which the retriever performs the search, shifting the relevance criterion to geographic proximity to ensure that the retrieved addresses are within a spatial range close to the query address.

[0077] In step three, the BERT encoder architecture is adopted as the base model E, the initialization of the retriever utilizes the parameters from the G2PTL text encoder, and BERT is subsequently fine-tuned on the geocoding task. Several fully connected (FC) layers are attached to the existing BERT structure, which convert the embeddings into geographic coordinates, i.e., longitude and latitude. The geocoding dataset of the present invention contains 200 million <address, coordinate> pairs. BERT is pre-trained, the added FC layer is randomly initialized, and a two-stage training strategy is adopted. In the first training cycle, the BERT parameters are frozen and focus on training the FC layer. For subsequent epochs, the entire neural network is trained, including BERT and FC layers. When retrieving addresses, the representation generated by BERT is used as a spatial embedding.

[0078] The specific implementation of the present invention is described below in conjunction with specific experimental results.

[0079] The address text, longitude and latitude, and site information used in the experiment are all provided by the logistics industry, totaling tens of millions of addresses, longitude and latitude, and site information.

[0080] Perform step 1, such as Figure 1, multi-instruction supervised fine-tuning for multiple tasks uses address completion, address rewriting, and address segmentation datasets based on logistics data mixed together to form the SFT dataset D. The process of generating text using a large language model using multi-instruction supervised fine-tuning (SFT) can be regarded as autoregressive sampling. In autoregressive language generation, one word is predicted at a time, and each prediction condition is based on the prompt and the previously generated words. Given a model input x and a standard output y, the training goal is to find a conditional probability p(y|x)=Π i=1 p(y i |y 0:i-1 ,x). Thus, the address rewriting task is performed to fine-tune all parameters of the large model;

[0081] Execute step 2, use unbiased target alignment based on reinforcement learning PPO, combine logistics package delivery data, apply address encoding, address decoding, semantic information and other information of the logistics system, design a reward function, and perform unbiased alignment training on the large model fine-tuned in step 1;

[0082] Execute step 3. For the trained large model, design prompt words to obtain highly relevant addresses as prompt words. The retriever that obtains the relevant degree re-fine-tunes the decoder part according to the logistics address service data of G2PTL to obtain the fine-tuned retriever. The fine-tuned retriever is used to encode the input rewritten address to obtain embedding. The t-SNE graph corresponding to the fine-tuned retriever is as follows: Figure 4 ,from Figure 4 It can be seen that the fine-tuned retriever has improved the association between address text and longitude and latitude, proving the effectiveness of fine-tuning. The retrieved addresses with similarities are input into the trained large model to obtain the final successfully rewritten address. Table 1 shows the accuracy of the address large model in different tasks.

[0083] Table 1. Accuracy of the large address model in different tasks

[0084]

[0085]

[0086] Table 1 shows that the fine-tuned large model is significantly better than the unfine-tuned and basic Bert address tasks in address completion tasks and address correction tasks.

[0087] Figure 5 and Figure 6 The recall rate of the pre-sorting system in Zhejiang Province increased by 10% and the average daily recovery volume of the pre-sorting system in Zhejiang Province reached more than 500 orders. Among them, the recovered addresses are cumulative, and the recovered orders will be refreshed in the address database. This proves the effectiveness of the method described in the present invention.

[0088] The method described in the present invention can achieve the following goals:

[0089] Correction of abnormal addresses: Existing address rewriting methods are usually optimized for specific types of errors or require frequent retraining to handle new address data. This paper proposes a framework based on retrieval-augmented large language model (LLM) that can uniformly handle multiple types of address errors, including missing administrative divisions, nested addresses, unofficial aliases, irrelevant words, and spelling errors.

[0090] Reduced retraining requirements: Traditional address rewriting methods often require retraining models to adapt to new address data, which is inefficient in practical applications. The present invention decouples knowledge storage and reasoning capabilities through retrieval-augmented generation (RAG) technology, so that updated answers can be generated without retraining when new knowledge emerges.

[0091] Improving the accuracy and efficiency of address rewriting: By fine-tuning a large language model through supervision and designing an unbiased target alignment framework, the present invention can run efficiently on real data streams across the country and significantly improve the accuracy of address rewriting.

[0092] Improvement of practical application effect: The application of the present invention in a real logistics system shows that it can effectively reduce the package rerouting events caused by abnormal addresses, thereby reducing the operating costs of logistics companies.

[0093] Example 1: Automatic address standardization

[0094] In the logistics system, the address rewriting method of the present invention can automatically convert the non-standard address entered by the user into a standardized format. This process recognizes and corrects errors in the address, such as spelling errors, missing administrative divisions, etc., by retrieving the enhanced large language model (LLM), thereby improving the accuracy and efficiency of package delivery.

[0095] Example 2: Abnormal address detection and correction

[0096] The present invention can detect and correct abnormal addresses in real time during the logistics distribution process. When the system identifies anomalies or inconsistencies in the address information, it uses reinforcement learning target alignment technology to automatically generate a rewritten address that is more consistent with the user's intention, thereby reducing the rerouting events of the package and reducing operating costs.

[0097] Example 3: Cross-region address translation

[0098] For multinational logistics companies, the present invention can effectively handle the differences in address formats between different countries and regions. Through multi-task multi-instruction supervision and fine-tuning, the system can understand and translate address information in different languages ​​and formats, ensuring that packages can pass international borders smoothly and be delivered accurately.

[0099] Example 4: Personalized address suggestion

[0100] On an e-commerce platform, when a user enters an address, the present invention can provide personalized address suggestions. By analyzing user historical data and current input, the system uses retrieval-augmented generation (RAG) technology to provide address suggestions that best meet user needs, improve user experience and reduce address input errors.

[0101] The present invention also proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the address rewriting method based on a large-scale language model when executing the computer program.

[0102] The present invention also proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the address rewriting method based on a large-scale language model.

[0103] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DRRAM). It should be noted that the memory of the method described in the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0104] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a high-density digital video disc (DVD)), or a semiconductor medium (eg, a solid state disc (SSD)).

[0105] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software. The steps of the method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor for execution, or a combination of hardware and software modules in a processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it is not described in detail here.

[0106] It should be noted that the processor in the embodiment of the present application can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiment can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor can be combined and performed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0107] The above is a detailed introduction to the address rewriting method based on a large-scale language model proposed in the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. An address rewriting method based on a large-scale language model, characterized in that: The method comprises the following steps: Step 1: Multi-task multi-instruction supervised fine-tuning uses the address completion, address rewriting and address segmentation datasets based on logistics data to form an SFT dataset D; the process of multi-instruction supervised fine-tuning SFT using a large-scale language model to generate text is regarded as autoregressive sampling. In autoregressive language generation, one word is predicted each time, and each prediction condition is based on the prompt and the previously generated words; given the model input x and the standard output y, the training goal is to find the conditional probability p(y|x=Π i=1 p(y i |y 0:i-1 ,x); thereby fine-tuning all parameters of the large-scale language model; Step 2: Use unbiased target alignment based on reinforcement learning PPO, combine the parcel delivery data of logistics, apply the address encoding, address decoding and semantic information of the logistics system, design the reward function, and perform unbiased alignment training on the large-scale language model fine-tuned in step 1; Step three, for the trained large-scale language model, design prompt words to obtain highly relevant addresses as prompt words, wherein the retriever that obtains the relevance refines the decoder part according to G2PTL using the logistics address service data to obtain the fine-tuned retriever, and uses the fine-tuned retriever to encode the input rewritten address to obtain embedding, retrieve similar addresses and input them into the trained large-scale language model obtained in step two to obtain the final successfully rewritten address.

2. The method according to claim 1, characterized in that In step one, a multi-task dataset is constructed: a dataset is constructed using address parsing, address entity prediction, and address rewriting, where the address parsing task involves the process of decomposing an address into its components, and the address parsing dataset contains 20 million <address, address component> pairs obtained from the logistics address service system; the LLM is fine-tuned on the address parsing task to generate addresses that conform to the standard hierarchical structure; finally, the address parsing, address entity prediction, and address rewriting datasets are mixed together to form the SFT dataset D.

3. The method according to claim 1, characterized in that In step 1, the training objective is expressed as: Where D is the SFT dataset, π is the large-scale language model, and θ is the parameter.

4. The method according to claim 1, characterized in that: In step 2, the reward function r consists of three scores: semantic score, reverse geocoding score and geocoding score; the semantic score seman(x,y): seman(x,y)=cos(f(x),f(y)) Where x is the original address, y is the rewritten address, and cos is the cosine similarity function. The formula is: f is a semantic embedding model; the pre-trained model BERT is selected as f; the rewritten address needs to be semantically close to the address obtained by reverse geocoding by passing the coordinates.

5. The method according to claim 4, characterized in that In step 2, the reverse geocoding score revgeo(y,c) formula is: revgeo(y,c)=cos(f(y),f(reverse(c))) Where y is the rewritten address, c is the successfully delivered coordinates, reverse is the reverse geocoding service, and f is the semantic embedding model.

6. The method according to claim 5, characterized in that In step 2, the geocoding score geo(y,c) formula is: k=dis(geocoding(y),c) where dis is the Euclidean distance function, geocode is the JD geocoding system that maps addresses to coordinates; and the geocoding scores are normalized to [0,1] by weights θ1 and θ2.

7. The method according to claim 1, characterized in that In step three, first, the retriever identifies and extracts a set of relevant addresses from the database, and then the large-scale language model generates the output retriever. The function of the retriever is to identify all addresses related to the input and determine their priorities; formally, the role of the retriever is encapsulated by the function M, which maps the query q and the database K to a subset K′, such that M:[q,K]→K′, where q is the address to be modified and K′ is a set of addresses related to q; the base model of the retriever is the encoder model E, which maps addresses to representations; then the similarity score between the query address q and each sample in K is calculated: s(p,q)=sim(E(q),E(p)),p∈K Where sin: is the similarity function; finally, the retriever returns the sample with the highest similarity, i.e. K′; the retriever encodes the text information into the embedding space, and performs the search in the embedding space, shifting the relevance criterion to geographical proximity to ensure that the retrieved address is located in a spatial range close to the query address.

8. The method according to claim 7, characterized in that In step 3, the BERT encoder architecture is adopted as the base model E, the retriever is initialized using the parameters from the G2PTL text encoder, and then BERT is fine-tuned on the geocoding task.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Process reward model training method and system

    CN120430424A

  • Address text correlation analysis method, system and equipment and storage medium

    CN120850955A