Logistics address data processing method and device, equipment and storage medium
Through the combination of sharded storage and distributed model, the problem of low efficiency in logistics address data processing is solved, and efficient logistics address data processing and target address selection is achieved.
Patent Information
- Application Number
- CN202311821301.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the processing efficiency of logistics addresses is low, mainly due to the complex address segmentation algorithm and large data volume, resulting in low processing efficiency.
A logistics address data processing method is proposed. By obtaining the initial logistics address data and slicing it, the pre-trained address segmentation model is read in the distributed file system, the data segmentation input model is used for address segmentation, and the slicing address data is obtained, and the address base table is generated based on the slicing address data to select the target address.
Through parallel processing and the use of distributed models, the processing efficiency of logistics address data is significantly improved, the complexity of address labeling is reduced, and efficient logistics address data processing is realized.
Smart Images

Figure CN120216499A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, apparatus, device, and storage medium for processing logistics address data. Background Art
[0002] With the rise of e-commerce and the development of global trade, the logistics industry plays an increasingly important role. Logistics refers to the entire process of delivering goods from producers to consumers, including transportation, warehousing, packaging, distribution, and other links of goods. And different group portrait characteristics are contained in the logistics address information.
[0003] Since the logistics addresses filled in by users when sending express deliveries have different personal styles, the logistics addresses cannot be directly used and need to be processed accordingly. In related technologies, the address segmentation algorithm is relatively complex, and the data volume of logistics addresses is generally large, so the efficiency of processing addresses is low. Summary of the Invention
[0004] The main objective of the embodiments of this application is to propose a method, apparatus, device, and storage medium for processing logistics address data, so as to improve the processing efficiency of logistics address data.
[0005] To achieve the above objective, a first aspect of the embodiments of this application proposes a method for processing logistics address data, including:
[0006] Obtain a plurality of initial logistics address data to be processed, and store the initial logistics address data in a data set in a sliced manner; the data set includes a plurality of data slices;
[0007] Read a pre-trained address segmentation model from a distributed file system, and broadcast the address segmentation model to the data set;
[0008] Input the initial logistics address data in the data slices into the address segmentation model for address segmentation to obtain segmented address data; the segmented address data includes a plurality of address segments, and each address segment is composed of a start address data, at least one middle address data, and an end address data;
[0009] Obtain an address base table according to the segmented address data, and select a target address based on the address base table.
[0010] In some embodiments, the training process of the address segmentation model includes the following steps:
[0011] Obtain an address training data set; the address training data set includes a plurality of training data, the training data includes a plurality of levels of training address segments, and each training address segment includes a corresponding segmentation label;
[0012] Input the corresponding training address segment and the segmentation label into the address segmentation model in the order of the said levels for prediction, to obtain the likelihood function value of the training address segment;
[0013] Maximize the likelihood function value, and adjust the model weights of the address segmentation model until the trained address segmentation model is obtained.
[0014] In some embodiments, the address segmentation model is a distributed probability model; the distributed probability model includes a plurality of the node sequences; each of the node sequences includes three prediction nodes; the step of inputting the training address segment and the segmentation label into the address segmentation model in the order of the said levels for prediction to obtain the likelihood function value of the training address segment includes:
[0015] Input the training address segment into the node sequence in the order of the said levels;
[0016] Successively obtain selection points from the prediction nodes in the node sequence;
[0017] Generate a prediction path according to the selection points, and obtain the prediction score of the prediction path;
[0018] Generate a label path and a corresponding label score based on the segmentation label;
[0019] Obtain the likelihood function value according to the label score and the prediction score.
[0020] In some embodiments, the step of obtaining the likelihood function value according to the label score and the prediction score includes:
[0021] Traverse the node sequence, generate a plurality of prediction paths, and calculate the prediction score of each prediction path;
[0022] Accumulate the prediction scores to obtain an accumulated value, and use the ratio of the label score to the accumulated value as the likelihood function value.
[0023] In some embodiments, the levels successively include: province level, city level, district level, street level and house number level.
[0024] In some embodiments, each level corresponds to a first number of the node sequences, and the training data segment is composed of a head address training data, a second number of intermediate address training data and a tail address training data; the step of inputting the training address segment into the node sequence in the order of the said levels includes:
[0025] If the difference between the first number and the second number is greater than two, then generate at least one supplementary intermediate address training data all of which are zero based on the difference;
[0026] Input the head address training data into the first node sequence.
[0027] Input the middle address training data and the supplementary middle address training data into the second node sequence in turn.
[0028] Input the tail address training data into the last node sequence.
[0029] In some embodiments, obtaining the address base table according to the segmented address data and selecting the target address based on the address base table includes:
[0030] Determine the address base table by using the segmented address data and the corresponding time information.
[0031] Screen the segmented address data in the address base table according to a preset area to obtain a plurality of first screening areas.
[0032] According to the time information of the segmented address data, screen the segmented address data located in a preset time period in the first screening areas as candidate address data.
[0033] Cluster the candidate address data, and obtain the target address according to the clustering result.
[0034] To achieve the above object, a second aspect of the embodiments of the present application proposes a logistics address data processing device, including:
[0035] Address data acquisition module: configured to acquire a plurality of initial logistics address data to be processed, and store the initial logistics address data in a data set in a segmented manner; the data set includes a plurality of data segments.
[0036] Model acquisition module: configured to read a pre-trained address segmentation model from a distributed file system, and broadcast the address segmentation model to the data set.
[0037] Address segmentation module: configured to input the initial logistics address data in the data segments into the address segmentation model for address segmentation to obtain segmented address data; the segmented address data includes a plurality of address segments, and each address segment is composed of a head address data, at least one middle address data, and a tail address data.
[0038] Target address selection module: configured to obtain an address base table according to the segmented address data, and select a target address based on the address base table.
[0039] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0040] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, which is a storage medium that stores a computer program. When the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0041] The logistics address data processing method, device, equipment, and storage medium proposed in the embodiments of the present application obtain multiple initial logistics address data to be processed, and perform sharded storage on the initial logistics address data to obtain multiple data shards. Read the pre-trained address segmentation model from the distributed file system, and broadcast the address segmentation model to each data shard. Then, input the initial logistics address data in the data shard into the address segmentation model for address segmentation to obtain segmented address data. Among them, the segmented address data includes multiple address segments, and each address segment is composed of a starting address data, at least one middle address data, and an ending address data. Finally, an address base table is obtained based on the segmented address data, and a target address is selected based on the address base table. In the embodiments of the present application, for the characteristic of a large amount of logistics address data, the distributed file system is first used for parallel processing of data, which greatly improves the data processing efficiency. In addition, when segmenting the logistics address, based on the address characteristics of the logistics address, it is segmented into multiple address segments, and at the same time, each address segment is divided into three parts. Through this annotation method, while retaining the address information, the complexity of address annotation is reduced, and the data processing efficiency of the logistics address is further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flowchart of the logistics address data processing method provided by the embodiments of the present application.
[0043] Figure 2 is a schematic diagram of the training process of the address segmentation model in the embodiments of the present application.
[0044] Figure 3 is a schematic diagram of the structure of the address segmentation model in the embodiments of the present application.
[0045] Figure 4 is a flowchart of step S220 provided by the embodiments of the present application.
[0046] Figure 5 is a flowchart of inputting the training address segments into the node sequence in hierarchical order provided by the embodiments of the present application.
[0047] Figure 6It is a schematic diagram of a node sequence in an embodiment of the present application.
[0048] Figure 7 It is a flowchart for obtaining a likelihood function value according to a label score and a prediction score provided by an embodiment of the present application.
[0049] Figure 8 It is a flowchart for obtaining an address base table according to segmented address data and selecting a target address based on the address base table provided by an embodiment of the present application.
[0050] Figure 9 It is a schematic diagram of a first screening area provided by an embodiment of the present application.
[0051] Figure 10 It is an overall schematic diagram of a logistics address data processing method provided by an embodiment of the present application.
[0052] Figure 11 It is a structural block diagram of a logistics address data processing device provided by another embodiment of the present application.
[0053] Figure 12 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0054] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0055] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the flowchart.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0057] First, several nouns involved in the present application are analyzed:
[0058] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.
[0059] Distributed File System (Hadoop Distributed File System, HDFS): Aims to provide reliable and high-capacity data storage and enable high-throughput access to data. HDFS can meet the needs of large-scale data processing and storage. It has the following characteristics: Distributed storage: HDFS splits file data into multiple blocks and stores these blocks distributively on multiple computer nodes in the cluster. This can achieve parallel reading and writing of data and high availability. Fault tolerance: HDFS provides fault tolerance by replicating file blocks in the cluster. Each file block will have multiple copies stored on different nodes to prevent data loss. High throughput: HDFS is suitable for batch reading and writing operations of large datasets. It achieves high-throughput data access through data locality and parallel processing. Scalability: HDFS can add new computer nodes to the cluster to increase storage capacity and processing power. It supports horizontal scaling, enabling the system to adapt to growing data demands.
[0060] With the rise of e-commerce and the development of global trade, the logistics industry plays an increasingly important role. Logistics refers to the entire process of delivering goods from producers to consumers, including transportation, warehousing, packaging, distribution, and other links of goods. And different group portrait characteristics are contained in logistics address information.
[0061] Since the logistics addresses filled in by users when sending express deliveries have different personal styles, the logistics addresses cannot be directly used and need to be processed accordingly. In related technologies, the address segmentation algorithm is relatively complex, and the data volume of logistics addresses is generally large. Therefore, the efficiency of processing addresses is relatively low.
[0062] Based on this, the embodiments of the present application provide a method, apparatus, device, and storage medium for processing logistics address data. In view of the characteristic of a large amount of logistics address data, the parallel processing of data is first carried out by using a distributed file system, which greatly improves the data processing efficiency. In addition, when segmenting the logistics address, based on the address characteristics of the logistics address, it is segmented into multiple address segments, and at the same time each address segment is divided into three parts. Through this annotation method, while retaining the address information, the complexity of address annotation is reduced, and the data processing efficiency of the logistics address is further improved.
[0063] The embodiments of the present application provide a method, apparatus, device, and storage medium for processing logistics address data, which will be specifically described through the following embodiments. First, the method for processing logistics address data in the embodiments of the present application will be described.
[0064] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, method, technology, and application system. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0065] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0066] The logistics address data processing method provided by the embodiments of the present application relates to the field of artificial intelligence technology. The logistics address data processing method provided by the embodiments of the present application can be applied to a terminal, or to a server, or can be a computer program running on a terminal or a server. For example, the computer program can be a native program or software module in an operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in an operating system to run, such as a client supporting logistics address data processing, or it can be a small program, that is, a program that only needs to be downloaded to a browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module or plug-in. Among them, the terminal communicates with the server through a network. The logistics address data processing method can be executed by the terminal or the server, or by the terminal and the server in cooperation.
[0067] In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer or a smart watch, etc. The server can be an independent server, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms; it can also be a service node in a blockchain system, and the service nodes in the blockchain system form a peer-to-peer (P2P, Peer To Peer, P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP) protocol. The terminal and the server can be connected through communication connection methods such as Bluetooth, Universal Serial Bus (USB) or network, and this embodiment does not make any restrictions here.
[0068] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0069] It should be noted that in each specific embodiment of this application, when it comes to relevant processing based on data related to the user's identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. In addition, when this application embodiment needs to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of this application embodiment will be obtained.
[0070] The following describes the logistics address data processing method in the embodiments of this application.
[0071] Figure 1 is an optional flowchart of the logistics address data processing method provided by the embodiments of this application. Figure 1 The method in may include but is not limited to steps S110 to S140. At the same time, it can be understood that this embodiment does not specifically limit the order of steps S110 to S140 in Figure 1 and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.
[0072] Step S110: Obtain multiple initial logistics address data to be processed and store the initial logistics address data in slices in a data set.
[0073] In one embodiment, the initial logistics address data refers to the logistics address-related data that needs to be processed, and the quantity thereof is large, usually in the tens of millions. Therefore, in order to improve the processing efficiency, the embodiments of the present application process the logistics address data in a parallel manner. For example, the initial logistics address data to be processed is fragmented and stored to obtain a plurality of data fragments, and these data fragments are processed in parallel.
[0074] In one embodiment, the initial logistics address data to be processed is stored on the distributed file system HDFS, and then based on the Apache Spark big data processing framework, each data fragment is sent to a Resilient Distributed Dataset (RDD). Among them, a plurality of partitions are defined in the distributed data set RDD, and each partition is distributed and stored on different nodes in the cluster, so that the data in different partitions can be processed and calculated in parallel. Therefore, the above data fragments are stored in the data set partitions. The data set includes a plurality of data fragments, and the number of data fragments can be the same as the number of partitions.
[0075] In one embodiment, when the data fragments perform data processing in each partition, it is necessary to cache the intermediate results in the calculation process of the distributed data set RDD for subsequent use. Therefore, through a persistence operation, Spark can cache the calculation results of the distributed data set RDD in the memory, thereby avoiding recalculation in subsequent use. Since the data volume of the initial logistics address data is large, multiple RDD persistence operations are required. In order to improve the data processing efficiency and ensure that the persisted data can be accommodated in the memory, the embodiments of the present application appropriately increase the storage parameter of the memory. For example, set: spark.storage.memoryFraction = 0.6, so as to avoid the problem that the memory is not enough for caching, resulting in data being written to the disk only, causing performance degradation.
[0076] In addition, during the calculation process of the distributed data set RDD, the Shuffle operation may be used to re-partition the initial logistics address data on different fragments, redistribute the data to different nodes, usually accompanied by reordering, merging or aggregation of the data. Therefore, in order to improve the data processing efficiency, the embodiments of the present application appropriately reduce the memory occupation ratio parameter of the Shuffle operation, so that during prediction, fewer Shuffle operations are generated, avoiding the situation that the memory is not enough when there is too much data during the Shuffle process and having to spill to the disk, resulting in performance degradation. For example, it can be set: spark.shuffle.memoryFraction = 0.2.
[0077] The above describes the relevant settings for parallel computing based on a distributed file system in the embodiments of the present application. Since the initial logistics address data may contain some non-standard logistics addresses, the embodiments of the present application also perform data preprocessing on the initial logistics address data. For example, if some address information such as "province" and "city" is missing in part of the initial logistics address data, it is supplemented. For example, if there are typos or spelling mistakes in part of the initial logistics address data, they are corrected.
[0078] Step S120: Read the pre-trained address segmentation model from the distributed file system and broadcast the address segmentation model to the data set.
[0079] In one embodiment, since the number of address training data sets used to train the address segmentation model may reach 1.6 billion, the Apache Spark big data processing framework needs to be used to annotate, train and deploy the model, and perform distributed inference on the data. Specifically, the setting parameters of the Apache Spark big data processing framework can be: set num-executors = 100, which is used to specify that the number of Executors started in the Apache Spark application is 100. Then set executor-memory = 20G, which is used to specify that the memory size that each Executor can use is 20G. Then set executor-cores = 4, which is used to specify that the number of CPU cores that each Executor can use is 4. Set conf spark.default.parallelism = 800, which is used to specify the default parallelism as 800. The parallelism indicates how many tasks are to be used to execute in parallel when performing data processing. By setting the above parameters, the embodiments of the present application can make full use of the resources of the Spark cluster for distributed training. After the address segmentation model is obtained after training, the model is deployed on the distributed file system HDFS. When address segmentation is required, the pre-trained address segmentation model is read from the distributed file system, and the address segmentation model is broadcast to the distributed data set RDD in a broadcast manner, so that the initial logistics address data in each data slice on the data set can use the trained address segmentation model for address segmentation.
[0080] In one embodiment, the address segmentation model is a distributed probability model (Conditional Random Fields, CRF). Refer to Figure 2 , Figure 2 is a schematic diagram of the training process of the address segmentation model in the embodiments of the present application, specifically including the following steps S210 to S230:
[0081] Step S210: Obtain the address training data set.
[0082] Among them, the address training data set includes multiple training data. The training data is obtained by annotating the logistics address data, and the obtained training address segments include multiple levels. The annotation information constitutes the segmentation label corresponding to each training address segment.
[0083] In one embodiment, the levels during annotation successively include: provincial level, municipal level, district level, street level, and house number level. That is, the logistics address data is split according to province, city, district, street, and house number. Among them, according to administrative divisions, some districts are at the same level as counties, and some streets are at the same level as towns. Therefore, for the counties that appear in some logistics address data, they can be split according to districts, and for the towns that appear, they can be split according to streets. Then, the data related to each level is used as a training address segment. For example, if the logistics address data is "No. 10, Lane t, opqrs Street, hijkmn District, defg City, abc Province", the annotated training address segments are respectively: provincial level (abc Province), municipal level (defg City), district level (hijkmn District), street level (opqrs Street), and house number level (Lane t, No. 10).
[0084] Then each training data segment is labeled into three parts. After annotation, the training data segment consists of a first address training data, a second number of intermediate address training data, and a last address training data.
[0085] In one embodiment, different training data segments include different labels. For example, the labels corresponding to the provincial level include: B_PRO, M_PRO, E_PRO, where B_PRO is used to label the first address training data, M_PRO is used to label the intermediate address training data, and E_PRO is used to label the last address training data. By analogy, the municipal level includes: B_CITY, M_CITY, E_CITY; the district level includes B_COUNTY, M_COUNTY, E_COUNTY; the street level includes B_TOWN, M_TOWN, E_TOWN; the house number level includes: B_DETA, M_DETA, E_DETA.
[0086] For example, the segmentation labels obtained from the above logistics address data are respectively:
[0087] {B_PRO|--|a, M_PRO|--|b, M_PRO|--|c, E_PRO|--|Province};
[0088] {B_CITY|--|d, M_CITY|--|e, M_CITY|--|f, M_CITY|--|g, E_CITY|--|City};
[0089] {B_COUNTY|--|h, M_COUNTY|--|i, M_COUNTY|--|j, M_COUNTY|--|k, M_COUNTY|--|m, M_COUNTY|--|n, E_COUNTY|--|District};
[0090] {B_TOWN|--|o, M_TOWN|--|p, M_TOWN|--|q, M_TOWN|--|r, M_TOWN|--|s, M_TOWN|--|Street, E_TOWN|--|Avenue};
[0091] {B_DETA|--|t, M_DETA|--|Lane, M_DETA|--|1, M_DETA|--|0, M_DETA|--|No.};
[0092] As can be seen from the above, the number of intermediate address training data in each training data segment is set according to the actual logistics address data.
[0093] Step S220: According to the order of levels, input the corresponding training address segment and segmentation label into the address segmentation model for prediction to obtain the likelihood function value of the training address segment.
[0094] In one embodiment, referring to Figure 3 , Figure 3 is a schematic structural diagram of the address segmentation model in the embodiment of the present application. In the figure, the address segmentation model includes multiple node sequences, and each node sequence includes 3 prediction nodes. Moreover, the node sequences can be grouped according to the levels of the training data segments, that is, each training data segment has a corresponding node sequence, and the number of node sequences in each group can be set by the longest number of the corresponding training data segments in the training dataset. For example, if the longest number of characters in the text of the training data segment corresponding to the "province level" in the training dataset is 9, then it is set that there are 9 node sequences in the group corresponding to the "province level" node sequence. And so on, the node sequences of each training data segment are set.
[0095] Then each node sequence includes 3 prediction nodes, and the prediction nodes in the node sequence of each training data segment are the same. For example, the three prediction nodes corresponding to the province level are: B_PRO, M_PRO, E_PRO; the three prediction nodes corresponding to the city level are: B_CITY, M_CITY, E_CITY; the three prediction nodes corresponding to the district level are: B_COUNTY, M_COUNTY, E_COUNTY; the three prediction nodes corresponding to the street level are: B_TOWN, M_TOWN, E_TOWN; the three prediction nodes corresponding to the house number level are: B_DETA, M_DETA, E_DETA.
[0096] Finally, the predicted nodes in each node sequence are connected to the predicted nodes in the adjacent node sequences. There are corresponding weights for the connection lines between different predicted nodes. If the weight is 0, it means that these two predicted nodes are not connected. It can be understood that the training process of the address segmentation model is the process of adjusting these weights.
[0097] Next, on the basis of Figure 3 , in one embodiment, referring to Figure 4 , Figure 4 is the flowchart of step S220 provided by the embodiment of the present application, which specifically includes the following steps S410 to step S450:
[0098] Step S410: Input the training address segments into the node sequences in the order of levels.
[0099] In one embodiment, since each training data segment corresponds to different groups of node sequences, each training data segment is input into the corresponding node sequence in the order of levels of province, city, district, street, and house number.
[0100] Considering that the number of characters contained in not every training data segment is the same as that of the training data segment with the maximum length, assuming that each level corresponds to the first number of node sequences, in one embodiment, referring to Figure 5 , Figure 5 is the flowchart of inputting the training address segments into the node sequences in the order of levels provided by the embodiment of the present application, which specifically includes the following steps S510 to step S540:
[0101] Step S510: If the difference between the first number and the second number is greater than two, at least one supplementary intermediate address training data all of which are zeros is generated based on the difference.
[0102] Step S520: Input the head address training data into the first node sequence;
[0103] Step S530: Input the intermediate address training data and the supplementary intermediate address training data into the second node sequence in turn.
[0104] Step S540: Input the tail address training data into the last node sequence.
[0105] Since the first node sequence corresponds to the head address training data and the last node sequence corresponds to the tail address training data, the number obtained by subtracting two from the first number is the number of node sequences corresponding to the intermediate address training data. At this time, if the difference between the first number and the second number is greater than two, that is, this number is more than the second number, it is necessary to fill the intermediate address training data so that the length of each training data segment is the same.
[0106] For example, the training data segment is specifically "wasde Province", the training data at the start address is "w", the training data at the middle address is "asde", and the training data at the end address is "Province". Assume that the first quantity of the node sequence corresponding to the Province level is 9, the second quantity is 4 at this time, and the difference between the first quantity and the second quantity is 5, which is greater than 2. It is necessary to generate 5 - 2 = 3 supplementary intermediate address training data all of which are zero. It can be understood that the supplementary intermediate address training data is only for aligning different training data segments and does not contain specific meanings. The supplementary intermediate address training data can be spliced behind, in front of, or interspersed in the intermediate address training data. For example, after filling, the intermediate address training data "asde" can be in different ways such as "asde000", "000asde", "a0s0d0e", or "00asde0", etc. This embodiment does not limit this.
[0107] Referring to Figure 6 , Figure 6 is a schematic diagram of the node sequence in an embodiment of the present application. In the figure, taking "wasde Province" as an example, after filling, the intermediate address training data is "asde000". First, "w" is sent to the 1st node sequence, then "a" is sent to the 2nd node sequence, "s" is sent to the 3rd node sequence, "d" is sent to the 4th node sequence, "e" is sent to the 5th node sequence, "0" is sent to the 6th node sequence, "0" is sent to the 7th node sequence, "0" is sent to the 8th node sequence, and "Province" is sent to the 9th node sequence.
[0108] Step S420: Sequentially obtain selection points from the nodes in the node sequence.
[0109] In one embodiment, the selection point is which label the input of the node sequence may be predicted to be during a prediction process. It can be understood that each node sequence will have a corresponding selection point during a prediction process.
[0110] Step S430: Generate a prediction path according to the selection point and obtain the prediction score of the prediction path.
[0111] Step S440: Generate a label path and the corresponding label score based on the segmentation label.
[0112] In one embodiment, referring to Figure 6 , the dashed line shows a prediction path, and there is at least one different selection point between different prediction paths. The solid line shows the label path corresponding to the segmentation label.
[0113] In one embodiment, the training address segment is represented as X = {x1, x2, x3, …, xn}, and the corresponding segmentation label is represented as Y = {y1, y2, y3, …, yn}. At this time, for each label path, there is a label score e S(X,Y) , which is expressed as:
[0114]
[0115] wherein, represents the initial probability that the predicted output of the i-th position node sequence is yi, and represents the transition probability from yi to yi+1. Both the initial probability and the transition probability here are obtained by processing the input sequence through feature extraction and a neural network. For example, a bidirectional LSTM network can be used. For different prediction paths calculate the prediction scores in the same manner as described above
[0116] Step S450: Obtain the likelihood function value according to the label score and the prediction score.
[0117] In one embodiment, referring to Figure 7 , Figure 7 is a flowchart for obtaining the likelihood function value according to the label score and the prediction score provided by an embodiment of the present application, specifically including the following steps S710 to step S720:
[0118] Step S710: Traverse the node sequence, generate multiple prediction paths, and calculate the prediction scores of each prediction path.
[0119] Step S720: Accumulate the prediction scores to obtain an accumulated value, and use the ratio of the label score to the accumulated value as the likelihood function value.
[0120] Among them, all possible prediction paths are traversed, the prediction scores corresponding to each prediction path are calculated, and then the prediction scores are accumulated to obtain an accumulated value. Finally, the ratio of the label score to the accumulated value is used as the likelihood function value. Therefore, the likelihood function value P(Y|X) is expressed as:
[0121]
[0122] In addition, the likelihood function value can also be expressed as a log-likelihood function:
[0123]
[0124] Step S230: Maximize the likelihood function value, adjust the model weights of the address segmentation model until a trained address segmentation model is obtained.
[0125] In one embodiment, the training objective of the address segmentation model is that for the input X, it can output Y with the highest probability. Therefore, the model weights of the address segmentation model can be adjusted to maximize the likelihood function value. After the iteration termination condition is met, the trained address segmentation model is obtained. The iteration termination conditions here include: the number of iterations reaches the preset number of iterations, or the model performance of the address segmentation model after iteration meets the preset performance requirements. This embodiment does not make specific limitations on the iteration termination conditions.
[0126] In one embodiment, after obtaining the trained address segmentation model, it is serially deployed to the distributed file system HDFS. Specifically, it includes the following steps: First, save the trained address segmentation model as a file. Tools or libraries can be used to save information such as model parameters, weights, and configurations as files. Then, use the corresponding Hadoop command-line tool or programming language API to connect to the distributed file system HDFS, and then use the hdfs dfs - mkdir command to create a directory in the distributed file system HDFS to store the model file of the serialized address segmentation model. Finally, upload the model file to the directory in the distributed file system HDFS. After the upload is completed, use the hdfs dfs - ls command to verify whether the file has been successfully uploaded to the specified directory in HDFS. After verification, it can be further used on the distributed file system HDFS. For example, broadcast the pre-trained address segmentation model in the distributed file system to the distributed data set RDD so that the address segmentation model can be used to perform parallel address segmentation prediction on data shards.
[0127] Step S130: Input the initial logistics address data in the data shard into the address segmentation model for address segmentation to obtain segmented address data.
[0128] Among them, the segmented address data includes multiple address segments, and each address segment is composed of a start address data, at least one middle address data, and an end address data.
[0129] It can be understood that when the initial logistics address data needs to be filled, there is also a corresponding likelihood function value for the supplementary intermediate address training data when calculating the likelihood function value, and the likelihood function value corresponds to the segmentation label. That is, for example, if the initial logistics address data is "abcde Province", after filling, it becomes "abcde000 Province", and the complete segmented address data is: {B_PRO|--|a, M_PRO|--|b, M_PRO|--|c, M_PRO|--|d, M_PRO|--|0, M_PRO|--|0, M_PRO|--|0, E_PRO|--|Province}. Additionally, the segmentation label of the likelihood function value of the supplementary intermediate address training data can also be set to null in the segmented address data: {B_PRO|--|a, M_PRO|--|b, M_PRO|--|c, M_PRO|--|d, M_PRO|--|null, M_PRO|--|null, M_PRO|--|null, E_PRO|--|Province}.
[0130] Step S140: Obtain an address base table based on the segmented address data, and select a target address based on the address base table.
[0131] In one embodiment, a large number of obtained segmented address data are written into a database table to obtain an address base table. Considering that the information at the street level may include more detailed: xx Community, xx Township, xx Road, etc., and the house number level may include: road number and community number, etc., so similar to the setting of the number of node sequences, the address base table includes all possible address columns, and only the segmented address data needs to be filled into the corresponding columns in the address base table. Assuming there are n segmented address data, the address base table is as shown in Table 1 below:
[0132] Province City District Street Community Road Road Number Residential Area Number Province a1 City b1 District c1 Road f1 No. 98 Province a1 City b1 District c2 Town d1 Road f2 No. 47 Province a2 City b2 District c3 Province a2 City b2 District c3 Province a3 City b3 District c4 Community e1 Province a3 City b3 District c4 Community e1 Province a3 City b3 District c4 Road f3 Province a3 City b3 County c5 Town d2 Road f4 38 No. 2 Province a4 City b4 County c6 Street d3 …
[0133] Each row in Table 1 corresponds to a segmented address data, where the blank places represent Null. After obtaining the address base table, user portraits can be obtained to get group portrait features, and different management strategies can be specified according to the group portrait features.
[0134] In one embodiment, referring to Figure 8 , Figure 8 is the flowchart of obtaining an address base table based on the segmented address data and selecting a target address provided by the embodiment of the present application, specifically including the following steps S810 to step S840:
[0135] Step S810: Use the segmented address data and the corresponding time information to determine the address base table.
[0136] In one embodiment, taking Table 1 as an example, the address base table is updated by adding the time information of each segmented address data. According to this time information, the acquisition time of the segmented address data can be obtained, and the acquisition time corresponds to the user's shipping time.
[0137] Step S820: Screen the segmented address data in the address base table according to a preset level to obtain a plurality of first screening regions.
[0138] In one embodiment, the preset region refers to the screening range. For example, when screening according to the provincial level, the segmented address data of the same province in the address base table is used as the first screening region, and different first screening regions correspond to different provinces. It can also be screened in the ways of city, district, street, community, road, road number, community number, etc. The preset region here is set according to the actual population portrait requirements.
[0139] Step S830: According to the time information of the segmented address data, screen the segmented address data located in the preset time period in the first screening region as candidate address data.
[0140] In one embodiment, assuming that the screening is carried out at the community level, refer to Figure 9 , Figure 9 is a schematic diagram of the first screening region provided by the embodiment of the present application. Figure 9 A plurality of first screening regions are schematically shown in it. The segmented address data in each first screening region is located in the same community, and the segmented address data is schematically shown by hollow circles. Then, a preset time is selected, for example, 19:00 - 23:00. According to the time information of the segmented address data in the first screening region, the segmented address data that meets the preset time is screened in each first screening region. The screened segmented address data constitutes the candidate address data, which is represented by black solid circles in the figure.
[0141] Step S840: Cluster the candidate address data, and obtain the target address according to the clustering result.
[0142] Then, the candidate address data can be clustered according to the division of geographical communities. Each clustering result corresponds to a target address, and the clustering result indicates the concentrated location of the candidate address data. In addition, the clustering results can also be screened according to the number of candidate address data in the clustering results. Only the clustering results with the number of candidate address data greater than the preset number are regarded as a target address. For example, refer to Figure 9, if the preset quantity is 10, then only two of the clustering results can be used as target addresses, such as Community A and Community B in the figure. Taking Community A as an example, it being used as a target address indicates that there are a relatively large number of outgoing parcels in Community A during the night period from 19:00 to 23:00. Community A can be circled as a community with the potential for night-time parcel collection. Then, night-time parcel collectors can be additionally assigned to these communities with the potential for night-time parcel collection, so as to be able to collect parcels in Community A in a timely manner, so as to avoid the loss of this part of the outgoing parcel volume due to staffing reasons.
[0143] Referring to Figure 10 , Figure 10 is the overall schematic diagram of the logistics address data processing method provided by the embodiment of the present application. First, data cleaning is performed on 1.6 billion pieces of original address corpus to remove unusable data. Then, the data after cleaning is labeled to obtain a training data set. Next, the address segmentation model is trained using the training data set under the Apache Spark framework. After training is completed, the trained address segmentation model is serialized and deployed to the distributed file system HDFS. When it is necessary to process daily logistics address data, the pre-trained address segmentation model is read from the distributed file system, and the address segmentation model is broadcast to the data set for address segmentation prediction to create or update the address base table. After verification, the logistics address data processing method of the embodiment of the present application has a prediction time of only 0.04 ms for an average piece of logistics address data, and the accuracy rate reaches 96%, greatly improving the efficiency of logistics address processing.
[0144] The technical solution provided by the embodiment of the present application obtains multiple initial logistics address data to be processed, and stores the initial logistics address data in slices to obtain multiple data slices. The pre-trained address segmentation model is read from the distributed file system, and the address segmentation model is broadcast to each data slice. Then, the initial logistics address data in the data slice is input into the address segmentation model for address segmentation to obtain segmented address data. Among them, the segmented address data includes multiple address segments, and each address segment is composed of a start address data, at least one middle address data, and an end address data. Finally, an address base table is obtained according to the segmented address data, and target addresses are selected based on the address base table. In the embodiment of the present application, aiming at the characteristic of a large amount of logistics address data, data parallel processing is first carried out using the distributed file system, which greatly improves the data processing efficiency. In addition, when segmenting the logistics address, based on the address characteristics of the logistics address, it is segmented into multiple address segments, and at the same time each address segment is divided into three parts. By this labeling method, while retaining the address information, the complexity of address labeling is reduced, further improving the data processing efficiency of the logistics address.
[0145] The embodiment of the present application also provides a logistics address data processing device, which can implement the above-mentioned logistics address data processing method. Referring to Figure 11, the device includes:
[0146] An address data acquisition module 1110: configured to acquire a plurality of initial logistics address data to be processed, and store the initial logistics address data in a data set in a fragmented manner; the data set includes a plurality of data fragments.
[0147] A model acquisition module 1120: configured to read a pre-trained address segmentation model from a distributed file system, and broadcast the address segmentation model to the data set.
[0148] An address segmentation module 1130: configured to input the initial logistics address data in the data fragment into the address segmentation model for address segmentation to obtain segmented address data; the segmented address data includes a plurality of address segments, and each address segment is composed of a starting address data, at least one intermediate address data, and an ending address data.
[0149] A target address selection module 1140: configured to obtain an address base table according to the segmented address data, and select a target address based on the address base table.
[0150] The specific implementation manner of the logistics address data processing device in this embodiment is basically the same as that of the above-mentioned logistics address data processing method, and will not be elaborated here.
[0151] An embodiment of the present application further provides an electronic device, including:
[0152] At least one memory;
[0153] At least one processor;
[0154] At least one program;
[0155] The program is stored in the memory, and the processor executes the at least one program to implement the above-mentioned logistics address data processing method of the present application. The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, abbreviated as PDA), an in-vehicle computer, etc.
[0156] Please refer to Figure 12 , Figure 12 which illustrates the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0157] The processor 1201 can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0158] The memory 1202 can be implemented in forms such as ROM (Read Only Memory), a static storage device, a dynamic storage device, or RAM (Random Access Memory). The memory 1202 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1202 and are called by the processor 1201 to execute the logistics address data processing method of the embodiments of the present application;
[0159] The input / output interface 1203 is used to implement information input and output;
[0160] The communication interface 1204 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and
[0161] The bus 1205 transmits information between various components of the device (such as the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204);
[0162] Among them, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 achieve communication connections with each other inside the device through the bus 1205.
[0163] The embodiments of the present application also provide a storage medium. The storage medium is a storage medium that stores a computer program, and when the computer program is executed by a processor, the above-mentioned logistics address data processing method is implemented.
[0164] As a non-transitory storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0165] In the method, apparatus, device, and storage medium for processing logistics address data provided by the embodiments of the present application, multiple initial logistics address data to be processed are obtained, and the initial logistics address data are sliced and stored to obtain multiple data slices. A pre-trained address segmentation model is read from a distributed file system, and the address segmentation model is broadcast to each data slice. Then, the initial logistics address data in the data slices are input into the address segmentation model for address segmentation to obtain segmented address data, where the segmented address data includes multiple address segments, and each address segment is composed of a start address data, at least one middle address data, and an end address data. Finally, an address base table is obtained according to the segmented address data, and a target address is selected based on the address base table. In the embodiments of the present application, aiming at the characteristic of a large amount of logistics address data, parallel processing of data is first performed using a distributed file system, which greatly improves the data processing efficiency. In addition, when segmenting the logistics address, based on the address characteristics of the logistics address, it is segmented into multiple address segments, and at the same time, each address segment is divided into three parts. By this annotation method, while retaining the address information, the complexity of address annotation is reduced, and the data processing efficiency of the logistics address is further improved.
[0166] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0167] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.
[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0169] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0170] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0171] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0172] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0173] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0174] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0175] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0176] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A method for processing logistics address data, characterized in that, Including: Obtain a plurality of initial logistics address data to be processed, and store the initial logistics address data in a data set in a sliced manner; The data set includes a plurality of data slices; Read a pre-trained address segmentation model from a distributed file system, and broadcast the address segmentation model to the data set; Input the initial logistics address data in the data slices into the address segmentation model for address segmentation to obtain segmented address data; the segmented address data includes a plurality of address segments, and each address segment is composed of a first address data, at least one intermediate address data, and a last address data; Obtain an address base table according to the segmented address data, and select a target address based on the address base table.
2. The logistics address data processing method according to claim 1, wherein The training process of the address segmentation model includes the following steps: Obtain an address training data set; the address training data set includes a plurality of training data, the training data includes a plurality of levels of training address segments, and each training address segment includes a corresponding segmentation label; Input the corresponding training address segment and the segmentation label into the address segmentation model in the order of the levels for prediction to obtain the likelihood function value of the training address segment; Maximize the likelihood function value, and adjust the model weights of the address segmentation model until the trained address segmentation model is obtained.
3. The logistics address data processing method according to claim 2, wherein The address segmentation model is a distributed probability model; the distributed probability model includes a plurality of the node sequences; each node sequence includes three prediction nodes; the inputting the training address segment and the segmentation label into the address segmentation model in the order of the levels for prediction to obtain the likelihood function value of the training address segment includes: Input the training address segment into the node sequence in the order of the levels; Successively obtain selection points from the prediction nodes in the node sequence; Generate a prediction path according to the selection points, and obtain the prediction score of the prediction path; Generate a label path and a corresponding label score based on the segmentation label; Obtain the likelihood function value according to the label score and the prediction score.
4. The logistics address data processing method according to claim 3, wherein The obtaining the likelihood function value according to the label score and the prediction score includes: Traverse the node sequence, generate a plurality of prediction paths, and calculate the prediction scores of each prediction path; Accumulate the prediction scores to obtain an accumulated value, and use the ratio of the label score to the accumulated value as the likelihood function value.
5. The logistics address data processing method according to claim 3, wherein The levels successively include: provincial level, municipal level, district level, street level, and house number level.
6. The logistics address data processing method according to claim 3, wherein Each level corresponds to a first number of the node sequences, and the training data segment is composed of a first address training data, a second number of intermediate address training data, and a last address training data; The inputting the training address segment into the node sequence in the order of the levels includes: If the difference between the first number and the second number is greater than two, generate at least one supplementary intermediate address training data all of which are zero based on the difference; Input the first address training data into the first node sequence; Input the intermediate address training data and the supplementary intermediate address training data into the second node sequence in sequence; Input the tail address training data into the last node sequence.
7. The logistics address data processing method according to any one of claims 1 to 6, characterized in that The obtaining of the address base table according to the segmented address data and the selection of the target address based on the address base table include: Determine the address base table by using the segmented address data and the corresponding time information; Screen the segmented address data in the address base table according to a preset area to obtain a plurality of first screening areas; According to the time information of the segmented address data, screen the segmented address data located in a preset time period in the first screening area as candidate address data; Cluster the candidate address data, and obtain the target address according to the clustering result.
8. A logistics address data processing device, characterized in that, Include: Address data acquisition module: used to acquire a plurality of initial logistics address data to be processed, and store the initial logistics address data in a data set in a segmented manner; The data set includes a plurality of data segments; Model acquisition module: used to read a pre-trained address segmentation model from a distributed file system, and broadcast the address segmentation model to the data set; Address segmentation module: used to input the initial logistics address data in the data segments into the address segmentation model for address segmentation to obtain segmented address data; the segmented address data includes a plurality of address segments, and each address segment is composed of a start address data, at least one intermediate address data, and a tail address data; Target address selection module: used to obtain an address base table according to the segmented address data, and select a target address based on the address base table.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the logistics address data processing method according to any one of claims 1 to 7.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the logistics address data processing method according to any one of claims 1 to 7.