Address search method and device based on deep learning, storage medium and server
By converting text addresses, spatial coordinates, surrounding POIs and street scene pictures into multimodal vectors, combined with semantic enhancement information, the problems of high computational complexity and low efficiency in traditional address search methods are solved, and efficient and accurate address query and matching are achieved, which is suitable for applications such as map navigation and logistics distribution.
Patent Information
- Application Number
- CN202510764124.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
When traditional text address search methods deal with complex and diverse address data, there are problems such as high computational complexity, low efficiency, and inability to effectively utilize the spatial topological relationship of addresses, which is difficult to meet the needs of modern applications.
By converting text addresses, spatial coordinates, peripheral POIs and street scene pictures into multimodal vectors, combining semantic enhancement information, multimodal vectors are generated for address search, and deep learning technology is used to improve the accuracy and intelligence of address matching.
It realizes more accurate and efficient address query, provides intuitive query results, supports address matching and personalized search in complex scenarios, and is suitable for map navigation, logistics and distribution and other fields.
Smart Images

Figure CN120277168A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of logistics information processing, and in particular, to an address search method, device, storage medium, and server based on deep learning. Background Art
[0002] In today's digital age, with the rapid development of fields such as e-commerce, logistics distribution, and geographic information system (GIS), the accurate processing and efficient utilization of address information have become increasingly important. As the key information for positioning and describing geographical locations, the accuracy and integrity of addresses directly affect the efficiency and quality of related services. However, when faced with increasingly complex and diverse address data, traditional text address search and processing methods have gradually revealed many technical problems and are difficult to meet the needs of modern applications.
[0003] Traditional text addresses are usually presented in natural language form. Different people may use different words, grammatical structures, and expressions when describing the same address. The city is developing and changing rapidly, with new roads and buildings emerging continuously, and the old address information may change accordingly. The update of traditional address databases often relies on manual collection and entry, which is not only inefficient but also prone to the situation of untimely data update. Traditional text addresses exist in an unstructured natural language form and lack a unified standard and format. Address information from different regions and sources has significant differences in expression methods, field orders, etc., which makes it difficult for computers to directly process and analyze them effectively. In traditional text address search, a string matching-based method is usually adopted, that is, the similarity between addresses is determined by comparing each character one by one. This method has a high computational complexity and low search efficiency when dealing with large-scale address data. Traditional address databases mainly focus on the text information of addresses and ignore the spatial topological relationships between addresses. In many practical applications, such as path planning and regional analysis, understanding the spatial location relationships between addresses is crucial. However, traditional text address search systems cannot directly provide the spatial topological information between addresses, which requires additional complex data processing and calculations when performing relevant spatial analysis, increasing the complexity and computational cost of the system.
[0004] In summary, traditional text address search and processing methods have many technical problems in semantic understanding, data update, structured processing, matching efficiency, and utilization of spatial relationships, and are difficult to meet the requirements of modern applications for accurate and efficient processing of address information. Summary of the Invention
[0005] The embodiments of this application provide an address search method, device, storage medium, and terminal device based on deep learning, which can solve the problems of inaccurate and incomplete text address search results in the prior art. The technical solutions are as follows: In a first aspect, an embodiment of the present application provides a method for address search based on deep learning. The method includes: Obtaining various multimodal data in a multimodal address database, where the multimodal data includes: text address, spatial coordinates, surrounding POIs, and street view pictures; Converting the text address into a text address vector, converting the spatial coordinates into a spatial vector, converting the surrounding POIs into POI vectors, and converting the street view pictures into street view picture vectors; Performing weighted summation on the text address vector, the spatial vector, the POI vector, and the street view picture vector to generate a multimodal vector; Dividing the text address into multiple entities, querying corresponding semantic enhancement information in a knowledge base using the entities, and encoding the queried semantic enhancement information to generate a semantic enhancement vector; Concatenating the semantic enhancement vector and the multimodal vector to generate a final address vector, and writing the final address vector into an address vector library; Querying in the address vector library according to query information input by a user on an interaction interface, and displaying a query result on the interaction interface.
[0006] In a second aspect, an embodiment of the present application provides a device for address search based on deep learning. The device includes: An obtaining unit, configured to obtain various multimodal data in a multimodal address database, where the multimodal data includes: text address, spatial coordinates, surrounding POIs, and street view pictures; A conversion unit, configured to convert the text address into a text address vector, convert the spatial coordinates into a spatial vector, convert the surrounding POIs into POI vectors, and convert the street view pictures into street view picture vectors; A generating unit, configured to perform weighted summation on the text address vector, the spatial vector, the POI vector, and the street view picture vector to generate a multimodal vector; divide the text address into multiple entities, query corresponding semantic enhancement information in a knowledge base using the entities, encode the queried semantic enhancement information to generate a semantic enhancement vector; concatenate the semantic enhancement vector and the multimodal vector to generate a final address vector, and write the final address vector into an address vector library; A query unit, configured to query in the address vector library according to query information input by a user on an interaction interface, and display a query result on the interaction interface.
[0007] In a third aspect, an embodiment of the present application provides a computer storage medium storing multiple instructions adapted to be loaded and executed by a processor to perform the above method steps.
[0008] In a fourth aspect, an embodiment of the present application provides a server, which may include: a processor and a memory; wherein, the memory stores a computer program adapted to be loaded and executed by the processor to perform the above method steps.
[0009] The beneficial effects brought by the technical solutions provided by some embodiments of the present application at least include: By integrating multi-modal data such as text addresses, spatial coordinates, surrounding POIs, and street view images, the address information can be represented more comprehensively and accurately, improving the richness and accuracy of address descriptions. Converting the multi-modal data into vector form facilitates mathematical operations and similarity comparisons, providing an efficient data basis for subsequent address queries and matches. By dividing the text address into multiple entities and querying the corresponding semantic enhancement information in the knowledge base, the semantic representation of the address is further enriched, improving the accuracy and intelligence of address matching. Generating a query text address vector based on the query text address input by the user and quickly querying the candidate address vector with the highest similarity in the address vector library realizes an efficient address query function. Displaying the text address, spatial coordinates, surrounding POIs, and street view images corresponding to the queried candidate address vector on the interaction interface provides the user with an intuitive and comprehensive query result, facilitating the user's selection and judgment. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 is a schematic diagram of the network architecture provided by the present application; Figure 2 is a schematic diagram of the process of the address search method based on deep learning provided by the present application; Figure 3 is a schematic diagram of the process of dynamically adjusting weights based on context information in the present application; Figure 4 is a schematic diagram of the structure of an address search device based on deep learning provided by the present application; Figure 5 is a schematic diagram of the computer storage medium provided by the present application; Figure 6It is a schematic structural diagram of a server provided by this application. Specific Embodiments
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0013] It should be noted that the address search method based on deep learning provided by this application is generally executed by a server. Correspondingly, the address search device based on deep learning is generally set in the server.
[0014] Figure 1 An exemplary network architecture is shown that can be applied to the address search method based on deep learning or the address search device based on deep learning of this application.
[0015] As Figure 1 shown, the network architecture may include: a terminal device 501 and a server 600. Communication can be carried out between the terminal device 501 and the server 600 through a network, and the network is a medium for providing a communication link between the above-mentioned various units. The network may include various types of wired communication links or wireless communication links. For example, the wired communication links include optical fibers, twisted pairs, or coaxial cables, etc., and the wireless communication links include Bluetooth communication links, Wireless-Fidelity (Wi-Fi) communication links, or microwave communication links, etc.
[0016] Among them, a vector address library is deployed in the server 600, the terminal device 501 displays an interaction interface, and the user inputs a query text in the interaction interface, and address search is performed in the vector address library based on the input query text.
[0017] The core concept of this application lies in: 1. Multimodal address vector search: Combine address text with other modal data (such as map coordinates, POI information, images, etc.) to construct a multimodal address vector. By fusing text, spatial, and visual information, the accuracy and richness of address search are improved. For example, when a user inputs an address, the system not only matches the text vector but also combines map coordinates and surrounding POI information to provide more accurate search results. This method is particularly suitable for complex scenarios, such as fuzzy address matching or cross-regional search.
[0018] 2. Dynamic context-aware search: Use context information (such as user historical search records, geographical location, time, etc.) to dynamically adjust the address vector search strategy. By introducing a context-aware model, the system can understand the user's real-time needs and provide personalized search results. For example, when a user searches for "Train Station in City A", the system gives priority to recommending nearby transportation hubs or related services according to the user's current location and historical preferences. This method improves the intelligence of the search and the user experience.
[0019] 3. Hierarchical Address Vector Search: Decompose the address vector by levels (such as country, province, city, street, etc.) to construct a hierarchical vector representation. During the search, first match the high-level vector (such as city), and then gradually refine to the low-level (such as street). This method can significantly reduce the search space and improve the search efficiency, especially suitable for large-scale address databases. At the same time, hierarchical search supports fuzzy matching, and even if the input address is incomplete, relevant results can be returned.
[0020] 4. Semantic-Enhanced Address Vector Search: Enhance the semantic representation of the address vector by introducing external knowledge bases (such as geographical information bases, administrative division bases). For example, fuse the associated information such as "Area A" with "Village A", "University A", etc. to generate a vector with richer semantics. During the search, the system can not only match the literal information but also understand the implicit semantics of the address to provide more relevant search results. This method is particularly suitable for complex queries, such as "addresses near the science and technology park".
[0021] 5. Real-Time Incremental Learning and Update: Introduce a real-time incremental learning mechanism into the address vector search system to dynamically update the address vector representation. For example, when new address data is added or old address information changes, the system can adjust the vector model in real time to ensure the timeliness and accuracy of search results. This method is applicable to scenarios where address data changes frequently, such as the logistics and food delivery industries, and can effectively cope with the challenges brought by address updates.
[0022] This technical solution is not only applicable to general address searches (such as map navigation, property queries), but also can be extended to scenarios such as logistics distribution, O2O services (such as food delivery, taxi-hailing), and smart cities, and can meet the business requirements of different fields through customized knowledge bases and weight configurations.
[0023] It should be noted that the terminal device 501 and the server 600 can be hardware or software. When the terminal device 501 and the server 600 are hardware, they can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the terminal device 501 and the server 600 are software, they can be implemented as multiple software or software modules (for example, used to provide distributed services), or as a single software or software module, and no specific limitation is made here.
[0024] Various communication client applications can be installed on the terminal device of this application, such as: video recording applications, video playback applications, voice interaction applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0025] The terminal device can be either hardware or software. When the terminal device is hardware, it can be various terminal devices with a display screen, including but not limited to smartphones, tablets, laptop portable computers, desktop computers, and so on. When the terminal device is software, it can be installed in the terminal devices listed above. It can be implemented as multiple software or software modules (for example, used to provide distributed services), or it can be implemented as a single software or software module, which is not specifically limited here.
[0026] When the terminal device is hardware, a display device and a camera can also be installed on it. The display device can be various devices that can implement the display function, and the camera is used to collect video streams. For example, the display device can be a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), etc. Users can use the display device on the terminal device to view information such as text, pictures, and videos displayed.
[0027] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0028] are only illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 2 Next, in combination with the attached Figure 1 drawings, the deep learning-based address search method provided by the embodiments of the present application will be introduced in detail. Among them, the deep learning-based address search device in the embodiments of the present application can be
[0029] the terminal device shown in Figure 2 . Please refer to Figure 2 , which is a schematic flowchart of a deep learning-based address search method provided by the embodiments of the present application. As shown, the method of the embodiments of the present application can include the following steps:
[0030] Among them, the server retrieves data from the multi-modal address database through a database interface (such as JDBC, ODBC, or RESTful API), supporting local deployment or cloud storage (such as MySQL, MongoDB, AWS S3). When retrieving data, the server dynamically generates a query statement according to preset conditions (such as data update time, data integrity), for example, filtering records updated within the last 30 days, or excluding records with empty key fields.
[0031] Text address: Check whether it contains garbled characters (such as non-Chinese and English characters), special symbols (such as!@#$%^&*), or the length exceeds the threshold (such as 200 characters). If it is abnormal, replace illegal characters through regular expressions or mark it as invalid.
[0032] Spatial coordinates: Verify whether the longitude and latitude are within a reasonable range (longitude [-180, 180], latitude [-90, 90]), and unify the coordinate system (such as converting the GCJ-02 Mars coordinate system to WGS84).
[0033] Surrounding POIs: Check whether the POI list is empty. If it is empty, fill in the default value (such as "unknown landmark") or mark it as missing; if the POI name contains rare words, split it through a word segmentation tool (such as IK Analyzer) and then supplement it by associating with the knowledge base.
[0034] Street view image: Verify whether the image file exists (such as checking whether the file path is valid), and whether the file format is supported (such as JPEG, PNG). If it is missing, replace it with the default placeholder image (such as a gray background image).
[0035] For example: The server retrieves a record from the database: Text address: "No. 1, Jiacun Street, Jia District, Jia City @#¥%" (including garbled characters) Spatial coordinates: 116.3184, 39.9642 (in WGS84 format, no conversion required) Surrounding POIs: [] (empty list) Street view image: changan_zhongcun_001.jpg (file exists) Cleaning result: The text address is corrected to "No. 1, Jiacun Street, Jia District, Jia City"; The POI list is filled with the default value "unknown landmark"; The verification of the image file passes, and the original path is retained.
[0036] S202. Convert the text address into a text address vector, convert the spatial coordinates into a spatial vector, convert the surrounding POIs into a POI vector, and convert the street view image into a street view image vector.
[0037] Among them, when the server converts each modality data into a high-dimensional vector, it is necessary to select an encoding model suitable for the scenario: Text address vector: Use a pre-trained language model (such as BERT, RoBERTa) to extract semantic features and output a 768-dimensional vector.
[0038] Spatial vector: Normalize the longitude and latitude to the range [0, 1] (e.g., through Min-Max scaling), or generate a binary hash value using GeoHash encoding (e.g., wx4g09c6).
[0039] POI vector: After tokenizing the POI list, generate word vectors through TF-IDF or Word2Vec, and then aggregate them (e.g., average pooling) into a fixed-dimensional vector (e.g., 300-dimensional).
[0040] Street view image vector: Use a convolutional neural network (e.g., ResNet-50, EfficientNet) to extract global features and output a 2048-dimensional vector.
[0041] For example: The text address "No. 1000, Jia Mingzhu (building name) Loop, Jia New Area, Jia City" is encoded into a 768-dimensional vector by BERT; The spatial coordinates (121.4737, 31.2304) are normalized to [0.674, 0.348]; The POI list ["Jia City Central Building"] generates two 300-dimensional vectors through Word2Vec, and after averaging, a 300-dimensional POI vector is obtained; The street view image Abc_001.jpg is extracted as a 2048-dimensional vector by ResNet-50.
[0042] S203. Perform weighted summation on the text address vector, the spatial vector, the POI vector, and the street view image vector to generate a multi-modal vector.
[0043] Among them, the server performs weighted summation on each modal vector, and the weights can be assigned according to the business priority (e.g., text 0.4, space 0.3, POI 0.2, image 0.1); or the contribution degree of each modality to the target task (e.g., address matching) can be inferred by training a classification model (e.g., SVM, XGBoost). After weighting, the vector needs to be L2-normalized to ensure that the vector distributions of different modalities are consistent.
[0044] For example: Assume the weights are text 0.4, space 0.3, POI 0.2, image 0.1, then: Multi-modal vector = 0.4 × text vector + 0.3 × spatial vector + 0.2 × POI vector + 0.1 × image vector, and after normalization, a 512-dimensional final vector is obtained.
[0045] S204. Divide the text address into multiple entities, query the corresponding semantic enhancement information in the knowledge base using the entities, and encode the queried semantic enhancement information to generate a semantic enhancement vector.
[0046] Among them, the server splits the text address into entities (such as "City A" and "New Area A"), and queries the associated information through a knowledge base (such as OpenStreetMap): Use a NER model (such as Stanford NER, LTP) to identify entity types (such as cities, regions, landmarks); query the knowledge base to obtain entity attributes (such as the area and population of "New Area A") or relationships (such as nearby subway stations), and convert the complemented semantic information into vectors through a graph neural network (such as GAT, GCN) or a knowledge embedding model (such as TransE, RotatE).
[0047] For example: The entity "New Area A, City A" is associated with the knowledge base information: "Population: 5.68 million" and "GDP: 1.6 trillion yuan"; the semantic enhanced vector is encoded into a 256-dimensional vector through GAT.
[0048] S205. Generate the final address vector by concatenating the semantic enhanced vector and the multi-modal vector, and write the final address vector into the address vector library.
[0049] Among them, the server concatenates the semantic enhanced vector and the multi-modal vector (such as directly connecting or fusing through an attention mechanism) to generate the final address vector. When writing the vector into the vector library, the index type needs to be selected: use the FlatIP index (inner product calculation) for exact matching; for approximate nearest neighbor (ANN): use the IVF Flat or HNSW index to accelerate large-scale queries.
[0050] For example: The dimension of the concatenated vector: 512 (multi-modal) + 256 (semantic enhanced) = 768 dimensions; store it in the Faiss vector library, and the index type is IVF1024, Flat.
[0051] S206. Query in the address vector library according to the query information input by the user on the interaction interface, and display the query results on the interaction interface.
[0052] Among them, the server receives the user's query (such as the text "near the Oriental Building" or a picture), converts it into a vector, and then searches in the vector library. When querying text, generate a vector through BERT; when querying a picture, extract the feature vector through ResNet-50. Use the inner product or cosine similarity to sort the candidate results. Filter the results according to business rules (such as spatial distance threshold, POI relevance). Sort the results in descending order of similarity, and display the text address, POI label, thumbnail, and distance information.
[0053] For example: The user inputs the text "Ping'anli" (the specific name of a small area of a certain location), and the server generates a query vector to find the top 3 addresses with the highest similarity in the vector library. The interface displays: "No. 1000, Ping'anli Ring Road, Jia New Area, Jia City (similarity 0.95, distance 500 meters)", "Jia City Central Building (similarity 0.90, distance 800 meters)", etc., along with a preview of the street view and a navigation link.
[0054] In some possible embodiments of the present application, the conversion of the text address into a text address vector includes: Decomposing the text address into a country address, a province / municipality address, a city / district address, and a street / house number address; Converting the addresses at each level into a country address vector, a province / municipality address vector, a city / district vector, and a street address vector according to a pre-trained PromptCSE-Base model, and splicing the address vectors at each level to obtain a text address vector.
[0055] Specifically, the server selects an address parsing tool according to the data source and the accuracy requirement. The address parsing tool can be: Regular expression: applicable to addresses with a standardized format (such as the "province-city-district-street" structure).
[0056] NLP model: such as an address parsing model based on BERT, which can process unstructured addresses.
[0057] Third-party API: such as the address parsing API of a certain map, which supports the splitting of complex addresses.
[0058] The server determines the country through keyword matching (such as "Country A") or geocoding.
[0059] Province / municipality address: identifying the name of the provincial administrative region (such as "Province A", "City A").
[0060] City / district address: extracting the name of the city or district (such as "Chang'an District").
[0061] Street / house number address: extracting the remaining part as the street and house number (such as "No. 103, Sports West Road").
[0062] It should be noted that if a certain level cannot be parsed (such as the province is missing in the address), a default value (such as "unknown province") may be filled or marked as missing data. For ambiguous addresses (such as "Zhongshan Road" may exist in multiple cities), context analysis or manual intervention is performed.
[0063] The server loads the pre-trained PromptCSE-Base model from the model repository, which may be stored in the local file system or cloud object storage. The decomposed address levels (country, province, city, street) are input into the model respectively, and the model generates the corresponding vectors through a multi-layer Transformer encoder. The vector dimension may be 768 or 1024, depending on the model structure. The vectors of the four levels are concatenated in a fixed order (such as country first and street second) to form a complete text address vector. The concatenation method may be direct connection or weighted connection (such as country vector weight 0.2, street vector weight 0.4). The concatenated vector is normalized (such as L2 normalization) to ensure that the vector length is consistent. Metadata (such as vector source, generation time) can be attached when storing vectors.
[0064] For example: Enter "No. 103, Tiyuxi Road, District A, City A": Country vector: [0.1, 0.3, ..., -0.2] (country A); Province vector: [0.2, -0.1, ..., 0.4] (Province A); City vector: [-0.3, 0.5, ..., 0.1] (Area A); Street vector: [0.4, -0.2, ..., 0.3] (No. 103, Tiyuxi Road); The dimension of the concatenated text address vector is the sum of the lengths of the four vectors (e.g. 768×4=3072).
[0065] In some embodiments of the present application, the process of converting the surrounding POIs to connected POIs includes: converting each POI included in the surrounding POIs into a corresponding vector, and averaging the converted vectors to obtain a POI vector.
[0066] Specifically, the server removes duplicates and standardizes the POI list (such as merging "KenKen" with "肯肯"), and filters invalid POIs (such as "unknown" and "other"). Use a word embedding model (such as Word2Vec, FastText) or a pre-trained language model (such as BERT) to convert each POI name into a vector. Take the arithmetic mean of all POI vectors to obtain the POI vector. If the POI list is empty, use the default vector (such as an all-zero vector) or skip this step. Smooth the POI vector (such as weighted averaging, assigning weights according to the importance of POI).
[0067] For example: POI list ["Jiacheng", "Jia Square", "Jiadong Station"]: "City A" vector: [0.5, -0.2, ..., 0.1]; "Vector of 'A Square': [0.3, 0.4, ..., -0.1]; "Vector of 'East Station of A': [-0.1, 0.6, ..., 0.2]; Average POI vector: [0.23, 0.27, ..., 0.07].
[0068] In some possible embodiments of the present application, the query information is a query text address; Among them, the query in the address vector library according to the query information input by the user on the interaction interface and the display of the query result on the interaction interface include: Obtain the query text address input by the user on the interaction interface; Decompose the query text address into a country address, a province / municipality address, a city / district address, and a street / house number address according to a preset rule; Call the pre-trained PromptCSE-Base model to convert the addresses at each decomposed level into corresponding address vectors respectively; Concatenate the address vectors at each level to obtain a text address vector; Perform weighted summation according to the default space vector, the default POI vector, the default street view image vector, and the concatenated text address vector to obtain a multi-modal vector; Divide the query text address into multiple entities, query corresponding semantic enhancement information in the knowledge base for each entity, and encode the semantic enhancement information to generate a semantic enhancement vector; Concatenate the generated semantic enhancement vector and the multi-modal vector obtained by weighted summation to obtain the query text address vector corresponding to the query text address; Query several candidate address vectors with the highest similarity in the address vector library according to the query text address vector; Display the text addresses, spatial coordinates, surrounding POIs, and street view images corresponding to the several candidate address vectors on the interaction interface.
[0069] Specifically, the server generates a query text address vector according to the query text address input by the user on the interaction interface.
[0070] Among them, clean the query address input by the user (such as removing punctuation marks and unifying case), and verify whether the query is empty or has an abnormal format.
[0071] Use the same PromptCSE-Base model as above to generate a vector representation of the query text. The vector dimension is the same as the final address vector generated above. Normalize or standardize the query vector to ensure comparability with the vectors in the address vector library.
[0072] In some embodiments of the present application, the process of generating the query text address vector includes: Obtain the query text address input by the user on the interaction interface; Decompose the query text address into a country address, a province / municipality directly under the Central Government address, a city / district address, and a street / house number address according to a preset rule; Call the pre-trained PromptCSE-Base model to convert the addresses at each decomposed level into corresponding address vectors respectively; Concatenate the address vectors at each level to obtain a text address vector; Perform weighted summation according to the default spatial vector, the default POI vector, the default street view image vector, and the concatenated text address vector to obtain a multi-modal vector; Divide the query text address into multiple entities, query corresponding semantic enhancement information in the knowledge base by using each entity, and encode the semantic enhancement information to generate a semantic enhancement vector; Concatenate the generated semantic enhancement vector and the multi-modal vector obtained by weighted summation to obtain the query text address vector corresponding to the query text address.
[0073] Among them, the server receives the text address (such as "No. 88, Jianguo Road, Jia District, Jia City") input by the user through an interaction interface (such as a Web page, a mobile APP, or an API interface).
[0074] Preprocess the input text, including: Format standardization: unify the case, remove redundant spaces or special symbols (such as considering "Chang'an City" and "Chang'an" as equivalent).
[0075] Anomaly detection: check whether the input is empty or contains illegal characters, and if there is an anomaly, return an error prompt.
[0076] The server decomposes the query text address into the following levels according to a preset address level rule (such as an administrative division coding table, a geographical dictionary): Country level: Extract the country name (such as "Country A").
[0077] Province / municipality directly under the Central Government level: Extract the name of the provincial administrative region (such as "City A").
[0078] City / district level: Extract the name of the municipal or district administrative region (such as "District A").
[0079] Street / house number level: Extract the street name and house number (such as "No. 88, Jianguo Road").
[0080] Fuzzy matching support: If the user input is incomplete (e.g., only "Jianguo Road, Jia District" is entered), the missing levels are supplemented according to the context or historical records (e.g., the default city is "Chang'an").
[0081] The server calls the pre-trained PromptCSE-Base model (a text encoding model based on contrastive learning) to convert each decomposed level of the address into a vector representation respectively: Each level of the address is used as an independent text input to the model (such as "Country A", "City A", "District A", "No. 88 Jianguo Road"). The model generates semantic vectors for each level (such as 128-dimensional or 256-dimensional), and each vector captures the semantic features of the corresponding level. The server concatenates the address vectors of each level into a single text address vector in a preset order. The concatenation rule can be: for example, country vector + province vector + city vector + street vector, and the vectors can be concatenated through a special delimiter (such as a vector of all zeros) or directly. L2 normalization is performed on the concatenated vectors to ensure that the vector norms are consistent and avoid numerical differences affecting subsequent calculations.
[0082] The server performs a weighted sum of the concatenated text address vector with the following default vectors to generate a multimodal vector: Default spatial vector: An embedding vector representing the geographical coordinates (such as longitude and latitude) of the address, generated through GeoHash or a deep learning model. The geographical coordinates can be manually input by the user or pre-stored fixed coordinates. For example, the fixed coordinates are empty.
[0083] Default POI vector: A semantic vector representing the points of interest (such as shopping malls, schools) near the address. The points of interest can be selected by the user from the POI database or pre-stored fixed POIs. For example, the POI is empty.
[0084] Default street view image vector: A visual feature vector representing the street view images around the address, obtained by encoding the street view images through a CNN model (such as ResNet). The surrounding street view images can be input by the user or be empty.
[0085] The weights of each vector can be adjusted through experiments (such as text vector weight 0.6, spatial vector weight 0.2, POI vector weight 0.1, street view vector weight 0.1) to balance the influence of different modalities.
[0086] The server divides the query text address into multiple entities (such as "City A", "District A", "No. 88 Jianguo Road") and uses a knowledge base (such as a geographical knowledge graph) to query the semantic enhancement information of each entity: The information types include entity categories (such as "administrative region", "street"), aliases (such as the alias of "Jia District" is "Dong'an"), and associated entities (such as the subway station near "Jianguo Road"). Encode the semantic enhancement information for each entity to generate a semantic enhancement vector: A pre-trained language model (such as BERT) can be used to encode the semantic text, or it can be mapped to a predefined semantic vector space through rules. Concatenate or weighted sum the semantic enhancement vectors of all entities to generate a global semantic enhancement vector.
[0087] The server concatenates the generated semantic enhancement vector with the multi-modal vector to obtain the final query text address vector: The semantic enhancement vector is directly concatenated to the end of the multi-modal vector, or the information of the two is dynamically fused through a gating mechanism. Normalize the final vector to ensure consistent numerical ranges.
[0088] The server converts the text address input by the user into a vector representation containing multi-modal information and semantic enhancement features, providing high-dimensional and high-semantic input features for subsequent address matching, and significantly improving the accuracy and efficiency of address parsing.
[0089] In some embodiments of the present application, the weights of the default space vector, the default POI vector, the default street view image vector, and the concatenated text address vector are dynamically adjusted according to the context information. See Figure 3 As shown, the flow diagram of the method for dynamically adjusting weights based on context information provided by the embodiments of the present application: S301. When the user inputs a query text address, collect the user's context information, where the context information includes: the user's historical search records, address location, and time.
[0090] S302. Convert the context information into a context vector.
[0091] S303. Input the context vector into a pre-trained context-aware model to calculate the weights of the default space vector, the default POI vector, the default street view image vector, and the concatenated text address vector.
[0092] Among them, the server extracts the user's most recent N (such as the most recent 10) address search records from the user behavior log, including the searched text address, the clicked address result, the search time, etc. Denoise the historical search records (such as removing duplicate or invalid records), and extract key features (such as search frequency, search result preference). Obtain the user's current geographical location (such as longitude and latitude coordinates) through the GPS, IP positioning, or Wi-Fi positioning of the user device. If the user does not authorize the positioning permission, use the common address in the historical record (such as home address, work address) as a substitute. Record the timestamp (accurate to the second) when the user initiates the query, and extract time features (such as weekdays / weekends, day / night, holidays, etc.).
[0093] For the text addresses in the historical search records, call the same text encoding model as in S210 (such as PromptCSE-Base) to generate address vectors. Perform one-hot encoding (One-Hot Encoding) or embedding encoding (Embedding) on the search result preferences (such as the types of POIs that users often click on) to generate preference vectors. Convert the user's geographical location (latitude and longitude) into GeoHash encoding (such as an 8-bit string), and generate spatial vectors through a pre-trained GeoHash embedding model. You can also directly use the latitude and longitude coordinates as input and map them to a low-dimensional vector space through a dense layer (Dense Layer). Split the timestamp into features such as hours, weeks, and months, and generate time vectors through an embedding layer (Embedding Layer). Perform one-hot encoding on special time markers such as holidays and concatenate them with the time vectors.
[0094] The server concatenates the historical search record vector, the address location vector, and the time vector into a single context vector: Concatenation rule: For example, historical search vector (256 dimensions) + address location vector (128 dimensions) + time vector (64 dimensions) = context vector (448 dimensions). If the dimensions of each vector are inconsistent, dimension mapping can be performed through a dense layer or zero vectors can be filled.
[0095] The server inputs the context vector into a pre-trained context-aware model (such as a lightweight model based on Transformer or a graph neural network): The context vector serves as the model input, and the model captures the associations between context features through self-attention mechanism (Self-Attention) or graph convolution (Graph Convolution).
[0096] The model outputs the following weight parameters: Default spatial vector weight: Used to adjust the contribution of the geographical coordinate vector in the multi-modal vector (such as 0.2).
[0097] Default POI vector weight: Used to adjust the contribution of the point of interest vector in the multi-modal vector (such as 0.15).
[0098] Default street view image vector weight: Used to adjust the contribution of the street view image vector in the multi-modal vector (such as 0.1).
[0099] Weight of the concatenated text address vector: Used to adjust the contribution of the text vector in the multi-modal vector (such as 0.55).
[0100] Weight constraint: The sum of all weights is 1, and normalization is performed through a Softmax layer.
[0101] The server converts user context information (historical search, geographical location, time) into vectors, generates dynamic weights for multi-modal vectors through a context-aware model, and finally generates a context-aware query text address vector. This process significantly improves the accuracy and personalization ability of address parsing, especially in complex scenarios (such as when users frequently search for addresses in the same area or during a specific time period).
[0102] Among them, the server calculates the cosine similarity or Euclidean distance between the query vector and all vectors in the address vector library. The similarity calculation may be parallelized to improve efficiency. The top N candidate vectors with the highest similarity are returned (such as Top-5). A threshold may be set during screening (such as similarity > 0.8). The candidate vectors are sorted in descending order of similarity.
[0103] S304. According to the calculated weights, perform weighted summation on the default space vector, default POI vector, default street view image vector, and the spliced text address vector to obtain a multi-modal vector.
[0104] Furthermore, the server supports hierarchical search. The specific process includes: If the hierarchical information has been marked when the query vector is generated (such as the vectors are spliced by level in S203), directly split the sub-vectors of each level.
[0105] If the query vector is an overall vector, it is dynamically segmented through a model or rules (such as using an attention mechanism to extract city-level features). Example: When querying "City A", the server splits its vector into: City-level vector: Corresponding to the semantic representation of "City A".
[0106] Street-level vector: Corresponding to the semantic representation of "Sports West Road, District A, City A".
[0107] High-level vector matching (city level): The server first uses the city-level vector to perform a coarse-grained search in the address vector library: Filter out all candidate addresses in the vector library that contain the city vector. Only calculate the cosine similarity between the query city vector and the city vectors in the candidate addresses, ignoring other levels. Retain the candidate addresses with a similarity higher than the preset threshold (such as 0.8), initially narrowing the search scope. This step can quickly filter out more than 90% of the irrelevant addresses, significantly reducing the subsequent calculation volume.
[0108] For the candidate addresses after high-level matching, the server further uses the street-level vector for refined search: calculate the similarity between the query street vector and the street vectors in the candidate addresses, and perform weighting in combination with the similarity at the city level (such as city weight 0.4 and street weight 0.6). If the street-level information is incomplete (such as the query is "A District" but the street is not specified), then reduce the street weight and increase the weight at the city or POI level. Example: When querying "A City", first match the addresses within "A District", and then refine to the vicinity of "Sports West Road".
[0109] For the candidate addresses after middle-level matching, the server further performs exact matching in combination with the house number or POI vector: If the query contains a POI (such as "A City Shopping Center"), then calculate the similarity between the query POI vector and the POI vectors in the candidate addresses. Use a string similarity algorithm (such as Levenshtein distance) for fuzzy matching of the house number, tolerating partial character differences (such as "No. 103" and "No. 105" may match). Weighted sum the similarity scores at each level to generate the final similarity ranking.
[0110] The server sorts the candidate addresses in descending order according to the comprehensive similarity and returns the Top-N results (such as Top-10).
[0111] Among them, retrieve the corresponding original data (text address, coordinates, POI, street view) from the database according to the candidate vector, and may also retrieve associated information (such as user reviews, business hours). Format the information into a list, map marker or card form, which may include address details, picture thumbnails, navigation buttons, etc. On the interactive interface displayed on the terminal device, the user is allowed to click on the candidate address to view more details (such as panoramic view, surrounding facilities), and the interactive interface also provides a feedback mechanism (such as "This address is incorrect"). After reviewing the feedback information, update the multi-modal address database and address vector database.
[0112] In some embodiments of the present application, the server can report address deviations, new addresses or temporary change information through the crowdsourcing client to update the multi-modal address database in real time. The specific process includes: The server receives data correction instructions from the crowdsourcing client (such as manual annotators or user feedback tools) through the crowdsourcing task platform or API interface. The instruction content includes the following information: Address identifier to be corrected: The ID that uniquely identifies the address to be modified (such as database primary key, address hash value). Correction field: Clearly indicate the field that needs to be modified (such as street name, house number, POI category, etc.). Corrected value: The corrected value provided by the crowdsourcing personnel (such as correcting "No. 88, Jianguo Road" to "No. 90, Jianguo Road"). Correction basis: Optional field, recording the source of the correction (such as user feedback, on-site verification record, satellite image comparison, etc.).
[0113] The server first performs automated verification on the correction instruction, including: Checking whether the instruction contains necessary fields and whether the field values meet the format requirements (e.g., the house number is a number). Verifying whether the corrected value is logically consistent with other fields of the address (such as the affiliated street and administrative division) (e.g., "No. 90, Jianguo Road" should be located in "District A" rather than "District B"). Checking whether the correction instruction conflicts with other unsubmitted correction instructions in the database (e.g., receiving two different correction requests for the same address simultaneously). For correction instructions that fail the automated review or are of high risk (such as those involving administrative division changes), submit them to the manual review queue. The reviewer checks the address information before and after the correction through a visualization tool (such as a map annotation interface) and verifies it with reference to external data sources (such as government public data and third-party map APIs). The instructions that pass the review are marked as "confirmed", and the instructions that fail the review are marked as "rejected" and the reason for rejection is returned (such as "The corrected value does not match the satellite image").
[0114] For the correction instructions that pass the review, the server first locks the record of the address to be modified in the database to prevent data inconsistency caused by concurrent modification. Directly modify the address text fields (such as street name and house number), and synchronously update the hierarchical representation of the address (such as the hierarchical relationship of city, district, and street). If the correction involves geographical location (such as the change of the house number resulting in the offset of longitude and latitude), recalculate the geographical coordinates of the address and update the GeoHash code or spatial index. Rematch the association relationship with nearby POIs (such as the bus stops near "No. 90, Jianguo Road" may change). If the corrected address corresponds to a new street view image (such as the change of the building appearance), trigger the street view image acquisition task and update the street view image vector. Archive the historical street view images and retain the change records for backtracking. Regenerate the semantic enhancement vector of the address (such as querying new entity association information through the knowledge base). Recode the text vector of the address (such as calling the PromptCSE-Base model to generate a new vector). According to the updated text vector, spatial vector, POI vector, and street view image vector, recalculate the weighted multi-modal vector (the weights follow the weights output by the most recent context-aware model, or trigger the weight update process).
[0115] Update the associated data of the address (such as the address information involved in the user's favorite records and historical search records). Cascade update other data that depends on this address (such as the logistics distribution range and merchant service area). Generate a version record for each modification of the address data, including the modification time, modifier, modification content, etc., to support data rollback and auditing.
[0116] The server returns the correction result to the crowdsourcing client: when successful, it returns "Correction successful" and the updated address details. When failed, it returns "Correction failed" and the specific reason (such as "Administrative division conflict"). For critical corrections (such as address changes related to public services), trigger a notification mechanism (such as email, SMS) to inform relevant stakeholders (such as merchants, government agencies).
[0117] The server can safely and efficiently respond to the data correction instructions of the crowdsourcing client, ensuring the accuracy and real-time nature of the multimodal address database. This process not only supports the rapid response to manual feedback but also guarantees data quality through strict auditing and version control mechanisms, and is applicable to scenarios that require high-precision geographic information.
[0118] In some embodiments of the present application, a full synchronization method is adopted to synchronize all multi-module data from the industry address library to the multimodal address database; When it is detected that the industry address library has data updates, an incremental synchronization method is adopted to synchronize the updated data in the industry address library to the multimodal address database; When it is detected that the multimodal address database has data updates, update the address vector library according to the updated data in the multimodal address database.
[0119] Among them, 1. Full synchronization stage.
[0120] When initializing the multimodal address database, or performing full data synchronization regularly (such as once a week) to ensure the initial consistency between the database and the industry address library. The server connects to the industry address library through a preset API interface or data file transfer protocol (such as FTP, SFTP) to obtain its multi-module data (such as basic address information, administrative divisions, POI data, spatial coordinates, etc.). Extract all module data from the industry address library and perform format conversion (such as converting XML format to JSON format) to ensure compatibility with the table structure of the multimodal address database. Batch-write the converted data into the multimodal address database, overwriting the original data (or clearing and reinserting). For large data volume scenarios, adopt a batch loading strategy to avoid single operation timeouts or out-of-memory errors. Record the start time, end time, processed data volume, and synchronization result (success / failure) of the full synchronization for subsequent auditing and troubleshooting.
[0121] 2. Incremental synchronization stage.
[0122] The server detects data updates through a polling mechanism (e.g., every 5 minutes) or change notifications provided by the industry address library (e.g., WebSocket push). If the industry address library supports change logs (such as an operation log table), the server directly queries the change records after the most recent synchronization. If change logs are not supported, the server identifies the changed content by comparing the current data with the data snapshot at the time of the previous synchronization (such as hash value comparison). Extract the changed data (including newly added, modified, and deleted records) and convert it into the format required by the multimodal address database.
[0123] Incremental data loading: Newly added data: Directly insert it into the multimodal address database.
[0124] Modified data: Update the corresponding records according to the primary key or unique identifier.
[0125] Deleted data: Delete the corresponding records from the multimodal address database according to the primary key or unique identifier (or mark them as the "deleted" status).
[0126] If there is a conflict between the incremental data and the unsynchronized local modifications in the multimodal address database (e.g., the same address is modified by both the crowdsourcing client and the industry address library), trigger the conflict resolution process: Give priority to retaining the data in the industry address library (regarded as the authoritative data source). Record the conflict event and notify the administrator for manual intervention. Update the synchronization progress schedule and record the latest timestamp or change ID of this incremental synchronization for the next synchronization.
[0127] 3. Address vector library update process.
[0128] Detect data updates in the multimodal address database (including initial updates after full synchronization, updates after incremental synchronization, or direct modifications by the crowdsourcing client). The server uses database triggers (such as INSERT / UPDATE / DELETE triggers) or change data capture (CDC) technology to capture the change records in the multimodal address database in real time. For newly added or modified address records, generate or update address vectors according to the following steps: Text vector generation: Call a pre-trained text encoding model (such as PromptCSE-Base) to convert the address text into a vector.
[0129] Spatial vector generation: Convert the geographical coordinates (latitude and longitude) of the address into a GeoHash code and generate a spatial vector through an embedding model.
[0130] POI and street view vector generation: Query the POI data and street view images associated with the address and generate corresponding semantic vectors and visual vectors respectively.
[0131] Multi-modal vector concatenation and weighting: According to the weights output by the context-aware model, the above vectors are weighted and summed to generate multi-modal vectors.
[0132] Semantic enhancement vector fusion: Combining the semantic enhancement information in the knowledge base to generate the final address vector.
[0133] Write the generated address vector into the address vector library and establish an association with the unique identifier (such as ID) of the corresponding address. For the deleted address records, delete the corresponding vectors from the address vector library (or mark them as "invalid" status). Scan the address vector library regularly (such as once a day) to check if there are vectors without associated addresses or invalid vectors and clean them up.
[0134] In this embodiment, through the combination of full synchronization and incremental synchronization, the server can efficiently and reliably maintain the consistency between the multi-modal address database and the industry address library. At the same time, by updating the address vector library in real time, the accuracy of the address matching and retrieval services is ensured.
[0135] Furthermore, the specific processes of full synchronization and incremental synchronization include: The server divides the data into multiple logical partitions (such as "Partition of City A" or "POI Data Partition") according to the table structure of the industry address library (such as by administrative division, address type, or data volume distribution). For large data tables (such as the address master table), hash partitioning (taking the modulus according to the hash value of the primary key) or range partitioning (by ID range) is used for partitioning to ensure the balance of data volume in each partition. Record the partitioning rules (such as partitioning fields, number of partitions) and partitioning boundaries (such as hash value range, ID range) into the metadata table for subsequent task scheduling and interruption recovery. The server creates an independent synchronization task (Job) for each data partition based on the xxjob distributed task scheduling framework, and the following parameters are configured in the task: Data source configuration: Connection information of the industry address library (such as database URL, username, password).
[0136] Target configuration: Connection information of the multi-modal address database and the target table corresponding to the partition.
[0137] Synchronization range: The data boundary of the current partition (such as hash value range, ID range).
[0138] The xxjob framework dynamically allocates tasks based on the number of partitions and server resources, supporting parallel execution (such as starting 10 partition tasks simultaneously). It manages task dependencies (such as subsequent tasks can only be started after the previous tasks are completed). During the execution of each synchronous task, progress information is reported to the xxjob monitoring center regularly (such as every 1000 records), including: the amount of data processed, the total amount of data, the processing speed (records / second). The current data position being processed (such as the ID or hash value of the last record). The server displays the progress bar, status (running / completed / failed), and log information of each partition task through the xxjob Web management interface. If a task is interrupted due to network failure, server downtime, or timeout, the xxjob framework automatically marks the task as "failed" and records the data position at the time of interruption (such as the ID of the last successfully processed record). When the task is restarted, the server reads the interrupted data position from the metadata table and configures the starting point of the synchronous task to this position, skipping the processed data to avoid duplicate synchronization. For failed tasks, the cause of the interruption is located through log analysis (such as database connection timeout, data format error), and retry or manual intervention is triggered.
[0139] The server deploys a canal client on the database server of the industry address library, disguising it as a MySQL slave (such as by simulating the MySQL replication protocol), and listens to the binlog of the master database in real time. Canal parses the new (INSERT), modified (UPDATE), and deleted (DELETE) operations in the binlog and extracts the changed data (such as row data, change timestamp, operation type).
[0140] Canal encapsulates the changed data into messages in JSON or Protobuf format, and the messages contain the following fields: Operation type: INSERT / UPDATE / DELETE. Data content: The complete row data after the change (for UPDATE operations, both the old data and the new data are included). Metadata: Table name, primary key value, change timestamp.
[0141] The server creates an independent kafka topic (such as address_db_changes) for incremental synchronization and configures multiple partitions (such as 8 partitions) to support high throughput.
[0142] Canal, as a producer, sends the change messages to this topic and selects a partition based on the table name or primary key hash value to ensure that change messages for the same address are sent to the same partition.
[0143] The server deploys a Kafka consumer group in the synchronization service of the multi-modal address database. The consumer group subscribes to the address_db_changes topic and configures the following parameters: Start multiple consumer instances according to the number of Kafka partitions to achieve parallel processing. Automatically manage offsets through the Kafka consumer group mechanism to ensure that messages are not lost and not processed repeatedly.
[0144] The consumer parses the change messages and performs the following operations according to the operation type: INSERT: Insert the new data into the multi-modal address database.
[0145] UPDATE: Update the corresponding records in the multi-modal address database according to the primary key.
[0146] DELETE: Delete the corresponding records in the multi-modal address database according to the primary key (or mark as "deleted").
[0147] For critical operations (such as batch UPDATE), adopt a transaction mechanism to ensure data consistency. For messages with synchronization failures, record them in the dead letter queue (DLQ) and trigger an alarm, supporting manual retry.
[0148] In this embodiment, by combining the full-volume synchronization partition and breakpoint resumption mechanism of xxjob, and the incremental synchronization real-time capture and push of canal + Kafka, the server can efficiently and reliably maintain the consistency between the multi-modal address database and the industry address database.
[0149] Beneficial effects of the embodiments of this application: By integrating multi-modal data such as text addresses, spatial coordinates, surrounding POIs, and street view pictures, it can represent address information more comprehensively and accurately, improving the richness and accuracy of address descriptions. Convert the multi-modal data into vector form, which is convenient for mathematical operations and similarity comparison, providing an efficient data basis for subsequent address queries and matches. By dividing the text address into multiple entities and querying the corresponding semantic enhancement information in the knowledge base, the semantic representation of the address is further enriched, improving the accuracy and intelligence of address matching. Generate a query text address vector according to the query text address input by the user, and quickly query the candidate address vector with the highest similarity in the address vector library, realizing an efficient address query function. Display the text address, spatial coordinates, surrounding POIs, and street view pictures corresponding to the queried candidate address vector on the interaction interface, providing the user with an intuitive and comprehensive query result, facilitating the user's selection and judgment.
[0150] In one or more possible embodiments, optimize the query process to improve the query speed.
[0151] Among them, the final address vector generated after concatenating the semantic enhancement vector and the multi-modal vector, and writing the final address vector into the address vector library includes: Calculate the hash value of the final address vector, and write the final address vector into the target node among multiple deployed nodes according to the hash value; Among them, querying in the address vector library according to the query information input by the user on the interaction interface, and displaying the query result on the interaction interface includes: Judge whether the query task is an exact query task or a fuzzy query task according to the query information input by the user on the interaction interface; If it is an exact query task, query in the cache server according to the exact query task; If it is hit in the cache server, display the query result returned by the cache server on the interaction interface; If it is not hit in the cache server, calculate the hash value according to the query information, determine the target node among the multiple nodes according to the calculated hash value, route the exact query task to the target node, receive the query result returned by the target node, and generate the final query result on the interaction interface; If it is a fuzzy query task, query in the cache server according to the fuzzy query task; If it is hit in the cache server, display the query result returned by the cache server on the interaction interface; If the cache server is not hit, broadcast the fuzzy query task to the multiple nodes, aggregate the query results returned by each node to generate the final query result, and display the final query result on the interaction interface.
[0152] Specifically, after the server generates the final address vector, it uses a consistent hashing algorithm (such as the Ketama algorithm) to calculate the hash value of the vector. The hash space is divided into multiple virtual nodes, and each physical node is responsible for managing a continuous hash ring area. According to the calculated hash value, the server determines the target node where the vector should be stored and writes the vector into the vector library of the node. This design can achieve uniform data distribution and support the minimum data migration when nodes are dynamically added or removed.
[0153] For example: Assume the hash space is 0 to 2^32 - 1, and servers A, B, and C are respectively responsible for the regions [0, 2^31), [2^31, 32^30), and [3230, 232 - 1] on the hash ring. If the hash value of a certain address vector is 500,000,000, it falls into the responsible region of server B and is finally routed to server B for storage.
[0154] The server determines the type of query task based on the query information input by the user (such as whether it contains wildcards, whether exact matching is required, etc.). An exact query task requires returning results that exactly match the query conditions, while a fuzzy query task allows returning partially matching or highly similar results. The judgment logic can be implemented through regular expression matching or query syntax parsing.
[0155] For example: when the user inputs "No. 88, Jianguo Road, Jia District, Jia City", the server recognizes it as an exact query task. When the user inputs "* Road, Jia District, Jia City", the server recognizes it as a fuzzy query task.
[0156] The server first queries in the cache server (such as Redis) to see if there is already a cached result for this exact query. If it hits, it directly returns the cached result; if it misses, it calculates the hash value of the query information, determines the target node according to the consistent hashing algorithm, and routes the query task to this node. After the target node executes the query, it returns the result to the requesting server, and the server then writes the result into the cache and returns it to the user.
[0157] For example: when the user queries the exact location information of "No. 88, Jianguo Road, Jia District, Jia City", the server does not find the query result in the cache. After calculating its hash value, it routes it to Server B. Server B queries the detailed information of this address in its vector library and returns it. The server writes the result into the cache and displays it to the user at the same time.
[0158] The server first queries the cached result of the fuzzy query in the cache server. If it hits, it directly returns; if it misses, it broadcasts the fuzzy query task to all nodes. Each node performs a similarity search in its vector library (such as using the IVFPQ index of FAISS) and returns results with relatively high similarity. The server aggregates the return results of each node (such as sorting by similarity and taking the top N), and displays the final result to the user.
[0159] For example: when the user queries the fuzzy location information of "* Road, Jia District, Jia City", the server does not find the result in the cache and broadcasts the query to Servers A, B, and C. Each node returns a list of addresses with relatively high similarity. The server aggregates them and sorts them by similarity, and displays results such as "No. 88, Jianguo Road" and "No. 90, Jianguo Road".
[0160] In this embodiment, the consistent hashing algorithm is used to achieve uniform data distribution and avoid overloading a single node. When nodes are added or removed, only a small amount of data needs to be migrated to ensure the high availability of the system. Exact queries preferentially hit the cache, reducing the calculation overhead; fuzzy queries achieve efficient similarity search through broadcasting and aggregation, balancing accuracy and performance. The system can dynamically add or remove nodes, and the hash ring is automatically adjusted without interrupting the service, adapting to changes in the business scale. The dual mechanisms of caching and node query ensure the consistency and real-time nature of the query results, improving the user experience.
[0161] In summary, the technical solution of this application improves the accuracy and richness of address descriptions through technical means such as multi-modal data fusion, vectorization processing, semantic enhancement, efficient query, and result visualization display, realizes an efficient address query function, and provides users with a more convenient and intelligent address service experience.
[0162] The following is an embodiment of the device of this application, which can be used to execute the method embodiment of this application. For details not disclosed in the device embodiment of this application, please refer to the method embodiment of this application.
[0163] Please refer to Figure 4 , which shows a schematic structural diagram of an address search device based on deep learning provided by an exemplary embodiment of this application, hereinafter referred to as device 4. This device 4 can be implemented as all or part of a server through software, hardware, or a combination of both. Device 4 includes: an acquisition unit 401, a conversion unit 402, a generation unit 403, and a query unit 404.
[0164] The acquisition unit is used to acquire various multi-modal data in the multi-modal address database, and the multi-modal data includes: text address, spatial coordinates, surrounding POIs, and street view pictures; The conversion unit is used to convert the text address into a text address vector, convert the spatial coordinates into a spatial vector, convert the surrounding POIs into POI vectors, and convert the street view pictures into street view picture vectors; The generation unit is used to perform weighted summation on the text address vector, the spatial vector, the POI vector, and the street view picture vector to generate a multi-modal vector; divide the text address into multiple entities, query corresponding semantic enhancement information in the knowledge base using the entities, and encode the queried semantic enhancement information to generate a semantic enhancement vector; splice the semantic enhancement vector and the multi-modal vector to generate a final address vector, and write the final address vector into the address vector library; The query unit is used to query in the address vector library according to the query information input by the user on the interaction interface, and display the query result on the interaction interface.
[0165] In one or more possible embodiments, the conversion of the text address into a text address vector includes: Decompose the text address into a national address, a provincial / municipal address, a city / district address, and a street / house number address; According to the pre-trained PromptCSE-Base model, convert the addresses at each level into a national address vector, a provincial / municipal address vector, a city / district vector, and a street address vector, and splice the address vectors at each level to obtain a text address vector.
[0166] In one or more possible embodiments, the query information is a query text address; Among them, the query in the address vector library according to the query information input by the user on the interaction interface and the display of the query result on the interaction interface include: Obtain the query text address input by the user on the interaction interface; Decompose the query text address into a country address, a province / municipality address, a city / district address, and a street / house number address according to a preset rule; Call the pre-trained PromptCSE-Base model to convert the decomposed addresses at each level into corresponding address vectors respectively; Concatenate the address vectors at each level to obtain a text address vector; Perform weighted summation according to the default spatial vector, the default POI vector, the default street view image vector, and the concatenated text address vector to obtain a multi-modal vector; Divide the query text address into multiple entities, query corresponding semantic enhancement information in the knowledge base using each entity, and encode the semantic enhancement information to generate a semantic enhancement vector; Concatenate the generated semantic enhancement vector and the multi-modal vector obtained by weighted summation to obtain the query text address vector corresponding to the query text address; Query several candidate address vectors with the highest similarity in the address vector library according to the query text address vector; Display the text addresses, spatial coordinates, surrounding POIs, and street view images corresponding to the several candidate address vectors on the interaction interface.
[0167] In one or more possible embodiments, the performing weighted summation according to the default spatial vector, the default POI vector, the default street view image vector, and the concatenated text address vector to obtain a multi-modal vector includes: When the user inputs a query text address, collect the user's context information, where the context information includes: the user's historical search records, address location, and time; Convert the context information into a context vector; Input the context vector into a pre-trained context-aware model to calculate the weights of the default spatial vector, the default POI vector, the default street view image vector, and the concatenated text address vector; Perform weighted summation on the default spatial vector, the default POI vector, the default street view image vector, and the concatenated text address vector according to the calculated weights to obtain a multi-modal vector.
[0168] In one or more possible embodiments, before converting the addresses at each level into national address vectors, provincial / municipal address vectors, city / district vectors, and street address vectors according to the pre-trained PromptCSE-Base model, the following steps are also included: Clean, segment, and standardize the addresses at each level.
[0169] In one or more possible embodiments, the following is also included: A correction unit for receiving a crowdsourcing client data correction instruction, and after passing the review of the data correction instruction, modifying the corresponding multimodal data in the multimodal address database based on the data correction instruction.
[0170] In one or more possible embodiments, the following is also included: A synchronization unit for synchronizing all multimodule data from the industry address library to the multimodal address database in a full synchronization manner; When it is detected that the industry address library has data updates, synchronize the updated data in the industry address library to the multimodal address database in an incremental synchronization manner; When it is detected that the multimodal address database has data updates, update the address vector library according to the updated data in the multimodal address database.
[0171] In one or more possible embodiments, during the full synchronization process, divide the industry address library into multiple data partitions based on xxjob technology, create a synchronization task for each data partition, and monitor the progress information of each synchronization task in real time; when a synchronization task is interrupted, obtain the interrupted data position and re-execute the synchronization operation based on the interrupted data position; During the incremental synchronization process, when it is detected by the canal middleware that the industry address library has new, modified, and deleted operations, push the updated data to kafka, and then use kafka to synchronize the updated data to the multimodal address database.
[0172] In one or more possible embodiments, the final address vector generated after splicing the semantic enhancement vector and the multimodal vector, and writing the final address vector into the address vector library, includes: Calculate the hash value of the final address vector, and write the final address vector to the target node among multiple deployed nodes according to the hash value; Among them, querying in the address vector library according to the query information input by the user on the interaction interface, and displaying the query result on the interaction interface, includes: Determine whether the query task is an exact query task or a fuzzy query task according to the query information input by the user on the interaction interface; If it is an exact query task, query in the cache server according to the exact query task; If it is hit in the cache server, display the query result returned by the cache server on the interaction interface; If it is not hit in the cache server, calculate the hash value according to the query information, determine the target node among the multiple nodes according to the calculated hash value, route the exact query task to the target node, receive the query result returned by the target node, and generate the final query result on the interaction interface; If it is a fuzzy query task, query in the cache server according to the fuzzy query task; If it is hit in the cache server, display the query result returned by the cache server on the interaction interface; If the cache server is not hit, broadcast the fuzzy query task to the multiple nodes, aggregate the query results returned by each node to generate the final query result, and display the final query result on the interaction interface.
[0173] It should be noted that when the device 4 provided in the above embodiment executes the address search method based on deep learning, only the above division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above functions. In addition, the address search device based on deep learning provided in the above embodiment and the embodiment of the address search method based on deep learning belong to the same concept, and the implementation process is detailed in the method embodiment, which will not be repeated here.
[0174] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0175] See Figure 5 As shown, it is a schematic diagram of a computer storage medium provided by an embodiment of the present application. The computer storage medium may store multiple instructions (that is, Figure 5 the computer program shown), and the instructions are suitable for being loaded and executed by the processor to perform the method steps of the above Figure 2 shown embodiment. The specific execution process can be seen in Figure 2 the specific description of the shown embodiment, which will not be repeated here.
[0176] The present application also provides a computer program product. The computer program product stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the address search method based on deep learning as described in the above embodiments.
[0177] Please refer to Figure 6 , which provides a schematic structural diagram of a server for an embodiment of the present application. As Figure 6 shown, the server 600 may include: at least one processor 601, at least one network interface 604, a user interface 603, a memory 605, and at least one communication bus 602.
[0178] Among them, the communication bus 602 is used to implement connection communication between these components.
[0179] Among them, the user interface 603 may include input units such as a mouse and a keyboard.
[0180] Among them, the network interface 604 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0181] Among them, the processor 601 may include one or more processing cores. The processor 601 connects various parts within the entire server 600 using various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 605, and by calling data stored in the memory 605, it performs various functions of the server 600 and processes data. Optionally, the processor 601 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 601 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 601 and may be implemented separately by a single chip.
[0182] Among them, the memory 605 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory 605 includes a non-transitory computer-readable storage medium. The memory 605 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 605 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 605 may also be at least one storage device located far from the aforementioned processor 601. As Figure 6 shown, the memory 605 as a computer storage medium may include an operating system, a network communication module, a user interface module, and application programs.
[0183] In Figure 6 the shown server 600, the user interface 603 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 601 can be used to call the application programs stored in the memory 605 and specifically execute the method as Figure 2 shown. The specific process can be referred to Figure 2 shown, and will not be elaborated here.
[0184] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0185] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A method for address search based on deep learning, characterized in that, Including: Obtain various multimodal data in the multimodal address database, where the multimodal data includes: text address, spatial coordinates, surrounding POIs, and street view pictures; Convert the text address into a text address vector, convert the spatial coordinates into a spatial vector, convert the surrounding POIs into POI vectors, and convert the street view pictures into street view picture vectors; Perform weighted summation on the text address vector, the spatial vector, the POI vector, and the street view picture vector to generate a multimodal vector; Divide the text address into multiple entities, query corresponding semantic enhancement information in the knowledge base using the entities, and encode the queried semantic enhancement information to generate a semantic enhancement vector; Concatenate the semantic enhancement vector and the multimodal vector to generate a final address vector, and write the final address vector into the address vector library; Query in the address vector library according to the query information input by the user on the interaction interface, and display the query result on the interaction interface.
2. The method according to claim 1, wherein The conversion of the text address into a text address vector includes: Decompose the text address into a national address, a provincial / municipal address, a city / district address, and a street / house number address; According to the pre-trained PromptCSE-Base model, convert the addresses at each level into a national address vector, a provincial / municipal address vector, a city / district vector, and a street address vector, and concatenate the address vectors at each level to obtain a text address vector.
3. The method according to claim 2, wherein The query information is a query text address; Among them, the query in the address vector library according to the query information input by the user on the interaction interface and the display of the query result on the interaction interface include: Obtain the query text address input by the user on the interaction interface; Decompose the query text address into a national address, a provincial / municipal address, a city / district address, and a street / house number address according to a preset rule; Call the pre-trained PromptCSE-Base model to convert the decomposed addresses at each level into corresponding address vectors respectively; Concatenate the address vectors at each level to obtain a text address vector; Perform weighted summation according to the default spatial vector, the default POI vector, the default street view picture vector, and the concatenated text address vector to obtain a multimodal vector; Divide the query text address into multiple entities, query corresponding semantic enhancement information in the knowledge base using each entity, and encode the semantic enhancement information to generate a semantic enhancement vector; Concatenate the generated semantic enhancement vector and the multimodal vector obtained by weighted summation to obtain the query text address vector corresponding to the query text address; Query several candidate address vectors with the highest similarity in the address vector library according to the query text address vector; Display the text address, spatial coordinates, surrounding POIs, and street view pictures corresponding to the several candidate address vectors on the interaction interface.
4. The method according to claim 3, characterized in that The weighted summation according to the default spatial vector, the default POI vector, the default street view picture vector, and the concatenated text address vector to obtain a multimodal vector includes: When the user inputs a query text address, collect the user's context information, where the context information includes: the user's historical search records, address location, and time; Convert the context information into a context vector; Input the context vector into a pre-trained context-aware model to calculate the weights of the default spatial vector, default POI vector, default street view image vector, and the concatenated text address vector; According to the calculated weights, perform weighted summation on the default spatial vector, default POI vector, default street view image vector, and the concatenated text address vector to obtain a multimodal vector.
5. The method according to claim 4, characterized in that It also includes: Receive a crowdsourcing client data correction instruction. After passing the review of the data correction instruction, modify the corresponding multimodal data in the multimodal address database based on the data correction instruction.
6. The method according to claim 5, wherein It also includes: Use the full synchronization method to synchronize all multi-module data from the industry address library to the multimodal address database; When it is detected that the industry address library has data updates, use the incremental synchronization method to synchronize the updated data in the industry address library to the multimodal address database; When it is detected that the multimodal address database has data updates, update the address vector library according to the updated data in the multimodal address database.
7. The method according to claim 6, wherein The final address vector generated by concatenating the semantic enhancement vector and the multimodal vector, and writing the final address vector into the address vector library, includes: Calculate the hash value of the final address vector, and write the final address vector to the target node among multiple deployed nodes according to the hash value; Among them, query in the address vector library according to the query information input by the user on the interaction interface, and display the query result on the interaction interface, including: Judge whether the query task is an exact query task or a fuzzy query task according to the query information input by the user on the interaction interface; If it is an exact query task, query in the cache server according to the exact query task; If it is hit in the cache server, display the query result returned by the cache server on the interaction interface; If it is not hit in the cache server, calculate the hash value according to the query information, determine the target node among the multiple nodes according to the calculated hash value, route the exact query task to the target node, receive the query result returned by the target node, and generate the final query result on the interaction interface; If it is a fuzzy query task, query in the cache server according to the fuzzy query task; If it is hit in the cache server, display the query result returned by the cache server on the interaction interface; If the cache server is not hit, broadcast the fuzzy query task to the multiple nodes, aggregate the query results returned by each node to generate the final query result, and display the final query result on the interaction interface.
8. An address search device based on deep learning, characterized in that, It includes: An acquisition unit for acquiring each multimodal data in the multimodal address database, where the multimodal data includes: text address, spatial coordinates, surrounding POIs, and street view images; A conversion unit for converting the text address into a text address vector, converting the spatial coordinates into a spatial vector, converting the surrounding POIs into POI vectors, and converting the street view images into street view image vectors; A generation unit for performing weighted summation on the text address vector, the spatial vector, the POI vector, and the street view image vector to generate a multi-modal vector; dividing the text address into multiple entities, querying corresponding semantic enhancement information in a knowledge base using the entities, encoding the queried semantic enhancement information to generate a semantic enhancement vector; concatenating the semantic enhancement vector and the multi-modal vector to generate a final address vector, and writing the final address vector into an address vector library; A query unit for querying in the address vector library according to query information input by a user in an interaction interface, and displaying a query result on the interaction interface.
9. A computer storage medium, characterized in that, The computer storage medium stores multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the method steps described in any one of claims 1 to 7.
10. A server, characterized in that, Comprising: A processor and a memory; wherein, the memory stores a computer program, and the computer program is suitable for being loaded and executed by the processor to perform the method steps described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-layer penetration query method and device for graph database
CN116244479A
Address retrieval method and device fusing text vector and longitude and latitude
CN116561449A
Method and system for optimizing man-machine conversation based on LLM model
CN117271735A
Information processing method and device, electronic equipment and storage medium
CN119026690A
Methods for determining a user's location using POI visibility inference
WO2013059734A1
Cited By
Method and system for replying petition document based on MCP protocol and large model
CN120832404A
Data co-processing method and system for multi-energy charging station, and program product
CN121350111A
Automobile channel intelligent question and answer method and system based on large language model
CN121524196A
Smart city-oriented multi-source heterogeneous address data intelligent modeling method and system
CN121962501A