Data construction method, device, computer storage medium and electronic device

By fuzzy matching the official name and address database information of the building, combined with the structured analysis of surrounding address information, the second name knowledge base of the building is built, and the data service accuracy problems caused by alias in the existing technology are solved, and more efficient and intelligent positioning and search functions are achieved.

CN119150977BActive Publication Date: 2025-06-03TAOBAO CHINA SOFTWARE +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411604498.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-06-03
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

In the prior art, when performing data processing or data services based on the official name of a building, there are problems of accuracy and limitations, especially because the data services are affected by the existence of alias.

Method used

By fuzzingly matching the first name of the target object with the address information in the address library, position coordinate information is determined, and a second name set is determined through the structured unit information of the surrounding address information, the second name knowledge base of the target object is constructed.

Benefits of technology

It improves the accuracy and accuracy of the position coordinate information of the target object, expands the name expression form of the target object, enhances the intelligence level of the search function, and ensures fast and accurate positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119150977B_ABST
    Figure CN119150977B_ABST
Patent Text Reader

Abstract

The present application discloses a data construction method, apparatus, computer storage medium, and electronic device. The method includes: matching the first name of the target object obtained with the address information in the address library to determine the position coordinate information of the target object; determining a set of second names corresponding to the target object in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected based on the position coordinate information; constructing a second name knowledge base of the target object according to the extended names of the second names determined by extending the second names in the set of second names; which can improve the intelligent level of the search function of the target object, ensure that it can quickly understand and respond to the diverse query needs of users, accelerate the search process of the target object, improve the user experience, and ensure that the position of the target object can be quickly and accurately located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data mining technology, and specifically relates to a data construction method and device. This application also relates to a building alias construction method and device, a computer storage medium, and an electronic device. Background Art

[0002] With the continuous development of Internet and computer technologies, more and more service requirements based on the Internet and computers are widely applied, such as navigation, delivery and other services. Therefore, positioning is particularly important because the accuracy of positioning determines the accuracy of services.

[0003] However, any kind of positioning needs to be based on a specific address or location. In real-life scenarios, buildings usually have a formal name, that is, a standard appellation, and there are also multiple other names different from the formal name, that is, aliases. Due to the existence of aliases, it will affect the data service or data processing process. Summary of the Invention

[0004] This application provides a data construction method to solve the limitations and accuracy problems caused in the prior art during the data processing or data service process based on the formal name.

[0005] This application provides a data construction method, including:

[0006] Matching the obtained first name of the target object with the address information in the address library to determine the position coordinate information of the target object;

[0007] Determining a set of second names corresponding to the target object in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected based on the position coordinate information;

[0008] Constructing a knowledge base of the second name of the target object according to the extended name of the second name determined by extending the second name in the set of second names.

[0009] In some embodiments, the matching the obtained first name of the target object with the address information in the address library to determine the position coordinate information of the target object includes:

[0010] Obtaining the first name of the target object according to the established information database of the target object;

[0011] Performing a fuzzy match between the first name and the address information in the established address library to determine the position coordinate information related to the first name.

[0012] In some embodiments, the fuzzy matching of the first name with the address information in the established address library to determine the position coordinate information related to the target object includes:

[0013] Fuzzy match the first name with the address information in the established address library to determine candidate address information related to the target object;

[0014] According to the latitude and longitude data corresponding to the candidate address information, select the median of the latitude and longitude data, or select the mode of the latitude and longitude data;

[0015] Determine the latitude and longitude data corresponding to the median or the mode as the position coordinate information corresponding to the target object.

[0016] In some embodiments, the determination of the set of second names corresponding to the target object in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected for the position coordinate information includes:

[0017] Select the surrounding address information of the target object according to the position coordinate information;

[0018] Analyze the surrounding address information to determine the structured unit information of the surrounding address information;

[0019] Determine the set of second names corresponding to the target object according to the structured unit information.

[0020] In some embodiments, the selection of the surrounding address information of the target object according to the position coordinate information includes:

[0021] Taking the position coordinate information as the center, perform grid coding processing on the search area within a preset range to determine the grid coding area;

[0022] Search for surrounding address information according to the grid coding area to obtain the surrounding address information of the target object.

[0023] In some embodiments, the analysis of the surrounding address information to determine the structured unit information of the surrounding address information includes:

[0024] Perform structured processing on the surrounding address information according to an address parsing tool to obtain the structured unit information constituting the surrounding address information.

[0025] In some embodiments, the construction of the second name knowledge base of the target object according to the extended name of the second name determined by the name extension of the second name in the set of second names includes:

[0026] Update the second name set according to the filtering process of the second names in the second name set;

[0027] Perform name expansion on the second names in the updated second name set to obtain the expanded names of the second names;

[0028] Construct the second name knowledge base of the target object according to the expanded names;

[0029] In some embodiments, the updating the second name set according to the filtering process of the second names in the second name set includes:

[0030] Filter the names whose occurrence frequencies do not meet the frequency filtering requirements according to the counted occurrence frequencies of the second names in the second name set, and / or filter the names that do not meet the similarity requirements according to the similarity between the second names in the second name set and the first name;

[0031] Update the second names in the second name set.

[0032] In some embodiments, the performing name expansion on the second names in the updated second name set to obtain the expanded names of the second names includes:

[0033] Perform name expansion on the second names in the updated second name set by using the label propagation method to obtain the expanded names.

[0034] In some embodiments, the constructing the second name knowledge base of the target object according to the expanded names of the second names determined by expanding the second names in the second name set includes:

[0035] Perform cleaning processing on the expanded names;

[0036] Construct the second name knowledge base of the target object according to the cleaned expanded names.

[0037] This application also provides a data construction device, including:

[0038] A first determination unit, configured to match the first name of the obtained target object with the address information in the address library to determine the position coordinate information of the target object;

[0039] A second determination unit, configured to determine the second name set corresponding to the target object in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected based on the position coordinate information;

[0040] A building block for constructing a second name knowledge base of the target object according to the extended name of the second name determined by the extension of the second name in the second name set.

[0041] This application also provides a method for constructing an alias of a building, including:

[0042] Match the obtained official name of the building with the address information in the address library to determine the location coordinate information of the building;

[0043] Determine the alias set corresponding to the building in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected based on the location coordinate information;

[0044] Construct a knowledge base of the alias of the building according to the extended name of the alias determined by the extension of the alias in the alias set.

[0045] This application also provides a computer storage medium, including a computer program, which, when running on an electronic device, enables the electronic device to execute the data construction method as described above, or execute the method for constructing an alias of a building as described above.

[0046] This application also provides an electronic device, including:

[0047] A processor;

[0048] A memory for storing a program for processing data generated by the electronic device, and when the program is read and executed by the processor, it executes the data construction method as described above, or executes the method for constructing an alias of a building as described above.

[0049] Compared with the prior art, this application has the following advantages:

[0050] A data construction method provided by this application can obtain the location coordinate information of a target object through fuzzy matching based on the first name of the target object and the address information in the address library, and further improve the accuracy and precision of the location coordinate information through the median or mode of longitude and latitude data. Then, the structured information of the surrounding address information obtained is used to determine the second name of the target object, so that in addition to having the first name, the target object can also have the expression form of the second name. In this process, the effectiveness and accuracy of the second name can be further improved through corresponding filtering methods. Furthermore, to ensure the comprehensiveness of the second name, the filtered second name can be expanded, and the expanded name is also used as the second name to establish the knowledge base of the second name of the target object, so as to improve the intelligent level of the search function of the target object in various subsequent scenario requirements, ensure that the diverse query requirements of users can be quickly understood and responded to, accelerate the search process of the target object, improve the user experience, and ensure that the location of the target object can be quickly and accurately located.

[0051] In a building alias construction method provided by this application, the building alias is mined using the receiving address information, and the alias knowledge base is constructed accordingly. Through the fuzzy matching of the official name of the building and the address, the longitude and latitude information is obtained, and the alias information that may be used to express the building is mined through the longitude and latitude information. Furthermore, the alias knowledge base can be established by widely collecting the adjacent receiving address information. In application scenarios such as real estate data processing, a more comprehensive and practical alias knowledge base is provided. Brief Description of the Drawings

[0052] Figure 1 is a flowchart of a data construction method provided by this application.

[0053] Figure 2 is a schematic structural diagram of a data construction device provided by this application.

[0054] Figure 3 is a flowchart of a building alias construction method provided by this application.

[0055] Figure 4 is a schematic structural diagram of a building alias construction device provided by this application.

[0056] Figure 5 is a schematic structural diagram of an electronic device provided by this application. Detailed Description of the Embodiments

[0057] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0058] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The description methods used in the present application and the appended claims, such as "a", "first", and "second", etc., are not limitations on the quantity or the order, but are used to distinguish information of the same type from each other.

[0059] Based on the above background technology, it can be seen that aliases other than the official name of a certain object are of great significance for data processing and data services. In the prior art, traditional alias recognition methods rely on keyword extraction and comparison technologies. By refining key elements in information descriptions, such as street names, prominent landmarks, and other relevant information, and pairing these keywords with existing community databases. This traditional method has intuitiveness and simplicity, but also has several limitations. For example, it cannot handle or inadequately handle situations such as abbreviations, spelling mistakes, or expression diversity, which may lead to problems such as omissions or incorrect matches in the recognition results.

[0060] Based on the above, the present application provides a data construction method, as Figure 1 shown, Figure 1 is a flowchart of a data construction method provided by the present application. The method includes:

[0061] Step S101: Match the obtained first name of the target object with the address information in the address library to determine the position coordinate information of the target object;

[0062] Step S102: Determine a set of second names corresponding to the target object in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected based on the position coordinate information;

[0063] Step S103: Construct a knowledge base of the second names of the target object according to the extended names of the second names determined by extending the second names in the set of second names.

[0064] The above steps S101 - S103 will be described in detail below in combination with specific implementation methods. Before the detailed description, some technical terms involved in the embodiments of the present application are explained.

[0065] Formal Name: It refers to a common name, official name, or standard name. In the real estate field, each property carries its unique identity and value. The formal name, as the standard appellation of the property in legal documents, official records, and market promotion, ensures the uniqueness and accuracy of the property, and is the basis for property ownership, transaction records, and administrative management. For example, the formal name of a certain community, the formal name of an office building, etc.

[0066] Alias: An additional name given by the market, the public, or a specific community based on the formal name. These aliases may stem from a prominent feature of the project, historical background, surrounding environment, or association with a well-known landmark, or sometimes because of a widely spread story or event. Aliases are often more vivid, easy to remember, and full of emotional color. They can quickly establish recognition and association among the target audience and become an important symbol of the uniqueness of the building.

[0067] LCS (Longest Common Subsequence): An algorithm used to measure string similarity, which evaluates the similarity of two sequences by finding the longest common subsequence between them.

[0068] LPA (Label Propagation Algorithm): A semi-supervised learning method aimed at classifying by propagating the labels of labeled neighbors to unlabeled data points, and is applicable to community detection and data clustering.

[0069] Regarding step S101: Match the first name of the obtained target object with the address information in the address library to determine the location coordinate information of the target object.

[0070] In step S101, the target object can be a target building, such as: residential community, apartment, office building, shopping mall, etc. Of course, it can also include some historical and cultural heritage buildings, tourist attraction buildings, etc. The first name is the formal name of the target object. The address library can be a pre-established database that stores address information corresponding to buildings. In this embodiment, the address library can be a receiving address library.

[0071] The purpose of step S101 is to determine the location coordinate information of the target object. The specific implementation process may include:

[0072] Step S101-1: Obtain the first name of the target object according to the established information database of the target object;

[0073] Step S101-2: Perform fuzzy matching between the first name and the address information in the established address library to determine the coordinate information related to the target object.

[0074] The information database in step S101-1 can be a pre-established database for storing information related to the target object. For example, it can store the official name of the target building, the construction time, the structural features, etc. Through the information database, the first name corresponding to the target object can be obtained. In this embodiment, the first name can also be understood as the official name, or it can be understood as the identification information of the target object.

[0075] In step S101-2, it can be achieved by performing fuzzy matching between the first name and a pre-established receiving address database, that is, performing fuzzy matching between the first name and the address information in the receiving address database. In this embodiment, common matching algorithms can be used for the fuzzy matching, such as the Levenshtein distance (edit distance), Jaccard similarity (Jaccard similarity), Cosine similarity (cosine similarity), TF-IDF algorithm, and n-gram model, etc. Among them, the Levenshtein distance is an algorithm based on editing operations, including insertion, deletion, and replacement operations. The smaller the edit distance, the more similar the two strings are. The Jaccard similarity is used to compare the similarity of two sets, and the similarity is measured by calculating the ratio of the intersection to the union. For text data, the text can be segmented into words or characters as the elements of the set. The Cosine similarity measures the angle between two vectors. Regarding the text as vectors, each dimension represents a word or a feature, and the similarity is measured by calculating the angle between the vectors. The TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is a text feature extraction method, which measures the importance of words by calculating the word frequency and the inverse document frequency, and can be used to calculate the similarity between texts or perform fuzzy matching. The n-gram model calculates the similarity between texts based on the frequency of consecutive subsequences. Common values of n include unigram (one-word segmentation), that is, n = 1, bigram (two-word segmentation), that is, n = 2, and trigram (three-word segmentation), that is, n = 3. Through fuzzy matching, the location coordinate information related to the target object can be obtained. For the receiving address database, since it stores a large amount of receiving address information, the address information related to the target object will be stored in the receiving address database. Therefore, multiple location coordinate information related to the target object can be obtained through fuzzy matching. In this embodiment, it can be filtered and judged based on whether the receiving address data contains the first name, that is, retaining the data similar to the first name.

[0076] To improve the accuracy and precision of the location coordinate information, in this embodiment, the specific process of step S101-2 can include:

[0077] Step S101-211: Perform fuzzy matching between the first name and the address information in the established address library to determine candidate address information related to the target object;

[0078] Step S101-212: According to the latitude and longitude data corresponding to the candidate address information, select the median of the latitude and longitude data, or select the mode of the latitude and longitude data;

[0079] Step S101-213: Determine the latitude and longitude data corresponding to the median or the mode as the position coordinate information corresponding to the target object.

[0080] Among them, when the mode is the latitude and longitude at which the frequency of occurrence is significantly higher than other latitude and longitude data, that is, when it becomes the mode, it is considered that this latitude and longitude can represent the relatively accurate position of the target object.

[0081] In addition to the method of selecting latitude and longitude data based on the above median or mode to determine the position coordinate information corresponding to the first name, it can also be determined by the mean value method. Therefore, the specific process of step S101-2 can also be implemented in the following way:

[0082] Step S101-221: Perform fuzzy matching between the first name and the address information in the established address library to determine candidate address information related to the target object;

[0083] Step S101-222: According to the latitude and longitude data corresponding to the candidate address information, select the average value of the latitude and longitude data;

[0084] Step S101-223: Determine the latitude and longitude data corresponding to the average value as the position coordinate information corresponding to the target object.

[0085] The above are at least two implementation methods for determining the position coordinate information of the target object. In fact, there can also be other implementation methods. For example: through empirical values, or by inputting the latitude and longitude data corresponding to the candidate address information into an AI intelligent large model to obtain the position coordinate information corresponding to the target object. Therefore, the specific implementation methods of step S101-2 are not limited to the two specific examples given above.

[0086] Regarding step S102: According to the analysis of the surrounding address information selected for the position coordinate information, determine the set of second names corresponding to the target object in the structured unit information of the surrounding address information.

[0087] The surrounding address information mentioned in step S102 may also be surrounding delivery address information, that is, based on the position coordinate information of the target object, the surrounding address information is selected. The structured unit information may be a constituent unit or part of the address information, such as: street, house number, name, POI (point of interest), AOI (area of interest), orientation, etc.

[0088] The purpose of step S102 is to: centrally determine a second name set corresponding to the target object according to the structured unit information of the surrounding address information. Of course, it can also be understood as a second name set corresponding to the first name. The specific implementation process may include:

[0089] Step S102-1: Select the surrounding address information of the target object according to the position coordinate information;

[0090] Step S102-2: Parse the surrounding address information to determine the structured unit information of the surrounding address information;

[0091] Step S102-3: Determine a second name set corresponding to the target object according to the structured unit information.

[0092] Among them, the specific implementation process of step S102-1 may include:

[0093] Step S102-11: Centering on the position coordinate information, perform grid coding processing on the search area within a preset range to determine the grid coding area; in this embodiment, the position coordinate of the target object can be used as the center, and the preset range is within a radius of 100 meters for searching.

[0094] Step S102-12: Search for surrounding address information according to the grid coding area to obtain the surrounding address information of the target object.

[0095] In this embodiment, to improve the search efficiency, grid coding processing can be performed on the determined search area, so that the preset range is a grid space, which is beneficial to improving the search efficiency and can also ensure the search accuracy and precision. It can still maintain efficient search, as well as search accuracy and precision in the face of a large amount of data. Of course, grid coding processing can also be performed on the entire surrounding area. To avoid resource waste, grid coding processing can be performed only on the preset range.

[0096] The specific implementation process of step S102-2 may include:

[0097] Step S102-21: Perform structured processing on the surrounding address information according to an address parsing tool to obtain the structured unit information constituting the surrounding address information.

[0098] In this embodiment, the address information can be deeply structured and parsed by methods such as Ptolemy. The parsing process can improve the automation level of data structured processing, and can also improve the flexibility and depth of subsequent analysis and processing. For example: political division information includes information such as province, city, county, and township; road network information includes road names, road numbers, road facilities, etc.; detailed address information includes community names, building numbers, house numbers, etc. Through the address parsing method, the delivery address text is parsed by using the text sequence annotation method. For example, No. 969, Wenyi West Road, Wuchang Street, Yuhang District, Hangzhou City, Zhejiang Province, a certain park can be parsed as province=Zhejiang Province, city=Hangzhou City, district=Yuhang District, town=Wuchang Street, road=Wenyi West Road, road_number=969, poi=a certain park.

[0099] The specific implementation process of step S102-3 may include:

[0100] Step S102-31: Count the structured unit information and select the unit information whose number of occurrences meets the selection requirement as the second name; specifically, the POI information or other structured unit information may be counted in the structured unit information, or some structured unit information may be mixed and then the unit information whose number of occurrences is counted is used as the second name;

[0101] Step S102-31: Establish the second name set according to the second name.

[0102] Regarding step S103: constructing a second name knowledge base of the target object according to the extended name of the second name determined by extending the second name in the second name set.

[0103] In step S103, the so-called extension can be understood as expanding the second name based on the recognition of the second name, thereby increasing the alias range corresponding to the target object, that is, the range of the second name, that is, the second name can include multiple (the alias can include multiple).

[0104] In order to ensure the accuracy and validity of the second name in the constructed second name knowledge base, the specific implementation process of step S103 may include:

[0105] Step S103-11: updating the second name set according to filtering processing of the second name in the second name set; in this embodiment, the filtering processing can also be understood as screening processing, that is, filtering out aliases that do not meet the selection requirements from the second name set, and the specific implementation process may include at least two methods:

[0106] Method 1: Filter the names whose occurrence frequencies do not meet the frequency filtering requirements according to the occurrence frequencies of the second names in the second name set. For example, by setting a first threshold, when the number of occurrences of a second name is less than the first threshold, it is excluded from the second name set.

[0107] Method 2: Filter the names that do not meet the similarity requirements according to the similarity between the second names in the second name set and the first name. The implementation process of Method 2 can be achieved through the Longest Common Subsequence (LCS), that is, calculate the similarity between the second name and the first name. For example, to calculate the Longest Common Subsequence (LCS) of "Binxing Home" and "Binxing Residential Quarter", the common characters and their order of the two strings can be found first.

[0108] String:

[0109] Let A = "Binxing Home" and B = "Binxing Residential Quarter". The common part of the two strings is "Binxing Home", and the longest common subsequence is 3. Combining the lengths of the two strings for coverage calculation, the final similarity is 2 * 3 / (4 + 4) = 0.75. Then, by setting a second threshold, it is determined whether the similarity is greater than or equal to the second threshold. If so, it is retained; otherwise, it is excluded. In some other embodiments, methods such as edit distance and cosine similarity can also be used to calculate the similarity between the second name and the first name.

[0110] The two methods can be executed sequentially or one of them can be selected. Of course, in this embodiment, for the sake of easy understanding, the occurrence frequency and similarity are used as examples for illustration, but it is not limited to the above two methods. For example, names that do not meet the requirements can also be excluded through confidence levels or empirical values, etc.

[0111] Step S103-12: Expand the second names in the updated second name set to obtain the extended names of the second names; the name expansion method given in this embodiment is the label propagation method (LPA: Label Propagation Algorithm) or the label propagation algorithm to implement name clustering, and further expand the name similarity text set, so as to extract more second names. For example: Based on the similarity obtained by the above-mentioned longest common subsequence (LCS), construct a target object name similarity graph model, where the nodes are target objects (such as: buildings), and edges of target objects are constructed when the similarity exceeds a certain threshold. For existing names, assign a unique label. For unknown names, their own IDs can be used as the initial label. Run LPA through the following steps: Iteratively update the labels: For the current node, view the neighbor nodes of the current node and update its own label according to the majority label of the neighbor nodes. Stop condition: Keep iterating until the labels no longer change or reach a certain number of iterations. After multiple iterations, obtain the final labels of each second name: Based on the final labels, cluster the nodes with the same label together, so as to determine the extended names in the second name set similar to the first name.

[0112] Step S103-13: Construct the second name knowledge base according to the extended names. To further improve the accuracy and precision of the second names, redundant information can be removed. Therefore, when constructing the second name knowledge base, it can include:

[0113] Step S103-131: Clean up the extended names; for example, remove abnormal information such as special characters carried in the extended names, and the special characters can be such as "-", "." and so on.

[0114] Step S103-132: Construct the second name knowledge base of the target object according to the cleaned extended names.

[0115] The above is a description of an embodiment of a data construction method provided by the present application. This method can obtain the location coordinate information of the target object by means of fuzzy matching based on the first name of the target object and the address information in the address library, and further improve the accuracy and precision of the location coordinate information by using the median or mode of the longitude and latitude data. Then, the structured information of the surrounding address information obtained is used to determine the second name of the target object, so that in addition to having the first name, the target object can also have the expression form of the second name. In this process, the effectiveness and accuracy of the second name can be further improved through corresponding filtering methods. Moreover, to ensure the comprehensiveness of the second name, the filtered second name can be expanded, and the expanded name can also be used as the second name to establish the second name knowledge base of the target object, so as to improve the intelligent level of the search function of the target object in various subsequent scenario requirements, ensure that it can quickly understand and respond to the diverse query requirements of users, accelerate the search process of the target object, enhance the user experience, and ensure that the location of the target object can be quickly and accurately located.

[0116] The above is a specific description of an embodiment of a data construction method provided by the present application. Corresponding to the embodiment of the data construction method provided above, the present application also discloses an embodiment of a data construction device. Please refer to Figure 2 , since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0117] As Figure 2 shown, Figure 2 is a schematic structural diagram of a data construction device provided by the present application. The device includes: a first determination unit 201, a second determination unit 202, and a construction unit 203;

[0118] The first determination unit 201 is configured to match the obtained first name of the target object with the address information in the address library to determine the location coordinate information of the target object;

[0119] The second determination unit 202 is configured to determine a set of second names corresponding to the target object in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected for the location coordinate information;

[0120] The construction unit 203 is configured to construct a second name knowledge base of the target object according to the extended names of the second names determined by expanding the second names in the set of second names.

[0121] Among them, the specific implementation process of the first determination unit 201 includes: an acquisition subunit and a determination subunit; the acquisition subunit is configured to acquire a first name of the target object according to the established information database of the target object; the determination subunit is configured to perform fuzzy matching on the first name with the address information in the established address library to determine the position coordinate information related to the first name.

[0122] The specific implementation process of the determination subunit may include: a matching subunit, a selection subunit, and a position determination subunit; the matching subunit is configured to perform fuzzy matching on the first name with the address information in the established address library to determine candidate address information related to the target object; the selection subunit is configured to select the median of the longitude and latitude data corresponding to the candidate address information, or select the mode of the longitude and latitude data; the position determination subunit is configured to determine the longitude and latitude data corresponding to the median or the mode as the position coordinate information corresponding to the target object.

[0123] For the specific content of the first determination unit 201, reference may be made to the relevant description of step S101 above, and details are not described herein again.

[0124] The specific implementation process of the second determination unit 202 may include: a selection subunit, a first determination subunit, and a second determination subunit; the selection subunit is configured to select the surrounding address information of the target object according to the position coordinate information; specifically, it may include: an encoding subunit, which is configured to perform grid encoding processing on the search area within a preset range centered on the position coordinate information to determine the grid encoding area; a search subunit, which is configured to search for the surrounding address information according to the grid encoding area to acquire the surrounding address information of the target object. The first determination subunit is configured to parse the surrounding address information to determine the structured unit information of the surrounding address information; specifically, it is configured to perform structured processing on the surrounding address information according to an address parsing tool to acquire the structured unit information constituting the surrounding address information. The second determination subunit is configured to determine a second name set corresponding to the target object according to the structured unit information.

[0125] For the specific content of the second determination unit 202, reference may be made to the relevant description of step S102 above, and details are not described herein again.

[0126] The specific implementation process of the third determination unit 203 may include: a filtering subunit, an obtaining subunit, and a constructing subunit; the filtering subunit is used to update the second name set according to the filtering process of the second names in the second name set; the specific implementation process may include: a first filtering subunit and / or a second filtering subunit, the first filtering subunit is used to filter the names whose occurrence frequency does not meet the frequency filtering requirement according to the counted occurrence frequency of the second names in the second name set; the second filtering subunit is used to filter the names that do not meet the similarity requirement according to the similarity between the second names in the second name set and the first name; the updating subunit is used to update the second names in the second name set. The obtaining subunit is used to perform name expansion on the second names in the updated second name set to obtain the extended names of the second names; specifically, it is used to perform name expansion on the second names in the updated second name set by using the label propagation method to obtain the extended names. The constructing subunit is used to construct the second name knowledge base according to the extended names; the specific implementation process may include: a cleaning subunit, which is used to perform cleaning processing on the extended names; the constructing subunit is specifically used to construct the second name knowledge base of the target object according to the extended names cleaned by the cleaning subunit.

[0127] The above is the description of an embodiment of a data construction device provided by the present application. The specific content of this device embodiment can refer to the content of the method embodiment, and will not be elaborated here in detail.

[0128] Based on the above content, the present application also provides a method for constructing building aliases, as Figure 3 shown, Figure 3 is the flowchart of a method for constructing building aliases provided by the present application. This method is mainly a description of the method for constructing building aliases. The building can be a community or a residential area, or an office building, etc. That is, the process of constructing the alias of the community according to the official name of the community. The specific implementation process may include:

[0129] Step S301: Match the obtained official name of the building with the address information in the address library to determine the location coordinate information of the building;

[0130] Step S302: Determine the alias set corresponding to the building in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected based on the location coordinate information;

[0131] Step S303: Construct the alias knowledge base of the building according to the extended names of the aliases determined by alias expansion in the alias set.

[0132] For the specific content of the above steps S301 to S303, reference can be made to the description of steps S101 to S103 in the above-mentioned data construction method embodiment, which will not be repeated here.

[0133] Based on the above, in the building alias construction method provided by this application, the building alias is mined using the delivery address information, and an alias knowledge base is constructed accordingly. Through the fuzzy matching of the official building name and the address, the latitude and longitude information is obtained. The alias information that may be used to express the building around is mined through the latitude and longitude information, and then the alias knowledge base can be established by widely collecting the adjacent delivery address information. In application scenarios such as real estate data processing, a more comprehensive and practical alias knowledge base is provided.

[0134] Correspondingly, this application also provides a building alias construction device, as Figure 4 shown, Figure 4 FIG. is a schematic structural diagram of a building alias construction device provided by this application. The device embodiment includes: a first determination unit 401, a second determination unit 402, and a construction unit 403.

[0135] The first determination unit 401 is configured to match the obtained official building name with the address information in the address library to determine the position coordinate information of the building;

[0136] The second determination unit 402 is configured to determine the alias set corresponding to the building in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected for the position coordinate information;

[0137] The construction unit 403 is configured to construct the alias knowledge base of the building according to the extended name of the alias determined by the alias extension in the alias set.

[0138] For the specific content of this building alias construction device, reference can also be made to the content of the above-mentioned data construction method embodiment and the content of the building alias construction method embodiment, which will not be repeated here.

[0139] Based on the above, this application also provides a computer storage medium, including a computer program. When the computer program runs on an electronic device, the electronic device is enabled to execute the relevant content involved in the data construction method as described above, or execute the relevant content involved in the building alias construction method as described above.

[0140] Based on the above, this application also provides an electronic device, as Figure 5 shown, Figure 5 FIG. is a schematic structural diagram of an electronic device provided by this application. The electronic device includes:

[0141] Processor 501;

[0142] A memory 502 for storing a program for processing data generated by an electronic device. When the program is read and executed by the processor, it executes the relevant content involved in the data construction method described above, or executes the relevant content involved in the building alias construction method described above.

[0143] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0144] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0145] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0146] 1. Computer-readable media include permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media such as modulated data signals and carrier waves.

[0147] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0148] Although the present application is disclosed above in preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be determined by the scope defined by the claims of the present application.

Claims

1. A data construction method, characterized in that: include: Performing fuzzy matching on the acquired first name of the target object and the address information in the address database to determine candidate address information related to the target object, and determining the location coordinate information of the target object according to the latitude and longitude data corresponding to the candidate address information; Determine, based on the analysis of the surrounding address information selected from the location coordinate information, a second name set corresponding to the target object in the structured unit information of the surrounding address information; including: counting the structured unit information, selecting unit information whose number of occurrences meets the selection requirement as the second name, and establishing the second name set based on the second name; A second name knowledge base of the target object is constructed according to an extended name of the second name determined by extending the second name in the second name set.

2. The data construction method according to claim 1, characterized in that: The step of fuzzily matching the acquired first name of the target object with the address information in the address database to determine candidate address information related to the target object, and determining the location coordinate information of the target object according to the latitude and longitude data corresponding to the candidate address information, includes: Acquire a first name of the target object according to the established information database of the target object; The first name is fuzzily matched with the address information in the established address database to determine the location coordinate information related to the target object.

3. The data construction method according to claim 2, characterized in that: The step of determining the location coordinate information of the target object according to the latitude and longitude data corresponding to the candidate address information includes: According to the longitude and latitude data corresponding to the candidate address information, selecting the median of the longitude and latitude data, or selecting the mode of the longitude and latitude data; The longitude and latitude data corresponding to the median or the mode is determined as the position coordinate information corresponding to the target object.

4. The data construction method according to claim 1, characterized in that: The step of determining, based on the analysis of the surrounding address information selected from the location coordinate information, a second name set corresponding to the target object in the structured unit information of the surrounding address information includes: Selecting the surrounding address information of the target object according to the location coordinate information; Parsing the peripheral address information to determine structured unit information of the peripheral address information; A second name set corresponding to the target object is determined according to the structured unit information.

5. The data construction method according to claim 4, characterized in that: The selecting the surrounding address information of the target object according to the position coordinate information includes: Taking the position coordinate information as the center, grid coding is performed on the search area within a preset range to determine a grid coding area; The surrounding address information is searched according to the grid coding area to obtain the surrounding address information of the target object.

6. The data construction method according to claim 4, characterized in that: The step of parsing the surrounding address information to determine the structured unit information of the surrounding address information includes: The peripheral address information is structured according to an address resolution tool to obtain structural unit information constituting the peripheral address information.

7. The data construction method according to claim 1, characterized in that: The step of constructing a second name knowledge base of the target object according to the extended name of the second name determined by extending the second name in the second name set includes: updating the second name set according to filtering the second names in the second name set; Performing name extension on the updated second name in the second name set to obtain an extended name of the second name; The second name knowledge base is constructed according to the extended name.

8. The data construction method according to claim 7, characterized in that: The updating of the second name set according to the filtering process of the second name in the second name set includes: According to the statistical occurrence frequency of the second name in the second name set, filtering the names whose occurrence frequency does not meet the frequency filtering requirement, and / or, according to the similarity between the second name in the second name set and the first name, filtering the names that do not meet the similarity requirement; The second name in the second name set is updated.

9. The data construction method according to claim 7 or 8, characterized in that: The step of extending the updated second name in the second name set to obtain an extended name of the second name includes: The name is extended according to the updated second name in the second name set by using a label propagation method to obtain the extended name.

10. The data construction method according to claim 1, characterized in that: The step of constructing a second name knowledge base of the target object according to the extended name of the second name determined by extending the second name in the second name set includes: Cleaning up the extension name; A second name knowledge base of the target object is constructed according to the cleaned extended name.

11. A data construction device, characterized in that: include: a first determining unit, configured to perform fuzzy matching on the first name of the acquired target object and the address information in the address library to determine candidate address information related to the target object, and determine the position coordinate information of the target object according to the latitude and longitude data corresponding to the candidate address information; A second determination unit is used to determine a second name set corresponding to the target object in the structured unit information of the surrounding address information according to the analysis of the surrounding address information selected from the location coordinate information; including: counting the structured unit information, selecting unit information whose number of occurrences meets the selection requirement as the second name, and establishing the second name set according to the second name; The construction unit is used to construct a second name knowledge base of the target object according to the extended name of the second name determined by extending the second name in the second name set.

12. A method for constructing a building alias, characterized in that: include: Perform fuzzy matching on the acquired formal name of the building and the address information in the address database to determine candidate address information related to the building, and determine the location coordinate information of the building according to the latitude and longitude data corresponding to the candidate address information; Determine, based on the analysis of the surrounding address information selected from the location coordinate information, a set of aliases corresponding to the building in the structured unit information of the surrounding address information; including: performing statistics on the structured unit information, selecting unit information whose number of occurrences meets the selection requirements as aliases, and establishing the alias set based on the aliases; An alias knowledge base of the building is constructed according to the extended name of the alias determined by extending the alias in the alias set.

13. A computer storage medium, characterized in that: It comprises a computer program, which, when executed on an electronic device, enables the electronic device to execute the data construction method as described in any one of claims 1 to 10, or the building alias construction method as described in claim 12.

14. An electronic device, characterized in that: include: processor; A memory for storing a program for processing data generated by an electronic device, wherein when the program is read and executed by the processor, the program executes the data construction method as described in any one of claims 1 to 10, or executes the building alias construction method as described in claim 12.

Citation Information

Patent Citations

  • Address information processing method and device, and address information display method and device

    CN113569564A

  • Object label determination method and device, electronic equipment and storage medium

    CN113761387A

  • Address matching method and device, computer equipment, storage medium and program product

    CN117076790A