House address analysis method, electronic equipment, storage medium and program product

By combining identification model and language model, the problem of inaccurate house address resolution in the existing technology is solved, and address analysis of multiple expression forms is realized, which can be divided into units, floors, room numbers and other levels in detail to improve the analysis effect.

CN120373285APending Publication Date: 2025-07-25KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510426478.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

It is difficult for the existing technology to accurately split and identify house addresses according to hierarchical structures, and the analysis effect is poor. Especially in areas with diverse expressions, the analysis rules are difficult to cover all possible address structures.

Method used

By identifying the first target hierarchy text in the specified house address hierarchy structure from the address text, determining the target expression form with a higher expression similarity than the threshold, and using the language model to analyze it with prompt words to achieve accurate splitting of the address.

Benefits of technology

It can analyze data such as units, floors, room numbers in detail, cover multiple expression forms, and realize the precise splitting and identification of house addresses according to hierarchical structure, improving analysis accuracy and iterability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373285A_ABST
    Figure CN120373285A_ABST
Patent Text Reader

Abstract

The invention provides a house address analysis method, electronic equipment, a readable storage medium and a computer program product, and the analysis method comprises the steps: firstly, recognizing a text corresponding to a first target hierarchy in a specified house address hierarchy structure from an address text of a geographic position where a house is located, and obtaining a first target text; wherein the house address hierarchical structure comprises N hierarchies, the first target hierarchy comprises n1 hierarchies in the N hierarchies, and N is greater than or equal to n1gt; the method comprises the following steps of: 1, determining a target expression form of which the similarity with an expression form of a first target text is higher than a similarity threshold value from a plurality of expression forms suitable for a first target level, and obtaining a cue word constructed based on the text applying the target expression form, and finally, taking information including the first target text and the cue words as input of a language model, and analyzing the first target text through the language model to obtain analyzed texts corresponding to the first target levels in the first target text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of computers, etc., and particularly relates to a method for parsing house addresses, an electronic device, a readable storage medium, and a computer program product. Background Art

[0002] House addresses are applied in many fields such as food delivery logistics and real estate transactions. House addresses include multiple different levels, such as levels of buildings, buildings, units, room numbers, etc.

[0003] Currently, the parsing of house addresses online mainly parses house addresses through set parsing rules to obtain the specific values of different levels in the house address. Since it is difficult for the parsing rules to cover all possible address structures, it is difficult to parse the specific values of each level in this way, and it is impossible to accurately split and identify the house address according to the hierarchical structure, resulting in a poor parsing effect. Summary of the Invention

[0004] The present disclosure provides a method for parsing house addresses, an electronic device, a readable storage medium, and a computer program product.

[0005] In a first aspect of the present disclosure, a method for parsing a house address is proposed, including: identifying, from the address text of the geographical location where the house is located, the text corresponding to the first target level in the specified hierarchical structure of the house address to obtain the first target text, where the specified hierarchical structure of the house address includes N levels, the first target level includes n1 levels among the N levels, N≥n1>1, and both N and n1 are positive integers; determining, from multiple expression forms applicable to the first target level, the target expression form whose similarity to the expression form of the first target text is higher than the similarity threshold, to obtain a prompt word constructed based on the text applying the target expression form, where at least one of the values of n1 corresponding to different expression forms and the names of the n1 levels is different; and using the information including the first target text and the prompt word as the input of a language model, and parsing the first target text through the language model to obtain the parsed text corresponding to each of the first target levels in the first target text.

[0006] According to some embodiments of the present disclosure, identifying text corresponding to a first target level in a specified housing address hierarchy from the address text of the geographical location where the house is located to obtain a first target text includes: identifying text corresponding to a second target level in the specified housing address hierarchy from the address text of the geographical location where the house is located through an identification model to obtain a second target text, where the second target level includes n2 levels out of the N levels, N = n1 + n2, and n2 ≥ 1; and determining the first target text from the remaining text in the address text excluding the second target text, where the first target text corresponds to the first target level in the housing address hierarchy.

[0007] According to some embodiments of the present disclosure, the first target level and the second target level are adjacent in the hierarchical order in the housing address hierarchy, and the lowest level in the second target level is higher than the highest level in the first target level.

[0008] According to some embodiments of the present disclosure, determining the first target text from the remaining text in the address text excluding the second target text includes: splitting the text located after the position of the second target text from the address text and using it as the first target text.

[0009] According to some embodiments of the present disclosure, the first target level includes, from high to low levels in sequence, a building number level, a unit level, a floor level, and a room number level.

[0010] According to some embodiments of the present disclosure, the second target level includes a real estate project level.

[0011] According to some embodiments of the present disclosure, before identifying the text corresponding to the first target level in the specified housing address hierarchy, the method further includes: fine-tuning the identification model that can be used for geographical text recognition using the annotated housing address text, where in the annotated housing address text, the annotated text only includes the text corresponding to the second target level.

[0012] According to some embodiments of the present disclosure, the annotation method is an annotation method based on named entity recognition.

[0013] According to some embodiments of the present disclosure, before identifying the text corresponding to the first target level in the specified housing address hierarchy, the method further includes: clustering the sample text containing the text corresponding to the first target level to obtain multiple expression categories.

[0014] According to some embodiments of the present disclosure, determining a target expression form whose similarity with the expression form of the first target text is higher than a similarity threshold from multiple expression forms applicable to the first target level includes: determining at least one target expression category with the highest similarity with the expression form of the first target text from the multiple expression categories, and using the expression forms included in the at least one target expression category as the target expression form.

[0015] According to some embodiments of the present disclosure, clustering sample texts containing the text corresponding to the first target level to obtain multiple expression categories includes: classifying the sample texts containing the text corresponding to the first target level according to the geographical regions corresponding to the sample texts to obtain sample text sets of multiple different geographical regions, where the regional scope of the geographical region is larger than the scope of the region corresponding to the highest level in the second target level; and respectively clustering the sample texts in each sample text set of the geographical region to obtain the expression categories corresponding to each geographical region.

[0016] According to some embodiments of the present disclosure, determining at least one target expression category with the highest similarity with the expression form of the first target text from the multiple expression categories includes: determining a same-region expression category corresponding to the geographical region where the location of the house is located from the expression categories corresponding to the geographical regions; and determining at least one target expression category with the highest similarity with the expression form of the first target text from the same-region expression category.

[0017] According to some embodiments of the present disclosure, after obtaining the multiple expression categories, the method further includes: determining at least one example text from the sample texts in each expression category, where the determined example text can represent all the expression forms included in the expression category; and constructing a prompt word for each expression category containing the corresponding example text and storing the constructed prompt words.

[0018] According to some embodiments of the present disclosure, obtaining a prompt word constructed based on the text applying the target expression form includes: determining the prompt word corresponding to the target expression category containing the target expression form from the stored prompt words.

[0019] A second aspect of the present disclosure provides an electronic device, including: a memory that stores execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the method according to any one of the above embodiments.

[0020] In a third aspect of the present disclosure, a readable storage medium is provided, in which a computer program is stored, and when the computer program is executed by a processor, it is used to implement the method described in any of the above embodiments.

[0021] In a fourth aspect of the present disclosure, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, it is used to implement the method described in any of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.

[0023] Figure 1 FIG. shows a schematic diagram of an application scenario of a method for parsing a house address according to some embodiments of the present disclosure.

[0024] Figures 2 - 7 FIG. shows a schematic diagram of the overall flow of a method M100 for parsing a house address according to some embodiments of the present disclosure.

[0025] Figure 8 FIG. is a schematic block diagram of a device for parsing a house address according to an embodiment of the present disclosure.

[0026] Figure 9 FIG. is a schematic block diagram of an electronic device 1000 according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The present disclosure will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant content and do not limit the present disclosure. Additionally, it should be noted that for the sake of convenience of description, only parts related to the present disclosure are shown in the drawings.

[0028] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the drawings and embodiments.

[0029] Unless otherwise specified, the exemplary embodiments / embodiments shown are understood to provide exemplary features of various details of some ways that can implement the technical concept of the present disclosure in practice. Therefore, unless otherwise specified, without departing from the technical concept of the present disclosure, the features of various embodiments / embodiments can be additionally combined, separated, interchanged, and / or rearranged.

[0030] The terms used in this document are for the purpose of describing specific embodiments and are not restrictive. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are also intended to include the plural forms. In addition, when the terms "comprise" and / or "include" and their variants are used in this specification, it is stated that the stated features, integers, steps, operations, components, assemblies and / or groups thereof exist, but do not preclude the existence or addition of one or more other features, integers, steps, operations, components, assemblies and / or groups thereof. It should also be noted that, as used herein, the terms "substantially", "about" and other similar terms are used as approximate terms and not as terms of degree, so they are used to explain the inherent deviations of measured values, calculated values and / or provided values that would be recognized by a person of ordinary skill in the art.

[0031] The parsing of house addresses is involved in a variety of scenarios. For example, when filling in the specific address for logistics distribution, or when entering the address of a housing unit in a database to build and update a building dictionary, or when verifying false and duplicate housing listings, and so on. In these scenarios, it is necessary to parse the house address according to the house address hierarchical structure to obtain specific values at different levels, such as the specific building number at the building level, the specific unit number at the unit level, the specific room number at the room number level, and so on.

[0032] The currently common way to parse a house address is to use a house address parsing interface to parse the house address according to the set parsing rules. This way can correctly parse data such as province, city, district, county, etc., but cannot parse in detail to data such as unit, floor, room number, etc. For example, the house address text A1 is "Building e in Community w, No. r-t-y", where w is the community name, e is the building number, and r, t, y are the unit number, floor and room number in sequence, and w, e, r, t, y are the specific values at different address levels. If the preset parsing rules include the exact same expression form as the above format, then the parsing can be correctly performed to obtain the specific values of the community, building, unit, floor, and room number, otherwise the parsing will fail.

[0033] The expression forms of house addresses adopted in different regions are different, that is, the expression forms of house addresses are diverse. Taking "No. 1-6-2, Building 24" in text A1 as an example, text A1 indicates that the address is located at No. 2, 6th floor, Unit 1, Building 24. Some regions may adopt the expression form of text A2, "Residential Building 24, 010602", and some regions may adopt the expression form of text A3, "24-010602". Even in some regions, the unit is not used to distinguish the address, and the expression forms are diverse. It is difficult to cover all expression forms only through parsing rules, so it is impossible to accurately split and identify some house addresses according to the hierarchical structure, and the parsing effect is poor.

[0034] To this end, the present disclosure proposes a method for parsing a house address.

[0035] Figure 1 The schematic diagram of the application scenario of the method for parsing a house address according to some embodiments of the present disclosure is shown. In this application scenario, a client 10 and a server 20 may be included.

[0036] The client 10 can communicate with the server 20 to perform data or instruction transceiver. In the present disclosure, both the client 10 and the server 20 each include at least one processor and at least one memory.

[0037] Exemplarily, the client 10 may be the terminal device of an enterprise employee, the server 20 may include the back-end server of the enterprise. The client 10 may parse the house address according to the specified house address hierarchy structure to obtain the specific values of each level, and then store the values of these levels into the house source database of the server 20, so as to add and maintain the building dictionary.

[0038] In Figure 1 The shapes and structures of the client 10 and the server 20 shown should not be construed as limiting the protection scope of the present disclosure. In the present disclosure, the "terminal device" may be different types of electronic devices. For example, the terminal device may be a tablet computer, a laptop computer or a desktop computer, etc. In addition, the back-end server of the server 20 may be a server with a physical form or a cloud server, and the present disclosure does not limit the type of the server.

[0039] Figure 2 The overall flow schematic diagram of the method M100 for parsing a house address according to some embodiments of the present disclosure is shown. As Figure 2 shown, the method includes step S110, step S120 and step S130. Among them, the method may be executed by an electronic device such as a computer.

[0040] S110, identify the text corresponding to the first target level in the specified house address hierarchy structure from the address text of the geographical location where the house is located, and obtain the first target text.

[0041] The address text contains the geographical location information where the house is located. The house address hierarchy structure includes N levels, N>1, that is, the hierarchy structure includes multiple different levels, such as building level, unit level, floor level and room number level, etc. Each level corresponds to only one specific value in the address text of a house.

[0042] Identifying the first target text from the address text can be achieved through an identification model. The identification model can be a pre-trained model for geographical text, such as the MGeo (Multi-modal multi-task Geographic) model that can be used for geographical text identification. Inputting the address text into the identification model can output the text corresponding to the first target level (the first target text).

[0043] The first target level includes n1 levels out of N levels, where N≥n1>1, that is, the first target level includes multiple different levels. For example, it can include multiple or all of the building number level, unit level, floor level, and room number level. It can be understood that both N and n1 are positive integers.

[0044] When N>n1, it means that the specified housing address level structure includes other levels in addition to the first target level. In addition to the first target text, the address text may also include other texts, and these other texts may correspond to other levels. Step S110 only needs to identify the first target text corresponding to the first target level. For example, the address text B1 can be "No. 24 Building, Unit 1-6-2, Happy Community", where the other text is "Happy Community", corresponding to the real estate level in the level structure, and the first target text G identified from B1 is "No. 24 Building, Unit 1-6-2".

[0045] When N=n1, it means that the specified housing address level structure only includes the first target level, but in addition to the first target text, the address text may also include other texts, and these other texts do not correspond to any level. For example, the address text B2 is "The address is No. 24 Building, Unit 1-6-2", where the other text is "The address is", which does not correspond to any level, and the first target text G identified from B2 is "No. 24 Building, Unit 1-6-2".

[0046] S120. Determine the target expression form whose similarity with the expression form of the first target text is higher than the similarity threshold from multiple expression forms applicable to the first target level, and obtain the prompt word constructed based on the text applying the target expression form.

[0047] The above-mentioned multiple expression forms applicable to the first target level are pre-configured and stored in a specified location. These existing expression forms can be represented by existing example texts. For example, the above texts A1, A2, and A3 are different expression forms of the same address and can all be used as example texts to represent one of the expression forms respectively.

[0048] The way to determine whether the expression forms are similar can be to calculate the cosine similarity between the vector of the first target text and the vectors of the existing example texts, and use the one or more example texts with a similarity higher than the similarity threshold or the highest similarity as the texts with similar expression forms, that is, the texts corresponding to the target expression forms; or calculate the Levenshtein distance between two texts, obtain the similarity through the distance, the lower the distance, the higher the similarity, and use the one or more example texts with a distance lower than a certain threshold or the closest distance as the texts with similar expression forms, that is, use the one or more example texts with a similarity higher than the similarity threshold or the highest similarity as the texts with similar expression forms, and obtain the texts corresponding to the target expression forms.

[0049] At least one of the value of n1 and the hierarchical names of the n1 levels corresponding to different expression forms is different. The hierarchical name is equivalent to the quantifier of the level, such as names like "building", "block", "unit", "floor", "number", etc. In the address texts under different expression forms, the hierarchical names used for some levels may be different. For example, different regions use "Building X", "Block X", and "Building X" respectively to represent the building level, so the hierarchical names of the building level in the address texts are different, that is, the expression ways are different. In addition, the number of levels included in different expression forms may also be different. For example, some real estate projects do not set units, so there is no unit level in the corresponding expression form. Compared with the address text with a unit level, the expression forms of the two are different.

[0050] For example, the expression form of the above first target text G is specifically: use "building" to express the building level, use the symbol "-" to connect the unit, floor, and room number, and use "number" to express the room number. Among the existing expression forms, there are C1, C2, and C3. Among them, C1 is specifically: use the symbol "-" to connect the building, unit, floor, and room number. The example text of C1 is "10-1-2-1", representing Building 10, Unit 1, Floor 2, Room 1. The expression form of C2 is the same as that of text G. The example text of C2 is "Building 6, 5-18-1802". C3 is specifically: use the symbol "-" to connect the building and the floor, there is no unit level, and the floor and the room number are directly adjacent. The example text of C3 is "7-1003". Then through step S120, it is obtained that the example texts of C1 and C2 have a higher similarity with text G, and the example text of C3 has a lower similarity with text G. Therefore, it can be determined that C1 and C2 are the target expression forms, and the example texts of C1 and C2 are obtained.

[0051] After obtaining the example text corresponding to the target expression form, a prompt can be generated based on the example text, that is, the prompt can be generated during the execution of step S120. The prompt can include the original text of the example text and the parsing result after the example text is correctly parsed. For example, the prompt can include the example text and parsing result of C1, and the example text and parsing result of C2. Exemplarily, the parsing result of C1 in the prompt can be: "Building 10 > Unit 1 > Floor 2 > No. 1", and the parsing result of C2 can be: "Building 6 > Unit 5 > Floor 18 > No. 1802". It can be understood that the ">" symbol in the parsing result is used to separate the hierarchical values of different levels, and the format of the parsing result can also use other delimiter symbols or other formats, which are not limited in this disclosure.

[0052] It can be understood that the correct parsing result of the example text is pre-configured and has a mapping relationship with the example text. After determining the example text, the correct parsing result of the example text can be obtained. In addition, the prompt can also be pre-generated and have a mapping relationship with the determined example text. For example, all possible combination methods of the existing example texts are determined in advance, and a corresponding prompt is generated for each combination method. When executing step S120, the corresponding prompt is obtained according to the determined example text.

[0053] S130, use the information including the first target text and the prompt as the input of the language model, and parse the first target text through the language model to obtain the parsed text corresponding to each of the first target levels in the first target text.

[0054] The language model can be a large language model. For example, an AIGC (Artificial Intelligence Generated Content) model can be used, such as the Qwen-14B model (a high-performance open-source model that supports multiple languages), etc., a pre-trained generative language model.

[0055] Input the prompt and "Building 24, Unit 1, No. 6, Floor 2" as the first target text G into the language model. The language model will parse text G according to the prompt of the prompt to obtain the correct parsed text. The number of parsed texts obtained is n1, and the n1 texts respectively correspond to n1 levels.

[0056] The following uses a specific house address text as an example for illustration. Assume that the address text is the above-mentioned text B1, "No. 24 Building, Unit 1, Room 602, Happy Community". The hierarchical structure of the house address includes 5 levels: the property, building, unit, floor, and room number, that is, N = 5. Among them, the first target level includes the remaining 4 levels except the property, that is, n1 = 4. First, through step S110, the text G corresponding to the first target level is identified from text B1, that is, "No. 24 Building, Unit 1, Room 602", and then the target expression forms C1 and C2 with a similar expression form to text G are determined, and the prompt words including C1 and C2, the example text, and the corresponding parsing results are obtained. The prompt words and text G are input into the Qwen-14B model, and the parsed text "Building 24 > Unit 1 > Floor 6 > Room 2" output by the Qwen-14B model is obtained.

[0057] According to the parsing method of the house address proposed by the embodiment of the present disclosure, by identifying the approximate expression form of the house address through the pre-constructed possible expression forms, and using the language model to normalize the parsing of the house address based on the approximate expression form, it is possible to parse in detail to data such as the unit, floor, and room number, and it can cover and be applicable to the parsing of house addresses in a variety of different expression forms, realizing the accurate splitting and identification of the house address according to the hierarchical structure, with good parsing effect, and it can completely replace the regular address parsing interface, improving the accuracy and iterability of address parsing.

[0058] Figure 3 The overall flow diagram of the parsing method M100 of the house address according to some other embodiments of the present disclosure is shown. Refer to Figure 3 , step S110 may specifically include the following steps S111 and S112.

[0059] S111, identify the text corresponding to the second target level in the specified hierarchical structure of the house address in the address text at the geographical location where the house is located through the identification model, and obtain the second target text. Among them, the second target level includes n2 levels among the N levels, N = n1 + n2, and n2 ≥ 1.

[0060] S112, determine the first target text from the remaining text in the address text except the second target text. Among them, the first target text corresponds to the first target level in the hierarchical structure of the house address.

[0061] The hierarchical structure of the house address can consist only of the second target level and the first target level. By first identifying the second target text from the address text and then identifying the first target text from the remaining text, the first target text can be identified in the address text. The address text can include three parts of text. The first part of the text corresponds to the second target level, the second part of the text corresponds to the first target level, and the third part of the text does not correspond to any hierarchical structure of the house address. The first part of the text is identified through step S111, and then the second part of the text is identified from the remaining text including the second part of the text and the third part of the text.

[0062] The second target level includes one or more different levels. For example, the second target level can only include the property level, or the second target level can successively include the city level, the urban area level, and the property level from high to low. The first target level can successively include the building level, the unit level, the floor level, and the room number level from high to low. The hierarchical order of the first target level and the second target level in the hierarchical structure of the house address can be adjacent, and the lowest level in the second target level can be higher than the highest level in the first target level.

[0063] Exemplarily, for example, the address text A4 is "The address is No. 1-6-2, Building 24, Happy Community", then the second target text corresponding to the second target level is identified from the text A4 through the MGeo model. The second target level only includes the property level, and the second target text is "Happy Community". At this time, the remaining text is "The address is No. 1-6-2, Building 24", and the first target text G is identified as "No. 1-6-2, Building 24". It can be understood that the identified second target text can also be used as one of the parsed texts obtained through parsing, that is, the parsed text includes the hierarchical value of the property level. Then, through step S111 and step S112, the address parsing problem is split into two parts. The first part is property parsing, and the second part is building, unit, floor, and room number parsing.

[0064] Figure 4 The overall flowchart of the house address parsing method M100 according to some other embodiments of the present disclosure is shown. Refer to Figure 4 , step S112 can specifically be: splitting the text located after the position of the second target text from the address text and using it as the first target text.

[0065] The address text can only include the first target text and the second target text, that is, it does not include the third part of the text that does not correspond to any hierarchical structure of the house address. At this time, identifying the second target text is equivalent to identifying the first target text.

[0066] For example, if the address text B1 is "No. 1-6-2, Building 24, Xingfu Community", the second target text H corresponding to the second target level is identified from the text A4 through the MGeo model. The second target level only includes the property level, and the second target text H is "Xingfu Community". At this time, the remaining text is "The address is No. 1-6-2, Building 24", and the first target text G identified therefrom is "No. 1-6-2, Building 24".

[0067] Continue to refer to Figure 4 , and the parsing method M100 may further include step S101. Step S101 is executed before step S110 and is completed before step S110 starts to be executed.

[0068] S101, fine-tune the recognition model that can be used for geographical text recognition by using the house address text with annotations. Among them, in the house address text with annotations, the annotated text only includes the text corresponding to the second target level.

[0069] The above recognition model that can be used for geographical text recognition can adopt the above MGeo model. The MGeo model can combine text information and geographical information, break through the traditional longitude and latitude limitations, and achieve accurate recognition of the specific shapes and adjacent inclusion relationships of geographical entities such as scenic spots and parks. The POI (Point of Interest) information in the address text is identified through the fine-tuned MGeo model. The POI information is the second target text including the property name. The above house address text with annotations can be the text in the text set for fine-tuning. By only annotating the second target text therein and performing fine-tuning, the MGeo model can better identify the second target text, and specifically improve the recognition effect of the second target level.

[0070] The annotation method can be the annotation method based on Named Entity Recognition (NER), such as the BIOES annotation method. The BIOES annotation method assigns a label to each word in the text to identify its position and role in the named entity. BIOES is the abbreviation of five tags, where B represents the start of a named entity; I represents the middle part of a named entity, that is, the part that does not belong to the beginning but is still inside the named entity; O represents that the word does not belong to any named entity; E represents the end part of a named entity; S represents a named entity composed of a single word alone, which is applicable to entities composed of only a single word. For example, for the text "No. 1-6-2, Building 24, Xingfu Community", the annotations of "Xing", "Fu", "Xiao", and "Qu" are B, I, I, and E in sequence, and "Xingfu Community" belongs to the POI content. The remaining content in this text is not annotated and does not belong to the POI content.

[0071] Figure 5 The overall flowchart of the parsing method M100 for the house address according to other embodiments of the present disclosure is shown. Refer to Figure 5 , the parsing method M100 may further include step S102. Step S102 is executed before step S110 and is completed before step S110 starts to be executed. The execution order between step S102 and step S101 can be arbitrary.

[0072] S102, clustering the sample texts containing the texts corresponding to the first target level to obtain multiple expression categories. Wherein, the expression forms of the sample texts included in each expression category are similar.

[0073] The sample texts are used to provide addresses in different expression forms. For example, texts A1, A2, and A3 can respectively provide three different expression forms of the same address. The clustering method can adopt a distance-based clustering algorithm, such as the K-means algorithm, a hierarchical clustering algorithm, a density-based clustering algorithm, etc. Specifically, the semantic vectors of each sample text can be calculated first, and then the K-means clustering algorithm is used to cluster the semantic vectors to obtain multiple expression categories, where the similarity between the semantic vectors in the same expression category satisfies the second similarity threshold. Or it can also be to calculate the Levenshtein distance between each sample text first, and then perform DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering or hierarchical clustering on the Levenshtein distance to obtain multiple expression categories. Among them, the smaller the value of the Levenshtein distance, the higher the similarity, so that the texts with high similarity are clustered into the same expression category.

[0074] One or more sample texts can be included in the same expression category. The expression forms of the sample texts in the same expression category are similar. For example, the semantic vector similarity between these sample texts is higher than the similarity threshold or the Levenshtein distance is lower than the distance threshold, while the semantic vector similarity between two sample texts belonging to different expression categories may be lower than the similarity threshold or the Levenshtein distance is higher than the distance threshold.

[0075] Correspondingly, step S120 may include step S121 and step S122.

[0076] S121, determining at least one target expression category with the highest similarity between the expression form and the first target text from the above multiple expression categories, and taking the expression forms included in the at least one target expression category as the target expression forms.

[0077] S122, obtain a prompt word constructed based on the text with the target expression form applied thereto.

[0078] The method for determining the target expression category of the first target text may be to calculate the vector similarity or Levenshtein distance between the first target text and the sample texts of each expression category. Then, for each expression category, determine the maximum value of the vector similarity or the minimum value of the Levenshtein distance between the sample texts included therein and the first target text, to obtain the maximum similarity or minimum distance between each expression category and the first target text. Then, use one or more expression categories with the highest maximum similarity or one or more expression categories with the lowest minimum distance as the target expression category.

[0079] Alternatively, it may also be that after obtaining each expression category, first calculate the central value of each expression category. For example, use the vector average value of the sample texts as the central value of the expression category; or for each expression category, calculate the sum of the cosine distances between each sample text in the expression category and other sample texts in the same expression category, and use the vector of the sample text with the smallest sum of cosine distances as the central value of the expression category. After obtaining the central value, it can be stored. Then, when determining the target expression category, calculate the vector similarity between the first target text and the central values of each expression category, and then screen out the target expression category through a specified threshold.

[0080] Then use the sample texts included in the target expression category as example texts to construct or obtain the corresponding prompt word.

[0081] Figure 6 Shows the overall flowchart of the house address parsing method M100 according to some other embodiments of the present disclosure. Refer to Figure 6 , step S102 may specifically include step S102a and step S102b.

[0082] S102a, classify the sample texts containing the text corresponding to the first target level according to the geographical regions corresponding to the sample texts, to obtain multiple sample text sets of different geographical regions. Among them. The regional scope of the geographical region is larger than the scope of the region corresponding to the highest level in the second target level.

[0083] S102b, perform clustering of the sample texts for each sample text set of each geographical region, to obtain the expression category corresponding to each geographical region.

[0084] The geographical area can be a city, whose area range is larger than that of a real estate project. There will be at least one real estate project (the second target level) in the city. The geographical area can also be a larger area than a city, such as a geographical region (including Northeast, North China, Central China, East China, South Central, Northwest, Southwest, etc.). First, classify the sample texts according to the geographical area. The sample texts belonging to the same geographical area are divided into the same category to form a sample text set, thus obtaining multiple sample text sets. Cluster each sample text set, and one or more expression categories are obtained after clustering each sample text set.

[0085] Correspondingly, step S121 can specifically include step S121a and step S121b.

[0086] S121a, determine the same-region expression category corresponding to the geographical area where the location of the house is located from the expression categories corresponding to the above geographical area.

[0087] S121b, determine at least one target expression category with the highest similarity between the expression form of the first target text and the expression forms in the same-region expression category, and use the expression forms included in the above at least one target expression category as the target expression form.

[0088] After identifying the first target text G through step S110, if the address of text G is in city R, first determine the expression category of city R, that is, the same-region expression category, through step S121a. If the sample texts corresponding to city R are clustered into 5 categories when clustering through step S102b before, then determine the target expression category of text G from these 5 expression categories, and obtain the sample texts in the target expression category and use them as example texts.

[0089] Since the address expression forms within the same geographical area may be relatively similar, or a certain expression form only exists in the same geographical area, clustering by geographical area and determining the target expression category with the highest text similarity or the smallest text distance from the geographical area can make the constructed prompt word closer to the target expression form, with higher accuracy and hit rate, and thus make the parsing result more accurate.

[0090] Figure 7 Shows the overall flowchart of the parsing method M100 for the house address in some other embodiments of the present disclosure. Refer to Figure 7 , the parsing method M100 can also include step S103 and step S104. Step S103 and step S104 are executed sequentially after completing step S102.

[0091] S103, determine at least one example text from the sample texts in each expression category. Among them, the determined example text can represent all the expression forms included in the expression category.

[0092] S104. Construct prompt words for each expression category that include corresponding example texts, and store the constructed prompt words.

[0093] Correspondingly, in step S120, the manner of obtaining the prompt words constructed based on the text to which the target expression form is applied may specifically be: determining the prompt words corresponding to the target expression category that includes the target expression form from the stored prompt words.

[0094] The prompt words may be constructed before starting to parse the house address text to be parsed, and after determining the target expression form, directly read the constructed appropriate prompt words and input them into the language model, which can improve the parsing efficiency. Specifically, each expression category may correspond to a prompt word, and the specific content of the prompt word may be constructed based on the sample text in the expression category. The prompt word of an expression category may include all the sample texts in that expression category. If the expression forms of some sample texts are the same, it may also only include some of the sample texts in that expression category, that is, for multiple sample texts with the same expression form, one of the sample texts may be selected to represent that expression form as the example text.

[0095] Based on any of the above embodiments, the present disclosure also provides an apparatus for parsing a house address. Figure 8 It is a structural schematic block diagram of an apparatus for parsing a house address according to an embodiment of the present disclosure. As Figure 8 shown, the apparatus for parsing a house address includes: a first target text recognition module 110, a prompt word acquisition module 120, and a house address parsing module 130.

[0096] The first target text recognition module 110 is configured to recognize, from the address text of the geographical location where the house is located, the text corresponding to the first target level in the specified house address hierarchical structure, and obtain the first target text. The house address hierarchical structure includes N levels, the first target level includes n1 levels among the N levels, N≥n1>1, and both N and n1 are positive integers.

[0097] The prompt word acquisition module 120 is configured to determine, from multiple expression forms applicable to the first target level, a target expression form whose similarity to the expression form of the first target text is higher than a similarity threshold, and obtain the prompt words constructed based on the text to which the target expression form is applied.

[0098] The housing address parsing module 130 uses the information including the first target text and the prompt word as the input of the language model, and parses the first target text through the language model to obtain the parsed text corresponding to each of the first target levels in the first target text. At least one of the values of n1 corresponding to different expression forms and the level names of the n1 levels is different.

[0099] The above-mentioned housing address parsing device may be in the form of computer software, and each module of the above-mentioned housing address parsing device may be implemented by computer software modules. The implementation processes of the functions and roles of each module in the above-mentioned device are specifically described in the implementation processes of the corresponding steps in the above-mentioned method, and will not be elaborated here.

[0100] The execution subject of the housing address parsing method in the specific implementation manner of the present disclosure may be an electronic device such as a computer or a server.

[0101] Therefore, based on any one of the above embodiments, the present disclosure further provides an electronic device, which can execute the housing address parsing method of any one of the above embodiments described in the present disclosure. Figure 9 It is a structural schematic diagram of an electronic device 1000 according to an embodiment of the present disclosure.

[0102] The hardware structure of the electronic device 1000 can be implemented by using a bus architecture. The bus architecture may include any number of interconnecting buses and bridges, depending on the specific application and overall design constraints of the hardware. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0103] The bus 1100 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only one connecting line is shown in this figure, but it does not mean that there is only one bus or one type of bus.

[0104] Among them, the processor 1200 may be a Central Processing Unit (CPU). The processor 1200 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips.

[0105] The memory 1300 can be used as a non-transitory computer-readable storage medium and can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions of the computer program in the embodiments of the present disclosure. The processor 1200 realizes the method for parsing the house address by running the non-transitory software programs, instructions, and modules stored in the memory 1300.

[0106] The memory 1300 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created by the processor 1200, such as address text, information on the house address hierarchy, the first target text, the second target text, data of the recognition model, data of the language model, sample text, prompt words, parsed text, etc. In addition, the memory 1300 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1300 may optionally include a memory remotely set relative to the processor 1200, and these remote memories may be connected to the processor 1200 through a network. Examples of the above networks include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0107] The present disclosure also provides a readable storage medium storing a computer program, which is used to implement the above method when executed by a processor. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.

[0108] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM, or other programmable devices.

[0109] The computer program or instructions can be stored in a readable storage medium, or transmitted from one readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any accessible medium or a data storage device such as a server or data center integrating one or more accessible media. The accessible medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.

[0110] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0111] This disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0112] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0114] In the description of this specification, the descriptions referring to the terms "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc., mean that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of this disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.

[0115] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0116] Those skilled in the art should understand that the above-described embodiments are merely for clearly illustrating the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or variations can be made based on the above disclosure, and these changes or variations are still within the scope of the present disclosure.

Claims

1. A method for parsing a house address, characterized in that, Including: Identifying, from the address text of the geographical location where the house is located, the text corresponding to the first target level in the specified house address hierarchy to obtain a first target text, where the specified house address hierarchy includes N levels, the first target level includes n1 levels among the N levels, N≥n1>1, and both N and n1 are positive integers; Determining, from multiple expression forms applicable to the first target level, a target expression form whose similarity to the expression form of the first target text is higher than a similarity threshold, and obtaining a prompt word constructed based on the text with the target expression form applied, where at least one of the values of n1 corresponding to different expression forms and the level names of the n1 levels is different; And Using the information including the first target text and the prompt word as the input of a language model, and parsing the first target text through the language model to obtain the parsed text corresponding to each of the first target levels in the first target text.

2. The parsing method of the house address according to claim 1, wherein Identifying, from the address text of the geographical location where the house is located, the text corresponding to the first target level in the specified house address hierarchy to obtain a first target text, including: Identifying, through an identification model, the text corresponding to the second target level in the specified house address hierarchy in the address text of the geographical location where the house is located to obtain a second target text, where the second target level includes n2 levels among the N levels, N=n1 + n2, and n2≥1; and Determining the first target text from the remaining text in the address text excluding the second target text, where the first target text corresponds to the first target level in the house address hierarchy.

3. The method for parsing a house address according to claim 2, wherein The first target level and the second target level are adjacent in the level order in the house address hierarchy, and the lowest level in the second target level is higher than the highest level in the first target level.

4. The parsing method of the house address according to claim 2 or 3, characterized in that, Determining the first target text from the remaining text in the address text excluding the second target text, including: Splitting the text located after the position of the second target text from the address text and using it as the first target text.

5. The parsing method of the housing address according to any one of claims 1-3, characterized in that The first target level sequentially includes, from high to low, a building number level, a unit level, a floor level, and a room number level.

6. The parsing method of the house address according to claim 2 or 3, characterized in that, The second target level includes a real estate project level.

7. The method for parsing a housing address according to any one of claims 1 to 6, characterized in that, Before identifying the text corresponding to the first target level in the specified house address hierarchy, the method further includes: Fine-tuning the identification model that can be used for geographical text identification with the house address text with annotations, where in the house address text with annotations, the annotated text only includes the text corresponding to the second target level; Optionally, the annotation method is an annotation method based on named entity recognition; Optionally, before identifying the text corresponding to the first target level in the specified house address hierarchy, the method further includes: Clustering the sample text containing the text corresponding to the first target level to obtain multiple expression categories; Optionally, determining a target expression form with a similarity higher than a similarity threshold between the expression forms applicable to the first target level and the expression form of the first target text includes: Determining at least one target expression category with the highest similarity between the expression forms in the multiple expression categories and the expression form of the first target text, and using the expression forms included in the at least one target expression category as the target expression form; Optionally, clustering sample texts containing the text corresponding to the first target level to obtain multiple expression categories, including: Classifying the sample texts containing the text corresponding to the first target level according to the geographical regions corresponding to the sample texts, to obtain multiple sets of sample texts for different geographical regions, where the regional scope of the geographical region is larger than the scope of the region corresponding to the highest level in the second target level; and Respectively clustering the sample texts in each set of sample texts for the geographical region to obtain the expression categories corresponding to each geographical region; Optionally, determining at least one target expression category with the highest similarity between the expression forms in the multiple expression categories and the expression form of the first target text includes: Determining the same-region expression category corresponding to the geographical region where the location of the house is located from the expression categories corresponding to the geographical regions; and Determining at least one target expression category with the highest similarity between the expression forms in the same-region expression category and the expression form of the first target text; Optionally, after obtaining the multiple expression categories, the method further includes: Determining at least one example text from the sample texts in each expression category, where the determined example text can represent all the expression forms included in the expression category; and Constructing a prompt word for each expression category containing the corresponding example text, and storing the constructed prompt word; Optionally, obtaining the prompt word constructed based on the text applying the target expression form includes: Determining the prompt word corresponding to the target expression category containing the target expression form from the stored prompt words.

8. An electronic device, characterized in that, Includes: A memory that stores execution instructions; And A processor that executes the execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1 to 7.

9. A readable storage medium, characterized in that, The computer program stored in the readable storage medium is used to implement the method according to any one of claims 1 to 7 when executed by a processor.

10. A computer program product, characterized in that, The computer program product includes a computer program, which is used to implement the method according to any one of claims 1 to 7 when executed by a processor.