Repeating a housing source determination method, device, computer device and storage device
Patent Information
- Application Number
- CN202610528585.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]基于此,本发明的目的是提供一种重复房源判定方法、装置、计算机装置及存储装置,以从根本上解决现有技术在阈值适应性、环境感知能力及实时处理能力方面均存在不足的问题
[0018]本发明实施例提供的重复房源判定方法,获取待判定房源的地址文本信息和地理坐标信息,基于地理坐标信息计算局部环境密度值,基于局部环境密度值生成去重半径阈值,基于地址文本信息、地理坐标信息和去重半径阈值实时输出重复判定结果。通过基于局部环境密度动态生成的去重半径阈值,使系统能够自适应高、低密度区域的不同判定需求,显著降低误判率,并依托在线实时处理架构,实现了在数据提交源头即时拦截重复房源,有效提升了数据治理的精准性与实时性。
Smart Images

Figure CN122838652A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of information judgment technology, and particularly relates to a method, apparatus, computer device, and storage device for determining duplicate housing listings. Background Technology
[0002] In the data governance and deduplication of listings in the online rental (or short-term rental / homestay) sector, efficiently and accurately identifying and merging duplicate listings with the same physical location is a key technical step to ensure the management goal of "real person, real house, real address" and improve data quality and regulatory effectiveness. Currently, the mainstream technologies for identifying duplicate listings in the industry are mainly divided into deduplication methods based on fixed spatial thresholds and processing schemes based on general clustering algorithms.
[0003] The deduplication technique based on a fixed spatial threshold uses GPS coordinates to determine spatial proximity. Specifically, the system pre-sets a globally uniform spatial distance threshold (such as 50 meters or 100 meters). When the Euclidean distance between the geographic coordinates of a newly submitted property and the coordinates of an existing property is less than this threshold, it is determined to be a duplicate property.
[0004] However, the aforementioned fixed threshold method has significant drawbacks in practical applications: Because it fails to consider the local environmental differences in housing distribution, a globally uniform threshold is difficult to adapt to the judgment needs of areas with different densities. In high-density apartment areas, with compact building structures and small distances between units, a fixed threshold can easily lead to different independent housing units being misjudged as duplicates because their coordinates fall within the same threshold range, causing confusion of occupant information and failure of trajectory tracking. In low-density villa areas or self-built housing areas, due to GPS signal drift and large building spacing, the coordinates collected from the same house multiple times may deviate by tens of meters. A fixed threshold can easily misjudge the same house as different housing units, leading to duplicate filing and wasted police resources. Furthermore, this method lacks the ability to perceive information such as surrounding building types and environmental context, and cannot distinguish whether coordinate overlap is due to GPS measurement errors or physical spatial overlap, exhibiting low intelligence and poor generalization ability.
[0005] Batch processing techniques based on general clustering algorithms are employed by some big data platforms, using general density clustering algorithms such as DBSCAN (Density-Based Spatial Clustering of Applications with Noise) for deduplication of housing listings. These methods scan the global dataset, dividing densely accessible spatial points into the same cluster, thereby achieving offline identification and batch cleaning of duplicate housing listings.
[0006] However, the aforementioned clustering methods are primarily applicable to offline data governance scenarios and are insufficient to meet the real-time management requirements of online rental services. On one hand, algorithms like DBSCAN have high computational complexity, typically requiring a full scan of historical datasets, resulting in long execution cycles and often employing a T+1 offline batch processing mode. This makes it impossible to perform immediate judgments during real-time processes such as property listing and check-in registration. On the other hand, the core parameters of these algorithms, such as neighborhood radius and minimum number of points, are usually fixed values, lacking the ability to adaptively adjust based on local environmental characteristics. This leads to significant differences in application performance across different cities and business districts, necessitating frequent manual parameter tuning. More importantly, the offline data separation mode results in duplicate data being stored in the database before cleaning, leaving dirty data still present in front-end displays and during police queries, failing to achieve the "source governance" management objective. Summary of the Invention
[0007] Based on this, the purpose of the present invention is to provide a method, apparatus, computer device and storage device for determining duplicate listings, so as to fundamentally solve the problems of the existing technology in terms of threshold adaptability, environmental perception ability and real-time processing ability.
[0008] This invention is implemented as follows: a method for determining duplicate listings is provided, comprising the following steps: Obtain the address text information and geographic coordinate information of the property to be judged; A dynamic search window is determined centered on the geographic coordinate information, and the local environmental density value within the dynamic search window is calculated. Based on the local environment density value, a deduplication radius threshold is dynamically generated using a monotonically decreasing function; and Based on the address text information, the geographic coordinate information, and the deduplication radius threshold, the property to be judged is compared with the properties in the historical property database in multiple dimensions, and the duplicate judgment result is output in real time.
[0009] In some implementations, the process of obtaining the address text information includes the following steps: Obtain the initial address text information; The initial address text information is semantically encoded using a pre-trained language model to extract contextual features; Address text elements are identified from the context features based on sequence labeling algorithms; The address text elements are formatted and normalized using an unsupervised learning algorithm to generate the address text information.
[0010] In some implementations, the step of dynamically generating the deduplication radius threshold based on the local environment density value using a monotonically decreasing function includes the following steps: The deduplication radius threshold is calculated based on the following formula: , in, The deduplication radius threshold is... The minimum deduplication radius threshold. Based on the deduplication radius threshold, The attenuation coefficient is... The local environmental density value is given.
[0011] In some implementations, the step of performing a multi-dimensional comparison between the property to be judged and the properties in the historical property database based on the address text information, the geographic coordinate information, and the deduplication radius threshold, and outputting the duplicate judgment result in real time, includes the following steps: Based on the geographic coordinate information of the property to be determined and the historical geographic coordinate information of the properties in the historical property database, the geographic distance is calculated; When the geographical location interval is less than the deduplication radius threshold, it is determined whether the address text information of the property to be judged is consistent with the historical address text information of the property in the historical property database. If they are consistent, it is determined to be a duplicate property, and the duplicate judgment result is output in real time. If they are inconsistent, a second judgment is made based on the building type attribute of the property to be judged, and the duplicate judgment result is output in real time.
[0012] Another embodiment of the present invention provides a device for determining duplicate listings, comprising: The information acquisition unit is used to acquire the address text information and geographic coordinate information of the property to be judged; The density calculation unit is used to determine a dynamic search window centered on the geographic coordinate information and calculate the local environmental density value within the dynamic search window. The threshold generation unit is used to dynamically generate a deduplication radius threshold based on the local environment density value using a monotonically decreasing function; and The result output unit is used to perform a multi-dimensional comparison between the property to be judged and the properties in the historical property database based on the address text information, the geographic coordinate information and the deduplication radius threshold, and output the duplicate judgment result in real time.
[0013] In some embodiments, the information acquisition unit includes: The information acquisition module is used to acquire the initial address text information; The feature extraction module is used to semantically encode the initial address text information using a pre-trained language model and extract contextual features; The feature recognition module is used to identify address text features from the context features based on a sequence labeling algorithm; The information generation module is used to perform format normalization processing on the address text elements through an unsupervised learning algorithm to generate the address text information.
[0014] In some embodiments, the threshold generation unit includes: The formula calculation module is used to calculate the deduplication radius threshold based on the following formula: , Where R is the deduplication radius threshold, Rmin is the minimum deduplication radius threshold, Rbase is the basic deduplication radius threshold, k is the attenuation coefficient, and D is the local environment density value.
[0015] In some implementations, the result output unit includes: The spacing calculation module is used to calculate the geographical distance based on the geographical coordinate information of the property to be judged and the historical geographical coordinate information of the properties in the historical property database. The result output module is used to determine whether the address text information of the property to be judged is consistent with the historical address text information of the property in the historical property database when the geographical location distance is less than the deduplication radius threshold. If they are consistent, the property is judged as a duplicate property and the duplicate judgment result is output in real time. If they are inconsistent, a second judgment is made based on the building type attribute of the property to be judged and the duplicate judgment result is output in real time.
[0016] Another embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the duplicate housing information determination method as described in any of the preceding embodiments.
[0017] Another embodiment of the present invention provides a storage medium storing a computer program that can be executed to implement the steps of the duplicate listing determination method as described in any of the preceding embodiments.
[0018] The duplicate listing determination method provided in this invention obtains the address text information and geographic coordinate information of the listing to be determined, calculates the local environmental density value based on the geographic coordinate information, generates a deduplication radius threshold based on the local environmental density value, and outputs the duplicate listing determination result in real time based on the address text information, geographic coordinate information, and deduplication radius threshold. By dynamically generating the deduplication radius threshold based on the local environmental density, the system can adapt to different determination requirements in high-density and low-density areas, significantly reducing the false judgment rate. Furthermore, relying on an online real-time processing architecture, it achieves instant interception of duplicate listings at the data submission source, effectively improving the accuracy and real-time performance of data governance. Attached Figure Description
[0019] Figure 1 This is a flowchart of the duplicate housing information determination method provided in the embodiments of the present invention; Figure 2 This is a structural block diagram of the duplicate housing information determination device provided in the embodiments of the present invention. Detailed Implementation
[0020] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0021] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0023] This invention provides a method for determining duplicate listings, as described in the following embodiments. Figure 1 This includes the following steps: S100. Obtain the address text information and geographic coordinate information of the property to be judged; S200. Determine a dynamic search window centered on the geographic coordinate information, and calculate the local environmental density value within the dynamic search window; S300. Based on the local environment density value, dynamically generate a deduplication radius threshold using a monotonically decreasing function; and S400. Based on the address text information, the geographic coordinate information, and the deduplication radius threshold, perform a multi-dimensional comparison between the property to be judged and the properties in the historical property database, and output the duplicate judgment result in real time.
[0024] First, when a user submits property information on the property listing or check-in registration interface, the system retrieves the address text information based on the user's input, such as "Room XX, No. XX, XX Road, XX District, XX City". Simultaneously, the system uses the map provider's application programming interface (API) to obtain the corresponding geographic coordinates based on the user's input. These geographic coordinates may include longitude and latitude coordinates. The property corresponding to the aforementioned information is the property to be assessed; in some implementations, this property is a rental property.
[0025] Next, a dynamic search window is defined centered on the acquired geographic coordinates (denoted as point P). In a preferred embodiment, the size of the dynamic search window can be preset according to the actual business scenario, for example, set as a square area of 100 meters × 100 meters, or a circular area with a radius of 50 meters or 100 meters. The center of the dynamic search window is dynamically determined by the geographic coordinates of the property to be judged each time, but its size can be a preset fixed value. Within this dynamic search window, the system counts the number of properties already stored in the historical property database. Furthermore, to improve the accuracy of density calculation, the system can also count environmental context information such as building outline coverage density and point of interest type distribution within the window. The system uses the kernel density estimation algorithm (KDE) to fuse and calculate the above statistical data to obtain the local environmental density value at the location of point P. The local environmental density value reflects the degree of property clustering in the area centered on the property to be judged: the larger the local environmental density value, the denser the properties in the area; the smaller the local environmental density value, the sparser the properties in the area.
[0026] The system then uses the calculated local environmental density value as input and calculates the deduplication radius threshold using a monotonically decreasing function. This monotonically decreasing function has the following characteristics: when the local environmental density value increases, the deduplication radius threshold R decreases accordingly; when the local environmental density value decreases, the deduplication radius threshold increases accordingly. In this way, the deduplication radius threshold is dynamically adjusted according to the local environmental density: in high-density areas (such as apartment buildings and commercial centers), the deduplication radius threshold automatically decreases to avoid misclassifying different rooms as the same property; in low-density areas (such as villa areas and self-built housing areas), the deduplication radius threshold automatically increases to avoid misclassifying the same property as different properties due to GPS signal drift.
[0027] Finally, the system comprehensively utilizes the acquired address text information and geographic coordinate information, as well as the generated deduplication radius threshold, to perform a one-by-one, multi-dimensional comparison between the property to be judged and each property in the historical property database. This comparison process considers both spatial distance and text consistency. Ultimately, the system generates and outputs the duplicate judgment result in real time, such as "judged as a duplicate property" or "judged as a non-duplicate property." If a property is judged as a duplicate, the system can further block the property submission operation and return a "Suspected duplicate, please verify" prompt message to the user interface. All judgment logs are stored in the audit database in real time for subsequent model optimization and data backtracking. In this embodiment, the historical property database is a property database generated based on past experience.
[0028] In this invention, the address text information and geographic coordinate information of the property to be judged are obtained. A dynamic search window is determined with the geographic coordinate information as the center and the local environmental density value is calculated. Then, based on the local environmental density value, a deduplication radius threshold is dynamically generated through a monotonically decreasing function. This achieves the technical effect of adaptive adjustment of the deduplication radius with the local environmental density, effectively solving the problem of high misjudgment rate caused by the inability of fixed spatial thresholds to adapt to density differences in the prior art. At the same time, the address text information, geographic coordinate information and dynamically generated deduplication radius threshold are combined for multi-dimensional comparison and the judgment result is output in real time, which significantly improves the real-time performance, accuracy and environmental adaptability of duplicate property judgment.
[0029] In some specific embodiments of this application, the process of obtaining the address text information includes the following steps: Obtain the initial address text information; The initial address text information is semantically encoded using a pre-trained language model to extract contextual features; Address text elements are identified from the context features based on sequence labeling algorithms; The address text elements are formatted and normalized using an unsupervised learning algorithm to generate the address text information.
[0030] First, the system receives the initial address text information input by the user (which may be included in the housing information described in the above embodiments). This initial address text information contains strings, which may contain various non-standard expressions. For example, the user may input different strings such as "Building 1, Room 301, XX Community", "1-301, XX Community", or "Room 301, Building 1, XX Community".
[0031] Next, the system inputs the acquired initial address text information into a pre-trained language model. In a preferred embodiment, this pre-trained language model uses the BERT (Bidirectional Encoder Representation from Transformers) model. The language model semantically encodes each character in the address text, captures the association features between the character and its context, and outputs the contextual features corresponding to each character. For example, for the address "XX Community, Building 1, Unit 301", the language model can learn contextual semantics such as the fact that "Community" is usually followed by building information and that the character "Building" has a strong association with the number "1".
[0032] The system then inputs the extracted context features into a sequence labeling layer. In a preferred embodiment, the sequence labeling layer uses a conditional random field algorithm. Based on the grammatical rules and transition probabilities of the context features, the sequence labeling algorithm optimally decodes the label category of each character, thereby accurately identifying core address text elements from the context features, including but not limited to the community name, building number, and room number. For example, for the address "XX Community, Building 1, Room 301", the sequence labeling algorithm can accurately label "XX Community" as the community name, "1" as the building number, "Building" as the building type suffix, and "301" as the room number.
[0033] Finally, the system inputs the identified address text elements into an unsupervised density estimation module. This module pre-learns the representation patterns of addresses in different regions, automatically standardizing address text elements of different formats through unsupervised learning algorithms without manual annotation. For example, "Building 1", "No. 1 Building", and "1 Tower" are standardized to "Building 1"; "Room 301", "3-301", and "Unit 301" are standardized to "Unit 301 Room". After standardization, the system generates standardized address text information with a uniform format for subsequent multi-dimensional comparisons. The unsupervised density estimation module enables incremental learning. When new regional address data is added, there is no need to modify the language model parameters; only the initial address text information for that region needs to be supplemented. The unsupervised density estimation module automatically relearns the regional features, ensuring the accuracy and adaptability of structured extraction.
[0034] This embodiment enables standardized processing of basic housing data, eliminating data noise, extracting structured information, and providing accurate and consistent data input for subsequent judgments, adapting to the differences in address representation across different cities and business districts. Specifically, it employs a pre-trained language model as a foundation, overlaid with a sequence labeling layer, responsible for accurately identifying core structured fields such as "community name, building number, and room number" in address text elements. Simultaneously, it introduces an unsupervised density estimation model to automatically learn the representation patterns of addresses in different regions, achieving adaptive adaptation.
[0035] In some specific embodiments of this application, the step of dynamically generating the deduplication radius threshold based on the local environment density value using a monotonically decreasing function includes the following steps: The deduplication radius threshold is calculated based on the following formula: , in, The deduplication radius threshold is... The minimum deduplication radius threshold. Based on the deduplication radius threshold, The attenuation coefficient is... The local environmental density value is given.
[0036] Minimum deduplication radius threshold Its function is to prevent duplicates in ultra-high-density areas (such as super high-rise office buildings and large commercial complexes) by setting a deduplication radius threshold. If compressed too small, it loses its basic deduplication capability; this parameter is set... The lower limit. For example, it can be... The preset distance is 5 meters or 8 meters to ensure that even in high-density areas, the system can still identify the same property whose coordinates have shifted due to GPS errors.
[0037] Basic deduplication radius threshold Represents the density value in the local environment The system uses a base deduplication radius when the ideal open area is zero (i.e., there are no historical properties around the property to be identified). This parameter can be preset according to the signal drift error range of the Global Positioning System in typical environments, for example, it can be preset to 30 meters or 50 meters.
[0038] denominator Achieved deduplication radius threshold With local environmental density value The function increases and monotonically decreases. Among them, the natural logarithm function... This makes the density value The radius decays more rapidly with smaller variations, while the density value... When the value is large, the attenuation rate gradually slows down, which aligns with the actual distribution patterns of housing units. Attenuation coefficient Used to control the sensitivity to decay: The larger the value, the faster the radius decreases with density; The smaller the value, the smoother the decay.
[0039] The calculation takes the calculated value and the minimum deduplication radius threshold. The larger of the two values is used to ensure the final deduplication radius threshold. Not less than the preset minimum deduplication radius threshold .
[0040] Through the above formula, this embodiment achieves the technical effect of smoothly and adaptively adjusting the deduplication radius threshold according to the local environmental density. It can adapt to different density regions without manual intervention, and significantly reduces the false merging rate in high-density regions and the false splitting rate in low-density regions.
[0041] In some specific embodiments of this application, the step of performing a multi-dimensional comparison between the property to be judged and the properties in the historical property database based on the address text information, the geographic coordinate information, and the deduplication radius threshold, and outputting the duplicate judgment result in real time, includes the following steps: Based on the geographic coordinate information of the property to be determined and the historical geographic coordinate information of the properties in the historical property database, the geographic distance is calculated; When the geographical location interval is less than the deduplication radius threshold, it is determined whether the address text information of the property to be judged is consistent with the historical address text information of the property in the historical property database. If they are consistent, it is determined to be a duplicate property, and the duplicate judgment result is output in real time. If they are inconsistent, a second judgment is made based on the building type attribute of the property to be judged, and the duplicate judgment result is output in real time.
[0042] The system reads the geographic coordinates of the properties to be evaluated. Historical geographic coordinates of a specific property in the historical property database The actual ground distance between the two is calculated using the Euclidean distance formula, thus obtaining the geographical distance. .
[0043] The calculated geographical distance Compare with the generated dynamic deduplication radius threshold R. This is a fast filtering step. > If the historical listing is not spatially adjacent to the listing to be compared, it is determined that they cannot be the same listing. The comparison with this listing ends, and the comparison continues with the next listing in the historical listing database. If, after traversing the entire historical database, all listings... All greater than If the property to be judged is determined to be a non-duplicate property, the result "non-duplicate" will be output in real time. If at least one historical property meets the condition... ≤ If so, proceed to the next step of in-depth verification.
[0044] For the selected geographic location spacing Less than the deduplication radius threshold For each candidate historical property, the system performs a deep comparison of the address text information. The core of the comparison is to check whether key identifier fields are consistent, especially whether the "building number" and "room number" match exactly. If they match exactly, it indicates that they describe the same room in the same building, and can be identified as a duplicate property. The system immediately outputs the duplicate determination result and the ID of the matched historical property in real time, and can trigger business interception logic.
[0045] If key fields in the address text information are inconsistent (e.g., different building numbers or room numbers), it cannot be simply determined that the address is unique. In this case, the system combines the building type attribute of the property for supplementary reasonableness verification as a reference for the final determination. For example: If the building type is a densely packed building such as a "high-rise apartment" or "commercial and residential building," the physical locations of different rooms within it are already close, and GPS signals may drift indoors. Therefore, inconsistencies in address text information are reasonable and can ultimately be determined as non-duplication.
[0046] If the building type is a sparse structure such as a "detached villa" or "self-built house," it typically occupies an entire plot of land. If the geographical distance between the two is... If the numbers are very small, but the address text information is completely different, it is very likely that it is a duplicate record of the same house due to different submissions. In this case, it can be judged as a high probability of duplication or marked as "requires manual verification".
[0047] In this embodiment, the entire cascading judgment process has a clear logic. First, it uses low-cost space calculation to quickly filter out most irrelevant records, and then performs precise text verification on a small number of candidate records. If necessary, attribute judgment is used as an auxiliary method. This ensures high accuracy while taking into account processing efficiency and meeting the requirements of real-time performance.
[0048] Another embodiment of the present invention provides a device for determining duplicate housing listings, with reference to Figure 2 ,include: The information acquisition unit 100 is used to acquire the address text information and geographic coordinate information of the property to be judged; Density calculation unit 200 is used to determine a dynamic search window centered on the geographic coordinate information and calculate the local environmental density value within the dynamic search window. Threshold generation unit 300 is used to dynamically generate a deduplication radius threshold based on the local environment density value using a monotonically decreasing function; and The result output unit 400 is used to perform a multi-dimensional comparison between the property to be judged and the properties in the historical property database based on the address text information, the geographic coordinate information and the deduplication radius threshold, and output the duplicate judgment result in real time.
[0049] First, when a user submits property information on the property listing or check-in registration interface, the system retrieves the address text information based on the user's input, such as "Room XX, No. XX, XX Road, XX District, XX City". Simultaneously, the system uses the map provider's application programming interface (API) to obtain the corresponding geographic coordinates based on the user's input. These geographic coordinates may include longitude and latitude coordinates. The property corresponding to the aforementioned information is the property to be assessed; in some implementations, this property is a rental property.
[0050] Next, a dynamic search window is defined centered on the acquired geographic coordinates (denoted as point P). In a preferred embodiment, the size of the dynamic search window can be preset according to the actual business scenario, for example, set as a square area of 100 meters × 100 meters, or a circular area with a radius of 50 meters or 100 meters. The center of the dynamic search window is dynamically determined by the geographic coordinates of the property to be judged each time, but its size can be a preset fixed value. Within this dynamic search window, the system counts the number of properties already stored in the historical property database. Furthermore, to improve the accuracy of density calculation, the system can also count environmental context information such as building outline coverage density and point of interest type distribution within the window. The system uses the kernel density estimation algorithm (KDE) to fuse and calculate the above statistical data to obtain the local environmental density value at the location of point P. The local environmental density value reflects the degree of property clustering in the area centered on the property to be judged: the larger the local environmental density value, the denser the properties in the area; the smaller the local environmental density value, the sparser the properties in the area.
[0051] The system then uses the calculated local environmental density value as input and calculates the deduplication radius threshold using a monotonically decreasing function. This monotonically decreasing function has the following characteristics: when the local environmental density value increases, the deduplication radius threshold R decreases accordingly; when the local environmental density value decreases, the deduplication radius threshold increases accordingly. In this way, the deduplication radius threshold is dynamically adjusted according to the local environmental density: in high-density areas (such as apartment buildings and commercial centers), the deduplication radius threshold automatically decreases to avoid misclassifying different rooms as the same property; in low-density areas (such as villa areas and self-built housing areas), the deduplication radius threshold automatically increases to avoid misclassifying the same property as different properties due to GPS signal drift.
[0052] Finally, the system comprehensively utilizes the acquired address text information and geographic coordinate information, as well as the generated deduplication radius threshold, to perform a one-by-one, multi-dimensional comparison between the property to be judged and each property in the historical property database. This comparison process considers both spatial distance and text consistency. Ultimately, the system generates and outputs the duplicate judgment result in real time, such as "judged as a duplicate property" or "judged as a non-duplicate property." If a property is judged as a duplicate, the system can further block the property submission operation and return a "Suspected duplicate, please verify" prompt message to the user interface. All judgment logs are stored in the audit database in real time for subsequent model optimization and data backtracking. In this embodiment, the historical property database is a property database generated based on past experience.
[0053] In this invention, the address text information and geographic coordinate information of the property to be judged are obtained. A dynamic search window is determined with the geographic coordinate information as the center and the local environmental density value is calculated. Then, based on the local environmental density value, a deduplication radius threshold is dynamically generated through a monotonically decreasing function. This achieves the technical effect of adaptive adjustment of the deduplication radius with the local environmental density, effectively solving the problem of high misjudgment rate caused by the inability of fixed spatial thresholds to adapt to density differences in the prior art. At the same time, the address text information, geographic coordinate information and dynamically generated deduplication radius threshold are combined for multi-dimensional comparison and the judgment result is output in real time, which significantly improves the real-time performance, accuracy and environmental adaptability of duplicate property judgment.
[0054] In some specific embodiments of this application, the information acquisition unit 100 includes: The information acquisition module is used to acquire the initial address text information; The feature extraction module is used to semantically encode the initial address text information using a pre-trained language model and extract contextual features; The feature recognition module is used to identify address text features from the context features based on a sequence labeling algorithm; The information generation module is used to perform format normalization processing on the address text elements through an unsupervised learning algorithm to generate the address text information.
[0055] First, the system receives the initial address text information input by the user (which may be included in the housing information described in the above embodiments). This initial address text information contains strings, which may contain various non-standard expressions. For example, the user may input different strings such as "Building 1, Room 301, XX Community", "1-301, XX Community", or "Room 301, Building 1, XX Community".
[0056] Next, the system inputs the acquired initial address text information into a pre-trained language model. In a preferred embodiment, this pre-trained language model uses the BERT (Bidirectional Encoder Representation from Transformers) model. The language model semantically encodes each character in the address text, captures the association features between the character and its context, and outputs the contextual features corresponding to each character. For example, for the address "XX Community, Building 1, Unit 301", the language model can learn contextual semantics such as the fact that "Community" is usually followed by building information and that the character "Building" has a strong association with the number "1".
[0057] The system then inputs the extracted context features into a sequence labeling layer. In a preferred embodiment, the sequence labeling layer uses a conditional random field algorithm. Based on the grammatical rules and transition probabilities of the context features, the sequence labeling algorithm optimally decodes the label category of each character, thereby accurately identifying core address text elements from the context features, including but not limited to the community name, building number, and room number. For example, for the address "XX Community, Building 1, Room 301", the sequence labeling algorithm can accurately label "XX Community" as the community name, "1" as the building number, "Building" as the building type suffix, and "301" as the room number.
[0058] Finally, the system inputs the identified address text elements into an unsupervised density estimation module. This module pre-learns the representation patterns of addresses in different regions, automatically standardizing address text elements of different formats through unsupervised learning algorithms without manual annotation. For example, "Building 1", "No. 1 Building", and "1 Tower" are standardized to "Building 1"; "Room 301", "3-301", and "Unit 301" are standardized to "Unit 301 Room". After standardization, the system generates standardized address text information with a uniform format for subsequent multi-dimensional comparisons. The unsupervised density estimation module enables incremental learning. When new regional address data is added, there is no need to modify the language model parameters; only the initial address text information for that region needs to be supplemented. The unsupervised density estimation module automatically relearns the regional features, ensuring the accuracy and adaptability of structured extraction.
[0059] This embodiment enables standardized processing of basic housing data, eliminating data noise, extracting structured information, and providing accurate and consistent data input for subsequent judgments, adapting to the differences in address representation across different cities and business districts. Specifically, it employs a pre-trained language model as a foundation, overlaid with a sequence labeling layer, responsible for accurately identifying core structured fields such as "community name, building number, and room number" in address text elements. Simultaneously, it introduces an unsupervised density estimation model to automatically learn the representation patterns of addresses in different regions, achieving adaptive adaptation.
[0060] In some specific embodiments of this application, the threshold generation unit 300 includes: The formula calculation module is used to calculate the deduplication radius threshold based on the following formula: , in, The deduplication radius threshold is... The minimum deduplication radius threshold. Based on the deduplication radius threshold, The attenuation coefficient is... The local environmental density value is given.
[0061] Minimum deduplication radius threshold Its function is to prevent duplicates in ultra-high-density areas (such as super high-rise office buildings and large commercial complexes) by setting a deduplication radius threshold. If compressed too small, it loses its basic deduplication capability; this parameter is set... The lower limit. For example, it can be... The preset distance is 5 meters or 8 meters to ensure that even in high-density areas, the system can still identify the same property whose coordinates have shifted due to GPS errors.
[0062] Basic deduplication radius threshold Represents the density value in the local environment The system uses a base deduplication radius when the ideal open area is zero (i.e., there are no historical properties around the property to be identified). This parameter can be preset according to the signal drift error range of the Global Positioning System in typical environments, for example, it can be preset to 30 meters or 50 meters.
[0063] denominator Achieved deduplication radius threshold With local environmental density value The function increases and monotonically decreases. Among them, the natural logarithm function... This makes the density value The radius decays more rapidly with smaller variations, while the density value... When the value is large, the attenuation rate gradually slows down, which aligns with the actual distribution patterns of housing units. Attenuation coefficient Used to control the sensitivity to decay: The larger the value, the faster the radius decreases with density; The smaller the value, the smoother the decay.
[0064] The calculation takes the calculated value and the minimum deduplication radius threshold. The larger of the two values is used to ensure the final deduplication radius threshold. Not less than the preset minimum deduplication radius threshold .
[0065] Through the above formula, this embodiment achieves the technical effect of smoothly and adaptively adjusting the deduplication radius threshold according to the local environmental density. It can adapt to different density regions without manual intervention, and significantly reduces the false merging rate in high-density regions and the false splitting rate in low-density regions.
[0066] In some specific embodiments of this application, the result output unit 400 includes: The spacing calculation module is used to calculate the geographical distance based on the geographical coordinate information of the property to be judged and the historical geographical coordinate information of the properties in the historical property database. The result output module is used to determine whether the address text information of the property to be judged is consistent with the historical address text information of the property in the historical property database when the geographical location distance is less than the deduplication radius threshold. If they are consistent, the property is judged as a duplicate property and the duplicate judgment result is output in real time. If they are inconsistent, a second judgment is made based on the building type attribute of the property to be judged and the duplicate judgment result is output in real time.
[0067] The system reads the geographic coordinates of the properties to be evaluated. Historical geographic coordinates of a specific property in the historical property database The actual ground distance between the two is calculated using the Euclidean distance formula, thus obtaining the geographical distance. .
[0068] The calculated geographical distance Compare with the generated dynamic deduplication radius threshold R. This is a fast filtering step. > If the historical listing is not spatially adjacent to the listing to be compared, it is determined that they cannot be the same listing. The comparison with this listing ends, and the comparison continues with the next listing in the historical listing database. If, after traversing the entire historical database, all listings... All greater than If the property to be judged is determined to be a non-duplicate property, the result "non-duplicate" will be output in real time. If at least one historical property meets the condition... ≤ If so, proceed to the next step of in-depth verification.
[0069] For the selected geographic location spacing Less than the deduplication radius threshold For each candidate historical property, the system performs a deep comparison of the address text information. The core of the comparison is to check whether key identifier fields are consistent, especially whether the "building number" and "room number" match exactly. If they match exactly, it indicates that they describe the same room in the same building, and can be identified as a duplicate property. The system immediately outputs the duplicate determination result and the ID of the matched historical property in real time, and can trigger business interception logic.
[0070] If key fields in the address text information are inconsistent (e.g., different building numbers or room numbers), it cannot be simply determined that the address is unique. In this case, the system combines the building type attribute of the property for supplementary reasonableness verification as a reference for the final determination. For example: If the building type is a densely packed building such as a "high-rise apartment" or "commercial and residential building," the physical locations of different rooms within it are already close, and GPS signals may drift indoors. Therefore, inconsistencies in address text information are reasonable and can ultimately be determined as non-duplication.
[0071] If the building type is a sparse structure such as a "detached villa" or "self-built house," it typically occupies an entire plot of land. If the geographical distance between the two is... If the numbers are very small, but the address text information is completely different, it is very likely that it is a duplicate record of the same house due to different submissions. In this case, it can be judged as a high probability of duplication or marked as "requires manual verification".
[0072] In this embodiment, the entire cascading judgment process has a clear logic. First, it uses low-cost space calculation to quickly filter out most irrelevant records, and then performs precise text verification on a small number of candidate records. If necessary, attribute judgment is used as an auxiliary method. This ensures high accuracy while taking into account processing efficiency and meeting the requirements of real-time performance.
[0073] This invention provides a computer device, which includes a processor. The processor executes a computer program stored in a memory to implement the steps of the duplicate housing information determination method described above.
[0074] The present invention also provides a storage medium storing a computer program (instructions) thereon, which, when executed by a processor, implements the steps of the duplicate housing information determination method described above.
[0075] For example, a computer program can be divided into one or more modules, one or more of which are stored in memory and executed by a processor to perform the present invention. One or more modules can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a computer device. For example, the computer program can be divided into the steps of the duplicate housing resource determination method provided in the above-described method embodiments.
[0076] Those skilled in the art will understand that the above description of the computer device is merely an example and does not constitute a limitation on the computer device. It may include more or fewer components than described above, or a combination of certain components, or different components, such as input / output devices, network access devices, buses, etc.
[0077] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the computer device, connecting various parts of the computer device via various interfaces and lines.
[0078] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as interface display function, interface interaction function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as map interface, selection interface, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0079] If the modules / units integrated into the computer device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, an electrical signal, and a software distribution medium, etc.
[0080] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0081] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A method for identifying duplicate listings, characterized in that, Includes the following steps: Obtain the address text information and geographic coordinate information of the property to be judged; A dynamic search window is determined centered on the geographic coordinate information, and the local environmental density value within the dynamic search window is calculated. Based on the local environment density value, a deduplication radius threshold is dynamically generated using a monotonically decreasing function; and Based on the address text information, the geographic coordinate information, and the deduplication radius threshold, the property to be judged is compared with the properties in the historical property database in multiple dimensions, and the duplicate judgment result is output in real time.
2. The method for determining duplicate listings according to claim 1, characterized in that, The process of obtaining the address text information includes the following steps: Obtain the initial address text information; The initial address text information is semantically encoded using a pre-trained language model to extract contextual features; Address text elements are identified from the context features based on sequence labeling algorithms; The address text elements are formatted and normalized using an unsupervised learning algorithm to generate the address text information.
3. The method for determining duplicate listings according to claim 1, characterized in that, The process of dynamically generating the deduplication radius threshold based on the local environment density value using a monotonically decreasing function includes the following steps: The deduplication radius threshold is calculated based on the following formula: , in, The deduplication radius threshold is... The minimum deduplication radius threshold. Based on the deduplication radius threshold, The attenuation coefficient is... The local environmental density value is given.
4. The method for determining duplicate listings according to claim 1, characterized in that, The process of comparing the property to be judged with properties in the historical property database based on the address text information, the geographic coordinate information, and the deduplication radius threshold, and outputting the duplicate judgment result in real time, includes the following steps: Based on the geographic coordinate information of the property to be determined and the historical geographic coordinate information of the properties in the historical property database, the geographic distance is calculated; When the geographical location interval is less than the deduplication radius threshold, it is determined whether the address text information of the property to be judged is consistent with the historical address text information of the property in the historical property database. If they are consistent, it is determined to be a duplicate property, and the duplicate judgment result is output in real time. If they are inconsistent, a second judgment is made based on the building type attribute of the property to be judged, and the duplicate judgment result is output in real time.
5. A device for determining duplicate housing listings, characterized in that, include: The information acquisition unit is used to acquire the address text information and geographic coordinate information of the property to be judged; The density calculation unit is used to determine a dynamic search window centered on the geographic coordinate information and calculate the local environmental density value within the dynamic search window. The threshold generation unit is used to dynamically generate a deduplication radius threshold based on the local environment density value using a monotonically decreasing function; and The result output unit is used to perform a multi-dimensional comparison between the property to be judged and the properties in the historical property database based on the address text information, the geographic coordinate information and the deduplication radius threshold, and output the duplicate judgment result in real time.
6. The duplicate housing information determination device according to claim 5, characterized in that, The information acquisition unit includes: The information acquisition module is used to acquire the initial address text information; The feature extraction module is used to semantically encode the initial address text information using a pre-trained language model and extract contextual features; The feature recognition module is used to identify address text features from the context features based on a sequence labeling algorithm; The information generation module is used to perform format normalization processing on the address text elements through an unsupervised learning algorithm to generate the address text information.
7. The duplicate housing information determination device according to claim 5, characterized in that, The threshold generation unit includes: The formula calculation module is used to calculate the deduplication radius threshold based on the following formula: , in, The deduplication radius threshold is... The minimum deduplication radius threshold. Based on the deduplication radius threshold, The attenuation coefficient is... The local environmental density value is given.
8. The duplicate housing information determination device according to claim 5, characterized in that, The result output unit includes: The spacing calculation module is used to calculate the geographical distance based on the geographical coordinate information of the property to be judged and the historical geographical coordinate information of the properties in the historical property database. The result output module is used to determine whether the address text information of the property to be judged is consistent with the historical address text information of the property in the historical property database when the geographical location distance is less than the deduplication radius threshold. If they are consistent, the property is judged as a duplicate property and the duplicate judgment result is output in real time. If they are inconsistent, a second judgment is made based on the building type attribute of the property to be judged and the duplicate judgment result is output in real time.
9. A computer device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the duplicate listing determination method as described in any one of claims 1 to 4.
10. A storage medium, characterized in that, The storage medium stores a computer program that can be executed to implement the steps of the duplicate listing determination method as described in any one of claims 1 to 4.