Site selection feature screening methods, devices, electronic equipment and storage media
By employing an automated site selection feature screening method that utilizes spatiotemporal filtering and temporal clustering analysis, the problems of low efficiency, high cost, and low accuracy in existing site selection technologies are solved, achieving more efficient and accurate site selection feature screening that is suitable for geographic location data processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2026-03-10
AI Technical Summary
Existing manual site selection methods are inefficient, costly, and inaccurate, and cannot effectively handle complex site selection feature types. In particular, when the number of features is large with the support of Internet data, it is difficult to effectively filter and reduce the dimensionality.
An automated site selection feature screening method is adopted. By acquiring site selection data, feature processing is performed, spatiotemporal filtering and temporal clustering analysis are conducted to screen out target site selection features that match the spatial differentiation relationship of the dependent variable. An open-source cluster computing framework is used for data decomposition and feature extraction, and feature selection is optimized by combining a geographic detector and K-Shape clustering algorithm.
It automates and improves the accuracy of the site selection process, reduces costs, is suitable for large-scale use by ordinary users, and improves the accuracy of site selection information.
Smart Images

Figure CN114329240B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to location data processing technology, and more particularly to a method, apparatus, electronic device, computer program product, and storage medium for site selection feature screening. Background Technology
[0002] In related technologies, geographical location has become a crucial factor influencing the operation of many industry outlets (such as the catering industry, logistics industry, server deployment, and point-of-interest advertising). Currently, the site selection method is typically manual, where site selection personnel conduct on-site surveys and combine this with their experience. However, this manual method is not only inefficient, time-consuming, and costly, but also suffers from low accuracy due to the limitations of human experience. Furthermore, the types of site selection characteristics are complex, mainly categorized into four types: customer characteristics, accessibility characteristics, competitive characteristics, and operational characteristics. With the support of actual internet data, the number of characteristics often reaches thousands. Therefore, effectively selecting and reducing the dimensionality of these characteristics is a key aspect of using AI technology for site selection. Summary of the Invention
[0003] In view of this, this application provides a site selection feature filtering method that can automatically filter site selection features to obtain target site selection features. Using the target site selection features, the site selection location that matches the target object can be determined, reducing the usage cost of the site selection process, which is beneficial for large-scale use by ordinary users, and at the same time, it can achieve the technical effect of more accurate site selection information.
[0004] The technical solution of this invention is implemented as follows:
[0005] This invention provides a location feature filtering method, including:
[0006] Acquire site selection data and perform feature processing on the site selection data to obtain initial site selection features;
[0007] The initial location features are subjected to spatiotemporal filtering, and initial location features that match the spatial differentiation relationship of the dependent variable are selected based on the time-series mean threshold.
[0008] Temporal clustering analysis is performed on the initial location features that match the spatial differentiation relationship of the dependent variable to filter the initial location features that match the spatiotemporal correlation, thereby obtaining the target location features, so as to determine the location matching the target object through the target location features.
[0009] This invention also provides a location feature filtering device, comprising:
[0010] The information transmission module is used to acquire site selection data and perform feature processing on the site selection data to obtain initial site selection features;
[0011] The information processing module is used to perform spatiotemporal filtering on the initial location features, and to filter the initial location features that match the spatial differentiation relationship of the dependent variable based on the time-series mean threshold.
[0012] The information processing module is used to perform temporal clustering analysis on the initial location features that match the spatial differentiation relationship of the dependent variable, filter the initial location features that match the spatiotemporal correlation, and obtain the target location features, so as to determine the location matching the target object through the target location features.
[0013] In the above scheme,
[0014] The information processing module is used to decompose the site selection data according to the time dimension using an open-source cluster computing framework to obtain site selection data in the time dimension.
[0015] The information processing module is used to perform feature extraction and feature normalization on the site selection data in the time dimension through the open-source cluster computing framework to obtain normalized initial site selection features.
[0016] The information processing module is used to perform feature deletion processing on the normalized initial location features based on the entropy value of the normalized initial location features to obtain the initial location features.
[0017] In the above scheme,
[0018] The information processing module is used to perform data transformation processing on the initial location features to determine the dependent and independent variables corresponding to the initial location features;
[0019] The information processing module is used to determine the correlation value between the dependent variable and the independent variable based on the dependent variable and the independent variable corresponding to the initial location characteristics.
[0020] The information processing module is used to calculate the time series mean of the correlation value, and filter the time series mean of the correlation value based on the time series mean threshold to obtain the time series mean of the correlation value that matches the time series mean threshold.
[0021] The information processing module is used to determine the initial location features that match the spatial differentiation relationship of the dependent variable based on the time series mean of the correlation values that match the time series mean threshold.
[0022] In the above scheme,
[0023] The information processing module is used to determine the correlation value corresponding to the initial location feature that matches the spatial differentiation relationship of the dependent variable;
[0024] The information processing module is used to determine the number of clusters that match the site selection data;
[0025] The information processing module is used to perform clustering processing on the association value according to the number of clusters, and obtain the clustering result of the association value;
[0026] The information processing module is used to filter initial location features that match the spatiotemporal correlation based on the clustering results of the correlation values, and obtain target location features.
[0027] In the above scheme,
[0028] The information processing module is used to acquire a set of interest point data to be processed;
[0029] The information processing module is used to combine the points of interest in the point of interest dataset to form corresponding point of interest sample pairs;
[0030] The information processing module is used to extract feature vectors corresponding to the interest point sample pairs by utilizing the target location features and through the feature combination network of the interest point selection model.
[0031] The information processing module is used to sort the feature vectors corresponding to the interest point samples through the sorting network of the interest point selection model, and determine the interest points that match the target location features.
[0032] In the above scheme,
[0033] The information processing module is used to acquire point-of-interest data from different data sources;
[0034] The information processing module is used to classify the data sources of the point of interest data;
[0035] The information processing module is used to determine the same point of interest in different data sources based on the classification results of the data sources of the point of interest according to the target location features.
[0036] The information processing module is used to aggregate interest point data belonging to the same interest point in order to obtain complete and detailed information about the interest point.
[0037] This invention also provides an electronic device, the electronic device comprising:
[0038] Memory, used to store executable instructions;
[0039] The processor, when executing executable instructions stored in the memory, implements a preceding addressing feature filtering method.
[0040] This invention also provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement a preceding addressing feature filtering method.
[0041] The embodiments of the present invention have the following beneficial effects:
[0042] This invention acquires site selection data and performs feature processing on the data to obtain initial site selection features. It then performs spatiotemporal filtering on these initial features, selecting those that match the spatial differentiation relationship of the dependent variable based on a time-series mean threshold. Finally, it performs time-series clustering analysis on these initial features to select those that match spatiotemporal correlation, obtaining target site selection features. This allows for the determination of a suitable location for a target object using these target site selection features. This reduces the cost of the site selection process, making it suitable for large-scale use by ordinary users. Furthermore, the automated filtering of site selection features yields more accurate target site selection features, resulting in more accurate site selection information. Attached Figure Description
[0043] Figure 1 This is a schematic diagram illustrating the application environment of the location feature filtering method provided in this embodiment of the invention;
[0044] Figure 2 This is a schematic diagram of the composition structure of the location feature screening device provided in an embodiment of the present invention;
[0045] Figure 3 This is an optional flowchart illustrating the location feature filtering method provided in an embodiment of the present invention.
[0046] Figure 4 This is a schematic diagram illustrating the calculation process of the correlation value between the dependent and independent variables in an embodiment of the present invention;
[0047] Figure 5 This is a schematic diagram of the k-Shape clustering process in an embodiment of the present invention;
[0048] Figure 6 A schematic diagram of an optional two-dimensional map display for the site selection feature filtering method provided in the embodiments of the present invention;
[0049] Figure 7 This is an optional flowchart illustrating the location feature filtering method provided in an embodiment of the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0052] In the implementation of this application, the collection and processing of relevant data should be strictly in accordance with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0053] Before providing a further detailed description of the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention will be explained, and the nouns and terms involved in the embodiments of the present invention shall be interpreted as follows.
[0054] 1) Responding to: used to indicate the conditions or states on which the operation is performed depends. When the conditions or states on which it depends are met, one or more operations can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.
[0055] 2) Location-Based Services (LBS): Location-based services (LBS) are location-related services provided by wireless operators to users. LBS utilizes various positioning technologies to determine the current location of a device and provides information resources and basic services to that device via the mobile internet. Users can first determine their spatial location using positioning technology, and then access location-related resources and information via the mobile internet. LBS services integrate multiple information technologies such as mobile communication, the internet, spatial positioning, location information, and big data. It utilizes mobile internet service platforms for data updates and interaction, enabling users to obtain relevant services through spatial positioning.
[0056] 3) Mobile Terminals: Mobile terminals, also known as mobile communication terminals, refer to computer devices that can be used while on the move, including mobile phones, laptops, tablets, and in-vehicle devices. With the development of networks and technologies towards greater broadband, the mobile communication industry is entering a true mobile information age. With the rapid development of integrated circuit technology, mobile terminals have acquired powerful processing capabilities, transforming from simple communication tools into comprehensive information processing platforms. Mobile terminals also offer a wide range of communication methods, including communication via GSM, CDMA, WCDMA, EDGE, 4G, and other wireless networks, as well as wireless LANs, Bluetooth, and infrared. Furthermore, mobile terminals integrate global navigation satellite system (GNSS) positioning chips for processing satellite signals and precise user positioning, currently widely used in location services; mobile terminals include devices with satellite positioning capabilities.
[0057] 4) Points of interest, a location attribute, can be information that can characterize a scene, such as identifiable buildings, areas (e.g., cities), landscapes (e.g., attractions), and third-party service entities (e.g., shops, restaurants, accommodations).
[0058] 5) Spark, a fast and versatile computing engine designed for large-scale data processing.
[0059] The site selection feature screening method provided in this application is described below, in which... Figure 1 This is a schematic diagram illustrating a usage scenario of the location feature filtering method provided in this embodiment of the invention. (See attached diagram.) Figure 1 The terminals (including terminals 10-1 and 10-2) are equipped with a client containing map information display software. Users can use the client to determine a suitable location matching the target object based on its location features, and the client will display this suitable location to the user. The terminals connect to a map server 200 via a network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both. Data transmission is achieved via a wireless link, enabling map information sharing between different terminals. The terminals (including terminals 10-1 and 10-2) can receive location data and perform feature processing on it to obtain initial location features. They then perform spatiotemporal filtering on these initial location features, selecting those that match the spatial differentiation relationship of the dependent variable based on a time-series mean threshold. Finally, they perform time-series cluster analysis on these initial location features to select those that match the spatiotemporal correlation, thus obtaining the target location features.
[0060] The structure of the location feature filtering device according to an embodiment of the present invention will be described in detail below. The location feature filtering device can be implemented in various forms, such as a dedicated terminal with terminal positioning function, or a server with terminal positioning function, for example, the preceding... Figure 1 Map server 200. Figure 2 This is a schematic diagram of the composition of the location feature screening device provided in an embodiment of the present invention. It can be understood that... Figure 2 Only an exemplary structure of the location feature screening device is shown, not the entire structure; it can be implemented as needed. Figure 2 The structure shown may be part or all of the structure.
[0061] The address feature filtering device provided in this embodiment of the invention includes at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. The various components in the address feature filtering device 20 are coupled together via a bus system 205. It can be understood that the bus system 205 is used to implement communication between these components. In addition to a data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 2 The general labeled all buses as Bus System 205.
[0062] The user interface 203 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.
[0063] It is understood that memory 202 can be volatile memory or non-volatile memory, or both. In this embodiment of the invention, memory 202 is capable of storing data to support the operation of a terminal (such as 10-1). Examples of this data include any computer programs used to operate on the terminal (such as 10-1), such as operating systems and applications. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications.
[0064] In some embodiments, the address feature filtering device provided in this invention can be implemented using a combination of hardware and software. For example, the question-answering model training device provided in this invention can be a processor in the form of a hardware decoding processor, programmed to execute the address feature filtering method provided in this invention. For instance, the processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0065] As an example of the address feature filtering device provided in this embodiment of the invention, which is implemented by combining software and hardware, the address feature filtering device provided in this embodiment of the invention can be directly embodied as a combination of software modules executed by processor 201. The software modules can be located in a storage medium, which is located in memory 202. Processor 201 reads the executable instructions included in the software modules in memory 202 and combines them with necessary hardware (e.g., including processor 201 and other components connected to bus 205) to complete the address feature filtering method provided in this embodiment of the invention.
[0066] As an example, processor 201 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0067] As an example of the hardware implementation of the address feature filtering device provided in the embodiments of the present invention, the device provided in the embodiments of the present invention can be directly executed by a processor 201 in the form of a hardware decoding processor. For example, it can be executed by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components to implement the address feature filtering method provided in the embodiments of the present invention.
[0068] In this embodiment of the invention, the memory 202 is used to store various types of data to support the operation of the address feature filtering device 20. Examples of such data include: any executable instructions for operation on the address feature filtering device 20, such as executable instructions that can be included in a program implementing the address feature filtering method of this embodiment of the invention.
[0069] In other embodiments, the location feature filtering device provided in this invention can be implemented in software. Figure 2 An address feature filtering device stored in memory 202 is shown. This device can be software in the form of programs and plug-ins, and includes a series of modules. As an example of a program stored in memory 202, it may include the address feature filtering device, which includes the following software modules: an information transmission module 2081 and an information processing module 2082. When the software modules in the address feature filtering device are read into RAM and executed by processor 201, the address feature filtering method provided in this embodiment of the invention will be implemented. The functions of each software module in the address feature filtering device will be further described below.
[0070] The information transmission module 2081 is used to acquire site selection data and perform feature processing on the site selection data to obtain initial site selection features.
[0071] The information processing module 2082 is used to perform spatiotemporal filtering on the initial location features and filter the initial location features that match the spatial differentiation relationship of the dependent variable based on the time series mean threshold.
[0072] The information processing module 2082 is used to perform temporal clustering analysis on the initial location features that match the spatial differentiation relationship of the dependent variable, filter the initial location features that match the spatiotemporal correlation, and obtain the target location features, so as to determine the location matching the target object through the target location features.
[0073] according to Figure 2 The electronic device shown, in one aspect of this application, also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform various embodiments and combinations of embodiments provided in the various optional implementations of the address feature filtering method described above.
[0074] Before introducing the site selection feature filtering method provided in this application, we will first introduce the site selection feature processing methods in related technologies. There are three main ways to obtain site selection features in related technologies:
[0075] 1) Wrapper method: This method treats feature selection as a search and optimization method. It divides features into different combinations, evaluates each combination, and compares them with other combinations. This treats feature selection as an optimization method, which can be solved using many optimization algorithms, such as genetic algorithms and artificial bee colony algorithms.
[0076] 2) Embedded method: In the process of determining the model, select attributes that are important for model training. Use methods such as decision tree algorithm, linear regression, RankNet ranking model, SVR, and grey relational analysis for feature processing.
[0077] 3) Filter method: This involves scoring features and then selecting features based on thresholds.
[0078] However, regardless of the method used, the addressing features are merely treated as a vector of numerical and textual attributes, without considering their spatiotemporal attributes. When selecting a site, the profile attributes include descriptive information such as the type of people depicted, as well as numerical information. Furthermore, these attributes are constantly changing spatiotemporally within the region. Ignoring these spatiotemporal attributes during feature selection leads to errors in feature importance calculation, negatively impacting the accuracy of feature selection and ultimately affecting the final site selection analysis results.
[0079] To address the aforementioned shortcomings, combined with Figure 2 The location feature filtering device shown illustrates the location feature filtering method provided in this embodiment of the invention. See also: Figure 3 , Figure 3 This is an optional flowchart illustrating the location feature filtering method provided in this embodiment of the invention. It can be understood that... Figure 3 The steps shown can be performed by various electronic devices that operate the site selection feature filtering device. These include, for example, dedicated terminals with site selection feature filtering devices, smartphones, smartwatches, and other electronic devices capable of receiving site selection data; or devices with satellite positioning capabilities. The dedicated terminal with the site selection feature filtering device can be a preceding... Figure 2 The electronic device with a location feature filtering device shown in the embodiment may also have a functional module with terminal positioning function. The following section addresses... Figure 3 The steps shown are explained.
[0080] Step 301: The site selection feature screening device acquires site selection data and performs feature processing on the site selection data to obtain initial site selection features.
[0081] In some embodiments of this invention, site selection data can be obtained using a grid index. This involves mapping the user-inputted area of interest onto multiple grids, obtaining the site data for each grid, and finally merging the site data from all grids. However, related grid indexing methods use an equally divided latitude and longitude grid to index site data across all dimensions. To balance query speed and data accuracy, a side length (i.e., scale) of 100 meters is typically chosen for the index grid. Site selection data refers to information such as population, economy, transportation, and environment existing within a certain area outline on a map. Population information includes, but is not limited to, headcount statistics, population profiles, and visitor flow statistics; economic information includes, but is not limited to, macroeconomics (GDP, etc.), industrial economics (GDP of primary, secondary, and tertiary industries and number of POIs, etc.), and industry economics (number and details of POIs for industries such as food, etc.); transportation information includes, but is not limited to, the number of transportation facilities, road conditions, and traffic conditions; and environmental information includes, but is not limited to, different information such as natural environment (green spaces and water systems) and human environment (public facilities).
[0082] After obtaining the site selection data, it is necessary to perform feature processing on the site selection data to obtain initial site selection features. Specifically, obtaining the site selection data and performing feature processing on the site selection data to obtain initial site selection features can be achieved in the following ways:
[0083] The site selection data is decomposed according to the time dimension using an open-source cluster computing framework to obtain site selection data in the time dimension. Feature extraction and feature normalization are performed on the site selection data in the time dimension using the open-source cluster computing framework to obtain normalized initial site selection features. Based on the entropy value of the normalized initial site selection features, feature deletion is performed on the normalized initial site selection features to obtain the initial site selection features. This involves utilizing a Web UI component to receive user-submitted parameters related to the open-source Spark computing framework and generating site selection data based on these parameters. Spark, as a fast and practical open-source computing framework, has wide applications in the field of massive user data processing, capable of efficiently scaling computing from one to thousands of nodes. To maximize flexibility, Spark supports operation on various cluster managers, such as the YARN Yet Another Resource Negotiator and the Mesos open-source distributed resource management framework. This allows for the construction of large-scale, low-latency data analysis applications to collect diverse data across various dimensions of the site selection data. By utilizing the entropy value of the normalized initial site selection features, feature deletion is performed on these features to obtain the initial site selection characteristics. These initial characteristics can then be used for urban market analysis, core area analysis, and cost analysis during the site selection process.
[0084] Step 302: The location feature screening device performs spatiotemporal filtering on the initial location features and filters the initial location features that match the spatial differentiation relationship of the dependent variable based on the time series mean threshold.
[0085] In some embodiments of the present invention, spatiotemporal filtering of initial location features is performed, and initial location features that match the spatial differentiation relationship of the dependent variable are selected based on the time-series mean threshold. This can be achieved in the following ways:
[0086] The initial site selection features are processed through data transformation to determine the dependent and independent variables corresponding to them. Based on these features, the correlation values are determined. The time-series mean of these correlation values is calculated, and a threshold is used to filter the time-series mean values to obtain those that match the threshold. Finally, the initial site selection features matching the spatial heterogeneity of the dependent variable are determined based on these time-series mean values. During the spatiotemporal filtering process, a geographic detector can be used. Geographic detectors are a novel statistical method for detecting spatial heterogeneity and revealing its underlying driving factors. When using a geographic detector, it is assumed that the study area is divided into several sub-regions. If the sum of the variances of the sub-regions is less than the total variance of the region, spatial heterogeneity exists; if the spatial distributions of two variables tend to be consistent, a statistical correlation exists between them. GeoDetector can effectively identify the interaction between multiple factors and geographical phenomena. It consists of four parts: factor detection, risk detection, ecological detection, and interaction detection. Factor detection can be used to measure the influence of factor X on variable Y. Interaction detection can be used to identify the explanatory power of variable Y under different factor interactions. Risk area detection can be used to determine whether there are significant differences between different regions. Ecological detection can be used to compare the differences in the influence of two factors on variable Y.
[0087] Referring to Table 1, when using a geographic detector to perform spatiotemporal filtering on the initial site selection features, the site selection data first needs to be processed into the format required for geographic detector analysis. Both the dependent and independent variables are converted to raster data; for the raster data of variable X, conversion is required monthly, followed by reclassification, for example, using a natural discontinuity classification method. For each month's data, it is converted to the format shown in Table 1 to perform data transformation on the initial site selection features, determining the corresponding dependent and independent variables.
[0088]
[0089] Table 1
[0090] As shown in Table 1, after performing data transformation on the initial location features and determining the dependent and independent variables corresponding to the initial location features, it is also necessary to determine the correlation value q between the dependent variable (Y) and the independent variable (X) based on the dependent and independent variables corresponding to the initial location features. Figure 4 This is a schematic diagram illustrating the calculation process of the correlation value between the dependent and independent variables in an embodiment of the present invention. Referring to Formulas 1 and 2, the q-statistic for each X variable (such as detailed features like population, POI, and transportation) and Y variable (such as profit) is calculated:
[0091] Formula 1
[0092] Formula 2
[0093] In Formulas 1 and 2 above, h = 1, …, L represents the stratification (strata) of variable Y or factor X, i.e., classification or partitioning; Nh and N are the number of units in stratum h and the entire region, respectively; σ²h and σ² are the variances of Y values in stratum h and the entire region, respectively. SSW and SST are the within-sum of squares and the total sum of squares, respectively, and the range of q is [0, 1]. A larger q value indicates a more pronounced spatial heterogeneity of Y; if the stratification is generated by the independent variable X, a larger q value indicates a stronger explanatory power of the independent variable X for attribute Y, and vice versa. In the extreme case, a q value of 1 indicates that factor X completely controls the spatial distribution of Y, and a q value of 0 indicates that factor X has no relationship with Y, and the q value represents that X explains 100*q% of Y.
[0094] Finally, the time series mean of the correlation values is calculated, and the time series mean of the correlation values is filtered based on the time series mean threshold to obtain the time series mean of the correlation values that match the time series mean threshold. Based on the time series mean of the correlation values that match the time series mean threshold, the initial location features that match the spatial differentiation relationship of the dependent variable are determined. Thus, the time series mean of the independent variable q values for all months of all years is calculated and sorted. The q values are filtered based on the time series mean threshold to filter out initial location features with low thresholds. This can filter out features that are obviously not very compatible with the spatial differentiation relationship of the dependent variable, and avoid the impact of initial location features that are not compatible with the spatial differentiation relationship of the dependent variable on the location selection process.
[0095] Step 303: The location feature screening device performs temporal clustering analysis on the initial location features that match the spatial differentiation relationship of the dependent variable, and screens the initial location features that match the spatiotemporal correlation to obtain the target location features.
[0096] Among them, the location that matches the target object can be determined by the target location characteristics.
[0097] In some embodiments of the present invention, temporal clustering analysis is performed on the initial location features that match the spatial differentiation relationship of the dependent variable, and the initial location features that match the spatiotemporal correlation are screened to obtain the target location features. This can be achieved in the following ways:
[0098] The process involves: determining the association values corresponding to initial site selection features that match the spatial heterogeneity of the dependent variable; determining the number of clusters matching the site selection data; clustering the association values according to the number of clusters to obtain the clustering results; and filtering initial site selection features that match spatiotemporal correlation based on the clustering results to obtain target site selection features. Since the geographic detector does not describe the temporal changes in spatial heterogeneity, temporal clustering analysis can be performed on the association values q to make them more suitable for describing the spatiotemporal causal influence of site selection features on the dependent variable and to remove features with high spatiotemporal correlation, thus making the site selection features more accurate. Specifically, when clustering the association values, K-Shape clustering can be used. K-Shape clustering is based on shape distance (SBD) and focuses on scaling and shift invariance. (Reference) Figure 5 , Figure 5 This is a schematic diagram of the k-Shape clustering process in an embodiment of the present invention. k-Shape has two main features: shape-based distance (SBD) and time-series shape extraction, specifically including the following steps:
[0099] Step 501: Sort the time series q-means of different features.
[0100] Step 502: Identify whether N adjacent features exist simultaneously in the same category of statistical clustering and shape clustering.
[0101] Where N is the number of features in the sliding window, with a minimum of 2.
[0102] Step 503: If it exists, keep only one feature; otherwise, return to step 502.
[0103] Therefore, statistical similarity and shape similarity can be obtained through q-value time series. If N features simultaneously exist in a certain cluster of statistical clustering and a certain cluster of shape clustering, it indicates that these N features may have similar spatial heterogeneous effects on the dependent variable in time and space, and have high correlation. In this case, redundant features can be reduced.
[0104] The site selection feature screening method provided in this application is based on artificial intelligence (AI). AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making functions.
[0105] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0106] In the embodiments of this application, the main artificial intelligence software technologies involved include the aforementioned speech processing technologies and machine learning. For example, it may involve Automatic Speech Recognition (ASR) technology in speech technology, including speech signal preprocessing, speech signal frequency analyzing, speech signal feature extraction, speech signal feature matching / recognition, and speech training.
[0107] For example, this could involve machine learning (ML), a multidisciplinary field encompassing probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning typically includes techniques such as deep learning, which includes artificial neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and deep neural networks (DNNs).
[0108] To reduce the number of target location features, the number of target location features can be controlled in any of the following ways to reduce the training cost of the neural network model.
[0109] 1) Reduce the number of samples by filtering through common features. Specifically, select the top K most common features, such as all features that have a probability greater than 0.001 in the target object. These features can characterize the user's common online behaviors.
[0110] 2) Features are extracted using clustering algorithms, specifically by using a length of... The vector fi represents each feature, where f is the number of target objects in the training data. i (j) represents the number of times user j contains feature i. All feature vectors are normalized, and then the K-means algorithm is applied to all feature vectors. The number of clusters is set according to the specific features. After clustering, a new length of is assigned to each new feature. vector The new feature is represented as: .
[0111] Where Cls(i) represents the category to which the original feature i belongs after clustering.
[0112] 3) Feature extraction via Locality Sensitive Hashing (LSH), specifically achieved by configuring a transformation matrix. Where d is the number of features in the original space, and k is the dimension of the lower-dimensional space. Multiplying the original features by the transformation matrix, i.e. This yields a new feature Y in the transformed low-dimensional space. Then, negative values in Y are replaced with zero, i.e. Features in the original space It can be used Approximate, where It is the Hamming distance between the local sensitive hashes of the two original features.
[0113] After obtaining the target location features through steps 301-303, a set of interest point (POI) data to be processed can be acquired. The POIs in the POI data set are combined to form corresponding POI sample pairs. Using the target location features, feature vectors corresponding to the POI sample pairs are extracted through the feature combination network of the POI selection model. The feature vectors corresponding to the POI sample pairs are sorted through the sorting network of the POI selection model to determine the POIs that match the target location features. In this invention, POIs refer to various public and service facilities used to provide public service products to citizens. Examples include POIs for educational facilities such as schools, kindergartens, and training institutions; POIs for medical and health facilities such as hospitals, clinics, and rehabilitation institutions; POIs for transportation facilities such as airports, train stations, and bus stations; POIs for sports facilities such as stadiums, swimming pools, and gyms; POIs for commercial and financial services such as shopping malls, cinemas, and banks; and POIs for social welfare and security facilities such as communication service centers and power supply bureaus.
[0114] In some embodiments of the present invention, after obtaining the target location features, point-of-interest (POI) data from different data sources can be acquired; the data sources of the POI data are classified; based on the target location features, the classification results of the POI data data sources are used to determine the same POI in the different data sources; and the POI data belonging to the same POI are aggregated to obtain complete detailed information about the POI. The representation of the POI data includes, but is not limited to: the name of the POI, the address of the POI, the contact number of the POI, the city information where the POI is located, and the latitude and longitude information of the POI. Since the data types of POI data from different data sources are not entirely the same, the technical solution shown in this embodiment can utilize the target location features to obtain detailed information about POIs of the same type belonging to the same POI. For example, by aggregating the data of POIs of the same type belonging to the same POI, all structured information in the detailed information about POIs belonging to the same POI can be obtained.
[0115] See Figure 6 , Figure 6This is an optional two-dimensional map display diagram of the site selection feature filtering method provided in the embodiment of the present invention. The displayed two-dimensional map includes various types of point of interest data, such as point of interest A, point of interest B, point of interest C, and point of interest D. These points of interest correspond to different site selection plots. When recommending a site selection location (any point of interest) that matches the target object to the user, it is necessary to filter the complex site selection features to reduce feature redundancy.
[0116] Figure 7 This is an optional flowchart illustrating the location feature filtering method provided in this embodiment of the invention. It can be understood that... Figure 7 The steps shown can be performed by various electronic devices that operate the site selection feature filtering device, such as dedicated terminals with site selection feature filtering devices, smartphones, smartwatches, and other electronic devices capable of receiving site selection data. The following section focuses on... Figure 7 The steps shown are explained.
[0117] Step 701: Target terminal location request, in response to the location request, obtain the location plot data.
[0118] Step 702: The target terminal performs feature processing on the selected site data to obtain initial site selection features.
[0119] Step 703: The target terminal performs spatiotemporal filtering on the initial location features and filters the initial location features that match the spatial differentiation relationship of the dependent variable based on the time-series mean threshold.
[0120] Step 704: The target terminal performs temporal clustering analysis on the initial location features that match the spatial differentiation relationship of the dependent variable, and filters out the initial location features that match the spatiotemporal correlation to obtain the target location features.
[0121] Step 705: The target terminal uses the target location features to extract feature vectors corresponding to the interest point sample pairs through the feature combination network of the interest point selection model.
[0122] Step 706: The target terminal sorts the feature vectors corresponding to the interest point samples through the sorting network of the interest point selection model to determine the interest points that match the target location features.
[0123] Beneficial technical effects:
[0124] This invention acquires site selection data and performs feature processing on the data to obtain initial site selection features. It then performs spatiotemporal filtering on these initial features, selecting those that match the spatial differentiation relationship of the dependent variable based on a time-series mean threshold. Finally, it performs time-series clustering analysis on these initial features to select those that match spatiotemporal correlation, obtaining target site selection features. This allows for the determination of a suitable location for a target object using these target site selection features. This reduces the cost of the site selection process, making it suitable for large-scale use by ordinary users. Furthermore, the automated filtering of site selection features yields more accurate target site selection features, resulting in more accurate site selection information.
[0125] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method of site selection feature screening, the method comprising: The method comprises: acquiring site plot data and performing feature processing on the site plot data to obtain initial site features; performing data conversion processing on the initial site features to determine dependent variables and independent variables corresponding to the initial site features; determining correlation relationship values of the dependent variables and the independent variables according to the dependent variables and the independent variables corresponding to the initial site features; calculating time series means of the correlation relationship values and screening the time series means of the correlation relationship values based on a time series mean threshold to obtain time series means of the correlation relationship values that match the time series mean threshold; determining initial site features that match a dependent variable spatial differentiation relationship according to the time series means of the correlation relationship values that match the time series mean threshold; performing clustering processing on the correlation relationship values corresponding to the initial site features according to a number of clusters that match the site plot data to obtain clustering results of the correlation relationship values; screening the initial site features that match spatiotemporal correlation according to the clustering results of the correlation relationship values to obtain target site features, so as to determine site locations that match target objects through the target site features.
2. The method of claim 1, wherein, The acquiring site plot data and performing feature processing on the site plot data to obtain initial site features comprises: performing data decomposition on the site plot data according to a time dimension through an open source cluster computing framework to obtain site plot data in the time dimension; performing feature extraction and feature normalization processing on the site plot data in the time dimension through the open source cluster computing framework to obtain normalized initial site features; performing feature deletion processing on the normalized initial site features based on entropy values of the normalized initial site features to obtain the initial site features.
3. The method of claim 1, wherein, The method further comprises: acquiring a set of interest point data to be processed; combining interest points in the set of interest point data to form corresponding interest point sample pairs; extracting feature vectors corresponding to the interest point sample pairs through a feature combination network of an interest point selection model by using the target site features; performing sorting processing on the feature vectors corresponding to the interest point sample pairs through a sorting network of the interest point selection model to determine interest points that match the target site features.
4. The method of claim 3, wherein, The method further comprises: acquiring interest point data in different data sources; classifying data sources of the interest point data; determining the same interest point in the different data sources based on a classification result of the data sources of the interest point according to the target site features; aggregating interest point data belonging to the same interest point to obtain complete detailed information of the interest point.
5. A site selection feature screening device characterized by, The apparatus comprises: an information transmission module configured to acquire site plot data and perform feature processing on the site plot data to obtain initial site features; The information processing module is configured to perform data conversion processing on the initial site selection feature, determine dependent variables and independent variables corresponding to the initial site selection feature, determine correlation values of the dependent variables and the independent variables according to the dependent variables and the independent variables corresponding to the initial site selection feature, calculate time series means of the correlation values, and perform screening on the time series means of the correlation values based on a time series mean threshold to obtain time series means of the correlation values that match the time series mean threshold, and determine initial site selection features that match a dependent variable spatial differentiation relationship according to the time series means of the correlation values that match the time series mean threshold. The information processing module is configured to perform clustering processing on the correlation values corresponding to the initial site selection feature according to a number of clusters that match the site block data, obtain a clustering result of the correlation values, and screen the initial site selection features that match a space-time correlation according to the clustering result of the correlation values to obtain target site selection features, so as to determine a site selection position that matches a target object by using the target site selection features.
6. The apparatus of claim 5, wherein, The information transmission module is further configured to: perform data decomposition on the site block data according to a time dimension by using an open source cluster computing framework to obtain site block data in the time dimension; perform feature extraction and feature normalization processing on the site block data in the time dimension by using the open source cluster computing framework to obtain normalized initial site selection features; perform feature deletion processing on the normalized initial site selection features based on entropy values of the normalized initial site selection features to obtain the initial site selection features.
7. The apparatus of claim 5, wherein, The information processing module is further configured to: obtain a set of point of interest data to be processed; combine points of interest in the set of point of interest data to form corresponding point of interest sample pairs; extract feature vectors corresponding to the point of interest sample pairs by using a feature combination network of a point of interest selection model through the target site selection features; perform sorting processing on the feature vectors corresponding to the point of interest sample pairs by using a sorting network of the point of interest selection model to determine points of interest that match the target site selection features.
8. The apparatus of claim 7, wherein, The information processing module is further configured to: obtain point of interest data in different data sources; classify data sources of the point of interest data; determine the same point of interest in the different data sources based on a classification result of the data sources of the point of interest through the target site selection features; aggregate point of interest data belonging to the same point of interest to obtain complete detailed information of the point of interest.
9. An electronic device, comprising: The electronic device includes: a memory configured to store executable instructions; a processor configured to execute the executable instructions stored in the memory to implement the site selection feature screening method in any one of claims 1 to 4.
10. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions are executed by the processor to implement the site selection feature screening method in any one of claims 1 to 4.
11. A computer-readable storage medium storing executable instructions, wherein the instructions, when executed by a processor, cause the processor to perform operations comprising: The executable instructions are executed by the processor to implement the site selection feature screening method in any one of claims 1 to 4.
Citation Information
Patent Citations
Site selection method and device
CN108984561A
Address selection method, device and non-transitory storage medium
CN110413722A