Lot Selection Method, Device, Electronic Device and Storage Medium
Through automated training model and similarity retrieval methods, the problem of low plot selection accuracy in the existing technology is solved, and more efficient plot selection is achieved, meeting user needs.
Patent Information
- Application Number
- CN202111057561.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-09-09
AI Technical Summary
The prior art is difficult to adapt to the basic data with huge amounts of data and the required data with small amounts of data for data mining, resulting in low plot selection accuracy.
By obtaining sample object data and preset plot image data, automate the model training, determine candidate plots, and reduce the plot selection results through similarity search to improve the accuracy.
It effectively improves the accuracy of plot selection, makes the selected plot more appropriate to user needs, and makes full use of the characteristics of plot image data and sample object data.
Smart Images

Figure CN113688299B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, this application relates to a plot selection method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] Plot selection refers to the process of selecting an address before construction. There are various factors to consider in plot selection, such as population density, traffic, existing buildings, etc., resulting in a large amount of basic data that needs to be considered during plot selection; in addition, plot selection also needs to be considered in combination with the development needs of the city or enterprise itself, that is, demand data with a relatively small amount also needs to be considered during plot selection.
[0003] However, in the related art, due to the inability to simultaneously adapt to the large amount of basic data and the relatively small amount of demand data for data mining, and unable to make full use of various data features, the accuracy of the final plot selection is very low. Summary of the Invention
[0004] The technical solutions provided in this application are intended to at least solve one of the above technical defects, especially the technical defect of low accuracy in plot selection. Among them, the technical solutions are as follows:
[0005] In the first aspect of this application, a plot selection method is provided, including:
[0006] Obtain sample object data related to plot selection, and based on the preset plot portrait data and the sample object data, determine candidate plots corresponding to the plot where the sample object is located;
[0007] Perform similarity retrieval on the candidate plots based on the preset plot portrait data to determine the target plot corresponding to the plot where the sample object is located among the candidate plots;
[0008] Among them, the plot portrait data includes portrait data corresponding to plots obtained by geographical space division.
[0009] In one embodiment, the determining, based on the preset plot portrait data and the sample object data, candidate plots corresponding to the plot where the sample object is located includes:
[0010] Automatically train the model based on the preset plot portrait data and the sample object data to obtain a target inference model;
[0011] Determine candidate plots corresponding to the plot where the sample object is located based on the target inference model;
[0012] Among them, the target inference model is an inference model constructed by selecting at least one model from multiple fusion models through semi-supervised automatic learning training.
[0013] In one embodiment, the automatic training of the model based on the preset plot portrait data and the sample object data to obtain the target inference model includes:
[0014] Automatically select features for the preset plot portrait data based on the sample object data to obtain the plot sample features of each plot;
[0015] Based on the training data determined by the correspondence between the plot where the sample object is located and the plot sample features, perform automatic training of the model to obtain the first inference model;
[0016] Infer the plot sample features of the unconfigured plots based on the first inference model, and merge the inference result data with a confidence level higher than the preset threshold with the training data to obtain the merged training data; the plot sample features of the unconfigured plots include plot features that do not have a correspondence with the plot where the sample object is located;
[0017] Use the merged training data for automatic training to obtain the target inference model;
[0018] Among them, automatic training includes automatically adjusting model parameters using classification model evaluation metrics.
[0019] In one embodiment, the automatic training of the model based on the training data determined by the correspondence between the plot where the sample object is located and the plot sample features to obtain the first inference model includes:
[0020] Determine the initial training data based on the correspondence between the plot where the sample object is located and the plot sample features;
[0021] Perform sampling processing on the initial training data to obtain the processed training data;
[0022] Based on the processed training data, perform automatic training of the model to obtain the first inference model.
[0023] In one embodiment, the use of the merged training data for automatic training to obtain the target inference model includes:
[0024] Use the merged training data for automatic training to obtain the second inference model;
[0025] Based on the received custom feature information and custom weight coefficients, adjust the merged training data;
[0026] Automatically train using the adjusted training data to obtain a target inference model.
[0027] In one embodiment, adjusting the merged training data based on the received custom feature information and custom weight coefficients includes:
[0028] Extract at least one piece of training feature information and its training weight coefficients used for training the second inference model and display them on the user interface of the client.
[0029] Based on the received custom feature information and custom weight coefficients, adjust the training feature information and its training weight coefficients to obtain adjusted training data.
[0030] In one embodiment, performing a similarity search on the candidate plots based on the preset plot portrait data to determine a target plot corresponding to the plot where the sample object is located among the candidate plots includes:
[0031] Obtain the basic features constructed based on the preset plot portrait data.
[0032] Extract the candidate features of the candidate plot.
[0033] Perform a feature similarity search on the basic features and the candidate features to determine a target plot corresponding to the plot where the sample object is located.
[0034] In one embodiment, constructing basic features based on the preset plot portrait data includes:
[0035] Determine the feature values of the preset metrics based on the preset plot portrait data.
[0036] Perform downsampling processing on the feature values to obtain the downsampled feature values.
[0037] Perform feature importance analysis based on the downsampled feature values to obtain basic features.
[0038] In one embodiment, determining the feature values of the preset metrics based on the preset plot portrait data includes at least one of the following:
[0039] At a set scale, determine the feature value of the neighborhood entropy based on the number of types of sample objects within the neighboring range of a preset radius centered on the location of the plot.
[0040] At a set scale, determine the feature value of the competition degree based on the number of sample objects of the same type within the neighboring range of a preset radius centered on the location of the plot.
[0041] Under a set scale, determine the characteristic value of the spatial correlation effect of the plot based on the number of set types of sample objects within the proximity range of a preset radius centered on the location of the plot and the attraction coefficient between the types of sample objects;
[0042] Under a set scale, determine the characteristic value of the transfer quality based on the number of user visits within the proximity range of a preset radius centered on the location of the plot and the transfer probability between the types of sample objects.
[0043] In one embodiment, the feature similarity retrieval of the basic features and candidate features to determine the target plot corresponding to the plot where the sample object is located includes:
[0044] Perform feature similarity retrieval on the basic features and candidate features to determine the first plot among the candidate plots;
[0045] Based on at least one of the number of transactions of the point of interest (POI) and the custom selection information, determine the target plot corresponding to the plot where the sample object is located among the first plots.
[0046] In one embodiment, before obtaining the sample object data related to plot selection, it further includes:
[0047] Respond to the plot selection request operation of the client;
[0048] After determining the target plot, it further includes:
[0049] Display the plot selection information corresponding to the target plot on the client;
[0050] The display of the plot selection information corresponding to the target plot includes: displaying the location of the target plot on the map interface based on a preset marking form.
[0051] In the second aspect of the present application, there is provided a plot selection device, including:
[0052] The first determination module is configured to obtain sample object data related to plot selection, and determine candidate plots corresponding to the plot where the sample object is located based on the preset plot portrait data and the sample object data;
[0053] The second determination module is configured to perform similarity retrieval on the candidate plots based on the preset plot portrait data to determine the target plot corresponding to the plot where the sample object is located among the candidate plots;
[0054] Among them, the plot portrait data includes the portrait data corresponding to the plots obtained by geographical space division.
[0055] In the third aspect of the present application, there is provided an electronic device, and the electronic device includes:
[0056] One or more processors;
[0057] A memory;
[0058] One or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute the method provided in the first aspect.
[0059] In the fourth aspect of the present application, there is provided a computer-readable storage medium for storing computer instructions, which when run on a computer, enable the computer to execute the method provided in the first aspect.
[0060] In the fifth aspect of the present application, there is provided a computer program product including a computer program or instructions, characterized in that when the computer program or instructions are executed by a processor, the steps of the method provided in the first aspect are implemented.
[0061] The beneficial effects brought by the technical solution provided in the present application are:
[0062] In the present application, when obtaining sample object data related to plot selection, first, based on the preset plot portrait data and the sample object data, candidate plots corresponding to the plot where the sample object is located are determined, and at this time, the number of candidate plots obtained is relatively large; on this basis, the present application also performs a similarity search on the candidate plots based on the preset plot portrait data to further narrow down the plot selection result and improve the accuracy of plot selection, so as to determine the target plot corresponding to the plot where the sample object is located among the candidate plots. Among them, the plot portrait data includes portrait data corresponding to plots obtained by geographical space division. The implementation of the solution of the present application can effectively mine the existing plot portrait data and the sample object data provided by the user, which is beneficial to improving the accuracy of plot selection and making the selected plot more in line with the user's needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.
[0064] Figure 1 It is a flowchart of a plot selection method provided by an embodiment of the present application;
[0065] Figure 2 It is a system flowchart of a plot selection method provided by an embodiment of the present application;
[0066] Figure 3 It is a flowchart of model automatic training in a plot selection method provided by an embodiment of the present application;
[0067] Figure 4 Flow chart of model automatic training in a plot selection method provided by an embodiment of the present application;
[0068] Figure 5 Flow chart of constructing basic features based on preset plot portrait data in a plot selection method provided by an embodiment of the present application;
[0069] Figure 6 Flow chart of similarity retrieval in a plot selection method provided by an embodiment of the present application;
[0070] Figure 7 Flow chart of selecting high-quality plots in a plot selection method provided by an embodiment of the present application;
[0071] Figure 8 System flow chart of another plot selection method provided by an embodiment of the present application;
[0072] Figure 9 Schematic diagram of an interaction environment applied to a plot selection method provided by an embodiment of the present application;
[0073] Figure 10 Schematic diagram of the structure of a plot selection device provided by an embodiment of the present application;
[0074] Figure 11 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0075] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.
[0076] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0077] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0078] The following explains the technologies and terms related to this application:
[0079] AI (Artificial Intelligence) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0080] In this application, directions such as machine learning / deep learning can be adopted.
[0081] Among them, ML (Machine Learning) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0082] In the plot selection method provided by the embodiments of this application, machine learning-related technologies can be used to process a small amount of sample object data, and deep learning-related technologies can be used to process a large amount of plot data.
[0083] Lot selection refers to the process of selecting a site before construction. There are various factors to consider in lot selection, such as population density, transportation, existing buildings, etc., which leads to a large amount of basic data to be considered during lot selection. In addition, lot selection also needs to be considered in combination with the development needs of the city or enterprise itself, that is, demand data with less data volume also needs to be considered during lot selection.
[0084] However, in the related art, due to the inability to simultaneously adapt to the large amount of basic data and the small amount of demand data for data mining, various data features cannot be fully utilized, resulting in very poor rationality of the finally selected lot. In addition, the main way to embed user knowledge in the related art is to directly specify weights for linear calculation. However, for massive data, similar simple processing cannot fully utilize the feature expression ability of the mined data and is also difficult to fit the scenario for use.
[0085] To solve at least one of the above problems, the present application provides a lot selection method, device, electronic device and computer-readable storage medium; it can be used in scenarios such as maps and vehicle networking. Specifically, it can make full use of the sample object data provided by the user and the existing lot portrait data, support automatic machine learning (AutoML) lot selection analysis for massive data, and can also integrate the user's experience. The algorithm framework proposed in the present application is designed to fit the scenario, which can effectively improve the rationality of lot selection.
[0086] The following will specifically describe the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0087] The present application provides a lot selection method in an embodiment, as Figure 1 shown, Figure 1 shows a schematic flowchart of a lot selection method provided by an embodiment of the present application. Among them, this method can be executed by any electronic device, such as a user terminal or a server. The user terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted device, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. However, the present application is not limited thereto.
[0088] First, the relevant content of the plot selection method provided by the embodiments of the present application will be described below.
[0089] Taking the selection of a plot by an enterprise to set up an offline store as an example, assuming a plot selection scenario for opening a store, the plot selection method provided by the embodiments of the present application can be used to select a plot suitable for opening a store based on the sample store data provided by the user. Among them, a plot can be used for certain application purposes, such as planning and plot selection, within a certain geographical space distribution range; in terms of shape, a plot can include regular plots and irregular plots. In the embodiments of the present application, a certain geographical space distribution range (such as a certain administrative region) can be divided into rectangular plots of a set length, and the plots obtained by the division are used as the basis for the implementation of the present application; however, the embodiments of the present application can also be applied to other plot planning methods, and the embodiments of the present application do not limit this.
[0090] Specifically, in combination with Figure 1 、 Figure 2 and Figure 8 , the following steps S101 - S102 included in this method will be described:
[0091] Step S101: Obtain sample object data related to plot selection, and determine candidate plots corresponding to the plot where the sample object is located based on the preset plot portrait data and the sample object data.
[0092] Among them, in the plot selection scenario, the user can upload the relevant data of some known sample objects to the plot selection platform, and the plot selection platform can fuse the data provided by the user with the existing plot data for processing, and screen out candidate plots corresponding to the plot where the sample object is located (that is, find candidate plots with higher relevance based on the corresponding relationship between the sample object and the plot where it is located).
[0093] Among them, the sample object data can include the attribute information of the sample object itself. For example, when the sample object is a clothing store, it can include the business data, store area, and store location (which can correspond to the plot location) of the clothing store. Optionally, the sample object data obtained in step S101 can be stored in units of sample objects, that is, one sample object corresponds to one piece of sample object data, and the amount of sample object data depends on the data that the user can provide.
[0094] Among them, the plot portrait data is data that, based on big data, abstracts each specific piece of information of the plot into tags, uses the tags to concretize the image of the plot, and is then used for analysis services. For example, the portrait data of plot 1 can include the geographical location, population information, traffic information, etc. within a certain range centered on the geographical location; the characteristics that the plot portrait data can include are diverse. Specifically, the plot portrait data includes the corresponding portrait data of the plot obtained by geographical space division.
[0095] Compared with the data volume of the plot portrait data, the sample object data belongs to the data to be analyzed with a relatively small data volume, and the plot portrait data belongs to the existing data with a huge data volume.
[0096] Specifically, the implementation of the above step S101 can be achieved through machine learning related technologies, and the specific implementation process will be described in subsequent embodiments.
[0097] Step S102: Perform a similarity search on the candidate plots based on the preset plot portrait data to determine the target plot corresponding to the plot where the sample object is located among the candidate plots.
[0098] In the embodiment of the present application, considering that the plot selection result obtained by screening based on step S101 is relatively rough, that is, the candidate plots can include many plots, and the accuracy of plot selection is not high. On this basis, the present application adopts step S102 to perform a similarity matching and filtering step based on the plot portrait data, perform a similarity search on the candidate plots, so as to further reduce the number of plots included in the plot selection result and improve the rationality and accuracy of plot selection.
[0099] Among them, the similarity search can be a feature search using a vector search algorithm, which is used to further optimize the candidate plots, that is, the number of candidate plots is greater than the number of target plots. Optionally, the target plot can include at least one plot.
[0100] Specifically, the implementation of the above step S102 can be achieved through deep learning related technologies, and the specific implementation process will be described in subsequent embodiments.
[0101] In a feasible embodiment, the plot selection method provided by the embodiment of the present application may also only include step S101, that is, output the candidate plots as the plot selection result; it may also only include step S102, that is, directly perform plot selection processing based on the plot portrait data without fusing the sample object data provided by the user.
[0102] The following describes the specific process of how to determine the candidate plots corresponding to the plot where the sample object is located.
[0103] In one embodiment, in step S101, based on the preset plot portrait data and the sample object data, determining candidate plots corresponding to the plot where the sample object is located includes steps A1 - A2:
[0104] Step A1: Based on the preset plot portrait data and the sample object data, perform automated training of the model to obtain a target inference model.
[0105] In the embodiments of the present application, considering that in some plot selection scenarios, after the user submits the sample object data, the plot selection algorithm is required to automatically give the plot selection result. However, in the related art, automatic learning and inference are not supported, and the plot selection algorithms in the related art are difficult to batch - replicate, resulting in operation and maintenance difficulties (for example, when the sample objects are catering stores and clothing stores, the learned plot selection algorithms are different. If different types need to be processed one by one, the computational complexity is very high and the operation and maintenance cost is very high). Therefore, the embodiments of the present application use AutoML (Automated Machine Learning) to solve this problem.
[0106] Among them, AutoML (Automated Machine Learning) is an end - to - end process automation that applies machine learning to real - world problems, and uses motorized training methods to train machine learning models and AI models. Specifically, as shown in Figure 3 and Figure 4 AutoML can achieve automation in three aspects: feature engineering, model construction, and parameter optimization.
[0107] Specifically, the preset plot portrait data and sample object data can be used to construct the training data (feature engineering) for the automatic learning and training of the inference model. Then, through automated training, a suitable model is selected for construction in a framework structure integrating multiple models, and finally, automated parameter tuning is achieved to complete the construction of the target inference model.
[0108] Step A2: Based on the target inference model, determine candidate plots corresponding to the plot where the sample object is located.
[0109] Specifically, the target inference model is a model trained based on the sample object data currently provided by the user and the existing plot portrait data. The target inference model can select relatively suitable candidate plots from the existing plots (known through the plot portrait data) based on the sample object data as the output of the model. For example, when the sample object data provided by the user is data of a catering store, the target inference model can screen out other relatively suitable candidate plots for opening a catering store.
[0110] In an embodiment of the present application, the target inference model may be an inference model constructed by selecting at least one model from a multi-fusion model through semi-supervised automatic learning training. The following describes the automatic learning training process of the target inference model.
[0111] In one embodiment, in step A1, based on the preset plot portrait data and the sample object data, automatic training of the model is performed to obtain the target inference model, including steps A11 - A14:
[0112] Step A11: Perform automatic feature selection on the preset plot portrait data based on the sample object data to obtain the plot sample features of each plot.
[0113] Specifically, step A11 belongs to the execution step of automatic feature engineering in AutoML. In the machine learning process, the human and time costs consumed by feature engineering are relatively high, and automatic feature engineering can be used to automate the operation and improve the efficiency of model training.
[0114] Among them, a feature is an abstract numerical transformation from a specific object to a numerically represented value, and feature engineering refers to the process of converting raw data into features. Features can be better expressed to the model, improving the accuracy of the model's data processing.
[0115] Optionally, the processing of automatic feature engineering may include feature completion, normalization, outlier processing, etc.
[0116] Among them, in the process of automatic feature selection, considering that thousands of features can be received based on the plot portrait data, and the number of features is much larger than the number of sample objects, and not all features are conducive to machine understanding and not all features are relevant to the modeling content, some features with higher importance can be selected and the plot sample features required for constructing the model in the present application can be adaptively generated. Optionally, if the sample object data includes the number of people flowing in store A, data with a relatively high correlation with the number of people flowing in store A can be selected from the plot portrait data to generate plot sample features. In some embodiments, automatic feature selection may also be performed without relying on the sample object data, that is, feature selection can be performed only on the plot portrait data; if the plot portrait data can be applied to multiple scenarios and adapt to the plot selection method of the present application, feature selection can be performed from the application perspective of plot selection.
[0117] Optionally, after automatic feature selection, the feature types in the plot sample features corresponding to each plot may be the same or different; it can be set according to actual needs.
[0118] Step A12: Based on the training data determined by the correspondence between the plot where the sample object is located and the plot sample features, perform automated training of the model to obtain the first inference model.
[0119] Specifically, based on the sample object data provided by the user, the plot where the sample object is located can be determined, and then the correspondence between the plot and the plot sample features can be determined. For example, plot B where sample object A is located has a correspondence with plot sample features C, D, and E; it can be understood that relative to plot sample features C, D, and E, there is a labeled plot B, and plot B has a mapping relationship with sample object A.
[0120] Among them, the correspondence between the plot where the sample object is located and the plot sample features can construct the training data of the positive samples, and then based on this positive sample training data, automated training of the model can be performed to obtain the first inference model. That is to say, the first inference model is trained based on the positive sample training data constructed from the sample object data provided by the user.
[0121] In one embodiment, in step A12, based on the training data determined by the correspondence between the plot where the sample object is located and the plot sample features, perform automated training of the model to obtain the first inference model, including the following steps A121 - A123:
[0122] Step A121: Determine the initial training data based on the correspondence between the plot where the sample object is located and the plot sample features;
[0123] Step A122: Perform sampling processing on the initial training data to obtain the processed training data;
[0124] Step A123: Based on the processed training data, perform automated training of the model to obtain the first inference model.
[0125] Specifically, considering that the amount of sample object data provided by the user is small, to further increase the proportion of the training data of the positive samples, the training data enhancement step can be executed, and the enhanced training data is used to perform automated training on the model. Among them, training data enhancement can use techniques such as Borderline - SMOT (Borderline Synthetic Minority Over - sampling Technique) for oversampling processing, which is mainly a refinement method for imbalanced datasets (such as in this application, the sample object data is less, while the amount of plot sample feature data is large).
[0126] Optionally, the sampling processing performed on the initial training data can be either up - sampling or down - sampling. Up - sampling processing can increase the amount of training data; down - sampling processing can extract the training data, thereby increasing the proportion of positive samples in the training data.
[0127] Step A13: Infer the plot sample features of the unconfigured plots based on the first inference model, and merge the inference result data with a confidence level higher than the preset threshold with the training data to obtain the merged training data.
[0128] Specifically, since the amount of sample object data provided by the user is small, there may be a large number of plot sample features of unconfigured plots, that is, training data without labels. For the training data without labels, the embodiments of the present application can use the first inference model for inference, that is, use the training data without labels as the input data of the first inference model, and the first inference model infers the plots corresponding to the relevant features.
[0129] Among them, considering that the inference accuracy of the first inference model is low, in the execution of subsequent steps, the inference result data with a confidence level higher than the preset threshold can be selected and merged with the training data (the data used to train the first inference model based on the sample object data) to obtain the merged training data.
[0130] Among them, the plot sample features of the unconfigured plots include plot features that do not have a corresponding relationship with the plot where the sample object is located.
[0131] Step A14: Use the merged training data for automated training to obtain the target inference model.
[0132] Specifically, the merged training data can be used to continue training the first inference model, or the initialized model can be retrained to obtain the final target inference model.
[0133] Optionally, the AutoML automatic training provided by the embodiments of the present application belongs to the process of model construction. It can be in the framework of multi-model fusion, make full use of the characteristics of various data, and select one or more suitable models to construct the inference model required by the present application. It can be understood that the algorithms used to construct the inference model may be different according to different plot selection requirements. For example, the target inference model can be a clustering model, a multi-classification model, etc.
[0134] Among them, the automated training includes automatically adjusting the model parameters using the classification model evaluation index.
[0135] Specifically, the automated training process includes the automated adjustment of model parameters. In the embodiments of the present application, a greedy search method based on the KS value can be used for automated parameter tuning. Among them, the KS value is an evaluation index used to distinguish the separation degree of predicted positive and negative samples in the model. The prediction result of each sample is a probability or a score range, from the smallest probability or lowest score to the largest probability or highest score, and the cumulative distribution of positive and negative samples; the KS value is the absolute value of the largest difference between the two distributions. The value range of the KS value is [0, 1]. The larger the KS value, the better the separation degree of positive and negative samples. Among them, greedy search is to select the output value with the largest probability in each step to obtain the decoded output sequence (that is, to give the search path corresponding to the search space); it can be understood that a label obtained by model inference can correspond to multiple paths. Therefore, the path with the largest probability is not equal to the path with the largest probability value of the final label. Specifically, when performing automated parameter tuning using the greedy search method based on the KS value, only a single parameter is considered for forward and backward search each time to obtain the optimal parameter.
[0136] In the embodiments of the present application, in combination with Figure 3 and Figure 4 as shown, through the implementation of steps A11 - A14, a semi-supervised automatic learning model training method is specifically adopted.
[0137] In a feasible embodiment, in step A14, the merged training data is used for automated training to obtain the target inference model, including steps A141 - A143:
[0138] Step A141: Use the merged training data for automated training to obtain the second inference model;
[0139] Step A142: Adjust the merged training data based on the received custom feature information and custom weight coefficients;
[0140] Step A143: Use the adjusted training data for automated training to obtain the target inference model.
[0141] As Figure 3As shown, after obtaining the second inference model through automated training based on the merged training data, the feature information (which can be the feature type) and its weight coefficients used during the training process of the second inference model can be extracted and presented to the user. If the user determines, based on their own needs, that the feature type and the corresponding weight coefficients need to be adjusted, the user can input custom feature information and the corresponding weight coefficients, and this process can involve adding and filtering relevant features in the sample plot feature set. That is, in the embodiments of the present application, the user can dynamically add or delete sample plot features based on actual experience, which can effectively improve the fit between the plot selection result and the user's own needs, thereby enhancing the user experience.
[0142] Specifically, after training the second inference model, at least one piece of training feature information used for training the second inference model and the training weight coefficient corresponding to the training feature information can be extracted and displayed on the user interface of the client. The user can consider whether to adjust the data of the training target inference model according to the plot selection requirements based on the displayed training feature information and the corresponding training weight coefficients, so as to obtain a plot selection result that better meets their own needs. Furthermore, the user can directly adjust the training feature information and its training weight coefficients, or input new feature information and the corresponding weight coefficients by themselves. The data generated during the user's operation on the user interface can be regarded as the corresponding custom feature information and custom weight coefficients. Based on this, the adjusted training data can be obtained.
[0143] The following describes the specific process of determining the target plot corresponding to the plot where the sample object is located.
[0144] In one embodiment, the similarity retrieval of the candidate plots based on the preset plot portrait data in step S102 to determine the target plot corresponding to the plot where the sample object is located among the candidate plots includes steps B1 - B3:
[0145] Step B1: Obtain the basic features constructed based on the preset plot portrait data.
[0146] Specifically, as Figure 5As shown, the existing plot portrait data in the embodiments of the present application can be constructed based on the basic data. Under the basic data, corresponding plot portrait data can be constructed for each divided plot. In addition, considering the adaptability of the plot portrait data to the plot selection method provided in the embodiments of the present application, in the construction of the feature data, in addition to portrait features such as the transportation, points of interest (POIs), and population of the plot, features facing plot selection also need to be constructed. Optionally, the feature data may include proximity entropy based on scale, competition degree, Jensen quality (spatial correlation effect of different plots), population transfer density, transfer quality, and land rent characteristics, etc. The embodiments of the present application can set a set of index systems for the actual application of plot selection, and determine the characteristic values corresponding to each index; it can be understood that step B1 is the execution step of feature extraction and preprocessing for the plot portrait data.
[0147] Among them, in a geographic information system, a POI can refer to a physical point such as a house, a store, a mailbox, or a bus stop.
[0148] In one embodiment, constructing basic features based on preset plot portrait data includes steps C1 - C3:
[0149] Step C1: Determine the characteristic values of preset indicators based on the preset plot portrait data.
[0150] Specifically, the present application can set the index system from the perspective of the actual application of plot selection, and consider the scale factor when dealing with the indicators of the involved types. Among them, when the sample object is a store, the type can refer to a clothing store, a restaurant, a toy store, etc.
[0151] In a feasible embodiment, determining the characteristic values of preset indicators based on the preset plot portrait data in step C1 includes at least one of the following steps C11 - C14:
[0152] Step C11: Determine the characteristic value of the proximity entropy based on the number of types of sample objects within the proximity range of a preset radius at the location of the plot under the set scale.
[0153] Specifically, the proximity entropy x l (r) can be calculated using the following formula (1):
[0154]
[0155] Among them, l represents the plot location, and r represents the preset radius; Nr(l,r) represents the number of types within the adjacent range corresponding to the r radius at the l location; it can be understood that at different plot scales, classification systems with different precisions are adopted. ST refers to all types at all specified scales.
[0156] Based on the above formula (1), it can be seen that the greater the adjacent entropy, the greater the type diversity of the area.
[0157] Step C12: At the set scale, determine the characteristic value of the competition degree based on the number of samples with the same type within the adjacent range of the preset radius at the location of the plot.
[0158] Specifically, the competition degree x l (r) can be calculated using the following formula (2):
[0159]
[0160] Among them, l represents the plot location, and r represents the preset radius; N Srl (l,r) refers to the number of the same types within a certain radius r at a specific scale. The types here are also associated with the scale, and different scales have corresponding classification systems.
[0161] Step C13: At the set scale, determine the characteristic value of the spatial correlation effect of the plot based on the number of set types of the sample objects within the adjacent range of the preset radius at the location of the plot and the attraction coefficient between the types of the sample objects.
[0162] Specifically, the Jensen quality x l (r) (spatial correlation effect of the plot) can be calculated using the following formula (3);
[0163]
[0164] Among them, l represents the plot location, and r represents the preset radius; refers to how many types r are observed on average at the location of the specified type rl at the specified scale. ST refers to the set of types at the specified scale. k p →rl refers to the attraction coefficient between types. rp →rl refers to the attraction coefficient between types.
[0165] Step C14: At the set scale, determine the characteristic value of the transfer quality based on the number of user visits within the adjacent range of the preset radius at the location of the plot and the transition probability between the types of the sample objects.
[0166] Specifically, the transfer quality index x l (r) can be calculated using the following formula (4):
[0167]
[0168] Among them, l represents the plot location, and r represents the preset radius; C p refers to the number of user visits to the specified location P; σ Srp→Srl refers to the transition probability between two types Srp and Srl at a certain scale.
[0169] Specifically, this indicator can be used to measure the user attraction of a specified location to other locations within a region at a certain scale. For example, the user attraction of specified location 1 (such as restaurant B) to other locations 2 and 3 (such as restaurants C and D) within the region at a certain scale A.
[0170] Step C2: Perform downsampling on the eigenvalue to obtain the downsampled eigenvalue.
[0171] Considering that the amount of data related to the features of the plot is very large, the embodiments of the present application can perform data dimensionality reduction on the features and reduce data sparsity.
[0172] Step C3: Perform feature importance analysis based on the downsampled eigenvalue to obtain the basic features.
[0173] Specifically, the execution of step C3 can be understood as the processing steps of feature importance analysis and feature screening. Algorithms such as xgboost (eXtreme Gradient Boosting) and linear regression can be used for feature importance analysis. In addition, the weighted analysis of feature weights can be combined with the geographical detector at the same time to perform feature screening and obtain the basic features.
[0174] Optionally, step C3 also involves feature deletion processing. When performing multiple rounds of feature deletion in the model applied to importance analysis, it can be processed through the KS index value and PSI index value of the model. If the index value does not change (the difference is small) or tends to change for the better, the corresponding feature can be deleted. Among them, the KS index value can be determined using the above KS value, and the PSI index value is the population stability index. The smaller the PSI value, the smaller the difference between the two distributions, indicating greater stability.
[0175] Optionally, as Figure 6 shown, the processing of steps C1 - C3 can correspond to the execution process of feature engineering and feature Embedding extraction.
[0176] In a feasible embodiment, the basic features applied to similarity matching filtering based on plot data may be different or the same as the plot sample features applied to automatic learning training for a small number of sample object data.
[0177] Step B2: Extract the candidate features of the candidate plot.
[0178] Specifically, the candidate features of the candidate plot can be extracted based on feature engineering and feature Embedding as described in Figure 6 , or can be obtained directly on the basis of the basic features by creating an index relationship between the basic features and the plot. Through index establishment, feature search can be carried out quickly, improving the efficiency of subsequent feature retrieval.
[0179] Step B3: Perform a feature similarity retrieval on the basic features and the candidate features to determine the target plot corresponding to the plot where the sample object is located.
[0180] Specifically, the similarity retrieval of the candidate plot can be performed based on deep learning and a vector engine, where the similarity retrieval of the candidate plot is achieved by performing a feature similarity retrieval on the basic features and the candidate features. The implementation of this step B3 is based on the basic features of the plot portrait data, and plot selection is performed among the candidate plots to obtain the target plot corresponding to the plot where the sample object is located.
[0181] As Figure 6 shown, after performing the feature retrieval, approximate plot data can be obtained, and this approximate plot data corresponds to the relevant data of the target plot.
[0182] In one embodiment, performing a feature similarity retrieval on the basic features and the candidate features in step B3 to determine the target plot corresponding to the plot where the sample object is located includes steps B31 - B32:
[0183] Step B31: Perform a feature similarity retrieval on the basic features and the candidate features to determine the first plot among the candidate plots.
[0184] Step B32: Based on at least one of the number of transactions of the point of interest (POI) and the custom selection information, determine the target plot corresponding to the plot where the sample object is located among the first plots.
[0185] Specifically, as Figure 7 shown, in order to further improve the rationality of plot selection and the degree of conformity with user needs in the embodiments of the present application, on the basis of the first plot obtained by performing similarity matching and filtering on the candidate plots, further filtering of the plot can be performed based on data such as the number of POI transactions and user-defined selection information, etc., to obtain the selected plot (target plot) after fine filtering.
[0186] Optionally, as Figure 7As shown, the model loads the Spark computing engine. Spark is a fast and general computing engine designed for large-scale data processing, which can be applied to complete various operations. In the embodiment of the present application, loading Spark can be applied to machine learning, that is, the processing of determining the first plot (which can correspond to the target plot in step B3) is completed by the model loading Spark.
[0187] Among them, the custom selection information can be set based on the user's personalized plot selection requirements, such as plot selection information like "not more than 300 meters away from the subway station" and "no competitor stores within 100 meters", that is, the first plot can be screened according to the user's needs and combined with geospatial analysis to obtain the selected plot (the final target plot).
[0188] In a feasible embodiment, before obtaining the sample object data related to plot selection in step S101, it further includes: responding to the plot selection request operation of the client.
[0189] Specifically, in the embodiment of the present application, it can be to obtain the sample object data in response to the plot selection request operation initiated by the user through the client.
[0190] Optionally, when the user initiates the plot selection request operation, the user can put forward plot selection requirements, such as restricting the administrative region (geospatial) of plot selection, the object type of plot selection (such as the specific type of a store when a certain enterprise wants to open a physical store), etc.
[0191] After determining the target plot in step S102, it further includes: displaying the plot selection information corresponding to the target plot on the client.
[0192] Among them, displaying the plot selection information corresponding to the target plot includes: displaying the location of the target plot on the map interface based on a preset marking form.
[0193] Specifically, the preset marking form can include displaying a location indication label at the location of the target plot, and displaying longitude and latitude information, plot size, plot orientation, etc. corresponding to the location of the target plot on the map interface.
[0194] Specifically, after determining the target plot corresponding to the plot where the sample object is located, plot selection information corresponding to the target plot can be displayed on the user interface of the client. Among them, the plot selection information may include the longitude and latitude coordinates of the plot, the size of the plot, the orientation of the plot, the feature information used to select the plot, etc. Among them, the display method of the plot selection information can be to use the form of a map to display the location of the target plot on the map through specific labels or effects. Among them, the data format transmitted to the front end of the client can adopt the json (JavaScript Object Notation, object representation method) format.
[0195] Optionally, in the embodiments of the present application, the target plot includes at least one piece. When there are more than one target plot, based on the geographical location where the user is currently located, the target plots can be sorted from near to far, and finally the sorted plot selection information items can be displayed.
[0196] The following combines Figure 8 and Figure 9 , and gives a feasible application example for the plot selection method provided in the embodiments of the present application.
[0197] Taking the user as enterprise A, which hopes to open a physical cinema in administrative region B as an example for illustration.
[0198] The user can upload data related to the cinema (sample object data) through the terminal 100. Among them, the data related to the cinema is not limited to the data of the cinema itself, but can also include the data of other objects in the vicinity of the cinema, such as the relevant data of the restaurant, toy store, etc. opened next to the cinema. Among them, although the user hopes to select a plot in administrative region B, the administrative region where the sample object is located in the sample object data uploaded by the user is not limited; because even the sample object data in other administrative regions has certain reference value for plot selection in administrative region B.
[0199] Based on the user's plot selection requirement (which can be the plot selection requirement triggered by submitting a plot selection request operation on the user interface of the terminal 100 after the user uploads data related to the cinema), existing plot portrait data can be obtained for plot selection. Among them, the existing plot portrait data includes the portrait data of plots obtained by geographical space division. In the embodiments of the present application, considering the user's hope to select a plot in administrative region B and the problem that the amount of plot portrait data is relatively large, only the plot portrait data corresponding to administrative region B can be obtained for plot selection processing.
[0200] Among them, such as Figure 8As shown, after receiving the data related to the cinema uploaded by the user, the plot sample features can be obtained (obtained in the feature engineering module), and then the training data can be constructed based on the sample object data uploaded by the user and the plot sample features. Additionally, considering that the amount of data in the sample object data uploaded by the user is small, and accordingly the amount of training data is also small, thus, the constructed training data can be subjected to data augmentation processing (oversampling processing) to perform AutoML automatic training based on the augmented training data to obtain the first inference model. Further, the training method adopted in the embodiments of this application is semi-supervised automatic learning. Therefore, the first inference model can be used to infer the plot sample features of the plots without configured plots, and then the inference result data with a confidence level higher than the preset threshold can be merged with the training data (which can be the augmented training data) to obtain the merged training data; then the merged training data is used for automatic training to obtain the second inference model. At this time, the feature types and corresponding weights used to construct the second inference model can be extracted and fed back to the user on the user interface of the client. If the user deems it necessary to adjust the feature types and weight coefficients in combination with their plot selection requirements, the user can input custom feature types and custom weight coefficients to further perform AutoML automatic training using the training data fine-tuned by the user, and finally obtain the target inference model. The target inference model is an inference model obtained by automatically training the model framework of the plot selection algorithm after the user submits data related to the cinema. At this time, within the general framework of the plot selection algorithm, the target inference model can determine candidate plots based on the data related to the cinema uploaded by the user. Among them, the candidate plots are determined based on the plot where the sample object uploaded by the user is located. The sample object can include one or multiple; at least one candidate plot can be determined for one plot where one sample object is located.
[0201] As Figure 8 shown, the candidate plots, as the model output, are the basic data for finally determining the selected plots. Among them, in the similarity matching filtering based on plot data, the basic data for processing is also the output of the target inference model, that is, the candidate plots. It can be understood that the first plot is determined by selecting plots from the candidate plots based on the plot portrait data.
[0202] Among them, the input data for similarity matching filtering based on plot data can be feature data extracted from plot portrait data, or can also be various feature values or basic features corresponding to preset indicators. In the process of similarity matching filtering based on plot data, first, feature extraction (including feature engineering between feature extractions) and preprocessing (such as index establishment) are performed on the plot portrait data; then, based on deep learning and a vector engine, similarity retrieval can be performed on candidate plots. Among them, the basic features and candidate features can be processed into basic feature vectors and candidate feature vectors. After performing feature retrieval, approximate plot data (data related to the first plot) can be obtained. It can be understood that the number of the first plots is less than the number of candidate plots, and the first plots can include at least one.
[0203] Such as Figure 8 shown, the relevant data of the first plot can be input into the model loaded with Spark for processing. This machine learning model can filter only the first plot to obtain a rough selection of plots, or can also filter in combination with the relevant data of candidate plots to obtain a rough selection of plots; in addition, after obtaining the rough selection of plots, fine filtering can also be performed based on the POI transaction times level and / or user-defined selection information to finally obtain a selected plot (target plot). Among them, the user-defined selection information can be input simultaneously when initiating a plot selection request, or can also be user-defined selection information input at any time point during the plot selection process. Specifically, the user-defined selection information can be "there cannot be other cinemas in the same business circle", "the distance from the subway station or bus stop cannot exceed 100 meters", "the pedestrian flow on non-working days is more than three times that on working days", etc.
[0204] After finally determining the target plot, the plot selection information corresponding to the target plot can be displayed on the user interface of the client. Such as star marking and displaying the geographical location of the target plot (in administrative region B) on the map interface.
[0205] Such as Figure 9 shown, the plot selection method provided by the embodiments of the present application can be executed on the terminal 100 or can also be executed on the server 200.
[0206] When the plot selection method is executed on the terminal 100, the preset plot portrait data can be stored in the server 200 or the database corresponding to the server 200. Furthermore, when the terminal 100 executes the plot selection method, it can obtain the plot portrait data from the server 200 through the network 300.
[0207] When the plot selection method is executed on the server 200, after the user uploads the sample object data from the terminal 100 to the server 200 via the network 300, the server 200 executes the plot selection method to determine the target plot and then feeds it back to the terminal 100 via the network 300, and finally displays the plot selection information corresponding to the target plot on the user interface of the terminal 100.
[0208] An embodiment of the present application provides a plot selection device, as Figure 10 shown. The plot selection device 900 may include: a first determination module 101 and a second determination module 102.
[0209] Among them, the first determination module 101 is used to obtain the sample object data related to plot selection, and determine the candidate plots corresponding to the plot where the sample object is located based on the preset plot portrait data and the sample object data; the second determination module 102 is used to perform similarity retrieval on the candidate plots based on the preset plot portrait data to determine the target plot corresponding to the plot where the sample object is located among the candidate plots; wherein, the plot portrait data includes the portrait data corresponding to the plots obtained by geographical space division.
[0210] In an embodiment, when the first determination module 101 is used to execute determining the candidate plots corresponding to the plot where the sample object is located based on the preset plot portrait data and the sample object data, it specifically is used for:
[0211] Automatically train the model based on the preset plot portrait data and the sample object data to obtain a target inference model;
[0212] Determine the candidate plots corresponding to the plot where the sample object is located based on the target inference model;
[0213] Among them, the target inference model is an inference model constructed by selecting at least one model from multiple fusion models through semi-supervised automatic learning training.
[0214] In an embodiment, when the first determination module 101 is used to execute automatically training the model based on the preset plot portrait data and the sample object data to obtain a target inference model, it specifically is used for:
[0215] Automatically perform feature selection on the preset plot portrait data based on the sample object data to obtain the plot sample features of each plot;
[0216] Automatically train the model based on the training data determined by the correspondence between the plot where the sample object is located and the plot sample features to obtain a first inference model;
[0217] Infer the plot sample features of the unconfigured plots based on the first inference model, and merge the inference result data with a confidence level higher than the preset threshold with the training data to obtain the merged training data; the plot sample features of the unconfigured plots include plot features that do not have a corresponding relationship with the plot where the sample object is located.
[0218] Use the merged training data for automated training to obtain a target inference model.
[0219] Among them, the automated training includes automatically adjusting model parameters using classification model evaluation metrics.
[0220] In one embodiment, when the first determination module 101 is used to perform automated training of the model based on the training data determined by the corresponding relationship between the plot where the sample object is located and the plot sample features to obtain the first inference model, it is specifically used for:
[0221] Determine the initial training data based on the corresponding relationship between the plot where the sample object is located and the plot sample features.
[0222] Perform sampling processing on the initial training data to obtain the processed training data.
[0223] Perform automated training of the model based on the processed training data to obtain the first inference model.
[0224] In one embodiment, when the first determination module 101 is used to perform automated training using the merged training data to obtain a target inference model, it is specifically used for:
[0225] Perform automated training using the merged training data to obtain a second inference model.
[0226] Adjust the merged training data based on the received custom feature information and custom weight coefficients.
[0227] Perform automated training using the adjusted training data to obtain a target inference model.
[0228] In one embodiment, when the first determination module 101 is used to perform adjusting the merged training data based on the received custom feature information and custom weight coefficients, it is specifically used for:
[0229] Extract at least one training feature information and its training weight coefficients used for training the second inference model and display them on the user interface of the client.
[0230] Adjust the training feature information and its training weight coefficients based on the received custom feature information and custom weight coefficients to obtain the adjusted training data.
[0231] In one embodiment, the second determination module 102 is configured to perform a similarity search on the candidate plots based on the preset plot portrait data to determine a target plot corresponding to the plot where the sample object is located among the candidate plots, including:
[0232] Obtain the basic features constructed based on the preset plot portrait data;
[0233] Extract the candidate features of the candidate plot;
[0234] Perform a feature similarity search on the basic features and the candidate features to determine a target plot corresponding to the plot where the sample object is located.
[0235] In one embodiment, constructing basic features based on the preset plot portrait data includes:
[0236] Determine the feature values of the preset metrics based on the preset plot portrait data;
[0237] Perform downsampling processing on the feature values to obtain the feature values after dimensionality reduction;
[0238] Perform feature importance analysis based on the feature values after dimensionality reduction to obtain basic features.
[0239] In one embodiment, determining the feature values of the preset metrics based on the preset plot portrait data includes at least one of the following:
[0240] At a set scale, determine the feature value of the proximity entropy based on the number of types of sample objects within the proximity range of a preset radius of the plot's location;
[0241] At a set scale, determine the feature value of the competition degree based on the number of sample objects of the same type within the proximity range of a preset radius of the plot's location;
[0242] At a set scale, determine the feature value of the plot space correlation effect based on the attraction coefficient between the number of preset types of sample objects within the proximity range of a preset radius of the plot's location and the types of sample objects;
[0243] At a set scale, determine the feature value of the transfer quality based on the number of user visits within the proximity range of a preset radius of the plot's location and the transition probability between the types of sample objects.
[0244] In one embodiment, the second determination module 102 is configured to perform a feature similarity search on the basic features and the candidate features to determine a target plot corresponding to the plot where the sample object is located, including:
[0245] Perform a feature similarity search on the basic features and the candidate features to determine a first plot among the candidate plots;
[0246] Determine a target plot corresponding to the plot where the sample object is located in the first plot based on at least one of the number of transactions of points of interest (POIs) and custom selection information.
[0247] In an embodiment, before the first determination module 101 executes to obtain sample object data related to plot selection, it is further configured to: respond to a plot selection request operation of the client;
[0248] After the second determination module 102 executes to determine the target plot, it is further configured to: display plot selection information corresponding to the target plot on the client;
[0249] The display of the plot selection information corresponding to the target plot includes: displaying the location of the target plot on the map interface based on a preset marking form.
[0250] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device in each embodiment of the present application correspond to the steps in the method in each embodiment of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details are not described herein again.
[0251] In an embodiment of the present application, an electronic device is provided. The electronic device includes: a memory and a processor; at least one program stored in the memory and configured to, when executed by the processor, compared with the prior art, achieve: in the present application, when obtaining sample object data related to plot selection, first determine candidate plots corresponding to the plot where the sample object is located based on preset plot portrait data and the sample object data. At this time, the number of candidate plots obtained is relatively large, and the plot selection result is relatively rough; on this basis, the present application further performs a similarity search on the candidate plots based on the preset plot portrait data to further narrow down the plot selection result and improve the accuracy of plot selection, so as to determine a target plot corresponding to the plot where the sample object is located among the candidate plots; wherein, the plot portrait data includes portrait data corresponding to plots obtained by geographical space division. The implementation of the solution of the present application can effectively mine the existing plot portrait data and the sample object data provided by the user, which is beneficial to improving the rationality of plot selection and making the selected plot more in line with the user's needs.
[0252] In an alternative embodiment, an electronic device is provided, as Figure 11 shown Figure 11The illustrated electronic device 1100 includes: a processor 1101 and a memory 1103. Among them, the processor 1101 and the memory 1103 are connected, such as connected through a bus 1102. Optionally, the electronic device 1100 may further include a transceiver 1104, and the transceiver 1104 can be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 1104 is not limited to one, and the structure of the electronic device 1100 does not constitute a limitation to the embodiments of the present application.
[0253] The processor 1101 can be a CPU (Central Processing Unit, central processing unit), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present application. The processor 1101 can also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0254] The bus 1102 may include a path for transmitting information between the above components. The bus 1102 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 1102 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0255] The memory 1103 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0256] The memory 1103 is used to store the application program code (computer program) for executing the solution of this application, and is controlled and executed by the processor 1101. The processor 1101 is used to execute the application program code stored in the memory 1103 to implement the content shown in the foregoing method embodiments.
[0257] Among them, the electronic device includes but is not limited to: smart phones, tablet computers, laptop computers, smart speakers, smart watches, vehicle-mounted devices, etc.
[0258] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the plot selection method provided in the above various optional implementation manners.
[0259] The embodiments of the present application provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When it runs on a computer, the computer can execute the corresponding content in the foregoing method embodiments.
[0260] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this text, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0261] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for plot selection, characterized in that, it includes: Determining candidate plots corresponding to the plot where the sample object is located through a target inference model; Performing similarity retrieval on the candidate plots based on preset plot portrait data to determine a target plot corresponding to the plot where the sample object is located among the candidate plots; Wherein, the plot portrait data includes portrait data corresponding to plots obtained by geographical space division; the target inference model is obtained through the following operations: Based on sample object data related to plot selection, performing automated feature selection on preset plot portrait data to obtain plot sample features of each plot; Based on training data determined by the corresponding relationship between the plot where the sample object is located and the plot sample features, performing automated training of the model to obtain a first inference model; Inferring the plot sample features of plots without configuration based on the first inference model, and merging the inference result data with a confidence level higher than a preset threshold with the training data to obtain merged training data; the plot sample features of plots without configuration include plot features that do not have a corresponding relationship with the plot where the sample object is located; Performing automated training using the merged training data to obtain a target inference model.
2. The method according to claim 1, characterized in that, The target inference model is an inference model constructed by selecting at least one model from multiple fusion models through semi-supervised automatic learning training.
3. The method according to claim 1, characterized in that, Automated training includes automatically adjusting model parameters using classification model evaluation metrics.
4. The method according to claim 3, characterized in that, The performing automated training of the model based on the training data determined by the corresponding relationship between the plot where the sample object is located and the plot sample features to obtain a first inference model includes: Determining initial training data based on the corresponding relationship between the plot where the sample object is located and the plot sample features; Performing sampling processing on the initial training data to obtain processed training data; Performing automated training of the model based on the processed training data to obtain a first inference model.
5. The method according to claim 3, characterized in that, The performing automated training using the merged training data to obtain a target inference model includes: Performing automated training using the merged training data to obtain a second inference model; Adjusting the merged training data based on received custom feature information and custom weight coefficients; Performing automated training using the adjusted training data to obtain a target inference model.
6. The method according to claim 5, characterized in that, The adjusting the merged training data based on received custom feature information and custom weight coefficients includes: Extracting at least one training feature information and its training weight coefficient used for training the second inference model and displaying them on the user interface of the client; Adjusting the training feature information and its training weight coefficient based on the received custom feature information and custom weight coefficients to obtain adjusted training data.
7. The method according to claim 1, characterized in that, Performing similarity retrieval on the candidate plots based on pre-set plot portrait data to determine a target plot corresponding to the plot where the sample object is located, including: Obtaining basic features constructed based on pre-set plot portrait data; Extracting candidate features of the candidate plots; Performing feature similarity retrieval on the basic features and the candidate features to determine a target plot corresponding to the plot where the sample object is located.
8. The method according to claim 7, wherein, Constructing basic features based on pre-set plot portrait data includes: Determining the feature values of pre-set indicators based on pre-set plot portrait data; Performing downsampling processing on the feature values to obtain downsampled feature values; Performing feature importance analysis based on the downsampled feature values to obtain basic features.
9. The method according to claim 8, wherein, Determining the feature values of pre-set indicators based on pre-set plot portrait data includes at least one of the following: At a set scale, determining the feature value of neighborhood entropy based on the number of types of sample objects within a neighboring range of a pre-set radius around the location of the plot; At a set scale, determining the feature value of competitiveness based on the number of sample objects of the same type within a neighboring range of a pre-set radius around the location of the plot; At a set scale, determining the feature value of the spatial correlation effect of the plot based on the attraction coefficient between the number of set types of sample objects within a neighboring range of a pre-set radius around the location of the plot and the types of sample objects; At a set scale, determining the feature value of transfer quality based on the number of user visits within a neighboring range of a pre-set radius around the location of the plot and the transfer probability between the types of sample objects.
10. The method according to claim 7, wherein, Performing feature similarity retrieval on the basic features and the candidate features to determine a target plot corresponding to the plot where the sample object is located includes: Performing feature similarity retrieval on the basic features and the candidate features to determine a first plot among the candidate plots; Determining a target plot corresponding to the plot where the sample object is located in the first plot based on at least one of the number of transactions of points of interest (POIs) and custom selection information.
11. The method according to claim 1, wherein, Before determining the candidate plots, it further includes: Responding to a plot selection request operation of the client; After determining the target plot, it further includes: Displaying plot selection information corresponding to the target plot on the client; The displaying of the plot selection information corresponding to the target plot includes: displaying the location of the target plot on the map interface based on a pre-set marking form.
12. A plot selection device, wherein, it includes: A first determination module, configured to determine candidate plots corresponding to the plot where the sample object is located through a target inference model; A second determination module, configured to perform similarity retrieval on the candidate plots based on pre-set plot portrait data to determine a target plot corresponding to the plot where the sample object is located among the candidate plots; wherein, the plot portrait data includes portrait data corresponding to plots obtained by geographical space division; the target inference model is trained through the following operations: Automatically select features for the preset plot portrait data based on the sample object data related to plot selection to obtain the plot sample features of each plot; Based on the training data determined by the correspondence between the plots where the sample objects are located and the plot sample features, perform automatic training of the model to obtain the first inference model; Based on the first inference model, infer the plot sample features of the plots without configured plots, and merge the inference result data with a confidence level higher than the preset threshold with the training data to obtain the merged training data; the plot sample features of the plots without configured plots include plot features that do not have a corresponding relationship with the plots where the sample objects are located; Use the merged training data for automatic training to obtain the target inference model.
13. The device according to claim 12, wherein, the target inference model is an inference model constructed by selecting at least one model from multiple fusion models through semi-supervised automatic learning training.
14. The device according to claim 12, wherein, automatic training includes automatically adjusting model parameters using classification model evaluation metrics.
15. The device according to claim 14, wherein, when the first determination module is used to perform automatic training of the model based on the training data determined by the correspondence between the plots where the sample objects are located and the plot sample features to obtain the first inference model, it specifically is used for: Determine the initial training data based on the correspondence between the plots where the sample objects are located and the plot sample features; Perform sampling processing on the initial training data to obtain the processed training data; Based on the processed training data, perform automatic training of the model to obtain the first inference model.
16. The device according to claim 14, wherein, when the first determination module is used to perform automatic training using the merged training data to obtain the target inference model, it specifically is used for: Perform automatic training using the merged training data to obtain the second inference model; Based on the received custom feature information and custom weight coefficients, adjust the merged training data; Perform automatic training using the adjusted training data to obtain the target inference model.
17. The device according to claim 16, wherein, when the first determination module is used to perform adjusting the merged training data based on the received custom feature information and custom weight coefficients, it specifically is used for: Extract at least one training feature information and its training weight coefficients used for training the second inference model and display them on the user interface of the client; Based on the received custom feature information and custom weight coefficients, adjust the training feature information and its training weight coefficients to obtain the adjusted training data.
18. The device according to claim 12, wherein, when the second determination module is used to perform similarity retrieval on the candidate plots based on the preset plot portrait data to determine the target plot corresponding to the plot where the sample object is located, it specifically is used for: Obtain the basic features constructed based on the preset plot portrait data; Extract the candidate features of the candidate plot; Perform a feature similarity search on the basic features and the candidate features to determine the target plot corresponding to the plot where the sample object is located.
19. The apparatus according to claim 18, wherein, when the second determination module is used to execute the construction of the basic features based on the preset plot portrait data, it is specifically used for: Determine the feature values of the preset indicators based on the preset plot portrait data; Perform downsampling processing on the feature values to obtain the dimension-reduced feature values; Perform feature importance analysis based on the dimension-reduced feature values to obtain the basic features.
20. The apparatus according to claim 19, wherein, when the second determination module is used to execute the determination of the feature values of the preset indicators based on the preset plot portrait data, it is specifically used for at least one of the following: At a set scale, determine the feature value of the neighborhood entropy based on the number of types of sample objects within the neighboring range of a preset radius around the location of the plot; At a set scale, determine the feature value of the competition degree based on the number of sample objects of the same type within the neighboring range of a preset radius around the location of the plot; At a set scale, determine the feature value of the spatial correlation effect of the plot based on the number of preset types of sample objects within the neighboring range of a preset radius around the location of the plot and the attraction coefficient between the types of sample objects; At a set scale, determine the feature value of the transfer quality based on the number of user visits within the neighboring range of a preset radius around the location of the plot and the transfer probability between the types of sample objects.
21. The apparatus according to claim 18, wherein, when the second determination module is used to execute the feature similarity search on the basic features and the candidate features to determine the target plot corresponding to the plot where the sample object is located, it is specifically used for: Perform a feature similarity search on the basic features and the candidate features to determine the first plot among the candidate plots; Based on at least one of the number of transactions of the point of interest (POI) and the custom selection information, determine the target plot corresponding to the plot where the sample object is located among the first plots.
22. The apparatus according to claim 12, wherein, before the first determination module executes the acquisition of the sample object data related to the plot selection, it is further used for: responding to the plot selection request operation of the client; after the second determination module executes the determination of the target plot, it is further used for: displaying the plot selection information corresponding to the target plot on the client; The display of the plot selection information corresponding to the target plot includes: displaying the location of the target plot on the map interface based on a preset marking form.
23. An electronic device, wherein, the electronic device includes: One or more processors; A memory; One or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured for: executing the method according to any one of claims 1 to 11.
24. A computer-readable storage medium, wherein, The computer storage medium is used to store computer instructions, which, when run on a computer, enable the computer to execute the method described in any one of claims 1 to 11 above.
25. A computer program product, comprising a computer program or instructions, wherein, when the computer program or instructions are executed by a processor, the steps of the method described in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Putting area determination method and device, model training method and storage medium
CN111222916A
Rapid store site selection method and device based on similarity extension, and storage medium
CN112308603A