Network site selection method and device

By combining the decision tree model with the reliability coefficient, bank branch addresses are automatically selected, solving the problems of poor accuracy, slow speed and high cost in existing branch site selection technologies, and achieving efficient branch site selection.

CN115965410BActive Publication Date: 2025-09-23INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310024931.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-09-23
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

In the existing technology, the accuracy of bank branch site selection is poor, the speed is slow and the cost is high, resulting in low efficiency of branch site selection.

Method used

By using the current feature information of multiple candidate network addresses and the input feature attributes of the trained decision tree model, classification is performed through the decision tree model to determine the type of candidate network points, and the target network point type is selected based on the reliability coefficient, reducing the dependence on consulting and investigation work and realizing automated site selection.

Benefits of technology

It improves the accuracy and speed of network site selection, reduces costs, reduces the time and financial investment in manual analysis and consulting investigations, and improves site selection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965410B_ABST
    Figure CN115965410B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for network site selection, particularly relating to the field of artificial intelligence. The method comprises: obtaining, based on current feature information corresponding to multiple candidate network site addresses and input feature attributes corresponding to multiple trained decision tree models, multiple current input feature information corresponding to the input feature attributes, and obtaining, based on the current input feature information and the corresponding trained decision tree models, multiple candidate network site types; obtaining, based on the model accuracy of the trained decision tree models corresponding to the candidate network site types, a reliability coefficient corresponding to the candidate network site types; determining, based on the reliability coefficient, a target network site type corresponding to the candidate network site address from the candidate network site types, and determining a final network site address from the multiple candidate network site addresses based on the target network site type. The present invention can improve the accuracy and speed of network site selection and reduce the cost of network site selection, thereby improving the efficiency of network site selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network site selection technology, in particular to the field of artificial intelligence, and more particularly to a network site selection method and device. Background Art

[0002] The address of a bank branch is closely related to its revenue and service quality. Therefore, in order to increase the corresponding revenue of the bank branch and better serve as many customers as possible, thereby improving the experience of the majority of customers, it is necessary to reasonably select a location for the bank branch.

[0003] In the prior art, branch site selection primarily relies on staff conducting tedious consultations and investigations, and analyzing the relevant information gathered during these consultations and investigations to determine the bank branch's location. However, since these consultations and investigations are time-consuming and expensive, and the analysis process relies on staff experience and is manual and time-consuming, the overall accuracy of branch site selection is low. This excessive time consumption also slows down the overall branch site selection process, and the high cost of branch site selection leads to high costs.

[0004] In summary, the existing technology has the problems of poor accuracy, slow speed and high cost of network site selection, which is not conducive to improving the efficiency of network site selection. Summary of the Invention

[0005] One object of the present invention is to provide a network site selection method to address the existing problems of poor accuracy, slow speed, and high cost in network site selection, which hinders the efficiency of network site selection. Another object of the present invention is to provide a network site selection device. Another object of the present invention is to provide a computer device. Yet another object of the present invention is to provide a readable medium.

[0006] In order to achieve the above objectives, one aspect of the present invention discloses a method for selecting a network site, the method comprising:

[0007] Based on current feature information corresponding to multiple candidate network point addresses and input feature attributes corresponding to multiple trained decision tree models, multiple current input feature information corresponding to the input feature attributes is obtained, and based on the current input feature information and the corresponding trained decision tree models, multiple candidate network point types are obtained;

[0008] Obtaining a reliability coefficient corresponding to the candidate network point type based on a model accuracy of a trained decision tree model corresponding to the candidate network point type;

[0009] Based on the reliability coefficient, a target network point type corresponding to the candidate network point address is determined from the candidate network point types, and based on the target network point type, a final network point address is determined from a plurality of candidate network point addresses.

[0010] Optionally, further including:

[0011] Before obtaining a plurality of current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the plurality of candidate network point addresses and the input feature attributes corresponding to the plurality of trained decision tree models,

[0012] Determining, based on a plurality of initial feature information preset in a plurality of historical network point feature information and the historical network point types corresponding to the initial feature information, the historical network point types corresponding to other historical network point feature information other than the initial feature information, wherein the historical network point types corresponding to the plurality of initial feature information are different from each other;

[0013] Based on the historical network point feature information, the corresponding historical network point type, and a plurality of preset input feature attributes corresponding to a plurality of preset untrained decision tree models, a plurality of samples to be divided corresponding to the untrained decision tree models are obtained, and based on a preset sample ratio, a plurality of training samples and a plurality of test samples among the plurality of samples to be divided are determined;

[0014] The untrained decision tree model is trained using the corresponding training samples to obtain a corresponding trained decision tree model, and the trained decision tree model is tested using the corresponding test samples to obtain a corresponding model accuracy.

[0015] Optionally, further including:

[0016] Before determining the historical network point types corresponding to other historical network point feature information other than the initial feature information based on the multiple initial feature information preset in the multiple historical network point feature information and the historical network point types corresponding to the initial feature information,

[0017] Data cleaning, data extraction and data standardization are performed on the initial historical feature information of multiple historical outlets to obtain historical outlet feature information corresponding to the historical outlets.

[0018] Optionally, further including:

[0019] Before determining the historical network point types corresponding to other historical network point feature information other than the initial feature information based on the multiple initial feature information preset in the multiple historical network point feature information and the historical network point types corresponding to the initial feature information,

[0020] Selecting a plurality of auxiliary feature information from a plurality of historical network point feature information, and determining a first Euclidean distance between each of the auxiliary feature information and a plurality of other historical network point feature information except the auxiliary feature information;

[0021] Based on the first Euclidean distance, other historical network point feature information other than the auxiliary feature information that is closest to the corresponding auxiliary feature information is determined as initial feature information corresponding to the auxiliary feature information.

[0022] Optionally, the determining, based on a plurality of initial feature information preset in the plurality of historical network point feature information and the historical network point types corresponding to the initial feature information, the historical network point types corresponding to other historical network point feature information except the initial feature information includes:

[0023] Using the initial feature information as cluster center feature information, and using other historical network point feature information except the initial feature information as feature information to be classified;

[0024] Determining a second Euclidean distance between each piece of feature information to be classified and each piece of cluster center feature information, and determining, based on the second Euclidean distance, the cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest cluster center feature information;

[0025] Based on the plurality of feature information to be classified having the same feature information of the nearest cluster center, a plurality of corresponding initial target clusters are obtained, and the historical network point type of the corresponding nearest cluster center feature information is used as the cluster type corresponding to the initial target cluster;

[0026] Repeat the clustering iteration step until there is a third Euclidean distance less than a preset distance threshold, wherein the clustering iteration step includes: based on the initial target cluster, obtaining corresponding intermediate cluster center feature information, and using the cluster type of the initial target cluster as the intermediate type of the corresponding intermediate cluster center feature information; using all the historical network feature information as feature information to be classified; determining the third Euclidean distance between each feature information to be classified and the intermediate cluster center feature information, and based on the third Euclidean distance, determining the intermediate cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest intermediate cluster center feature information; obtaining intermediate target clusters based on multiple feature information to be classified that are identical to the corresponding nearest intermediate cluster center feature information, and using the intermediate type of the corresponding nearest intermediate cluster center feature information as the cluster type of the intermediate target cluster; using the intermediate target cluster as the initial target cluster;

[0027] The cluster types of the plurality of intermediate target clusters are used as the historical network point types corresponding to other historical network point feature information corresponding to the intermediate target cluster except the initial feature information.

[0028] Optionally, obtaining corresponding intermediate cluster center feature information based on the initial target clustering includes:

[0029] Based on all the feature information to be classified included in the initial target cluster, obtaining mean feature information corresponding to the initial target cluster;

[0030] The mean feature information is used as the intermediate cluster center feature information.

[0031] Optionally, obtaining a plurality of to-be-classified samples corresponding to the untrained decision tree model based on the historical network point feature information, the corresponding historical network point type, and a plurality of preset input feature attributes corresponding to the preset plurality of untrained decision tree models includes:

[0032] Based on the feature parameters corresponding to the input feature attributes in the historical network point feature information, forming an input sample of the historical network point feature information corresponding to the untrained decision tree model, and using the corresponding historical network point type as the corresponding output sample;

[0033] Based on the input samples and the corresponding output samples, the corresponding samples to be divided are formed.

[0034] Optionally, further including:

[0035] Before obtaining a plurality of current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the plurality of candidate network point addresses and the input feature attributes corresponding to the plurality of trained decision tree models,

[0036] Data cleaning, data extraction and data standardization are performed on the initial current feature information corresponding to the multiple candidate network point addresses to obtain the current feature information corresponding to the candidate network point addresses.

[0037] Optionally, the obtaining of multiple current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the multiple candidate network point addresses and the input feature attributes corresponding to the multiple trained decision tree models includes:

[0038] Based on the feature parameters in the current feature information corresponding to the input feature attributes, current input feature information of the trained decision tree model corresponding to the current feature information is formed.

[0039] Optionally, obtaining the reliability coefficient corresponding to the candidate network point type based on the model accuracy of the trained decision tree model corresponding to the candidate network point type includes:

[0040] The model accuracy rates of multiple trained decision tree models corresponding to the candidate network point type are superimposed to obtain the reliability coefficient corresponding to the candidate network point type.

[0041] Optionally, determining the target network point type corresponding to the to-be-selected network point address from the candidate network point types based on the reliability coefficient includes:

[0042] The candidate network point type corresponding to the largest reliability coefficient is determined as the target network point type.

[0043] In order to achieve the above objectives, another aspect of the present invention discloses a network site selection device, the device comprising:

[0044] a type prediction module, configured to obtain, based on current feature information corresponding to multiple candidate network point addresses and input feature attributes corresponding to multiple trained decision tree models, multiple current input feature information corresponding to the input feature attributes, and obtain multiple candidate network point types based on the current input feature information and the corresponding trained decision tree models;

[0045] a reliability determination module, configured to obtain a reliability coefficient corresponding to the candidate network point type based on a model accuracy of a trained decision tree model corresponding to the candidate network point type;

[0046] The network point selection module is used to determine the target network point type corresponding to the candidate network point address from the candidate network point types based on the reliability coefficient, and determine the final network point address from multiple candidate network point addresses based on the target network point type.

[0047] The present invention also discloses a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method described above is implemented when the processor executes the program.

[0048] The present invention also discloses a computer-readable medium on which a computer program is stored. When the program is executed by a processor, the method described above is implemented.

[0049] The network site selection method and device provided by the present invention obtains multiple current input feature information corresponding to the input feature attributes based on the current feature information corresponding to multiple network site addresses to be selected and the input feature attributes corresponding to multiple trained decision tree models. It can form current input feature information corresponding to the input format supported by each of the multiple trained decision tree models based on the information of the actual network site addresses to be selected, thereby greatly improving the compatibility between the model input and the trained decision tree model, and thus improving the calculation accuracy and speed of the subsequent trained decision tree model, thereby greatly improving the accuracy and speed of the overall network site selection, and also indirectly leading to the training of the decision tree model. There is no need to use the information of all attributes in the corresponding samples for training and testing (because when the model is used later, the current input feature information corresponding to the input attribute type will be formed according to the input attribute type supported by the model, so there is no need to overemphasize the attributes of the model input sample during the training process). Instead, according to actual needs, only part of the attribute information can be extracted for different decision tree models (the attribute parts corresponding to different decision trees can be different or the same) for training and testing. On the basis of indirectly improving the flexibility of training, it also indirectly improves the training speed and reduces the training cost, thereby further indirectly improving the speed of overall network site selection and reducing the cost of overall network site selection. By obtaining multiple alternative network site types based on the current input feature information and the corresponding trained decision tree model, it is possible to quickly obtain multiple alternative network site types by relying on the advantage that the decision tree model is suitable for classification, and the alternative network site types output by different decision tree models all have relatively high accuracy, thereby greatly improving the accuracy and speed of overall network site selection.

[0050] By obtaining the reliability coefficient corresponding to the candidate network point type based on the model accuracy of the trained decision tree model corresponding to the candidate network point type, it is possible to improve the traditional voting mechanism of the random forest. Not only is the number of decision tree models corresponding to each candidate network point type output as the basis for determining the reliability of the output result (candidate network point type), but the calculation accuracy of the decision tree model corresponding to each candidate network point type is also closely considered. Since the calculation accuracy of the model is also closely related to the reliability of its output, it is possible to greatly improve the accuracy of the reliability coefficient for determining the reliability of each candidate network point type, thereby greatly improving the subsequent determination of the network type of each candidate network point address. The accuracy of the model is improved, thereby greatly improving the accuracy of the overall network site selection. Moreover, it also indirectly leads to the fact that when training the decision tree model, there is no need to re-select samples and re-train the decision trees that have been trained but have too low calculation accuracy (because in subsequent use, it is not important whether the accuracy of a single decision tree meets the standard, but only the calculation accuracy of its test is used as the parameter for subsequent prediction of the target network type), thereby indirectly greatly reducing the time consumption of training the model and thus greatly improving the speed of training the model, and also indirectly greatly reducing the capital investment in the sampling, training and testing processes and thus greatly reducing the cost of training the model, and further indirectly improving the speed of overall network site selection and reducing the cost of overall network site selection.

[0051] By determining the target network point type corresponding to the candidate network point address from the candidate network point types based on the reliability coefficient, and determining the final network point address from multiple candidate network point addresses based on the target network point type, it is possible to fully use the reliability coefficient that accurately measures the reliability of each alternative result as a basis to accurately select a better alternative result as the prediction result corresponding to the candidate network point address, that is, the target network point type, thereby improving the accuracy of determining the target network point type and being conducive to accurately selecting a better candidate network point address as the final network point address based on the target network point type of each candidate network point address, thereby greatly improving the accuracy of the overall network site selection.

[0052] The network site selection method and device provided by the present invention can perform relevant site selection predictions based on readily available and collected characteristic information of actual candidate network sites, significantly reducing reliance on tedious consultation and investigation work. This naturally reduces the additional time and financial investment incurred by such work, contributing to faster and lower costs of network site selection. Furthermore, the network site selection method of the present invention can be automated, significantly reducing reliance on manual analysis processes, thereby significantly improving the accuracy and speed of overall network site selection and reducing labor-related costs.

[0053] In summary, the network site selection method and device provided by the present invention can improve the accuracy and speed of network site selection and reduce the cost of network site selection, thereby improving the efficiency of network site selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0055] Figure 1 A schematic diagram showing a flow chart of a method for selecting a network site according to an embodiment of the present invention is shown;

[0056] Figure 2 A schematic diagram showing an optional step of model preparation work according to an embodiment of the present invention is shown;

[0057] Figure 3 A schematic diagram showing an optional step of obtaining multiple pieces of current input feature information corresponding to input feature attributes according to an embodiment of the present invention is shown;

[0058] Figure 4 A schematic diagram showing an optional step of obtaining a reliability coefficient according to an embodiment of the present invention is shown;

[0059] Figure 5 A schematic diagram showing an optional step of determining a target network point type according to an embodiment of the present invention;

[0060] Figure 6 A schematic diagram of a module of a network site selection device according to an embodiment of the present invention is shown;

[0061] Figure 7 A schematic diagram showing the structure of a computer device suitable for implementing an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0063] The terms “first,” “second,” etc. used herein do not particularly refer to an order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.

[0064] The words “include,” “including,” “have,” “contain,” etc. used in this document are open-ended terms, meaning including but not limited to.

[0065] As used herein, "and / or" includes any and all combinations of the items mentioned.

[0066] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of the present invention are in compliance with the relevant provisions of national laws and regulations.

[0067] It should be noted that the network site selection method and device disclosed in this application can be used in the field of network site selection technology, and can also be used in any field other than the field of network site selection technology. The application field of the network site selection method and device disclosed in this application is not limited.

[0068] The embodiment of the present invention discloses a method for selecting a network site. Figure 1 As shown, the method specifically includes the following steps:

[0069] S101: Based on current feature information corresponding to multiple candidate network point addresses and input feature attributes corresponding to multiple trained decision tree models, multiple current input feature information corresponding to the input feature attributes are obtained, and based on the current input feature information and the corresponding trained decision tree models, multiple candidate network point types are obtained.

[0070] S102: Obtaining a reliability coefficient corresponding to the candidate network point type based on the model accuracy of the trained decision tree model corresponding to the candidate network point type.

[0071] S103: Based on the reliability coefficient, determine a target network point type corresponding to the candidate network point address from the candidate network point types, and based on the target network point type, determine a final network point address from a plurality of candidate network point addresses.

[0072] For example, a candidate network address corresponds to a piece of current feature information, a trained decision tree model corresponds to a set of input feature attributes (a set of input feature attributes includes multiple input feature attributes), a piece of current feature information corresponds to multiple pieces of current input feature information, a piece of current input feature information is used to input into a corresponding trained decision tree model, a trained decision tree model outputs an alternative network type for a candidate network address, a candidate network address corresponds to multiple alternative network types, a trained decision tree model corresponds to a model accuracy, an alternative network type for a candidate network address corresponds to a reliability coefficient, and a candidate network type corresponds to a target network type. It should be noted that the relevant corresponding relationships can be determined by those skilled in the art based on actual circumstances, and the above description is only for example and does not constitute a limitation.

[0073] For example, the specific form of the feature information in the embodiments of the present invention may be, but is not limited to, a vector or matrix including feature values ​​corresponding to multiple feature attributes (the data form is usually digital, with one feature value corresponding to each attribute). The feature values ​​may be obtained by, but are not limited to, mapping the content of the corresponding feature attributes (generally based on preset mapping relationship information) or numerical processing. For example, for a feature attribute "the number of residents within a radius of x kilometers", different corresponding feature values ​​may be obtained when the population is in different preset population ranges. For another example, for a feature attribute "whether it is a commercial area", different corresponding feature values ​​may be obtained when the corresponding content is "yes commercial area" and "not commercial area". It should be noted that the specific form and source of the feature information can be determined by those skilled in the art based on actual circumstances. The above description is only an example and does not constitute a limitation.

[0074] For example, the characteristic attributes of the characteristic information in the embodiment of the present invention include but are not limited to the number of residents within a radius of x kilometers (which can be obtained by but is not limited to investigating statistics or querying relevant population density and then multiplying by the corresponding area), the number of companies / units within a radius of x kilometers (which can be obtained by but is not limited to investigating statistics or querying relevant information), the average house price / rental price within a radius of x kilometers, the number of working people within a radius of x kilometers (which can be obtained by but is not limited to investigating statistics or querying information), the number of parking spaces within a radius of x kilometers (which can be obtained by but is not limited to investigating statistics or querying information), the number of hotels within a radius of x kilometers (which can be obtained by but is not limited to investigating statistics or querying information), and the number of hotels within a radius of x kilometers (which can be obtained by but is not limited to investigating statistics or querying information). The number of shopping centers within a radius of x kilometers (which can be obtained by but not limited to survey statistics or information query, etc.), the number of restaurants within a radius of x kilometers (which can be obtained by but not limited to survey statistics or information query, etc.), whether it is a commercial area, the number of office buildings within a radius of x kilometers (which can be obtained by but not limited to survey statistics or information query, etc.), the number of similar outlets within a radius of x kilometers (which can be directly queried), the average willingness of people to participate in business within a radius of x kilometers (which can be obtained by surveying and inquiring them under the conditions of authorization of relevant personnel), and address location information (such as longitude and latitude, etc.). Among them, the set of input feature attributes corresponding to the decision tree model is a part of all feature attributes of the feature information. It should be noted that the specific type of feature attributes can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation.

[0075] For example, the collection of multiple trained decision tree models may also be called, but not limited to, a random forest.

[0076] For example, the branch types in the embodiments of the present invention may include, but are not limited to, "good business performance," "relatively good business performance," "average business performance," "poor business performance," and "low business performance." The nature of the branch type may include, but is not limited to, types that represent the branch's business performance or superiority. It should be noted that the specific content and nature of the branch type can be determined by those skilled in the art based on actual circumstances, and the above description is merely illustrative and does not constitute a limitation.

[0077] Exemplarily, the method of obtaining multiple candidate network point types based on the current input feature information and the corresponding trained decision tree model may be, but is not limited to, inputting the corresponding current input feature information into the corresponding trained decision tree model for computational processing to obtain the candidate network point types output by the corresponding trained decision tree model (since for each candidate network point address, one trained decision tree model outputs one candidate network point type, and there are multiple trained decision tree models, there can be multiple candidate network point types). It should be noted that the specific implementation method of obtaining multiple candidate network point types based on the current input feature information and the corresponding trained decision tree model can be determined by those skilled in the art based on actual circumstances, and the above description is merely an example and does not constitute a limitation thereto.

[0078] For example, determining the final network address from multiple candidate network addresses based on the target network type may include, but is not limited to, selecting one of the candidate network addresses corresponding to the expected target network type as the final network address. For example, selecting one of the candidate network addresses corresponding to the target network type of "good business performance" as the final network address. It should be noted that the specific implementation method for determining the final network address from multiple candidate network addresses based on the target network type can be determined by those skilled in the art based on actual circumstances, and the above description is merely an example and does not constitute a limitation.

[0079] For example, among the multiple candidate outlet types corresponding to a candidate outlet address, some may be the same. For example, some trained decision tree models all output the outlet type result of "good business performance".

[0080] Among them, determining the Euclidean distance between two pieces of feature information on the basis of clear feature information (which may be in the form of but not limited to feature vectors or feature matrices, specifically including the eigenvalues ​​corresponding to each attribute) is a conventional technical means in this field. Therefore, the various Euclidean distances involved in the embodiments of the present invention, their solution and calculation processes and principles will not be repeated.

[0081] The network site selection method and device provided by the present invention obtains multiple current input feature information corresponding to the input feature attributes based on the current feature information corresponding to multiple network site addresses to be selected and the input feature attributes corresponding to multiple trained decision tree models. It can form current input feature information corresponding to the input format supported by each of the multiple trained decision tree models based on the information of the actual network site addresses to be selected, thereby greatly improving the compatibility between the model input and the trained decision tree model, and thus improving the calculation accuracy and speed of the subsequent trained decision tree model, thereby greatly improving the accuracy and speed of the overall network site selection, and also indirectly leading to the training of the decision tree model. There is no need to use the information of all attributes in the corresponding samples for training and testing (because when the model is used later, the current input feature information corresponding to the input attribute type will be formed according to the input attribute type supported by the model, so there is no need to overemphasize the attributes of the model input sample during the training process). Instead, according to actual needs, only part of the attribute information can be extracted for different decision tree models (the attribute parts corresponding to different decision trees can be different or the same) for training and testing. On the basis of indirectly improving the flexibility of training, it also indirectly improves the training speed and reduces the training cost, thereby further indirectly improving the speed of overall network site selection and reducing the cost of overall network site selection. By obtaining multiple alternative network site types based on the current input feature information and the corresponding trained decision tree model, it is possible to quickly obtain multiple alternative network site types by relying on the advantage that the decision tree model is suitable for classification, and the alternative network site types output by different decision tree models all have relatively high accuracy, thereby greatly improving the accuracy and speed of overall network site selection.

[0082] By obtaining the reliability coefficient corresponding to the candidate network point type based on the model accuracy of the trained decision tree model corresponding to the candidate network point type, it is possible to improve the traditional voting mechanism of the random forest. Not only is the number of decision tree models corresponding to each candidate network point type output as the basis for determining the reliability of the output result (candidate network point type), but the calculation accuracy of the decision tree model corresponding to each candidate network point type is also closely considered. Since the calculation accuracy of the model is also closely related to the reliability of its output, it is possible to greatly improve the accuracy of the reliability coefficient for determining the reliability of each candidate network point type, thereby greatly improving the subsequent determination of the network type of each candidate network point address. The accuracy of the model is improved, thereby greatly improving the accuracy of the overall network site selection. Moreover, it also indirectly leads to the fact that when training the decision tree model, there is no need to re-select samples and re-train the decision trees that have been trained but have too low calculation accuracy (because in subsequent use, it is not important whether the accuracy of a single decision tree meets the standard, but only the calculation accuracy of its test is used as the parameter for subsequent prediction of the target network type), thereby indirectly greatly reducing the time consumption of training the model and thus greatly improving the speed of training the model, and also indirectly greatly reducing the capital investment in the sampling, training and testing processes and thus greatly reducing the cost of training the model, and further indirectly improving the speed of overall network site selection and reducing the cost of overall network site selection.

[0083] By determining the target network point type corresponding to the candidate network point address from the candidate network point types based on the reliability coefficient, and determining the final network point address from multiple candidate network point addresses based on the target network point type, it is possible to fully use the reliability coefficient that accurately measures the reliability of each alternative result as a basis to accurately select a better alternative result as the prediction result corresponding to the candidate network point address, that is, the target network point type, thereby improving the accuracy of determining the target network point type and being conducive to accurately selecting a better candidate network point address as the final network point address based on the target network point type of each candidate network point address, thereby greatly improving the accuracy of the overall network site selection.

[0084] The network site selection method and device provided by the present invention can perform relevant site selection predictions based on readily available and collected characteristic information of actual candidate network sites, significantly reducing reliance on tedious consultation and investigation work. This naturally reduces the additional time and financial investment incurred by such work, contributing to faster and lower costs of network site selection. Furthermore, the network site selection method of the present invention can be automated, significantly reducing reliance on manual analysis processes, thereby significantly improving the accuracy and speed of overall network site selection and reducing labor-related costs.

[0085] In summary, the network site selection method and device provided by the present invention can improve the accuracy and speed of network site selection and reduce the cost of network site selection, thereby improving the efficiency of network site selection.

[0086] In an optional embodiment, if Figure 2 As shown, further comprising the following steps:

[0087] S201: Before obtaining multiple current input feature information corresponding to the input feature attributes based on current feature information corresponding to multiple candidate network address addresses and input feature attributes corresponding to multiple trained decision tree models, determine the historical network type corresponding to other historical network feature information other than the initial feature information based on multiple initial feature information preset in the multiple historical network feature information and the historical network type corresponding to the initial feature information, wherein the historical network types corresponding to the multiple initial feature information are different from each other.

[0088] S202: Based on the historical network feature information, the corresponding historical network type and the preset multiple input feature attributes corresponding to the preset multiple untrained decision tree models, a plurality of samples to be divided corresponding to the untrained decision tree models are obtained, and based on a preset sample ratio, a plurality of training samples and test samples among the multiple samples to be divided are determined.

[0089] S203: Using the corresponding training samples to train the untrained decision tree model to obtain a corresponding trained decision tree model, and using the corresponding test samples to test the trained decision tree model to obtain a corresponding model accuracy.

[0090] Exemplarily, one historical network point feature information corresponds to one historical network point or historical network point address, one initial feature information corresponds to one historical network point type (wherein, one initial feature information is one of multiple historical network point feature information, and some of the multiple historical network point feature information corresponds to the initial feature information), and other historical network point feature information other than the initial feature information also corresponds to one historical network point type. In this way, each historical network point feature information can correspond to a historical network point type later (some historical network point feature information may correspond to the same historical network point type, but the initial feature information corresponds to different historical network point types), and an untrained decision tree model corresponds to multiple samples to be divided (preferably, each untrained decision tree model corresponds to samples to be divided, covering all historical network point feature information that has been determined for training and testing (one sample to be divided corresponds to one historical network point feature information), but the sets of input feature attributes taken are different). It should be noted that the relevant corresponding relationship can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation to this.

[0091] Exemplarily, the number of the initial feature information may be consistent with the number of preset values ​​of the outlet type. For example, when the value range of the outlet type includes "good business situation", "relatively good business situation", "average business situation", "poor business situation" and "bad business situation", the number of initial feature information is 5, and the historical outlet types corresponding to the five initial feature information are "good business situation", "relatively good business situation", "average business situation", "poor business situation" and "bad business situation". The determination of the historical outlet type of the initial feature information can be determined by relevant staff after research and analysis. It should be noted that the number and nature of the initial feature information can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation.

[0092] Exemplarily, the sample ratio may be, but is not limited to, 7:3 (70% for training samples and 30% for test samples). For example, if 300 pieces of historical site feature information are selected for training the test model, then for each untrained decision tree model, there are 300 corresponding samples to be divided. Then, for each untrained decision tree model, 210 samples are used as training samples for training, and the remaining 90 samples are used as test samples for testing (the feature attribute sets of samples corresponding to different untrained decision tree models are generally different. For example, the input feature attributes corresponding to an untrained decision tree model are A, B, C, D..., and the input feature attributes corresponding to another untrained decision tree model may be B, F, G, H...). It should be noted that the specific value of the sample ratio can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation.

[0093] For example, using training samples to train a model and using test samples to test the trained model to obtain the accuracy is a conventional technical means in this field and will not be described in detail here.

[0094] Through the above steps, when preparing decision tree models, only partial attribute information (the attribute portions corresponding to different decision trees can be different or the same) can be extracted for training and testing. This not only improves training flexibility, but also increases training speed and reduces training costs, thereby indirectly improving the speed and reducing the cost of overall network site selection. Furthermore, when training decision tree models, there is no need to reselect samples and retrain decision trees that have already been trained but have low computational accuracy. This indirectly significantly reduces the time spent on model training and thus significantly improves the speed of model training. It also indirectly significantly reduces the capital investment in the sampling, training, and testing processes and thus significantly reduces the cost of model training, further indirectly improving the speed and reducing the cost of overall network site selection. By accurately and quickly providing the trained decision tree models and model accuracy required for the network site selection process, the above steps provide sufficient and excellent preparation for the main network site selection process, facilitate the smooth progress of the network site selection process, and thus help improve the efficiency of network site selection.

[0095] In an optional embodiment, further comprising:

[0096] Before determining the historical network point types corresponding to other historical network point feature information other than the initial feature information based on the multiple initial feature information preset in the multiple historical network point feature information and the historical network point types corresponding to the initial feature information,

[0097] Data cleaning, data extraction and data standardization are performed on the initial historical feature information of multiple historical outlets to obtain historical outlet feature information corresponding to the historical outlets.

[0098] For example, the data cleaning method may include, but is not limited to, replacing or deleting abnormal data in the feature information using cleaning methods such as spline interpolation and linear regression. The data extraction method may include, but is not limited to, performing a dimensionality reduction operation on attribute variables with strong correlation. For example, if the feature information contains two attribute variables, namely, longitude and latitude and the cell to which the attribute belongs, since the embodiment of the present invention does not focus on the cell to which the attribute belongs, and the two attributes, longitude and latitude and the cell to which the attribute belongs, are essentially the same (both represent geographic location features) and are highly correlated, the variable of the cell to which the attribute belongs is deleted (so that subsequent related attributes and element types do not include the cell to which the attribute belongs) to complete the dimensionality reduction operation. The data normalization method may include, but is not limited to, converting the relevant data into various appropriate formats. For example, if the eigenvalue of a certain attribute is 876, and the format of the eigenvalue in the feature information requires normalization, the eigenvalue 876 is normalized to 0.876. It should be noted that the specific implementation methods of data cleaning, data extraction, and data normalization can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation.

[0099] For example, one historical network point corresponds to one initial historical feature information, and one initial historical feature information corresponds to one historical network point feature information. It should be noted that the relevant correspondence can be determined by those skilled in the art according to actual conditions, and the above description is only an example and does not constitute a limitation.

[0100] Exemplarily, the acquisition and processing of the initial historical feature information may be implemented through, but not limited to, a corresponding big data platform, such as, but not limited to, a Hadoop big data platform.

[0101] Through the above steps, the initial characteristic information of the historical network points can be corrected and simplified, so that the related calculations and processing operations in the subsequent steps are more concise and accurate, and the efficiency of the overall network point site selection is effectively improved.

[0102] In an optional embodiment, further comprising:

[0103] Before determining the historical network point types corresponding to other historical network point feature information other than the initial feature information based on the multiple initial feature information preset in the multiple historical network point feature information and the historical network point types corresponding to the initial feature information,

[0104] Selecting a plurality of auxiliary feature information from a plurality of historical network point feature information, and determining a first Euclidean distance between each of the auxiliary feature information and a plurality of other historical network point feature information except the auxiliary feature information;

[0105] Based on the first Euclidean distance, other historical network point feature information other than the auxiliary feature information that is closest to the corresponding auxiliary feature information is determined as initial feature information corresponding to the auxiliary feature information.

[0106] Exemplarily, the selection of multiple auxiliary feature information from multiple historical outlet feature information may include, but is not limited to, random selection of multiple auxiliary feature information from multiple historical outlet feature information. The number of auxiliary feature information may be consistent with the number of preset possible values ​​for the outlet type. For example, if the value range of the outlet type includes "good business conditions," "relatively good business conditions," "average business conditions," "poor business conditions," and "terrible business conditions," the number of auxiliary feature information is 5. It should be noted that the specific implementation of selecting multiple auxiliary feature information from multiple historical outlet feature information can be determined by those skilled in the art based on actual circumstances, and the above description is merely an example and does not constitute a limitation.

[0107] For example, the Euclidean distance in the embodiment of the present invention may also be called but not limited to Euclidean distance.

[0108] Exemplarily, the determining of the first Euclidean distance between each of the auxiliary feature information and a plurality of other historical network point feature information other than the auxiliary feature information may be, but is not limited to, determining the first Euclidean distance between each of the auxiliary feature information and other auxiliary feature information other than the auxiliary feature information and other historical network point feature information other than the auxiliary feature information. For example, the current auxiliary feature information is A, the other auxiliary feature information is B and C, and the other historical network point feature information other than the auxiliary feature information is D, E, and F. Then, the first Euclidean distance includes the Euclidean distance between the auxiliary feature information A and the auxiliary feature information B, the Euclidean distance between the auxiliary feature information A and the auxiliary feature information C, the Euclidean distance between the auxiliary feature information A and the historical network point feature information D, and the Euclidean distance between the auxiliary feature information A and the historical network point feature information The Euclidean distance between the auxiliary feature information and the historical network point feature information is determined, and the Euclidean distance between the auxiliary feature information and the historical network point feature information is determined. The same is true for the auxiliary feature information B and the auxiliary feature information C. Alternatively, the first Euclidean distance between each auxiliary feature information and other historical network point feature information other than the auxiliary feature information is determined. For example, if the current auxiliary feature information is A, the other auxiliary feature information is B and C, and the other historical network point feature information other than the auxiliary feature information is D, E, and F, then the first Euclidean distance includes the Euclidean distance between the auxiliary feature information A and the historical network point feature information D, the Euclidean distance between the auxiliary feature information A and the historical network point feature information E, and the Euclidean distance between the auxiliary feature information A and the historical network point feature information F. The same is true for the auxiliary feature information B and the auxiliary feature information C. It should be noted that the specific implementation method for determining the first Euclidean distance between each of the auxiliary feature information and multiple historical network point feature information other than the auxiliary feature information can be determined by those skilled in the art according to actual circumstances. The above description is only an example and does not constitute a limitation.

[0109] Exemplarily, based on the first Euclidean distance, determining the historical network point feature information other than the auxiliary feature information that is closest to the corresponding auxiliary feature information as the initial feature information corresponding to the auxiliary feature information may include, but is not limited to, determining the historical network point feature information other than the auxiliary feature information with the smallest first Euclidean distance (which may include, but is not limited to, other auxiliary feature information other than the current auxiliary feature information and other historical network point feature information that is not auxiliary feature information, or other historical network point feature information that only includes non-auxiliary feature information) as the initial feature information corresponding to the auxiliary feature information. For example, if the current auxiliary feature information is A, the other auxiliary feature information is B and C, and the other historical network point feature information that is not auxiliary feature information is D, E, and F, and the first Euclidean distance between auxiliary feature information A and other historical network point feature information D is the smallest, then historical network point feature information D is determined as the initial feature information corresponding to auxiliary feature information A. One piece of auxiliary feature information corresponds to one piece of initial feature information. It should be noted that the specific implementation method of determining other historical network feature information other than the auxiliary feature information that is closest to the corresponding auxiliary feature information based on the first Euclidean distance as the initial feature information corresponding to the auxiliary feature information can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation to this.

[0110] Through the above steps, the determined initial feature information is less likely to have extreme values ​​(because the auxiliary feature information has been used to participate in a distance calculation and the feature information is selected as the initial feature information based on this, the relevant feature values ​​of the determined initial feature information can be kept as close as possible to the feature boundary values ​​of the overall feature information, and the selected initial feature information is indeed historical network feature information, which is beneficial to improving the data authenticity of subsequent calculations). This can reduce the extremes of the historical types corresponding to other historical network feature information determined based on the initial feature information and thus reduce errors, so that the quality of the initial feature information used as the initial clustering center for determining the historical type is relatively high, thereby improving the convergence speed, accuracy and adaptability of the subsequent determination of the historical types corresponding to other historical network feature information, and thus helping to improve the speed and accuracy of the overall network site selection.

[0111] In an optional embodiment, determining the historical network point type corresponding to other historical network point feature information other than the initial feature information based on the multiple initial feature information preset in the multiple historical network point feature information and the historical network point type corresponding to the initial feature information includes:

[0112] Using the initial feature information as cluster center feature information, and using other historical network point feature information except the initial feature information as feature information to be classified;

[0113] Determining a second Euclidean distance between each piece of feature information to be classified and each piece of cluster center feature information, and determining, based on the second Euclidean distance, the cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest cluster center feature information;

[0114] Based on the plurality of feature information to be classified having the same feature information of the nearest cluster center, a plurality of corresponding initial target clusters are obtained, and the historical network point type of the corresponding nearest cluster center feature information is used as the cluster type corresponding to the initial target cluster;

[0115] Repeat the clustering iteration step until there is a third Euclidean distance less than a preset distance threshold, wherein the clustering iteration step includes: based on the initial target cluster, obtaining corresponding intermediate cluster center feature information, and using the cluster type of the initial target cluster as the intermediate type of the corresponding intermediate cluster center feature information; using all the historical network feature information as feature information to be classified; determining the third Euclidean distance between each feature information to be classified and the intermediate cluster center feature information, and based on the third Euclidean distance, determining the intermediate cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest intermediate cluster center feature information; obtaining intermediate target clusters based on multiple feature information to be classified that are identical to the corresponding nearest intermediate cluster center feature information, and using the intermediate type of the corresponding nearest intermediate cluster center feature information as the cluster type of the intermediate target cluster; using the intermediate target cluster as the initial target cluster;

[0116] The cluster types of the plurality of intermediate target clusters are used as the historical network point types corresponding to other historical network point feature information corresponding to the intermediate target cluster except the initial feature information.

[0117] Exemplarily, the initial feature information is used as cluster center feature information, and other historical network feature information other than the initial feature information is used as feature information to be classified. Examples include the following:

[0118] The initial feature information includes feature information A, feature information B, and feature information C. The other historical network feature information except the initial feature information includes feature information D, feature information E, feature information F, feature information G, feature information H, and feature information I. The feature information to be classified includes feature information D, feature information E, feature information F, feature information G, feature information H, and feature information I. The cluster center feature information includes feature information A, feature information B, and feature information C.

[0119] It should be noted that the specific implementation method of using the initial feature information as the cluster center feature information and using other historical network feature information other than the initial feature information as the feature information to be classified can be determined by those skilled in the art based on actual conditions. The above description is only an example and does not constitute a limitation to this.

[0120] Exemplarily, the second Euclidean distance between each piece of feature information to be classified and each piece of cluster center feature information is determined as follows:

[0121] There is cluster center feature information A, cluster center feature information B and cluster center feature information C, feature information to be classified D, feature information to be classified E, feature information to be classified F, feature information to be classified G, feature information to be classified H and feature information to be classified I, then the second Euclidean distance includes but is not limited to the Euclidean distance between feature information to be classified D and cluster center feature information A, the Euclidean distance between feature information to be classified D and cluster center feature information B, the Euclidean distance between feature information to be classified D and cluster center feature information C, the Euclidean distance between feature information to be classified E and cluster center feature information A, the Euclidean distance between feature information to be classified E and cluster center feature information B, the Euclidean distance between feature information to be classified E and cluster center feature information C, the Euclidean distance between feature information to be classified F and cluster center feature information A, the Euclidean distance between feature information to be classified F and cluster center feature information The Euclidean distance between the class feature information F and the cluster center feature information B, the Euclidean distance between the feature information to be classified F and the cluster center feature information C, the Euclidean distance between the feature information to be classified G and the cluster center feature information A, the Euclidean distance between the feature information to be classified G and the cluster center feature information B, the Euclidean distance between the feature information to be classified G and the cluster center feature information C, the Euclidean distance between the feature information to be classified H and the cluster center feature information A, the Euclidean distance between the feature information to be classified H and the cluster center feature information B, the Euclidean distance between the feature information to be classified H and the cluster center feature information C, the Euclidean distance between the feature information to be classified I and the cluster center feature information A, the Euclidean distance between the feature information to be classified I and the cluster center feature information B, and the Euclidean distance between the feature information to be classified I and the cluster center feature information C.

[0122] It should be noted that the specific implementation method for determining the second Euclidean distance between each of the feature information to be classified and each of the cluster center feature information can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation thereto.

[0123] Exemplarily, the cluster center feature information closest to the corresponding feature information to be classified based on the second Euclidean distance is determined as the corresponding nearest cluster center feature information. This can be, but is not limited to, using the cluster center feature information with the smallest corresponding second Euclidean distance as the nearest cluster center feature information of the corresponding feature information to be classified. For example, for feature information to be classified I, among cluster center feature information A, cluster center feature information B, and cluster center feature information C, the second Euclidean distance between cluster center feature information A and feature information to be classified I is the smallest, then cluster center feature information A is used as the nearest cluster center feature information corresponding to feature information to be classified I. It should be noted that the specific implementation method for determining the cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest cluster center feature information based on the second Euclidean distance can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation thereto.

[0124] Exemplarily, the obtaining of the corresponding multiple initial target clusters based on the multiple feature information to be classified that are the same as the corresponding nearest cluster center feature information can be, but is not limited to, clustering each nearest cluster center feature information and the multiple feature information to be classified corresponding to the nearest cluster center feature information to obtain the initial target cluster corresponding to the nearest cluster center feature information, wherein one nearest cluster center feature information corresponds to one initial target cluster. For example, the nearest cluster center feature information corresponding to feature information to be classified D and feature information to be classified E is cluster center feature information A, the nearest cluster center feature information corresponding to feature information to be classified F and feature information to be classified G is cluster center feature information B, and the nearest cluster center feature information corresponding to feature information to be classified H and feature information to be classified I is cluster center feature information C. Then, feature information to be classified D, feature information to be classified E, and cluster center feature information A are clustered to obtain an initial target cluster A, feature information to be classified F, feature information to be classified G, and cluster center feature information B are clustered to obtain another initial target cluster B, and feature information to be classified H, feature information to be classified I, and cluster center feature information C are clustered to obtain another initial target cluster C. It should be noted that the specific implementation method for obtaining multiple corresponding initial target clusters based on multiple pieces of feature information to be classified that have the same corresponding nearest cluster center feature information can be determined by those skilled in the art according to actual circumstances. The above description is only an example and does not constitute a limitation thereto.

[0125] Exemplarily, the historical network point type of the corresponding recent cluster center feature information is used as the cluster type corresponding to the initial target cluster, as follows:

[0126] The cluster center feature information A corresponds to an initial target cluster A, and the cluster center feature information A is the initial feature information in the historical network feature information. Its historical network type is known and its historical network type is "good business conditions", so the cluster type of the initial target cluster A is set to "good business conditions".

[0127] It should be noted that the specific implementation method of using the historical network type of the corresponding recent cluster center feature information as the cluster type corresponding to the initial target cluster can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation to this.

[0128] For example, the distance threshold may be determined by those skilled in the art according to actual conditions, and the embodiments of the present invention do not limit this.

[0129] Exemplarily, the number of intermediate cluster center feature information is the same as the number of initial target clusters, the number of possible values ​​of the node type, the number of initial feature information and the number of cluster center feature information, and the intermediate types corresponding to different cluster center feature information are also different.

[0130] Exemplarily, the determining of the third Euclidean distance between each piece of feature information to be classified and the intermediate cluster center feature information may be, but is not limited to, determining the third Euclidean distance between each piece of feature information to be classified and each piece of intermediate cluster center feature information. For example, if there are five different pieces of feature information to be classified and three different pieces of intermediate cluster center feature information, then the number of third Euclidean distances may be 3×5=15. It should be noted that the specific implementation of determining the third Euclidean distance between each piece of feature information to be classified and the intermediate cluster center feature information may be determined by those skilled in the art based on actual circumstances, and the above description is for illustrative purposes only and does not constitute a limitation thereto.

[0131] Exemplarily, the determination of the intermediate cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest intermediate cluster center feature information based on the third Euclidean distance may be, but is not limited to, taking the intermediate cluster center feature information with the smallest corresponding third Euclidean distance as the nearest intermediate cluster center feature information corresponding to the feature information to be classified. For example, for a certain feature information to be classified A, among the intermediate cluster center feature information A, the intermediate cluster center feature information B, and the intermediate cluster center feature information C, the intermediate cluster center feature information A has the smallest third Euclidean distance with the feature information to be classified A, then the intermediate cluster center feature information A is taken as the nearest intermediate cluster center feature information corresponding to the feature information to be classified A. It should be noted that the specific implementation method for determining the intermediate cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest intermediate cluster center feature information based on the third Euclidean distance can be determined by those skilled in the art according to actual circumstances. The above description is only an example and does not constitute a limitation thereto.

[0132] Exemplarily, the intermediate target clusters are obtained based on the multiple feature information to be classified that have the same nearest intermediate cluster center feature information. This can be, but is not limited to, clustering the multiple feature information to be classified corresponding to each nearest intermediate cluster center feature information to obtain the intermediate target clusters, wherein one nearest intermediate cluster center feature information corresponds to one intermediate target cluster. For example, if feature information to be classified D and feature information to be classified E correspond to nearest intermediate cluster center feature information A, feature information to be classified F and feature information to be classified G correspond to nearest intermediate cluster center feature information B, and feature information to be classified H and feature information to be classified I correspond to nearest intermediate cluster center feature information C, then feature information to be classified D and feature information to be classified E are clustered to obtain an intermediate target cluster A, feature information to be classified F and feature information to be classified G are clustered to obtain another intermediate target cluster B, and feature information to be classified H and feature information to be classified I are clustered to obtain another intermediate target cluster C. It should be noted that the specific implementation method for obtaining the intermediate target clustering based on multiple pieces of feature information to be classified that have the same feature information corresponding to the nearest intermediate cluster center can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation to this.

[0133] Exemplarily, the intermediate type of the corresponding nearest intermediate cluster center feature information is used as the cluster type of the intermediate target cluster, as follows:

[0134] The most recent intermediate cluster center feature information A corresponds to an intermediate target cluster A, and the historical network type of the most recent intermediate cluster center feature information A is "good business conditions", then the cluster type of the intermediate target cluster A is set to "good business conditions".

[0135] It should be noted that the specific implementation method of using the intermediate type of the corresponding nearest intermediate cluster center feature information as the cluster type of the intermediate target cluster can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation to this.

[0136] Exemplarily, the clustering types of the multiple intermediate target clusters are used as the historical network point types corresponding to other historical network point feature information corresponding to the intermediate target cluster in addition to the initial feature information. It can be, but is not limited to, the step of repeatedly performing clustering iterations until there are multiple clustering types of the intermediate target clusters after a third Euclidean distance less than a preset distance threshold (one intermediate target cluster corresponds to one cluster type) as the historical network point types corresponding to other historical network point feature information included in the intermediate target cluster in addition to the initial feature information. For example, when the clustering iteration step is no longer executed, the intermediate target clusters obtained include intermediate target cluster A (cluster type is "good business situation"), intermediate target cluster B (cluster type is "relatively good business situation"), intermediate target cluster C (cluster type is "medium business situation"), intermediate target cluster D (cluster type is "poor business situation") and intermediate target cluster E (cluster type is "poor business situation"), and the historical network point feature information corresponding to the intermediate target cluster A in addition to the initial feature information includes feature information A, the historical network point feature information corresponding to the intermediate target cluster B in addition to the initial feature information includes feature information B and feature information C, and the historical network point feature information corresponding to the intermediate target cluster C in addition to the initial feature information includes feature information B and feature information C. The historical network point characteristic information other than the initial characteristic information includes characteristic information D and characteristic information E, the historical network point characteristic information other than the initial characteristic information corresponding to the intermediate target cluster D includes characteristic information F and characteristic information G, and the historical network point characteristic information other than the initial characteristic information corresponding to the intermediate target cluster E includes characteristic information H. Then, the historical network point type corresponding to characteristic information A is "good business conditions", the historical network point type corresponding to characteristic information B and characteristic information C is "relatively good business conditions", the historical network point type corresponding to characteristic information D and characteristic information E is "medium business conditions", the historical network point type corresponding to characteristic information F and characteristic information G is "poor business conditions", and the historical network point type corresponding to characteristic information H is "poor business conditions". It should be noted that the specific implementation method of using the cluster types of multiple intermediate target clusters as the historical network point types corresponding to the historical network point characteristic information other than the initial characteristic information corresponding to the intermediate target cluster can be determined by those skilled in the art according to actual circumstances. The above description is only an example and does not constitute a limitation.

[0137] Through the above steps, it is possible to perform clustering iterations based on the relevant principles of the K-means clustering algorithm and ultimately obtain accurate and stable clusters of historical network point feature information that can significantly differentiate and characterize different network point types, thereby automatically determining the network point types corresponding to the historical network point feature information (and its corresponding historical network points) other than the initial feature information. In this way, it is possible to achieve the goal of not manually labeling the types of many historical network point feature information, but to automatically determine them, thereby significantly improving the speed and accuracy of overall sample formation and significantly reducing the relevant labor costs, thereby significantly improving the speed and accuracy of overall network point location selection and significantly reducing the relevant costs. Moreover, in the above steps, the iteration is not stopped when the cluster center does not change at all, but when there is a corresponding center distance less than a threshold (which can also represent a relatively stable clustering characteristic). At this time, the determined cluster can still meet the relevant accuracy requirements, and on this basis, the number of iterations can be reduced, thereby further reducing the corresponding time, achieving an improvement on the basis of the K-means clustering algorithm, thereby further indirectly improving the speed of sample formation and thus indirectly further improving the speed of overall network point location selection.

[0138] In an optional embodiment, the obtaining corresponding intermediate cluster center feature information based on the initial target clustering includes:

[0139] Based on all the feature information to be classified included in the initial target cluster, obtaining mean feature information corresponding to the initial target cluster;

[0140] The mean feature information is used as the intermediate cluster center feature information.

[0141] Exemplarily, the method of obtaining the mean feature information corresponding to the initial target cluster based on all the feature information to be classified included in the initial target cluster may be, but is not limited to, superimposing all the feature information to be classified included in the initial target cluster to obtain the sum feature information, and then dividing the sum feature information by the number of all the feature information to be classified included in the initial target cluster to obtain the mean feature information. Since the feature information is generally in the form of a vector or matrix including feature values ​​of multiple attributes, it can participate in related addition, subtraction, multiplication and division operations. It should be noted that the specific implementation method of obtaining the mean feature information corresponding to the initial target cluster based on all the feature information to be classified included in the initial target cluster can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation thereto.

[0142] Through the above steps, the intermediate cluster center feature information can be made consistent with the overall average of the initial target cluster, so that the intermediate cluster center feature information is indeed the center or center of gravity of the corresponding initial target cluster, which improves the accuracy of obtaining the intermediate cluster center feature information.

[0143] The accuracy of the characteristic information of the cluster center is improved, thereby improving the accuracy of the relevant iteration, and further improving the accuracy of determining the historical network point type corresponding to the historical network point characteristic information based on the relevant 5 iteration steps and thus improving the accuracy of the overall network point location selection.

[0144] In an optional embodiment, the obtaining of multiple to-be-classified samples corresponding to the untrained decision tree model based on the historical network point feature information, the corresponding historical network point type, and multiple preset input feature attributes corresponding to the preset multiple untrained decision tree models includes: 0. forming input samples corresponding to the untrained decision tree model of the historical network point feature information based on feature parameters corresponding to the input feature attributes in the historical network point feature information, and using the corresponding historical network point type as the corresponding output sample;

[0145] Based on the input samples and the corresponding output samples, the corresponding samples to be divided are formed.

[0146] Exemplarily, the feature parameters corresponding to the input feature attributes in the historical network point feature information are used to form the input samples of the historical network point feature information corresponding to the untrained decision tree model, and the corresponding historical network point type is used as the corresponding output sample. This can be, but is not limited to, integrating the feature parameters of multiple feature attributes that are the same as the input feature attributes in the historical network point feature information to form the input samples of the historical network point feature information corresponding to the untrained decision tree model, and using the historical network point feature information as the corresponding output sample.

[0147] The corresponding historical outlet type is used as the output sample corresponding to the input sample, wherein one historical outlet feature information 0 and an untrained decision tree model together correspond to one input sample. For example, a historical outlet feature information has 30 feature attributes, each feature attribute corresponds to a feature parameter (for example, the average house price attribute corresponds to a specific average house price feature value, etc.), and an untrained decision tree model supports the use of parameters of 15 of the above 30 feature attributes as input for operation processing, then the parameters of the 15 feature attributes in the historical outlet feature information are used as the input.

[0148] The characteristic parameters corresponding to the characteristics of the historical network point feature information are integrated to obtain an input sample corresponding to the untrained decision tree model, and the historical network point type of the historical network point feature information is used as the corresponding output sample. It should be noted that the specific implementation method for forming the input sample corresponding to the untrained decision tree model of the historical network point feature information based on the characteristic parameters corresponding to the input characteristic attributes in the historical network point feature information, and using the corresponding historical network point type as the corresponding output sample can be determined by those skilled in the art based on actual circumstances. The above description is only an example and does not constitute a limitation thereto.

[0149] Exemplarily, the multiple input feature attributes corresponding to each untrained decision tree model may be, but are not limited to, randomly selected from all feature attributes corresponding to the historical network point feature information or obtained based on the properties of the untrained decision tree model, and the sets of input feature attributes corresponding to different untrained decision tree models are generally different. For example, if the historical network point feature information has 30 feature attributes, then for a certain untrained decision tree model A, 15 of the 30 feature attributes in the historical network point feature information can be randomly selected as the multiple input feature attributes of the untrained decision tree model A. It should be noted that the specific source of the input feature attributes can be determined by those skilled in the art based on actual circumstances, and the above description is only for example and does not constitute a limitation thereto.

[0150] Exemplarily, a sample to be divided corresponds to an input sample and an output sample corresponding to the input sample. Specifically, for example, there are 300 historical network feature information for training and testing, then for each untrained decision tree model, there are 300 samples to be divided (corresponding to 300 input samples and 300 output samples). For another example, the untrained decision tree model A corresponds to 300 samples to be divided, and the untrained decision tree model B corresponds to 300 samples to be divided. However, the feature attribute set of the 300 samples to be divided corresponding to the untrained decision tree model A and the feature attribute set of the 300 samples to be divided corresponding to the untrained decision tree model B are likely to be different. It should be noted that the relevant corresponding relationship can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation to this.

[0151] Through the above steps, the corresponding samples to be divided can more comprehensively cover multiple historical network feature information that characterize the relevant characteristics of actual network points, and can also make the samples to be divided accurately compatible with the corresponding untrained decision tree model in attribute format, thereby improving the speed and accuracy of subsequent model training and testing, and further improving the speed and accuracy of the overall network site selection.

[0152] In an optional embodiment, further comprising:

[0153] Before obtaining a plurality of current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the plurality of candidate network point addresses and the input feature attributes corresponding to the plurality of trained decision tree models,

[0154] Data cleaning, data extraction and data standardization are performed on the initial current feature information corresponding to the multiple candidate network point addresses to obtain the current feature information corresponding to the candidate network point addresses.

[0155] For example, the data cleaning method may include, but is not limited to, replacing or deleting abnormal data in the feature information using cleaning methods such as spline interpolation and linear regression. The data extraction method may include, but is not limited to, performing a dimensionality reduction operation on attribute variables with strong correlation. For example, if the feature information contains two attribute variables, namely, longitude and latitude and the cell to which the attribute belongs, since the embodiment of the present invention does not focus on the cell to which the attribute belongs, and the two attributes, longitude and latitude and the cell to which the attribute belongs, are essentially the same (both represent geographic location features) and are highly correlated, the variable of the cell to which the attribute belongs is deleted (so that subsequent related attributes and element types do not include the cell to which the attribute belongs) to complete the dimensionality reduction operation. The data normalization method may include, but is not limited to, converting the relevant data into various appropriate formats. For example, if the eigenvalue of a certain attribute is 876, and the format of the eigenvalue in the feature information requires normalization, the eigenvalue 876 is normalized to 0.876. It should be noted that the specific implementation methods of data cleaning, data extraction, and data normalization can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation.

[0156] For example, one candidate network point address corresponds to one initial current feature information, and one initial current feature information corresponds to one current feature information. It should be noted that the relevant correspondence can be determined by those skilled in the art according to actual conditions, and the above description is only an example and does not constitute a limitation thereto.

[0157] Exemplarily, the acquisition and processing of the initial current feature information may be implemented through, but not limited to, a corresponding big data platform, for example, through, but not limited to, a Hadoop big data platform.

[0158] Through the above steps, the initial current feature information of the address of the selected network point can be corrected and simplified, so that the related calculation and processing operations in the subsequent steps are more concise and accurate, and the efficiency of the overall network point site selection is effectively improved.

[0159] In an optional embodiment, if Figure 3As shown, the method of obtaining multiple current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the multiple candidate network addresses and the input feature attributes corresponding to the multiple trained decision tree models includes the following steps:

[0160] S301: Based on the feature parameters corresponding to the input feature attributes in the current feature information, forming current input feature information corresponding to the trained decision tree model.

[0161] Exemplarily, the current input feature information of the trained decision tree model corresponding to the current feature information is formed based on the feature parameters corresponding to the input feature attributes in the current feature information. It can be, but is not limited to, integrating feature parameters of multiple feature attributes that are the same as the input feature attributes in the current feature information to form the current input feature information of the trained decision tree model corresponding to the current feature information, wherein one current feature information and one trained decision tree model jointly correspond to one current input feature information. For example, a current feature information has 30 feature attributes, each of which corresponds to a feature parameter, and a trained decision tree model supports the use of parameters of 15 of the 30 feature attributes as input for computational processing. The feature parameters corresponding to the 15 feature attributes in the current feature information are integrated to obtain a current input feature information corresponding to the trained decision tree model. The attribute sets of the current input feature information corresponding to different trained decision tree models may be different. For example, the current input feature information corresponding to a trained decision tree model is (A, B, C, D, ..., O), and the current input feature information corresponding to another trained decision tree model is (B, C, H, G, ..., P). It should be noted that the specific implementation method for forming the current input feature information corresponding to the trained decision tree model based on the feature parameters corresponding to the input feature attributes in the current feature information can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation.

[0162] Through the above steps, the current input feature information input into the corresponding model can be made more compatible with the input format supported by the model, thereby improving the calculation speed and accuracy of the trained model, thereby making the speed and accuracy of outputting the addresses of the selected outlets higher, and thereby improving the speed and accuracy of the overall outlet site selection.

[0163] In an optional embodiment, if Figure 4 As shown, obtaining the reliability coefficient corresponding to the candidate network point type based on the model accuracy of the trained decision tree model corresponding to the candidate network point type includes the following steps:

[0164] S401: superimposing the model accuracy rates of multiple trained decision tree models corresponding to the candidate network point type to obtain a reliability coefficient corresponding to the candidate network point type.

[0165] For example, since for a candidate outlet address, some different trained decision tree models may output the same candidate outlet type, a certain candidate outlet can correspond to multiple trained decision tree models. For example, for a candidate outlet address, the candidate outlet types output by trained decision tree model A and trained decision tree model B are both "good business conditions", while the candidate outlet types output by trained decision tree model C and trained decision tree model D are both "relatively good business conditions", and the candidate outlet type output by trained decision tree model E is "medium business conditions". Then, the trained decision tree models corresponding to the candidate outlet type "good business conditions" are trained decision tree model A and trained decision tree model B, the trained decision tree models corresponding to the candidate outlet type "relatively good business conditions" are trained decision tree model C and trained decision tree model D, and the trained decision tree model corresponding to the candidate outlet type "medium business conditions" is trained decision tree model E. It should be noted that the specific correspondence between the candidate network types and the trained decision tree models and the corresponding origins can be determined by those skilled in the art according to actual conditions. The above description is only an example and does not constitute a limitation thereto.

[0166] For example, one candidate network point type corresponds to one reliability coefficient, and the step S401 has the following example:

[0167] The trained decision tree models corresponding to the alternative outlet type "good business conditions" are trained decision tree model A and trained decision tree model B, and the model accuracy of trained decision tree model A is 0.7, while the model accuracy of trained decision tree model B is 0.65. Therefore, the reliability coefficient corresponding to the alternative outlet type "good business conditions" is 0.7+0.65=1.35.

[0168] It should be noted that the specific implementation of step S401 can be determined by those skilled in the art according to actual conditions, and the above description is only an example and does not constitute a limitation thereto.

[0169] Through the above steps, the reliability coefficient can comprehensively represent the number of votes of the decision tree that outputs the corresponding classification and the accuracy of the decision tree. Therefore, the reliability coefficient can closely correspond to the degree of possibility that the address of the candidate network point does meet the corresponding alternative network point, thereby improving the accuracy of the reliability coefficient, thereby improving the accuracy of determining the target network point type of the address of the candidate network point, and further improving the accuracy of the overall network site selection.

[0170] In an optional embodiment, if Figure 5 As shown, the step of determining the target network point type corresponding to the to-be-selected network point address from the candidate network point types based on the reliability coefficient includes the following steps:

[0171] S501: Determine the candidate network point type corresponding to the largest reliability coefficient as the target network point type.

[0172] For example, the step S501 may be as follows:

[0173] For a certain candidate outlet address, the reliability coefficient of the corresponding alternative outlet type "good business conditions" is 81.25, the reliability coefficient of the corresponding alternative outlet type "relatively good business conditions" is 50.15, the reliability coefficient of the corresponding alternative outlet type "medium business conditions" is 60.75, the reliability coefficient of the corresponding alternative outlet type "poor business conditions" is 30.85, and the reliability coefficient of the corresponding alternative outlet type "poor business conditions" is 15.50. The alternative outlet type "good business conditions" corresponding to the largest reliability coefficient 81.25 is determined as the target outlet type of the candidate outlet address.

[0174] It should be noted that the specific implementation of step S501 can be determined by those skilled in the art according to actual conditions, and the above description is only an example and does not constitute a limitation thereto.

[0175] Through the above steps, the type among multiple candidate network point types that is most likely to be consistent with the actual address of the network point to be selected can be determined as the target network point type, thereby improving the accuracy of determining the target network point type and further improving the accuracy of the overall network point site selection.

[0176] Based on the same principle, the embodiment of the present invention discloses a network site selection device 600, such as Figure 6 As shown, the network site selection device 600 includes:

[0177] The type prediction module 601 is configured to obtain, based on current feature information corresponding to multiple candidate network point addresses and input feature attributes corresponding to multiple trained decision tree models, multiple current input feature information corresponding to the input feature attributes, and obtain multiple candidate network point types based on the current input feature information and the corresponding trained decision tree models;

[0178] A reliability determination module 602 is configured to obtain a reliability coefficient corresponding to the candidate network point type based on the model accuracy of the trained decision tree model corresponding to the candidate network point type;

[0179] The network point selection module 603 is configured to determine a target network point type corresponding to the candidate network point address from the candidate network point types based on the reliability coefficient, and determine a final network point address from multiple candidate network point addresses based on the target network point type.

[0180] In an optional embodiment, the system further includes a preparation module for:

[0181] Before obtaining a plurality of current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the plurality of candidate network point addresses and the input feature attributes corresponding to the plurality of trained decision tree models,

[0182] Determining, based on a plurality of initial feature information preset in a plurality of historical network point feature information and the historical network point types corresponding to the initial feature information, the historical network point types corresponding to other historical network point feature information other than the initial feature information, wherein the historical network point types corresponding to the plurality of initial feature information are different from each other;

[0183] Based on the historical network point feature information, the corresponding historical network point type, and a plurality of preset input feature attributes corresponding to a plurality of preset untrained decision tree models, a plurality of samples to be divided corresponding to the untrained decision tree models are obtained, and based on a preset sample ratio, a plurality of training samples and a plurality of test samples among the plurality of samples to be divided are determined;

[0184] The untrained decision tree model is trained using the corresponding training samples to obtain a corresponding trained decision tree model, and the trained decision tree model is tested using the corresponding test samples to obtain a corresponding model accuracy.

[0185] In an optional embodiment, the system further includes a historical data preprocessing module for:

[0186] Before determining the historical network point types corresponding to other historical network point feature information other than the initial feature information based on the multiple initial feature information preset in the multiple historical network point feature information and the historical network point types corresponding to the initial feature information,

[0187] Data cleaning, data extraction and data standardization are performed on the initial historical feature information of multiple historical outlets to obtain historical outlet feature information corresponding to the historical outlets.

[0188] In an optional embodiment, the system further includes an initial feature information determination module, configured to:

[0189] Before determining the historical network point types corresponding to other historical network point feature information other than the initial feature information based on the multiple initial feature information preset in the multiple historical network point feature information and the historical network point types corresponding to the initial feature information,

[0190] Selecting a plurality of auxiliary feature information from a plurality of historical network point feature information, and determining a first Euclidean distance between each of the auxiliary feature information and a plurality of other historical network point feature information except the auxiliary feature information;

[0191] Based on the first Euclidean distance, other historical network point feature information other than the auxiliary feature information that is closest to the corresponding auxiliary feature information is determined as initial feature information corresponding to the auxiliary feature information.

[0192] In an optional embodiment, the preparation module is used to:

[0193] Using the initial feature information as cluster center feature information, and using other historical network point feature information except the initial feature information as feature information to be classified;

[0194] Determining a second Euclidean distance between each piece of feature information to be classified and each piece of cluster center feature information, and determining, based on the second Euclidean distance, the cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest cluster center feature information;

[0195] Based on the plurality of feature information to be classified having the same feature information of the nearest cluster center, a plurality of corresponding initial target clusters are obtained, and the historical network point type of the corresponding nearest cluster center feature information is used as the cluster type corresponding to the initial target cluster;

[0196] Repeat the clustering iteration step until there is a third Euclidean distance less than a preset distance threshold, wherein the clustering iteration step includes: based on the initial target cluster, obtaining corresponding intermediate cluster center feature information, and using the cluster type of the initial target cluster as the intermediate type of the corresponding intermediate cluster center feature information; using all the historical network feature information as feature information to be classified; determining the third Euclidean distance between each feature information to be classified and the intermediate cluster center feature information, and based on the third Euclidean distance, determining the intermediate cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest intermediate cluster center feature information; obtaining intermediate target clusters based on multiple feature information to be classified that are identical to the corresponding nearest intermediate cluster center feature information, and using the intermediate type of the corresponding nearest intermediate cluster center feature information as the cluster type of the intermediate target cluster; using the intermediate target cluster as the initial target cluster;

[0197] The cluster types of the plurality of intermediate target clusters are used as the historical network point types corresponding to other historical network point feature information corresponding to the intermediate target cluster except the initial feature information.

[0198] In an optional embodiment, the preparation module is used to:

[0199] Based on all the feature information to be classified included in the initial target cluster, obtaining mean feature information corresponding to the initial target cluster;

[0200] The mean feature information is used as the intermediate cluster center feature information.

[0201] In an optional embodiment, the preparation module is used to:

[0202] Based on the feature parameters corresponding to the input feature attributes in the historical network point feature information, forming an input sample of the historical network point feature information corresponding to the untrained decision tree model, and using the corresponding historical network point type as the corresponding output sample;

[0203] Based on the input samples and the corresponding output samples, the corresponding samples to be divided are formed.

[0204] In an optional embodiment, the system further includes a current data preprocessing module for:

[0205] Before obtaining a plurality of current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the plurality of candidate network point addresses and the input feature attributes corresponding to the plurality of trained decision tree models,

[0206] Data cleaning, data extraction and data standardization are performed on the initial current feature information corresponding to the multiple candidate network point addresses to obtain the current feature information corresponding to the candidate network point addresses.

[0207] In an optional embodiment, the type prediction module 601 is configured to:

[0208] Based on the feature parameters in the current feature information corresponding to the input feature attributes, current input feature information of the trained decision tree model corresponding to the current feature information is formed.

[0209] In an optional implementation, the reliability determination module 602 is configured to:

[0210] The model accuracy rates of multiple trained decision tree models corresponding to the candidate network point type are superimposed to obtain the reliability coefficient corresponding to the candidate network point type.

[0211] In an optional implementation, the network site selection module 603 is configured to:

[0212] The candidate network point type corresponding to the largest reliability coefficient is determined as the target network point type.

[0213] Since the principle of solving the problem by the network site selection device 600 is similar to that of the above method, the implementation of the network site selection device 600 can refer to the implementation of the above method, which will not be repeated here.

[0214] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer device. Specifically, the computer device may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0215] In a typical example, a computer device specifically includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method described above is implemented.

[0216] Reference below Figure 7 , which shows a schematic structural diagram of a computer device 700 suitable for implementing an embodiment of the present application.

[0217] like Figure 7 As shown, computer device 700 includes a central processing unit (CPU) 701, which can perform various appropriate tasks and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. Various programs and data required for the operation of system 700 are also stored in RAM 703. CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0218] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, and the like; an output section 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 708 including devices such as a hard disk; and a communication section 709 including a network interface card such as a LAN card or a modem. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. Removable media 711, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 710 as needed, so that computer programs read therefrom can be installed in the storage section 708 as needed.

[0219] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709 and / or installed from removable media 711.

[0220] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0221] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0222] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0223] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 The functions specified in one or more processes and / or block diagrams and 5 or more blocks.

[0224] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 process or processes and / or

[0225] or box Figure 1 The steps for the function specified in one or more boxes.

[0226] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, product, or apparatus. In the absence of further limitations, the elements defined by the phrase "comprises a..." do not preclude the presence of other identical elements in the process, method, product, or apparatus that includes the elements.

[0227] Those skilled in the art should understand that the embodiments of the present application may be provided as methods, systems, or computer program products.

[0228] Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0229] This application may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In distributed computing environments, program modules may be located in both local and remote computer storage media, including storage devices.

[0230] 5 The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments may be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.

[0231] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for selecting a network site, characterized in that: include: Based on current feature information corresponding to multiple candidate network point addresses and input feature attributes corresponding to multiple trained decision tree models, multiple current input feature information corresponding to the input feature attributes is obtained, and based on the current input feature information and the corresponding trained decision tree models, multiple candidate network point types are obtained; Obtaining a reliability coefficient corresponding to the candidate network point type based on a model accuracy of a trained decision tree model corresponding to the candidate network point type; Based on the reliability coefficient, determining a target network point type corresponding to the candidate network point address from the candidate network point types, and based on the target network point type, determining a final network point address from a plurality of candidate network point addresses; The method further comprises: Before obtaining a plurality of current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the plurality of candidate network point addresses and the input feature attributes corresponding to the plurality of trained decision tree models, Determining, based on a plurality of initial feature information preset in a plurality of historical network point feature information and the historical network point types corresponding to the initial feature information, the historical network point types corresponding to other historical network point feature information other than the initial feature information, wherein the historical network point types corresponding to the plurality of initial feature information are different from each other; Based on the historical network point feature information, the corresponding historical network point type, and a plurality of preset input feature attributes corresponding to a plurality of preset untrained decision tree models, a plurality of samples to be divided corresponding to the untrained decision tree models are obtained, and based on a preset sample ratio, a plurality of training samples and a plurality of test samples among the plurality of samples to be divided are determined; Using the corresponding training samples to train the untrained decision tree model to obtain a corresponding trained decision tree model, and using the corresponding test samples to test the trained decision tree model to obtain a corresponding model accuracy; The determining, based on a plurality of initial feature information preset in the plurality of historical network point feature information and the historical network point types corresponding to the initial feature information, of the historical network point feature information other than the initial feature information includes: Using the initial feature information as cluster center feature information, and using other historical network point feature information except the initial feature information as feature information to be classified; Determining a second Euclidean distance between each piece of feature information to be classified and each piece of cluster center feature information, and determining, based on the second Euclidean distance, the cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest cluster center feature information; Based on the plurality of feature information to be classified having the same feature information of the nearest cluster center, a plurality of corresponding initial target clusters are obtained, and the historical network point type of the corresponding nearest cluster center feature information is used as the cluster type corresponding to the initial target cluster; Repeat the clustering iteration step until there is a third Euclidean distance less than a preset distance threshold, wherein the clustering iteration step includes: based on the initial target cluster, obtaining corresponding intermediate cluster center feature information, and using the cluster type of the initial target cluster as the intermediate type of the corresponding intermediate cluster center feature information; using all the historical network feature information as feature information to be classified; determining the third Euclidean distance between each feature information to be classified and the intermediate cluster center feature information, and based on the third Euclidean distance, determining the intermediate cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest intermediate cluster center feature information; obtaining intermediate target clusters based on multiple feature information to be classified that are identical to the corresponding nearest intermediate cluster center feature information, and using the intermediate type of the corresponding nearest intermediate cluster center feature information as the cluster type of the intermediate target cluster; using the intermediate target cluster as the initial target cluster; The cluster types of the plurality of intermediate target clusters are used as the historical network point types corresponding to other historical network point feature information corresponding to the intermediate target cluster except the initial feature information.

2. The method according to claim 1, characterized in that Further including: Before determining the historical network point types corresponding to other historical network point feature information other than the initial feature information based on the multiple initial feature information preset in the multiple historical network point feature information and the historical network point types corresponding to the initial feature information, Data cleaning, data extraction and data standardization are performed on the initial historical feature information of multiple historical outlets to obtain historical outlet feature information corresponding to the historical outlets.

3. The method according to claim 1, characterized in that Further including: Before determining the historical network point types corresponding to other historical network point feature information other than the initial feature information based on the multiple initial feature information preset in the multiple historical network point feature information and the historical network point types corresponding to the initial feature information, Selecting a plurality of auxiliary feature information from a plurality of historical network point feature information, and determining a first Euclidean distance between each of the auxiliary feature information and a plurality of other historical network point feature information except the auxiliary feature information; Based on the first Euclidean distance, other historical network point feature information other than the auxiliary feature information that is closest to the corresponding auxiliary feature information is determined as initial feature information corresponding to the auxiliary feature information.

4. The method according to claim 1, wherein The obtaining of corresponding intermediate cluster center feature information based on the initial target clustering includes: Based on all the feature information to be classified included in the initial target cluster, obtaining mean feature information corresponding to the initial target cluster; The mean feature information is used as the intermediate cluster center feature information.

5. The method according to claim 1, wherein The step of obtaining a plurality of to-be-classified samples corresponding to the untrained decision tree models based on the historical network point feature information, the corresponding historical network point type, and a plurality of preset input feature attributes corresponding to the preset plurality of untrained decision tree models includes: Based on the feature parameters corresponding to the input feature attributes in the historical network point feature information, forming an input sample of the historical network point feature information corresponding to the untrained decision tree model, and using the corresponding historical network point type as the corresponding output sample; Based on the input samples and the corresponding output samples, the corresponding samples to be divided are formed.

6. The method according to claim 1, characterized in that Further including: Before obtaining a plurality of current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the plurality of candidate network point addresses and the input feature attributes corresponding to the plurality of trained decision tree models, Data cleaning, data extraction and data standardization are performed on the initial current feature information corresponding to the multiple candidate network point addresses to obtain the current feature information corresponding to the candidate network point addresses.

7. The method according to claim 1, characterized in that The step of obtaining a plurality of current input feature information corresponding to the input feature attributes based on the current feature information corresponding to the plurality of candidate network point addresses and the input feature attributes corresponding to the plurality of trained decision tree models includes: Based on the feature parameters in the current feature information corresponding to the input feature attributes, current input feature information of the trained decision tree model corresponding to the current feature information is formed.

8. The method according to claim 1, characterized in that The obtaining of the reliability coefficient corresponding to the candidate network point type based on the model accuracy of the trained decision tree model corresponding to the candidate network point type includes: The model accuracy rates of multiple trained decision tree models corresponding to the candidate network point type are superimposed to obtain the reliability coefficient corresponding to the candidate network point type.

9. The method according to claim 1, characterized in that The determining, based on the reliability coefficient, the target network point type corresponding to the to-be-selected network point address from the candidate network point types includes: The candidate network point type corresponding to the largest reliability coefficient is determined as the target network point type.

10. A network site selection device, characterized in that: include: a type prediction module, configured to obtain, based on current feature information corresponding to multiple candidate network point addresses and input feature attributes corresponding to multiple trained decision tree models, multiple current input feature information corresponding to the input feature attributes, and obtain multiple candidate network point types based on the current input feature information and the corresponding trained decision tree models; a reliability determination module, configured to obtain a reliability coefficient corresponding to the candidate network point type based on a model accuracy of a trained decision tree model corresponding to the candidate network point type; a network point selection module, configured to determine, based on the reliability coefficient, a target network point type corresponding to the candidate network point address from the candidate network point types, and determine a final network point address from a plurality of candidate network point addresses based on the target network point type; The device further includes: a preparation module for determining, based on multiple initial feature information preset in multiple historical network point feature information and historical network point types corresponding to the initial feature information, historical network point types corresponding to other historical network point feature information other than the initial feature information, before obtaining multiple current input feature information corresponding to the input feature attributes based on current feature information corresponding to multiple candidate network point addresses and input feature attributes corresponding to multiple trained decision tree models, wherein the historical network point types corresponding to the multiple initial feature information are different from each other; obtaining multiple samples to be divided corresponding to the untrained decision tree model based on the historical network point feature information, the corresponding historical network point types, and the multiple input feature attributes preset corresponding to multiple untrained decision tree models, and determining multiple training samples and test samples among the multiple samples to be divided based on a preset sample ratio; training the untrained decision tree model using the corresponding training samples to obtain a corresponding trained decision tree model, and testing the trained decision tree model using the corresponding test samples to obtain a corresponding model accuracy; The preparation module is specifically used to use the initial feature information as cluster center feature information, and use other historical network point feature information except the initial feature information as feature information to be classified; determine the second Euclidean distance between each feature information to be classified and each cluster center feature information, and based on the second Euclidean distance, determine the cluster center feature information closest to the corresponding feature information to be classified as the corresponding nearest cluster center feature information; obtain corresponding multiple initial target clusters based on multiple feature information to be classified that are the same as the corresponding nearest cluster center feature information, and use the historical network point type of the corresponding nearest cluster center feature information as the cluster type corresponding to the initial target cluster; repeatedly perform the clustering iteration step until there is a third Euclidean distance less than a preset distance threshold, wherein the clustering iteration step includes: based on the initial target cluster, obtaining the corresponding intermediate cluster center feature information, And the cluster type of the initial target cluster is used as the intermediate type of the corresponding intermediate cluster center feature information; all the historical network feature information are used as feature information to be classified; the third Euclidean distance between each feature information to be classified and the intermediate cluster center feature information is determined, and based on the third Euclidean distance, the intermediate cluster center feature information closest to the corresponding feature information to be classified is determined as the corresponding nearest intermediate cluster center feature information; based on multiple feature information to be classified that are the same as the corresponding nearest intermediate cluster center feature information, intermediate target clusters are obtained, and the intermediate type of the corresponding nearest intermediate cluster center feature information is used as the cluster type of the intermediate target cluster; the intermediate target cluster is used as the initial target cluster; the cluster type of multiple intermediate target clusters is used as the historical network point type corresponding to other historical network point feature information corresponding to the intermediate target cluster except the initial feature information.

11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 9 is implemented.

12. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Decision tree model generation method and data recommendation method based on decision tree model

    CN114418035A

  • Branch site selection method and device, storage medium and equipment

    CN115392759A