Potential user identification method and device, electronic equipment and storage medium
By acquiring and filtering user data and using pre-trained recognition models for prediction processing, the problem of low accuracy of potential user recognition in the prior art is solved, and more efficient potential user recognition is achieved.
Patent Information
- Application Number
- CN202311515177.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
The existing potential user identification methods have the problem of low recognition accuracy.
By obtaining the associated user data of the user data to be identified and the target product, similarity filtering is performed to obtain candidate potential users, and predict the candidate users using a pre-trained identification model to determine the potential users of the target product.
Effectively identify potential users of target products and improve the accuracy of potential users' identification.
Smart Images

Figure CN120013573A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing technology, and in particular to a potential user identification method, device, electronic device and storage medium. Background Art
[0002] With the development of big data technology, more and more big data analysis technologies are being applied to various industries.
[0003] In areas such as real estate, current methods for identifying potential users include: identifying potential users based on business rules, and the scope of using data is limited.
[0004] In the process of implementing the present invention, it is found that there are at least the following technical problems in the prior art: the existing potential user identification method has the problem of low potential user identification accuracy. Summary of the invention
[0005] The present invention provides a potential user identification method, device, electronic device and storage medium to effectively identify potential users of a target product and improve the accuracy of potential user identification of the target product.
[0006] According to one aspect of the present invention, a potential user identification method is provided, comprising:
[0007] Obtain the user data to be identified and the associated user data of the target product;
[0008] Performing similarity screening on the to-be-identified user data based on the associated user data to obtain candidate potential users of the target product;
[0009] The candidate potential users are predicted based on a pre-trained recognition model to obtain a prediction result, and potential users of the target product are determined from among the candidate potential users based on the prediction result.
[0010] According to another aspect of the present invention, there is provided a potential user identification device, comprising:
[0011] A data acquisition module is used to acquire the user data to be identified and the associated user data of the target product;
[0012] A data screening module, used to perform similarity screening on the to-be-identified user data based on the associated user data to obtain candidate potential users of the target product;
[0013] The potential user determination module is used to perform prediction processing on the candidate potential users based on a pre-trained recognition model to obtain prediction results, and determine the potential users of the target product from among the candidate potential users based on the prediction results.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the potential user identification method described in any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the potential user identification method described in any embodiment of the present invention when executed.
[0019] The technical solution of the embodiment of the present invention obtains the user data to be identified and the associated user data of the target product, and then performs similarity screening on the user data to be identified based on the associated user data to obtain candidate potential users of the target product, so as to filter out some non-compliant user data and reduce the amount of user data, and then performs prediction processing on the candidate potential users based on a pre-trained recognition model to obtain prediction results, and determines the potential users of the target product from among the candidate potential users based on the prediction results, so as to effectively identify the potential users of the target product and improve the accuracy of identifying the potential users of the target product.
[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 is a flow chart of a potential user identification method provided by an embodiment of the present invention;
[0023] Figure 2 is a flow chart of a potential user identification method provided by an embodiment of the present invention;
[0024] Figure 3is a flow chart of a potential user identification method provided by an embodiment of the present invention;
[0025] Figure 4 is a flow chart of a potential user identification method provided by an embodiment of the present invention;
[0026] Figure 5 is a flow chart of a potential user identification method provided by an embodiment of the present invention;
[0027] Figure 6 is a flow chart of a method for identifying potential real estate users provided by an embodiment of the present invention;
[0028] Figure 7 is a flow chart of a data preprocessing method provided by an embodiment of the present invention;
[0029] Figure 8 is a schematic structural diagram of a potential user identification device provided by an embodiment of the present invention;
[0030] Fig. 9 It is a schematic diagram of the structure of an electronic device for implementing the potential user identification method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.
[0033] Figure 11 is a flow chart of a potential user identification method provided by an embodiment of the present invention. This embodiment is applicable to the case of performing potential user identification on user data. The method can be executed by a potential user identification device. The potential user identification device can be implemented in the form of hardware and / or software. The potential user identification device can be configured in a terminal and / or a server. Figure 1 As shown, the method includes:
[0034] S110: Acquire the user data to be identified and the associated user data of the target product.
[0035] In this embodiment, the user data to be identified refers to the data to be used for potential user identification. For example, the user data to be identified can be multiple user data in the application platform, or multiple user data in the site, etc. For example, the user to be identified can include but is not limited to registered users of the application platform, browsing users or concerned users of the application platform, etc. The associated user of the target product can be a user associated with the target object. The association with the target object can be an interactive operation with the display information of the target object or a transaction with the target object, wherein the interactive operation is not limited to browsing, evaluating, forwarding and other interactive operations on the display information of the target object. Correspondingly, the associated user of the target object includes but is not limited to the browsing user, historical transaction user, concerned user, and purchase user of the target object. User data can be user attribute data authorized by the user for potential user analysis, such as attribute data including at least one attribute item pre-set, where the attribute item can be set according to demand and is not limited. The target product refers to the product to be recommended or traded, such as the target product can be a product such as real estate, automobile, etc. The associated user data refers to the user data associated with the target product. Exemplarily, when the target product is real estate, the associated user data can be one or more user data associated with the real estate.
[0036] Specifically, the user data to be identified and the associated user data of the target product can be obtained from a preset storage location of the electronic device, or they can be retrieved from other devices connected to the electronic device or cloud devices, without limitation here.
[0037] S120: Perform similarity screening on the to-be-identified user data based on the associated user data to obtain candidate potential users of the target product.
[0038] In this embodiment, the candidate potential user refers to the user matched in the user to be identified after similarity screening, which can form a candidate user set of potential users. It can be understood that through similarity screening, the user data with low similarity to the user data associated with the target product in the user data to be identified can be filtered out, and the user data with high similarity to the user data associated with the target product in the user data to be identified can be retained, and the screened user data to be identified is used as the data of the candidate potential user, that is, the candidate potential user is determined based on the user identifier in the screened user data to be identified, so as to realize the mining of potential users. The candidate potential user here can be identified by the user identifier in the user data, and the user identifier can be an identifier of user registration information, account information or a string, etc.
[0039] Exemplarily, the similarity screening method may include but is not limited to a method based on distance measurement and a method based on similarity measurement, wherein the method based on distance measurement may be Euclidean distance, and the method based on similarity measurement may be cosine similarity, etc.
[0040] S130: performing prediction processing on the candidate potential users based on a pre-trained recognition model to obtain a prediction result, and determining potential users of the target product from among the candidate potential users based on the prediction result.
[0041] In this embodiment, the pre-trained recognition model refers to a network model that can be used to identify potential users of the target product. Potential users refer to users who have potential transaction opportunities for the target product.
[0042] Specifically, the user data of the candidate potential users can be used as input data of the recognition model, and then the user data of the candidate potential users can be input into the pre-trained recognition model, and then the recognition model makes predictions based on the candidate potential users and outputs prediction results, which can characterize the probability data of each user being a potential user of the target product. Furthermore, accurate potential users of the target product can be screened out from the candidate potential users based on the prediction results, for example, the candidate potential users whose probability data in the prediction results is greater than a threshold can be determined as potential users.
[0043] The recognition model can be trained in advance with a large amount of user sample data. In the trained recognition model, the user sample data is pre-extracted for features, and the recognition model is trained based on the extracted feature information. By continuously adjusting the model parameters, the distance deviation between the output result of the recognition model and the user sample data label is gradually reduced and stabilized, thereby obtaining a trained recognition model.
[0044] Based on the above embodiments, optionally, after determining the potential users of the target product, the method may further include: sending the target product data to the electronic devices of the potential users according to the potential users of the target product, so as to achieve accurate reach of the target product data. The target product data may be information in the form of pictures, texts, etc.
[0045] The technical solution of this embodiment obtains the associated user data of the user to be identified and the target product, and then performs similarity screening on the user data to be identified based on the associated user data to obtain candidate potential users of the target product, so as to filter out some non-compliant user data and reduce the amount of user data, and then performs prediction processing on the candidate potential users based on the pre-trained recognition model to obtain prediction results, and determines the potential users of the target product from among the candidate potential users based on the prediction results, so as to effectively identify the potential users of the target product and improve the accuracy of identifying the potential users of the target product.
[0046] Figure 2 A flow chart of a potential user identification method provided in an embodiment of the present invention, the method of this embodiment can be combined with the various optional schemes in the potential user identification method provided in the above embodiments. The potential user identification method provided in this embodiment is further optimized. Optionally, the similarity screening of the user data to be identified based on the associated user data to obtain candidate potential users of the target product includes: for each of the associated user data, determining the similarity data between the associated user data and each of the user data to be identified, and determining the candidate potential users matching the associated user data based on the similarity data; deduplicating the candidate potential users matching each of the associated user data to obtain candidate potential users of the target product.
[0047] like Figure 2 As shown, the method includes:
[0048] S210: Acquire the user data to be identified and the associated user data of the target product.
[0049] S220: For each of the associated user data, determine similarity data between the associated user data and each of the to-be-identified user data, and determine candidate potential users matching the associated user data based on the similarity data.
[0050] In this embodiment, the number of the user data to be identified and the associated user data can be multiple, in other words, the user data to be identified and the associated user data of multiple users can be obtained. The similarity data can be used to characterize the similarity between the associated user data and the user data to be identified.
[0051] Specifically, similarity data between each associated user data and each to-be-identified user data may be determined, and then the to-be-identified user data may be screened according to each similarity data to obtain a plurality of candidate potential users matching the associated user data.
[0052] S230: De-duplicate candidate potential users that match the associated user data to obtain candidate potential users of the target product.
[0053] It should be noted that there may be duplicate user data among the candidate potential users that match the associated user data. This embodiment deduplicates the candidate potential users that match the associated user data, deletes the duplicate user data, thereby avoiding the occurrence of data duplication and redundancy, and reduces the amount of user data, which can reduce the occupancy of device resources.
[0054] S240: performing prediction processing on the candidate potential users based on the pre-trained recognition model to obtain prediction results, and determining potential users of the target product from among the candidate potential users based on the prediction results.
[0055] Based on the above embodiments, optionally, determining candidate potential users that match the associated user data based on the similarity data includes: sorting each user data to be identified based on the similarity data, and determining the user data to be identified in a preset sorting range as candidate potential users that match the associated user data based on the sorting result.
[0056] The similarity data may be a similarity value. For example, the similarity data may be a value ranging from 0 to 1. It is understandable that a larger similarity value indicates a higher similarity.
[0057] For example, taking the real estate scenario as an example, the user data to be identified may be the user data within the shopping website, and the associated user data may be the user data within the real estate channel of the shopping website. The amount of user data within the shopping website is much larger than the user data within the real estate channel of the shopping website. Specifically, the user data within the shopping website may be sorted according to the similarity value, and the user data within the shopping website with a similarity value between 0.5 and 1 may be determined as candidate potential users that match the user data within the real estate channel, that is, candidate user data with potential purchase intentions may be obtained.
[0058] The technical solution of the embodiment of the present invention determines, for each associated user data, similarity data between the associated user data and each user data to be identified, and then determines candidate potential users similar to the associated user data based on the similarity data; further, deduplication processing is performed on the candidate potential users that match the associated user data, and duplicate user data is deleted, thereby avoiding the occurrence of data duplication and redundancy, and reducing the amount of user data, which can reduce the occupation of device resources.
[0059] Figure 3 A flowchart of a potential user identification method provided in an embodiment of the present invention, the method of this embodiment can be combined with each optional scheme in the potential user identification method provided in the above embodiment. The potential user identification method provided in this embodiment is further optimized. Optionally, the method also includes: reading a pre-configuration file, the pre-configuration file includes a potential user screening condition; based on the potential user screening condition, the user data to be identified or the candidate potential user is screened.
[0060] like Figure 3 As shown, the method includes:
[0061] S310: Obtain the user data to be identified and the associated user data of the target product.
[0062] S320: Perform similarity screening on the to-be-identified user data based on the associated user data to obtain candidate potential users of the target product.
[0063] S330: Read a pre-configuration file, wherein the pre-configuration file includes a potential user screening condition; and screen the to-be-identified user data or the candidate potential users based on the potential user screening condition.
[0064] S340: Perform prediction processing on the candidate potential users based on the pre-trained recognition model to obtain prediction results, and determine potential users of the target product from among the candidate potential users based on the prediction results.
[0065] In this embodiment, the potential user screening condition may be a pre-configured potential user screening condition, which may be used to screen potential users and may be stored in a pre-configuration file.
[0066] Exemplarily, the potential user screening condition may be to retain the data of users within a preset age range, or the potential user screening condition may be to delete the data of users who have not logged into the application for more than a preset number of days, or the potential user screening condition may be to retain the data of users with excellent credit ratings, etc.
[0067] Specifically, after obtaining candidate potential users for the target product, the candidate potential users can be screened based on the potential user screening conditions; or, before obtaining candidate potential users for the target product, the user data to be identified can be screened based on the potential user screening conditions to filter out user data that does not meet the potential user screening conditions, retain user data that meets the potential user screening conditions, and realize the mining of potential user-related data.
[0068] Based on the above embodiments, optionally, the associated user data includes transaction user data of the target product and / or browsing user data of the target product; the method also includes: determining the transaction user data as positive sample data, and determining the non-transaction browsing user data as negative sample data; and obtaining the recognition model based on the training of the positive sample data and the negative sample data.
[0069] The transaction user data refers to the relevant data of users who have completed the target product transaction. The target product browsing user data refers to the relevant data of users who browse the target product.
[0070] For example, taking the real estate scenario as an example, the transaction user data of the target product can be the user data of those who have completed the real estate transaction, and the browsing user data of the target product can be the user data of those who have browsed the real estate information but have not completed the real estate transaction. Specifically, the user data of those who have completed the real estate transaction can be determined as positive sample data, and the user data of those who have browsed the real estate information but have not completed the real estate transaction can be determined as negative sample data, and then the initial network model is trained, adjusted and optimized based on the positive sample data and the negative sample data to obtain a recognition model. Optionally, the initial network model can be a prediction model such as the LGBM (Light Gradient Boosting Machine) model.
[0071] The technical solution of the embodiment of the present invention reads a pre-configuration file, which includes potential user screening conditions, and then screens the user data to be identified or the candidate potential users based on the potential user screening conditions to filter out user data that does not meet the potential user screening conditions and retains user data that meets the potential user screening conditions, thereby realizing the mining of potential user-related data.
[0072] Figure 4A flowchart of a potential user identification method provided in an embodiment of the present invention, the method of this embodiment can be combined with the various optional schemes in the potential user identification method provided in the above embodiment. The potential user identification method provided in this embodiment is further optimized. Optionally, after obtaining the user data to be identified and the associated user data of the target product, it also includes: preprocessing the user data to be identified and the associated user data, and the preprocessing includes one or more of the following: feature selection of the user data to be identified and the associated user data; filling in missing values in the user data to be identified and the associated user data; outlier processing of the user data to be identified and the associated user data; deriving features in the user data to be identified and the associated user data; normalizing the feature values in the user data to be identified and the associated user data; encoding the feature values in the user data to be identified and the associated user data.
[0073] like Figure 4 As shown, the method includes:
[0074] S410: Acquire the user data to be identified and the associated user data of the target product.
[0075] S420: Preprocess the to-be-identified user data and the associated user data to obtain preprocessed to-be-identified user data and associated user data.
[0076] It is understandable that preprocessing the to-be-identified user data and the associated user data can improve data quality and reduce the workload and working time of subsequent data mining.
[0077] In this embodiment, the preprocessing includes one or more of the following: feature selection for the user data to be identified and the associated user data; filling in missing values in the user data to be identified and the associated user data; outlier processing for the user data to be identified and the associated user data; deriving features in the user data to be identified and the associated user data; normalizing feature values in the user data to be identified and the associated user data; and encoding feature values in the user data to be identified and the associated user data.
[0078] Specifically, features whose missing value ratios in the to-be-identified user data and the associated user data meet a preset threshold, or features whose feature values are single values, can be eliminated. Missing values in the to-be-identified user data and the associated user data can be filled in according to the data type. Outlier processing can be performed on the to-be-identified user data and the associated user data using a preset multiple standard deviation. Features in the to-be-identified user data and the associated user data can be derived, such as adding a variable called "user activity within 30 days". Feature values in the to-be-identified user data and the associated user data can be normalized. Feature values in the to-be-identified user data and the associated user data can be encoded using one-hot, label, or other methods.
[0079] On the basis of the above embodiments, optionally, feature selection is performed on the user data to be identified and the associated user data, including one or more of the following: eliminating features whose missing value ratios in the user data to be identified and the associated user data meet a preset threshold; eliminating features whose feature values in the user data to be identified and the associated user data are single values; obtaining feature importance data output during the training of the recognition model, and performing feature screening on the user data to be identified and the associated user data based on the feature importance data; obtaining a feature profile of the target product, and performing feature screening on the user data to be identified and the associated user data based on the feature profile, wherein the feature profile includes features associated with the target product.
[0080] In some embodiments, features whose missing value ratio in the to-be-identified user data and the associated user data is greater than a preset threshold may be eliminated. The preset threshold may be set according to data processing requirements to complete feature selection.
[0081] In some embodiments, features whose feature values in the to-be-identified user data and the associated user data are single values may also be eliminated to complete feature selection.
[0082] In some embodiments, feature importance data output during the recognition model training process can also be obtained, and feature screening can be performed on the user data to be identified and the associated user data based on the feature importance data, wherein the feature importance data can be used to characterize the importance of the data. It can be understood that by performing feature screening on the user data to be identified and the associated user data through the feature importance data, the user data to be identified and the associated user data with high importance can be retained. Exemplarily, the feature importance data output during the recognition model training process can be the importance value of the user's city level feature. If the importance value of the user's city level feature is greater than a preset judgment threshold, the data corresponding to the user is retained, otherwise it is removed.
[0083] In some embodiments, a feature configuration file of the target product may also be obtained, and feature screening may be performed on the user data to be identified and the associated user data based on the feature configuration file, wherein the feature configuration file includes associated features with the target product, wherein the associated features may be pre-configured associated features of the target product, and may be used to perform feature screening on the user data to be identified and the associated user data. Exemplarily, the associated features may be the location of the area where the property is located, and if the location of the area where the property is located is area A, then the user data of the user data to be identified and the associated user data whose residence is area A is retained.
[0084] S430: Perform similarity screening based on the pre-processed to-be-identified user data and the associated user data to obtain candidate potential users of the target product.
[0085] S440: Perform prediction processing on the candidate potential users based on the pre-trained recognition model to obtain prediction results, and determine potential users of the target product from among the candidate potential users based on the prediction results.
[0086] The technical solution of the embodiment of the present invention can improve data quality and reduce the workload and working time of subsequent data mining by preprocessing the user data to be identified and the associated user data.
[0087] Figure 5 This is a flow chart of a potential user identification method provided by an embodiment of the present invention. The method of this embodiment is a preferred example of the above embodiments. In this embodiment, the target product is real estate, the user data to be identified is the user portrait label data in the site, and the associated user data of the target product is the user portrait label data in the real estate channel. Figure 6 A flowchart of a method for identifying potential real estate users provided by an embodiment of the present invention, the method for identifying potential real estate users comprising:
[0088] S510: Collect user portrait label data in the real estate channel and user portrait label data in the website, and pre-process the user portrait label data in the real estate channel and the user portrait label data in the website respectively.
[0089] The preprocessing methods include one or more of the following: feature selection, missing value processing, outlier processing, adding derivative variables, data reduction and feature encoding, etc. The order of the preprocessing methods is not limited here. Figure 7 A flow chart of a data preprocessing method provided by an embodiment of the present invention. Specifically, the user portrait label data in the real estate channel and the user portrait label data in the station can be processed in sequence by feature selection, missing value processing, outlier processing, adding derivative variables, data reduction and feature coding to obtain the preprocessed user portrait label data in the real estate channel and the user portrait label data in the station.
[0090] In this embodiment, a feature library can be constructed based on the user portrait tag data in the real estate channel and the user portrait tag data in the site. The feature library can include but is not limited to information such as the user's age, gender, education, occupation, city area, purchasing power, etc.
[0091] In this embodiment, the user portrait label data in the station is the data obtained by filtering out the user portrait label data in the real estate channel. The data volume is large, and direct processing will cause memory overflow. To solve this problem, this embodiment will collect the user portrait label data in the real estate channel and the user portrait label data in the station and split them into multiple files, store them in a distributed file system (Hadoop Distributed File System, HDFS), and then read each file in the distributed file system for processing.
[0092] S520. The preprocessed user portrait label data in the real estate channel and the user portrait label data in the site are processed by a cosine similarity model based on Faiss to obtain a similarity value, and the user portrait label data in the site are sorted in descending order according to the similarity value to filter out the users in the site whose similarity value is greater than a threshold.
[0093] In this embodiment, the pre-processed user portrait label data in the real estate channel and the user portrait label data in the site can be processed in batches, and a GPU can be used for accelerated processing to improve the similarity calculation speed.
[0094] S530: By using an expert business method, the user data in the site with a similarity value greater than a threshold can be filtered to obtain candidate potential users of the property.
[0095] In this embodiment, the potential user screening condition can be set by the expert business method. For example, the potential user screening condition can be to retain the data of users within a preset age range, or the potential user screening condition can also be to delete the data of users who have not logged into the application for more than a preset number of days, or the potential user screening condition can also be to retain the data of users with excellent credit ratings, etc.
[0096] S540: taking the candidate potential users of the property as input data of the LGBM recognition model, the LGBM recognition model performs prediction processing on the candidate potential users of the property to obtain a prediction result, and determines the potential users of the property from among the candidate potential users of the property based on the prediction result.
[0097] The prediction result includes the high potential probability value of each candidate potential user. Exemplarily, the candidate potential users of the property can be sorted in descending order according to the high potential probability value of the property, and the user data corresponding to the first k candidate potential users are determined as potential users of the property, where k>0.
[0098] In this embodiment, the training data of the LGBM recognition model can be the historical user data in the real estate channel, and the model effect evaluation index can be evaluated by indicators such as Accuracy and F1, where Accuracy represents the ratio of the number of correctly predicted samples to the total number of predicted samples, and F1 represents the harmonic mean of precision and recall. The precision is the proportion of correct predictions in the set of all predicted positive samples, and the recall is the proportion of correct predictions in all positive samples.
[0099] After obtaining potential users of the property, the recognition model can be iteratively optimized based on the feedback results of the property sales, thereby improving the model's prediction accuracy.
[0100] The technical solution of the embodiment of the present invention collects user portrait label data in the real estate channel and user portrait label data in the station, and pre-processes the user portrait label data in the real estate channel and the user portrait label data in the station respectively; then, the pre-processed user portrait label data in the real estate channel and the user portrait label data in the station are processed through a cosine similarity model based on Faiss to obtain similarity values, and the user portrait label data in the station are sorted in descending order according to the similarity values to screen out the users in the station whose similarity values are greater than a threshold; then, the user data in the station whose similarity values are greater than a threshold can be filtered through an expert business method to obtain candidate potential users of the real estate; then, the candidate potential users of the real estate are used as input data of the recognition model, and the recognition model predicts the candidate potential users of the real estate to obtain prediction results, and based on the prediction results, accurate potential users of the real estate are determined among the candidate potential users of the real estate.
[0101] Figure 8 Schematic diagram of a potential user identification device provided by an embodiment of the present invention. Figure 8 As shown, the device comprises:
[0102] The data acquisition module 610 is used to acquire the user data to be identified and the associated user data of the target product;
[0103] A data screening module 620 is used to perform similarity screening on the to-be-identified user data based on the associated user data to obtain candidate potential users of the target product;
[0104] The potential user determination module 630 is used to perform prediction processing on the candidate potential users based on the pre-trained recognition model to obtain prediction results, and determine the potential users of the target product from among the candidate potential users based on the prediction results.
[0105] The technical solution of this embodiment obtains the associated user data of the user to be identified and the target product, and then performs similarity screening on the user data to be identified based on the associated user data to obtain candidate potential users of the target product, so as to filter out some non-compliant user data and reduce the amount of user data, and then performs prediction processing on the candidate potential users based on the pre-trained recognition model to obtain prediction results, and determines the potential users of the target product from among the candidate potential users based on the prediction results, so as to effectively identify the potential users of the target product and improve the accuracy of identifying the potential users of the target product.
[0106] Based on the above embodiment, optionally, the data screening module 620 includes:
[0107] A candidate potential user matching unit, configured to determine, for each of the associated user data, similarity data between the associated user data and each of the user data to be identified, and determine a candidate potential user matching the associated user data based on the similarity data;
[0108] The candidate potential user deduplication unit is used to perform deduplication processing on the candidate potential users that match the associated user data to obtain the candidate potential users of the target product.
[0109] Based on the above embodiment, optionally, the candidate potential user matching unit is further used to:
[0110] The user data to be identified are sorted based on the similarity data, and the user data to be identified within a preset sorting range are determined as candidate potential users matching the associated user data based on the sorting result.
[0111] Based on the above embodiment, optionally, the device further includes:
[0112] A pre-configuration file reading module, used to read a pre-configuration file, wherein the pre-configuration file includes a potential user screening condition;
[0113] The potential user screening module is used to screen the user data to be identified or the candidate potential users based on the potential user screening conditions.
[0114] Based on the above embodiment, optionally, the device further includes:
[0115] A data preprocessing module is used to preprocess the to-be-identified user data and the associated user data, wherein the preprocessing includes one or more of the following:
[0116] Performing feature selection on the to-be-identified user data and the associated user data;
[0117] Filling missing values in the to-be-identified user data and the associated user data;
[0118] Performing outlier processing on the to-be-identified user data and the associated user data;
[0119] Performing derivative processing on the features in the to-be-identified user data and the associated user data;
[0120] Normalizing the feature values in the to-be-identified user data and the associated user data;
[0121] Encoding is performed on the feature values in the to-be-identified user data and the associated user data.
[0122] Based on the above embodiment, optionally, the performing feature selection on the to-be-identified user data and the associated user data includes one or more of the following:
[0123] Eliminate features in which the ratio of missing values in the to-be-identified user data and the associated user data meets a preset threshold;
[0124] Eliminate features whose feature values are single values in the to-be-identified user data and the associated user data;
[0125] Acquire feature importance data output during the recognition model training process, and perform feature screening on the user data to be recognized and the associated user data based on the feature importance data;
[0126] A feature configuration file of the target product is obtained, and feature screening is performed on the to-be-identified user data and the associated user data based on the feature configuration file, wherein the feature configuration file includes features associated with the target product.
[0127] The potential user identification device provided in the embodiment of the present invention can execute the potential user identification method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0128] Fig. 9 1 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0129] like Fig. 9 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0130] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0131] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs the various methods and processes described above, such as a potential user identification method.
[0132] In some embodiments, the potential user identification method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the potential user identification method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the potential user identification method in any other appropriate manner (e.g., by means of firmware).
[0133] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0134] The computer programs for implementing the potential user identification method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0135] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute a potential user identification method, the method comprising:
[0136] Obtain the user data to be identified and the associated user data of the target product;
[0137] Performing similarity screening on the to-be-identified user data based on the associated user data to obtain candidate potential users of the target product;
[0138] The candidate potential users are predicted based on a pre-trained recognition model to obtain a prediction result, and potential users of the target product are determined from among the candidate potential users based on the prediction result.
[0139] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0141] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0142] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0143] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0144] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A potential user identification method, characterized in that: include: Obtain the user data to be identified and the associated user data of the target product; Performing similarity screening on the to-be-identified user data based on the associated user data to obtain candidate potential users of the target product; The candidate potential users are predicted based on a pre-trained recognition model to obtain a prediction result, and potential users of the target product are determined from among the candidate potential users based on the prediction result.
2. The method according to claim 1, characterized in that The similarity screening of the to-be-identified user data based on the associated user data to obtain candidate potential users of the target product includes: For each of the associated user data, determining similarity data between the associated user data and each of the user data to be identified, and determining candidate potential users matching the associated user data based on the similarity data; Deduplication is performed on candidate potential users that match the associated user data to obtain candidate potential users of the target product.
3. The method according to claim 2, characterized in that The determining, based on the similarity data, candidate potential users matching the associated user data comprises: The user data to be identified are sorted based on the similarity data, and the user data to be identified within a preset sorting range are determined as candidate potential users matching the associated user data based on the sorting result.
4. The method according to claim 1, characterized in that: The method further comprises: Reading a pre-configuration file, wherein the pre-configuration file includes a potential user screening condition; The to-be-identified user data or the candidate potential users are screened based on the potential user screening condition.
5. The method according to claim 1, characterized in that After obtaining the user data to be identified and the associated user data of the target product, it also includes: Preprocessing the to-be-identified user data and the associated user data, wherein the preprocessing includes one or more of the following: Performing feature selection on the to-be-identified user data and the associated user data; Filling missing values in the to-be-identified user data and the associated user data; Performing outlier processing on the to-be-identified user data and the associated user data; Performing derivative processing on the features in the to-be-identified user data and the associated user data; Normalizing the feature values in the to-be-identified user data and the associated user data; Encoding is performed on the feature values in the to-be-identified user data and the associated user data.
6. The method according to claim 5, characterized in that The performing feature selection on the to-be-identified user data and the associated user data includes one or more of the following: Eliminate features in which the ratio of missing values in the to-be-identified user data and the associated user data meets a preset threshold; Eliminate features whose feature values are single values in the to-be-identified user data and the associated user data; Acquire feature importance data output during the recognition model training process, and perform feature screening on the user data to be recognized and the associated user data based on the feature importance data; A feature configuration file of the target product is obtained, and feature screening is performed on the to-be-identified user data and the associated user data based on the feature configuration file, wherein the feature configuration file includes features associated with the target product.
7. The method according to claim 1, characterized in that The associated user data includes transaction user data of the target product and / or browsing user data of the target product; The method further comprises: Determine the transaction user data as positive sample data, and determine the non-transaction browsing user data as negative sample data; The recognition model is obtained by training based on the positive sample data and the negative sample data.
8. A potential user identification device, characterized in that: include: A data acquisition module is used to acquire the user data to be identified and the associated user data of the target product; A data screening module, used to perform similarity screening on the to-be-identified user data based on the associated user data to obtain candidate potential users of the target product; The potential user determination module is used to perform prediction processing on the candidate potential users based on a pre-trained recognition model to obtain prediction results, and determine the potential users of the target product from among the candidate potential users based on the prediction results.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the potential user identification method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the potential user identification method according to any one of claims 1 to 7 when executed.