User selection method and device, equipment, storage medium and program product
By analyzing the influence of user characteristics on intention, user instances are generated, thereby selecting potential users from users with low intention. This solves the problem of potential users being omitted in the existing technology and achieves more in-depth and comprehensive user selection.
Patent Information
- Application Number
- CN202510307598.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-09-23
AI Technical Summary
In the user selection process, the existing technology fails to fully analyze other characteristics of non-target users, resulting in the omission of potential users and the inability to select potential users from users with low intention.
The influence of intention is analyzed based on user characteristics. By generating user instances and determining the influence of characteristics, potential users are selected from the user group with low intention.
By enriching feature diversity and deeply analyzing the impact of user characteristics on intention, we can avoid omissions in the user selection process and achieve effective selection of potential users.
Smart Images

Figure CN120689087A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a user selection method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] In related technologies, machine learning models are widely used in the stratification and decision-making processes of various users. For example, the user's purchasing behavior is predicted based on the machine learning model to obtain the user's purchase intention tendency score. Then, users whose purchase intention tendency score exceeds the threshold are regarded as target users, and all other users are regarded as non-target users. Summary of the Invention
[0003] Embodiments of the present application provide a user selection method, apparatus, electronic device, and computer-readable storage medium, which can select potential users from users with low intention.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] This embodiment of the present application provides a user selection method, the method comprising:
[0006] Dividing the users into a first user group and a second user group based on their intentions for the first event, wherein the intention of a first user in the first user group is greater than or equal to an intention threshold, and the intention of a second user in the second user group is less than the intention threshold;
[0007] generating a first user instance based on a first feature of a first user in the first user group, and generating a second user instance based on a second feature of a second user in the second user group;
[0008] Determining a first degree of influence of the first feature on the intentionality based on the first user instance, and determining a second degree of influence of the second feature on the intentionality based on the second user instance;
[0009] Based on the first influence level and the second influence level, a third user whose influence level is greater than an influence level threshold is selected from the second user group.
[0010] An embodiment of the present application provides a user selection device, the device comprising:
[0011] a grouping module, configured to divide the users into a first user group and a second user group based on the users' intentions for the first event, wherein the intentions of first users in the first user group are greater than or equal to an intention threshold, and the intentions of second users in the second user group are less than the intention threshold;
[0012] a generating module, configured to generate a first user instance based on a first feature of a first user in the first user group, and to generate a second user instance based on a second feature of a second user in the second user group;
[0013] a determination module, configured to determine a first degree of influence of the first feature on the intentionality based on the first user instance, and to determine a second degree of influence of the second feature on the intentionality based on the second user instance;
[0014] The selection module is configured to select a third user whose influence degree is greater than an influence degree threshold from the second user group based on the first influence degree and the second influence degree.
[0015] An embodiment of the present application provides an electronic device, including:
[0016] Memory for storing computer-executable instructions or computer programs;
[0017] The processor is used to implement the user selection method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0018] An embodiment of the present application provides a computer-readable storage medium, which stores computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the user selection method provided by the embodiment of the present application.
[0019] The embodiments of the present application provide a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the user selection method provided in the embodiments of the present application.
[0020] The embodiments of the present application have the following beneficial effects:
[0021] After grouping users based on their intentions for the first event to obtain a first user group whose intentions are greater than or equal to a threshold intention threshold and a second user group whose intentions are less than the threshold intention threshold, a first user instance is determined based on the first feature, and a second user instance is determined based on the second feature. Then, based on the first user instance, a first influence of the first feature on the intention is determined, and based on the second user instance, a second influence of the second feature on the intention is determined. Then, based on the first influence and the second influence, a third user whose influence is greater than the threshold influence is selected from the second user group. In this way, compared with a solution that selects users based only on intention, multiple user instances are generated based on the features of different users, which enriches the diversity of features, thereby facilitating the determination of the influence of the features on the intention. Then, for users with low intention, user selection is performed from users with low intention based on the influence. In this way, for users with low intention, user selection can be performed in combination with the influence of the features on the intention, making the user selection process more in-depth and comprehensive, avoiding the omission of users who should be selected, and achieving the effect of selecting potential users from users with low intention. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 1 is a schematic diagram of the architecture of the user selection system 100 provided in an embodiment of the present application;
[0023] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0024] Figure 3 Schematic diagram of the user selection method provided in the embodiment of the present application;
[0025] Figure 4 Schematic diagram of the model structure of the intentionality model provided in the embodiment of the present application;
[0026] Figure 5 is a schematic diagram of a process for determining the first feature and the second feature provided in an embodiment of the present application;
[0027] Figure 6 Schematic diagram of the model structure of the discrimination model provided in the embodiment of the present application;
[0028] Figure 7 This is a flowchart of determining the first impact level provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0030] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0031] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0033] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0034] 1) Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics.
[0035] 2) Client, also known as the user end, refers to the program corresponding to the server that provides local services to users. Except for some applications that can only run locally, it is generally installed on an ordinary client and needs to cooperate with the server to run. That is, there must be corresponding servers and service programs in the network to provide corresponding services. In this way, a specific communication connection needs to be established between the client and the server to ensure the normal operation of the application.
[0036] During the research process, the inventors found that when promoting to users, the related technology would first predict the user's intention, then select users with high intention as target users and users with low intention as target users, so as to promote only to the target users;
[0037] However, during the prediction process, some features have a high degree of influence on the prediction results, which is equivalent to the model only using these important features to predict the user, but the user has other features. Although other features have a lower impact on the prediction results, they may affect the user's actual behavior. Therefore, when predicting a user as a non-target user based on important features, other non-important features are not fully analyzed. Therefore, potential users may also exist among the non-target users; in this way, the target user selection scheme in the related art may cause the omission of target users, and it is impossible to select potential users from the non-target users.
[0038] Based on this, an embodiment of the present application provides a user selection method. For users with low intention, that is, non-target users, after determining the degree of influence of the user's characteristics on the intention, the user is selected from the users with low intention based on the degree of influence. In this way, for users with low intention, the degree of influence of the characteristics on the intention can be combined to make the selection, making the user selection process more in-depth and comprehensive, avoiding the situation of missing users who should be selected, and thus selecting potential users from users with low intention.
[0039] See also Figure 1 , Figure 1 This is an architectural diagram of a user selection system 100 provided in an embodiment of the present application. To implement an application scenario of user selection, a terminal (terminal 400 is shown as an example) is connected to a server 200 via a network 300. The network 300 may be a wide area network or a local area network, or a combination of the two. The terminal 400 is used for the user to use a client 401, which is displayed on a display interface (display interface 401-1 is shown as an example). The terminal 400 and the server 200 are connected to each other via a wired or wireless network.
[0040] The terminal 400 is used to send the user's intention for the first event to the server 200;
[0041] The server 200 is used to receive the user's intention for the first event; based on the user's intention for the first event, divide the users into a first user group and a second user group, the intention of the first user in the first user group is greater than or equal to the intention threshold, and the intention of the second user in the second user group is less than the intention threshold; based on the first feature of the first user in the first user group, generate a first user instance, and based on the second feature of the second user in the second user group, generate a second user instance; based on the first user instance, determine a first influence degree of the first feature on the intention degree, and based on the second user instance, determine a second influence degree of the second feature on the intention degree; based on the first influence degree and the second influence degree, select a third user from the second user group whose influence degree is greater than the influence degree threshold.
[0042] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Deliver Network), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a set-top box, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, and a mobile device (for example, a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device, a smart speaker, and a smart watch), etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0043] See also Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. In practical applications, the electronic device can be Figure 1 The server 200 or terminal 400 shown, see Figure 2 , Figure 2 The electronic device shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .
[0044] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0045] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0046] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0047] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0048] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0049] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0050] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB).
[0051] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0052] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.
[0053] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 A user selection device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a grouping module 4551, a generation module 4552, a determination module 4553, and a selection module 4554. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0054] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the user selection device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the user selection method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0055] In some embodiments, the terminal or server can implement the user selection method provided in the embodiment of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP, a web browser APP; it can also be a small program, that is, a program that can be run only by downloading it into a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plug-in.
[0056] Based on the above description of the user selection system and electronic device provided by the embodiment of the present application, the user selection method provided by the embodiment of the present application is described below. In actual implementation, the user selection method provided by the embodiment of the present application can be implemented by the terminal or the server alone, or by the terminal and the server in collaboration, so that Figure 1 The example of the server 200 executing the user selection method provided in the embodiment of the present application alone is used for explanation. Figure 3 , Figure 3 This is a flow chart of the user selection method provided by the embodiment of the present application. Figure 3 The steps shown will be described.
[0057] In step 101, the server divides users into a first user group and a second user group based on the users' intentions for a first event, wherein the intention of the first user in the first user group is greater than or equal to the intention threshold, and the intention of the second user in the second user group is less than the intention threshold.
[0058] In actual implementation, the first event may be a user purchasing a certain product or a user downloading a certain APP, etc., which is not limited in the embodiments of the present application; and the user's intention degree for the first event refers to the probability that the user will execute the first event. If the user's intention degree for the first event is 1, that is, the probability that the user will execute the first event is 1, it means that the user will definitely execute the first event. If the user's intention degree for the first event is 0, that is, the probability that the user will execute the first event is 0, it means that the user will definitely not execute the first event.
[0059] The intentions of multiple users for the first event can be predicted by the trained intention model; specifically, see Figure 4 , Figure 4 This is a schematic diagram of the model structure of the intentionality model provided in the embodiment of the present application, based on Figure 4 The intention model includes a feature extraction layer and an intention prediction layer. Therefore, the process of obtaining the user's intention for the first event can be, for each user, inputting the user's attribute information into the intention model, and extracting the user's features based on the user's attribute information through the feature extraction layer of the intention model to obtain user features; through the intention prediction layer of the intention model, predicting the user's intention for the first event based on the user features to obtain the predicted user's intention for the first event; wherein the user's attribute information can be the user's age, gender, occupation, number of likes within 30 days, number of searches within 30 days, etc.
[0060] In actual implementation, after determining the user's intention for the first event, multiple users are grouped based on each user's intention and intention threshold to obtain a first user group and a second user group. Specifically, for each user, the user's intention is compared with the intention threshold. When the comparison result indicates that the user's intention is greater than or equal to the intention threshold, the user is regarded as the first user; when the comparison result indicates that the user's intention is less than the intention threshold, the user is regarded as the second user; wherein the intention threshold is pre-set, and this embodiment of the present application does not limit this.
[0061] Step 102 : generating a first user instance based on a first feature of a first user in a first user group, and generating a second user instance based on a second feature of a second user in a second user group.
[0062] In actual implementation, based on the first feature of the first user in the first user group, a first user instance is generated, that is, based on the first features of different first users, multiple user instances are generated. Before generating the first user instance based on the first feature of the first user in the first user group, see Figure 5 , Figure 5 This is a flow chart of determining the first and second features provided by the embodiment of the present application, based on Figure 5 Based on the first feature of the first user in the first user group, before generating the first user instance, the first feature and the second feature can be determined through the following steps.
[0063] In step 201, the server performs importance analysis on the user's features to obtain the importance of each feature.
[0064] It should be noted that the importance analysis of user features here refers to the importance analysis of each user's features, and the importance of each feature refers to the impact of the corresponding feature on the prediction result when predicting whether the user's intention reaches the intention threshold. For example, if the confidence of the prediction result decreases significantly when predicting whether the user's intention reaches the intention threshold based on other features that do not include a certain feature, it means that the feature is very important to the prediction process of predicting whether the user's intention reaches the intention threshold, that is, it is highly important, such as exceeding the pre-set importance threshold.
[0065] In actual implementation, the process of performing importance analysis on the user's features and obtaining the importance of each feature can be, for each of the user's multiple features, performing the following processing to obtain the importance of the feature: removing the corresponding feature from the user's multiple features to obtain a third feature; based on the third feature, predicting the probability that the user's intention exceeds the intention threshold; and determining the importance of the feature based on the probability.
[0066] It should be noted that for each feature, the feature is removed from the user's multiple features to obtain a third feature. Here, the third feature refers to the other features in the multiple features except the feature; then, based on the third feature, the probability of the user's intention exceeding the intention threshold is predicted; and the process of determining the importance of the feature based on the probability can be to determine the predicted intention of the user for the first event, and determine the user's label based on whether the user's intention exceeds the intention threshold; the absolute value of the difference between the above predicted probability and the user's label is used as the importance of the corresponding feature.
[0067] It should be noted that the user's label is used to indicate whether the user's intention exceeds the intention threshold, that is, to indicate the probability that the user's intention exceeds the intention threshold; wherein, since the user's intention has been determined, whether the user's intention exceeds the intention threshold has also been determined, that is, the probability that the user's intention exceeds the intention threshold has also been determined, so the user's label is 1 or 0. When the user's intention exceeds the intention threshold, that is, the probability that the user's intention exceeds the intention threshold is 1, at this time, the user's label is 1; when the user's intention does not exceed the intention threshold, that is, the probability that the user's intention exceeds the intention threshold is 0, at this time, the user's label is 0.
[0068] Based on this, if the probability that the user's intention exceeds the intention threshold is predicted, the greater the absolute value of the difference with the user's label, the less accurate the prediction result is. That is, if the prediction result of the user is less accurate based on other features other than a certain feature (equivalent to the less accurate prediction result obtained by predicting the user not based on the feature), the higher the importance of the corresponding feature.
[0069] In this way, after removing the features and using the remaining features to predict and obtain the prediction results, the importance of the features is determined based on the difference between the predicted results and the actual results, and the impact of feature removal on the model prediction accuracy is quantified, that is, the importance of the features is quantified, which makes it easier to sort the importance of the features and determine the important features.
[0070] In actual implementation, the importance of user features is analyzed and the importance of each feature is obtained through a discriminant model. The discriminant model is used to determine whether the user's intention exceeds the intention threshold, that is, to determine the probability that the user's intention exceeds the intention threshold. At the same time, the discriminant model has the same variables as the intention model described above, that is, it processes the same type of features. For example, see Figure 6 , Figure 6 This is a schematic diagram of the model structure of the discrimination model provided in the embodiment of the present application, based on Figure 6The discriminant model includes a feature extraction layer and a prediction layer. Therefore, the process of judging the probability that the user's intention exceeds the intention threshold can be, for each user, inputting the user's attribute information into the discriminant model, and extracting the user's features based on the user's attribute information through the feature extraction layer of the discriminant model to obtain user features; through the prediction layer of the discriminant model, based on the user features, predicting the probability that the user's intention exceeds the intention threshold, and obtaining the probability that the user's intention exceeds the intention threshold; wherein the user's attribute information can also be the user's age, gender, occupation, number of likes within 30 days, number of searches within 30 days, etc.
[0071] It should be noted that after removing a feature, use all the remaining features to retrain the same discriminant model, and then compare the performance changes of the two discriminant models before and after removing the feature. If the model performance decreases significantly after removing the corresponding feature, it indicates that the feature is very important for model prediction.
[0072] Step 202 : Based on the importance of each feature, select features whose importance is greater than or equal to an importance threshold from the user's features as important features.
[0073] It should be noted that the importance threshold here can be set, and the embodiments of the present application do not limit this; for example, after determining the importance of each feature, multiple features are sorted based on the importance, and based on the sorting results, a target number of features such as the top 20% are selected. Here, the importance threshold can be the minimum importance among the importance corresponding to the target number of features.
[0074] It should be noted that the important features here refer to important feature types. Therefore, for different users, the important features are the same, that is, the important feature types are the same, but the values of the important features are different, that is, the feature values under the important feature types are different; for example, the user's feature types include the user's age, gender, occupation, number of likes within 30 days, and number of searches within 30 days. Here, the important features can be the number of likes and searches of the user within 30 days, and the corresponding feature values are the specific number of likes and searches of the user within 30 days.
[0075] Step 203: Determine a first feature of the first user and a second feature of the second user from the important features.
[0076] It should be noted that, since the important features are selected based on the features of the users, determining the first feature of the first user and the second feature of the second user from the important features is equivalent to using the important feature of the first user as the first feature and the important feature of the second user as the second feature; at the same time, the number of the first features can be one or more, and the number of the second features can also be one or more;
[0077] Continuing with the above example, the user's feature types include the user's age, gender, occupation, number of likes within 30 days, and number of searches within 30 days. The important features here can be the number of likes and searches within 30 days. Then the first feature of the first user is the number of likes and searches of the first user within 30 days; and the second feature of the second user is the number of likes and searches of the second user within 30 days.
[0078] In this way, important features are selected from multiple features of the user, and then a subsequent user selection process is performed based on the important features. This can not only improve the efficiency of the subsequent user selection process, but also improve the accuracy of the selected users.
[0079] In actual implementation, for the process of generating a first user instance based on the first feature of the first user in the first user group, there are multiple ways to generate the first user instance based on the first feature of the first user in the first user group. Next, taking several of them as examples, the process of generating the first user instance based on the first feature of the first user in the first user group is explained.
[0080] In some embodiments, there are multiple first features, and different first features have different feature types; thus, the process of generating a first user instance based on the first feature of the first user in the first user group can be to perform multiple selection processes for the first feature of the first user to obtain multiple combined features; wherein, for each selection process, the following operations are performed to obtain a combined feature: for each feature type, the first feature of each first user under the feature type is determined, and a first feature is selected from the determined multiple first features; the first features selected under each feature type are combined to obtain a combined feature; wherein the feature type corresponding to the combined feature is the same as the feature type of the first feature; and the first user instance is determined based on the multiple combined features.
[0081] It should be noted that the number of first features is one or more. When the number of first features is multiple, different first features correspond to different feature types. For each feature type, the first feature of each first user under the feature type is determined, thereby determining the feature set under each feature type. One feature type corresponds to one feature set, and one feature set includes the feature values of all first users under this feature type; then, in each selection process, a first feature is randomly selected from each feature set, thereby using the selected multiple first features as a combined feature. At the same time, since multiple selection processes are performed, multiple combined features are obtained; wherein, the feature type corresponding to the combined feature is the same as the feature type of the first feature, which means that the number of feature types corresponding to the features included in the combined feature is the same as the number of feature types of the first feature, that is, the number of features included in the combined feature is the same as the number of first features.
[0082] For example, the feature type of the first feature includes the number of likes and search numbers of the user within 30 days. There are 3 first users. The number of likes of the first user A within 30 days is 20, and the number of searches within 30 days is 18; the number of likes of the first user B within 30 days is 1, and the number of searches within 30 days is 1; the number of likes of the first user C within 30 days is 5, and the number of searches within 30 days is 0. Then, for each feature type, the first feature of each first user under the feature type is determined, that is, the number of likes within 30 days includes 20, 1, and 5, and the number of searches within 30 days is 18, 1, and 0. Then, one first feature is selected from the multiple first features determined under different feature types, and the first feature selected under each feature type is selected. The first feature is taken and combined to obtain the combined feature, and 9 combined features can be obtained, namely, {number of likes within 30 days = 20, number of searches within 30 days = 18}, {number of likes within 30 days = 20, number of searches within 30 days = 1}, {number of likes within 30 days = 20, number of searches within 30 days = 0}, {number of likes within 30 days = 1, number of searches within 30 days = 1}, {number of likes within 30 days = 1, number of searches within 30 days = 18}, {number of likes within 30 days = 1, number of searches within 30 days = 0}, {number of likes within 30 days 5, number of searches within 30 days = 18}, {number of likes within 30 days = 30 days = 1, number of searches within 30 days = 0}, {number of likes within 30 days 5, number of searches within 30 days = 1}, {number of likes within 30 days 5, number of searches within 30 days = 0}.
[0083] It should be noted that the number of selection processes can be pre-set, as long as it is less than P n times, where P is the number of first users and n is the number of first features.
[0084] In actual implementation, in the process of determining the first user instance based on multiple combined features, since the combined feature is generated based on the first feature, and the first user includes other features in addition to the first feature, therefore, in the process of determining the first user instance based on the multiple combined features, for each first user, starting from the first combined feature in the multiple combined features, the first feature of the first user is replaced with the corresponding combined feature in sequence, completing one replacement, and based on the feature formed after the replacement (equivalent to the splicing feature of the corresponding combined feature and the above-mentioned other features), a first user instance is determined. In this way, the replacement process is performed as many times as there are combined features, thereby obtaining multiple first user instances;
[0085] Among them, for each first user, the number of combined features is the same as the number of corresponding first user instances; and for the process of determining the first user instance based on multiple combined features, the first user instance determined here includes the first user instances corresponding to multiple first users. For example, if k first user instances are generated for each first user based on k combined features, then the process of determining the first user instance based on multiple combined features is to use the k first user instances corresponding to multiple first users as the first user instance determined above, where k is an integer greater than 1.
[0086] It should be noted that the user instance is not an actual user, but a virtual user, which is equivalent to indicating a new set of features. For example, as mentioned above, for each first user, starting from the first combined feature of multiple combined features, the first feature of the first user is replaced with the corresponding combined feature in turn, thereby obtaining multiple first user instances. Since there are multiple combined features, after replacing the first features respectively, a combined feature and the remaining features will form a new set of features, and this new set of features will be regarded as a user instance.
[0087] Continuing with the above example, the user's feature types include the user's age, gender, occupation, number of likes within 30 days, and number of searches within 30 days. The first feature is the number of likes within 30 days and the number of searches within 30 days. There are 9 combined features and the number of first users is 3. For the first user A, the first user A's age of 18, gender of female, and occupation of student are combined with these 9 combined features to form 9 new feature groups, namely {age = 18, gender = female, occupation = student, number of likes within 30 days = 20, number of searches within 30 days = 18}, {age = 18, gender = female, occupation = student, number of likes within 30 days = 20, number of searches within 30 days = 1}, {age = 18, gender = female, occupation = student, number of likes within 30 days = 20, number of searches within 30 days = 0}, {age = : {Age = 18, Gender = Female, Occupation = Student, Number of Likes in 30 Days = 1, Number of Searches in 30 Days = 1}, {Age = 18, Gender = Female, Occupation = Student, Number of Likes in 30 Days = 1, Number of Searches in 30 Days = 18}, {Age = 18, Gender = Female, Occupation = Student, Number of Likes in 30 Days = 1, Number of Searches in 30 Days = 0}, {Age = 18, Gender = Female, Occupation = Student, Number of Likes in 30 Days = 5, Number of Searches in 30 Days = 18}, {Age = 18, Gender = Female, Occupation = Student, Number of Likes in 30 Days = 5, Number of Searches in 30 Days = 1}, {Age = 18, Gender = Female, Occupation = Student, Number of Likes in 30 Days = 5, Number of Searches in 30 Days = 0}, thus taking these 9 groups of new features as the 9 first user instances corresponding to the first user A;
[0088] Each first user can correspond to 9 first user instances, and there are 3 first users here, so the final number of first user instances is 27, that is, the number of multiple first user instances generated here based on the first features of different first users is 27.
[0089] Similarly, the process of generating a second user instance based on the second feature of the second user in the second user group is similar to the process of generating a first user instance based on the first feature of the first user in the first user group described above, and is not repeated here in the embodiments of the present application.
[0090] In this way, a feature is randomly selected under different feature types, and the selected features are combined to obtain combined features. Then, based on the combined features, multiple user instances are generated. In this way, by combining different types of features, the diversity of user features is increased, which makes it easier to determine the degree of influence of features on intentionality based on user instances in the subsequent process.
[0091] In other embodiments, the number of first features is n, where n is a positive integer, and thus the process of generating a first user instance based on the first feature of the first user in the first user group may be, for each first user, performing k update processes on the first feature of the first user to obtain k first user instances; wherein, for each update process, the following operations are performed to obtain a first user instance: the m first features of the first user are updated to m new first features to obtain multiple first user instances, where m is a positive integer less than or equal to n, and k is an integer greater than 1.
[0092] Among them, updating the first feature refers to updating the feature value of the first feature, and the m corresponding to different update processes may be different or the same. When the m corresponding to different update processes are the same, the feature values of the first feature of the first user instance obtained by different update processes are different, which is equivalent to the new first features obtained by different update processes being different. At the same time, when the m corresponding to different update processes are the same, m is equal to n; among them, when the m corresponding to different update processes are different, in multiple update processes, the m corresponding to some update processes may be the same, and the m corresponding to other update processes may be different. This is not limited in the embodiments of the present application.
[0093] For example, the features of a first user are {age = 18, gender = female, occupation = student, number of likes within 30 days = 20, number of searches within 30 days = 18}, and the first features are age, number of likes within 30 days, and number of searches within 30 days, that is, n is 3. Here, if m is the same, the three first features of the first user need to be updated, that is, the age, number of likes within 30 days, and number of searches within 30 days of the first user are updated respectively. For example, the features of the first user {age = 18, gender = female, occupation = student, number of likes within 30 days = 20, number of searches within 30 days = 18} are updated to {age = 20, gender = female, occupation = student, number of likes within 30 days = 18, number of searches within 30 days = 1}. In this way, k update processes are performed to obtain k groups of new features corresponding to the first user, that is, to obtain k first user instances;
[0094] If m is different, then in the first update process, one of the first characteristics of the first user can be updated, that is, the number of likes of the first user within 30 days can be updated respectively. For example, the characteristics of the first user are {age = 18, gender = female, occupation = student, number of likes within 30 days = 20, number of searches within 30 days = 18}, which are updated to {age = 18, gender = female, occupation = student, number of likes within 30 days = 18, number of searches within 30 days = 18}; in the second update process, two of the first characteristics of the first user can be updated, that is, the number of likes and the number of searches within 30 days of the first user can be updated respectively. For example, the characteristics of the first user are {age = 18, gender = female, occupation = student, number of likes within 30 days = 20, 3 0 days = 18}, updated to {age = 18, gender = female, occupation = student, number of likes within 30 days = 18, number of searches within 30 days = 1}; in the third update process, the three first features of the first user can be updated, that is, the age of the first user, the number of likes within 30 days and the number of searches within 30 days are updated respectively. For example, the features of the first user are {age = 18, gender = female, occupation = student, number of likes within 30 days = 20, number of searches within 30 days = 18}, updated to {age = 20, gender = female, occupation = student, number of likes within 30 days = 18, number of searches within 30 days = 1}; perform k update processes in this way, thereby obtaining k groups of new features corresponding to the first user, that is, obtaining k first user instances.
[0095] At the same time, each first user can obtain k first user instances. For each first user, after obtaining k first user instances, the k first user instances corresponding to the plurality of first users are used as the first user instances determined above.
[0096] Similarly, the process of generating a second user instance based on the second feature of the second user in the second user group is similar to the process of generating a first user instance based on the first feature of the first user in the first user group described above, and is not repeated here in the embodiments of the present application.
[0097] In this way, for each first user, the first feature of the first user is updated multiple times, thereby generating multiple user instances. In this way, by updating the features, the diversity of user features is increased, which facilitates determining the degree of influence of the features on the intention based on the user instances in the subsequent process.
[0098] Step 103: Determine a first degree of influence of the first feature on the intentionality based on the first user instance, and determine a second degree of influence of the second feature on the intentionality based on the second user instance.
[0099] It should be noted that the degree of influence of a feature on intention refers to the correlation between the feature and the user's decision (that is, whether the user executes the first event) when determining the user's intention, that is, the weight of the corresponding feature when predicting the intention; for example, the higher the degree of influence of a feature on intention, the greater the impact of the feature change on the prediction result, which is equivalent to the prediction result of the intention being more dependent on the corresponding feature.
[0100] In actual implementation, multiple first user instances can be generated for each first user, and the multiple first user instances here are multiple first user instances corresponding to multiple first users, that is, a user instance set generated based on the multiple first user instances corresponding to each first user, so as to determine the first influence of the first feature on the intentionality based on the first user instance, that is, based on the multiple first user instances corresponding to each first user, determine the influence of the first feature of the corresponding first user on the intentionality of the first user; wherein each first user corresponds to a first influence degree, that is, there is a one-to-one correspondence between the first user and the first influence degree;
[0101] Similarly, multiple second user instances can be generated for each second user, and the multiple second user instances here are multiple second user instances corresponding to multiple second users, that is, a user instance set generated based on the multiple second user instances corresponding to each second user, so that the second influence of the second feature on the intention is determined based on the multiple second user instances, that is, based on the multiple second user instances corresponding to each second user, the influence of the second feature of the corresponding second user on the intention of the second user is determined; wherein, each second user corresponds to a second influence degree, that is, there is a one-to-one correspondence between the second user and the second influence degree.
[0102] In actual implementation, see Figure 7 , Figure 7 This is a flow chart of determining the first impact level provided by the embodiment of the present application, based on Figure 7 The number of the first user instances is multiple, so the process of determining the first influence degree of the first feature on the intention degree based on the first user instance can be achieved through the following steps.
[0103] Step 1031 : For each first user instance, determine the intention of the first user instance with respect to the first event.
[0104] In actual implementation, as described above, the intention of the first user instance for the first event is also the probability that the first user instance executes the first event, and the intention of the first user instance for the first event can be predicted by the trained intention model. This embodiment of the present application will not go into details.
[0105] Step 1032 : averaging the intentional degrees of the multiple first user instances to obtain an average intentional degree of the multiple first user instances.
[0106] In actual implementation, the process of averaging the intentional degrees of multiple first user instances to obtain the average intentional degree of multiple first user instances can be to sum the intentional degrees of multiple first user instances to obtain the sum of the intentional degrees; and use the ratio of the sum of the intentional degrees to the number of first user instances as the average intentional degree.
[0107] Step 1033 : Determine the dispersion degree of the intentional degrees of the plurality of first user instances relative to the average intentional degree, and use the dispersion degree as the first influence degree.
[0108] In actual implementation, the process of determining the dispersion of the intentional degrees of multiple first user instances relative to the average intentional degree may be: subtracting the intentional degree of each first user instance from the average intentional degree to obtain an intentional degree difference; summing the squares of the intentional degree differences of each first user instance to obtain a sum of the intentional degrees; determining a ratio of the sum of the intentional degrees to the number of first user instances, and performing square root processing on the ratio to obtain the dispersion degree, that is:
[0109]
[0110] Among them, σ U is the degree of dispersion, n is the number of first user instances, S D(i) is the intention of the first user instance i, S v is the average intention.
[0111] It should be noted that, since the first features of different first user instances are different, that is, the feature values of the first features of different first user instances are different, if the dispersion of the intentions determined based on different first features relative to the average intention is very high, then it means that the difference in intentions determined based on different first features will be very large, that is, the difference in the first features has a great impact on the intentions, and thus the degree of influence of the first feature on the intentions will also be very large, which is equivalent to reflecting the size of the influence of the first feature on the intentions based on the different impacts on the fluctuations of the intentions. Therefore, the dispersion of the intentions of multiple first user instances relative to the average intention can be used as the first degree of influence.
[0112] Similarly, based on the second user instance, the process of determining the second degree of influence of the second feature on the intention degree may be, for each second user instance, determining the intention degree of the second user instance for the first event; averaging the intention degrees of multiple second user instances to obtain the average intention degree of multiple second user instances; determining the degree of dispersion of the intention degrees of multiple second user instances relative to the corresponding average intention degree, and using the degree of dispersion as the second degree of influence; wherein, the process of determining the degree of dispersion of the intention degrees of multiple second user instances relative to the average intention degree is similar to the process of determining the degree of dispersion of the intention degrees of multiple first user instances relative to the average intention degree as described above, and this embodiment of the present application will not be elaborated on.
[0113] In this way, the degree of influence of the features on the intentionality is determined based on the fluctuation of the intentionality caused by the differences in the features, and the influence of the features on the intentionality is quantified, thereby improving the accuracy of determining the degree of influence of the corresponding features on the intentionality.
[0114] Step 104 : Based on the first influence level and the second influence level, select a third user from the second user group whose influence level is greater than an influence level threshold.
[0115] In actual implementation, as described above, there is a one-to-one correspondence between the first degree of influence and the first user, and there is a one-to-one correspondence between the second degree of influence and the second user; thus, based on the first degree of influence and the second degree of influence, the process of selecting a third user whose degree of influence is greater than the influence degree threshold from the second user group can be, for example, summing the first degree of influence of each first user to obtain the sum of the influence degrees; determining the ratio of the sum of the influence degrees to the number of first users, and determining the influence degree threshold based on the ratio; and selecting a second user whose corresponding second degree of influence is greater than the influence degree threshold from the second users as the third user.
[0116] It should be noted that the process of determining the impact threshold based on the ratio may be to multiply the ratio by a preset first coefficient to obtain the impact threshold, that is:
[0117]
[0118] Where ψ is the influence threshold, Q is the number of first users, σ i is the first influence degree of the i-th first user, α is a preset first coefficient, and here, the first coefficient generally takes a value between 0 and 1, and here it can be 0.7.
[0119] In actual implementation, after determining the influence degree threshold, the second influence degree of each second user is compared with the influence degree threshold to obtain a comparison result; from multiple comparison results, the comparison result whose corresponding second influence degree is greater than the influence degree threshold is selected, and the second user corresponding to the selected comparison result is used as the third user.
[0120] It should be noted that if the second influence level of the second user is greater, it means that the second user's intention is more affected by the fluctuation of the important feature, that is, the second feature, that is, the other remaining features (that is, the features other than the second feature) have a relatively weaker decisive effect on the second user's intention. At the same time, the influence level threshold is determined based on the first influence level of the first user. Therefore, the influence level threshold can be considered as the average influence level of the important feature, that is, the first feature, on the first user. If the second influence level of the third user is greater than the influence level threshold, it means that the influence level of the important feature on the third user's intention is greater than the influence level of the important feature on the intention levels of most first users, that is, the influence level of the other remaining features on the third user's intention is less than the influence level of the other remaining features on the intention levels of most first users.
[0121] Since the first user has been identified as the target user based on the important features and other remaining features, it can be considered that relative to the first user, the important features and other remaining features have been fully analyzed and utilized, and the influence of the other remaining features on the third user's intention is less than the influence of the other remaining features on the intention of most first users. Therefore, it can be considered that for the third user, the other remaining features have not been fully analyzed and utilized; in short, the influence of the important features on the third user is too large, that is, greater than the influence threshold, which is equivalent to obtaining the third user's intention only by relying on the important features. Therefore, the other remaining features have no influence on the third user's intention, so it can be considered that the other remaining features have not been fully analyzed and utilized; based on this, the third user can be regarded as a potential user, so that the user can be further accurately identified in the subsequent process to determine whether the user will execute the first event.
[0122] In this way, after determining the influence level threshold based on the first influence level of the first user, a second user whose second influence level is greater than the influence level threshold is selected from the second users as the third user, that is, the potential user. In this way, not only can users be selected comprehensively and the omission of target users be avoided, but potential users can also be located more simply and effectively, reducing the waste of promotion resources and improving the conversion rate and return on investment of promotion activities.
[0123] In some embodiments, as described above, the user's intention for the first event is predicted by a trained intention model; thus, based on the first influence level and the second influence level, after selecting a third user whose influence level is greater than the influence level threshold from the second user group, it is also possible to determine the execution result of the third user for the first event, and the execution result is used to indicate whether the third user executes the first event; based on the execution result, determine the label of the third user, and the label is used to indicate the probability of the third user executing the first event; and train the intention model again based on the label and the third user's intention.
[0124] It should be noted that the execution result of the third user for the first event refers to whether the third user has executed the first event. For example, after the third user is selected, promotional information such as advertisements of the first event can be sent to the third user. When feedback information indicating that the third user has executed the first event is received, the execution result indicating that the third user has executed the first event is determined; when no feedback information indicating that the third user has executed the first event is received within a target period, such as one month, the execution result indicating that the third user has not executed the first event is determined; or, the staff can conduct real-life promotion for the third user, such as telephone calls, and then obtain the promotion results, i.e., the execution results, uploaded by the staff for the third user. The promotion results are used to indicate that the third user has executed the first event, or to indicate that the third user has not executed the first event.
[0125] In actual implementation, after determining the execution result of the third user for the first event, in the process of determining the label of the third user, if the third user executes the first event, the probability of the third user executing the first event is determined to be 1, that is, the label of the third user is determined to be 1; if the third user does not execute the first event, the probability of the third user executing the first event is determined to be 0, that is, the label of the third user is determined to be 0;
[0126] Then, based on the label and the intention of the third user, the process of re-training the intention model is carried out, that is, obtaining the difference between the label and the intention of the third user, and re-training the intention model based on the difference; wherein, the intention of the third user is obtained by predicting the third user based on the intention model. As described above, the embodiment of the present application will not be elaborated here.
[0127] It should be noted that, since the intention model predicts the intention of the third user for the first time, the predicted intention is less than the intention threshold, that is, the third user is predicted to not perform the first event, but the third user is a potential user. Therefore, based on whether the third user performs the first event, the label of the third user is determined, and then the intention model is trained again based on the label and the intention of the third user; in this way, by using the actual behavior of the third user (whether to perform the first event) as a label to train the model again, not only the prediction accuracy of the model is further improved, but also more accurate and specific behavior data are used to train the model again, and the generalization ability of the model can be enhanced.
[0128] In the above embodiment of the present application, after grouping users based on their intentions for the first event to obtain a first user group whose intentions are greater than or equal to the intention threshold and a second user group whose intentions are less than the intention threshold, a first user instance is determined based on the first feature, and a second user instance is determined based on the second feature. Then, based on the first user instance, a first influence of the first feature on the intention is determined, and based on the second user instance, a second influence of the second feature on the intention is determined. Thus, based on the first influence and the second influence, a third user whose influence is greater than the influence threshold is selected from the second user group. In this way, compared to the scheme of selecting users based only on intention, multiple user instances are generated based on the characteristics of different users, which enriches the diversity of features, thereby facilitating the determination of the influence of the characteristics on the intention. Then, for users with low intention, based on the influence, users are selected from users with low intention. In this way, for users with low intention, user selection can be performed in combination with the influence of the characteristics on the intention, making the user selection process more in-depth and comprehensive, avoiding the omission of users who should be selected, and achieving the effect of selecting potential users from users with low intention.
[0129] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0130] During the research process, it was discovered that machine learning models are widely used in various user stratification and decision-making processes. For example, machine learning models are used to predict user purchasing behavior and obtain their purchase intention scores. Users with purchase intention scores exceeding a threshold are then designated as target users, while all other users are designated as non-target users. Promotion operations are then performed only on the target users. Therefore, only the target users have actual performance data on promotion operations. Since non-target users were not promoted, there is no performance data on these users. Therefore, the model cannot continuously adjust and correct prediction biases by collecting actual performance data. Consequently, because the model has not fully explored some users, it assigns these users scores below the threshold, resulting in the system rejecting them. The model is then unable to continue learning based on the actual performance data of these rejected users until it has fully explored them.
[0131] In practical applications, commonly used methods for achieving a good balance between exploration and exploitation include greedy-based methods and upper confidence bound-based methods. However, these methods do not describe how to distinguish between rejected samples that are truly underexplored and those that have been fully explored by the model but rejected due to low decision scores. They often assume that all rejected samples are underexplored and in the "exploration state." This assumption is often false in real-world problems. Samples that are already in the "exploration state" but rejected simply due to low decision scores may account for a significant proportion of all rejected samples and cannot be ignored. Consequently, continuous learning schemes based on the assumption that all rejected samples represent "exploration state" samples may not only be inefficient, but also, due to the high proportion of "exploitation state" samples among rejected samples, even though the model has spent a lot of effort to re-explore them, it may never fully explore the samples that truly need to be explored.
[0132] Based on this, an embodiment of the present application provides a user selection method, which selects users whose model scores are too low due to insufficient model exploration in the past from rejected users, so as to complete the continued exploration of these users for whom the model exploration was insufficient, that is, it is possible to select potential users from the rejected users.
[0133] In actual implementation, assuming that in a promotion scenario, the strategy currently being implemented is to use a machine learning model M 0 意向度 Intention model: Each user in the entire marketable population Q (multiple users) is first scored for their intention to purchase a product (the first event). Users whose intention exceeds a pre-set threshold (intention threshold) will receive further promotional actions.
[0134] Among them, suppose that the set of users whose intention degree exceeds the threshold in Q is Q + (first user group), and the user set below the threshold is Q - (Second user group), the goal is to - Find a sub-user group Q that the current model M has not fully explored - 探索 (third user), then in the future you can - 探索 The user in designs an exploration strategy, such as - 探索 Take a random sample and then take the sum of Q for each person in the sample. + The same promotion action as the user in the Q is then used to retrain the model based on the feedback to complete the Q - 探索 Full exploration of potential customer groups.
[0135] The following is how to - Find Q - 探索 Provide explanation.
[0136] The first step is to train a linear machine learning discriminant model M 1 判别 Model (discriminant model), where the variables of this model are used and model M 0 意向度 The same variables in the model, where M 1 判别 The model is used to identify the probability that the user's intention exceeds the intention threshold, that is, to indicate whether the user's intention exceeds the intention threshold.
[0137] The second step is to take the model M 1 判别 Perform univariate importance analysis on each variable feature in the variables in M. 1 判别 Remove the single variable feature X from the variable set; then, use all the remaining variable features except X to retrain the same linear discriminant model M -X 判别 , and then compare the model M before and after removing the variable X 1 判别 and Model M -X 判别 If the model performance drops significantly after removing the X variable, it indicates that the variable is very important for the model prediction. The degree of performance drop can be used to quantify the importance of each variable. Assuming that the model M 1判别 There are n variables X1, X2, ..., X n , their corresponding importance values (importance) are Z x1 , Z x2 ,…,Z xn .
[0138] The third step is to put the 20% variables (important features) with the highest importance values calculated in the previous step into the set R. Then R contains [n*20%] variables in total. Assume that they are {X1, X2, ..., X n*20%}; At the same time, let M 1 判别 The set of all remaining variables except the variables in set R is ~R. + Randomly sample 100 users from the dataset (first user group) and get different value combinations of variables in the set R, denoted as {D {1} , D {2} ,…,D {100}};
[0139] For example, R = {number of likes in 30 days, number of searches in 30 days}, ~R = {user gender, user age, user occupation}, then {D {1} ={30-day likes = 20, 30-day searches = 18}D {2} ={30-day likes = 1, 30-day searches = 1}, ...D {100} ={30-day likes = 5, 30-day searches = 0}.
[0140] Step 4: Through Q + The data in the dataset is used to determine the threshold of the disturbance score (the threshold of the degree of influence). The specific method is to + Each user U (first user) in the dataset, for example, first user U = {user gender = male, user age = 43, user occupation = employee, number of likes in 30 days = 10, number of searches in 30 days = 12}, according to {D {1} , D {2} ,…,D {100}}, expand it to generate 100 new user instances (first user instances), for example, generate instance U D(1) ={user gender = male, user age = 43, user occupation = employee, number of likes in 30 days = 20, number of searches in 30 days = 18}, generate instance U D(2) ={user gender = male, user age = 43, user occupation = employee, number of likes in 30 days = 1, number of searches in 30 days = 1}; ...; Generate instance U D(100)={user gender = male, user age = 43, user occupation = employee, 30-day likes = 5, 30-day search count = 0}; this is equivalent to replacing the user's 30-day likes and 30-day search count with the instance D extracted above, respectively, to obtain 100 new user instances;
[0141] Then, for each user U, 100 instances U are generated. D(1) , U D(2) ,…,U D(100) , call model M 0 意向度 For each generated instance, a corresponding intention degree is predicted, i.e., s D(1) , s D(2) ,…,s D(100) , and then calculate the dispersion of this set of scores for this user U, that is:
[0142]
[0143] Next, based on the above formula (2), the disturbance score threshold is determined.
[0144] Step 5: From Q - (Second user group) determines Q - 探索 Sub-customer groups. Specifically, for Q - Each user U (second user) in the data set, according to {D {1} , D {2} ,…,D {100}}, expand it to generate 100 new user instances (second user instances); then for each instance U generated D(1) , U D(2) ,…,U D(100) , call model M 0 意向度 For each generated instance, a corresponding intention degree is predicted, i.e., s D(1) , s D(2) ,…,s D(100) , combined with the above formula (3), calculate the dispersion degree of this set of scores for this user U.
[0145] It should be noted that σ U The larger the value of , the more the user's intention score is affected by the fluctuation of the features in set R. That is, the other remaining features (features in set ~R) play a relatively low role in determining the user's intention score, which is equivalent to these features in set ~R not being fully "utilized". Therefore, the user should be marked as a user who needs to be "explored".
[0146] Based on this, if σU If the value is greater than the disturbance threshold, user U is placed in Q - 探索 In the sub-customer group, because in this case, among all the features of the current user U, the features in set ~R are better than those in set R in M 0 意向度 The degree of determination of user intention in the model is greater than that of Q + The average level of all known users in the "exploitation" state is lower, so they should be classified as "exploration" users.
[0147] In the above embodiment of the present application, after grouping users based on their intentions for the first event to obtain a first user group whose intentions are greater than or equal to the intention threshold and a second user group whose intentions are less than the intention threshold, a first user instance is determined based on the first feature, and a second user instance is determined based on the second feature. Then, based on the first user instance, a first influence of the first feature on the intention is determined, and based on the second user instance, a second influence of the second feature on the intention is determined. Thus, based on the first influence and the second influence, a third user whose influence is greater than the influence threshold is selected from the second user group. In this way, compared to the scheme of selecting users based only on intention, multiple user instances are generated based on the characteristics of different users, which enriches the diversity of features, thereby facilitating the determination of the influence of the characteristics on the intention. Then, for users with low intention, based on the influence, users are selected from users with low intention. In this way, for users with low intention, user selection can be performed in combination with the influence of the characteristics on the intention, making the user selection process more in-depth and comprehensive, avoiding the omission of users who should be selected, and achieving the effect of selecting potential users from users with low intention.
[0148] The following continues to describe the exemplary structure of the user selection device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the user selection device 455 of the memory 450 may include:
[0149] a grouping module 4551 configured to divide the users into a first user group and a second user group based on the users' intentions for the first event, wherein the intention of a first user in the first user group is greater than or equal to an intention threshold, and the intention of a second user in the second user group is less than the intention threshold;
[0150] a generating module 4552 configured to generate a first user instance based on a first feature of a first user in the first user group, and to generate a second user instance based on a second feature of a second user in the second user group;
[0151] a determination module 4553 configured to determine, based on the first user instance, a first degree of influence of the first feature on the intentionality, and to determine, based on the second user instance, a second degree of influence of the second feature on the intentionality;
[0152] The selection module 4554 is configured to select, from the second user group, a third user whose influence level is greater than an influence level threshold based on the first influence level and the second influence level.
[0153] In some embodiments, the device also includes an analysis module, which is used to perform importance analysis on the user's features to obtain the importance of each feature; based on the importance of each feature, select features whose importance is greater than or equal to an importance threshold from the user's features as important features; and determine the first feature of the first user and the second feature of the second user from the important features.
[0154] In some embodiments, the analysis module is further used to perform the following processing for each of the multiple features of the user to obtain the importance of the feature: remove the corresponding feature from the multiple features of the user to obtain a third feature; based on the third feature, predict the probability that the user's intention exceeds the intention threshold; based on the probability, determine the importance of the feature.
[0155] In some embodiments, the number of the first features is multiple, and the feature types of different first features are different; the generation module 4552 is also used to perform multiple selection processes for the first features of the first user to obtain multiple combined features; wherein, for each selection process, the following operations are performed to obtain a combined feature: for each feature type, the first feature of each first user under the feature type is determined, and one first feature is selected from the multiple determined first features; the first features selected under each feature type are combined to obtain a combined feature; wherein the feature type corresponding to the combined feature is the same as the feature type of the first feature; based on the multiple combined features, the first user instance is determined.
[0156] In some embodiments, the number of the first features is n, where n is a positive integer. The generation module 4552 is further used to perform k update processes on the first features of each first user to obtain k first user instances; wherein, for each update process, the following operations are performed to obtain a first user instance: the m first features of the first user are updated to m new first features to obtain a first user instance, where m is a positive integer less than or equal to n, and k is an integer greater than 1.
[0157] In some embodiments, there are multiple first user instances, and the determination module 4553 is further used to determine, for each first user instance, the intention of the first user instance towards the first event; average the intentions of multiple first user instances to obtain the average intention of the multiple first user instances; determine the degree of dispersion of the intentions of the multiple first user instances relative to the average intention, and use the dispersion degree as the first influence degree.
[0158] In some embodiments, the determination module 4553 is further used to calculate the difference between the intentionality of each first user instance and the average intentionality to obtain an intentionality difference; sum the squares of the intentionality differences of each first user instance to obtain a sum of intentionality; determine the ratio of the sum of intentionality to the number of first user instances, and perform square root processing on the ratio to obtain the degree of dispersion.
[0159] In some embodiments, there is a one-to-one correspondence between the first influence degree and the first user, and there is a one-to-one correspondence between the second influence degree and the second user; the selection module 4554 is further used to sum the first influence degree of each first user to obtain the sum of the influence degrees; determine the ratio of the sum of the influence degrees to the number of the first users, and determine the influence degree threshold based on the ratio; from the second users, select a second user whose corresponding second influence degree is greater than the influence degree threshold as the third user.
[0160] In some embodiments, the user's intention for the first event is predicted by a trained intention model; the device also includes a training module, which is used to determine the execution result of the third user for the first event, and the execution result is used to indicate whether the third user executes the first event; based on the execution result, the label of the third user is determined, and the label is used to indicate the probability of the third user executing the first event; based on the label and the intention of the third user, the intention model is trained again.
[0161] The embodiment of the present application provides a computer program product, which includes computer executable instructions or a computer program, which is stored in a computer-readable storage medium. A processor of an electronic device reads the computer executable instructions or the computer program from the computer-readable storage medium, and the processor executes the computer executable instructions or the computer program, so that the electronic device performs the user selection method described above in the embodiment of the present application, for example, Figure 3 User selection method shown.
[0162] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the user selection method provided by the embodiment of the present application, for example, Figure 3 User selection method shown.
[0163] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0164] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0165] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0166] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0167] It should be noted that in the embodiments of the present application, when it comes to obtaining user attribute information and other related data, when the embodiments of the present application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0168] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A user selection method, characterized in that: The method comprises: Dividing the users into a first user group and a second user group based on their intentions for the first event, wherein the intention of a first user in the first user group is greater than or equal to an intention threshold, and the intention of a second user in the second user group is less than the intention threshold; generating a first user instance based on a first feature of a first user in the first user group, and generating a second user instance based on a second feature of a second user in the second user group; Determining a first degree of influence of the first feature on the intentionality based on the first user instance, and determining a second degree of influence of the second feature on the intentionality based on the second user instance; Based on the first influence level and the second influence level, a third user whose influence level is greater than an influence level threshold is selected from the second user group.
2. The method according to claim 1, characterized in that Before generating the first user instance based on the first feature of the first user in the first user group, the method further includes: Performing importance analysis on the user's features to obtain the importance of each feature; Based on the importance of each of the features, selecting features whose importance is greater than or equal to an importance threshold from the features of the user as important features; A first feature of the first user and a second feature of the second user are determined from the important features.
3. The method according to claim 2, characterized in that The importance analysis of the user's features to obtain the importance of each feature includes: For each of the multiple features of the user, perform the following processing to obtain the importance of the feature: removing the corresponding feature from the multiple features of the user to obtain a third feature; Based on the third feature, predicting a probability that the user's intention exceeds the intention threshold; Based on the probability, the importance of the feature is determined.
4. The method according to claim 1, wherein There are multiple first features, and different first features have different feature types; and generating a first user instance based on the first feature of the first user in the first user group includes: Performing multiple selection processes on the first feature of the first user to obtain multiple combined features; For each selection process, the following operations are performed to obtain a combined feature: For each of the feature types, determining a first feature of each of the first users under the feature type, and selecting one first feature from the multiple determined first features; combining the first features selected under each of the feature types to obtain a combined feature; wherein the feature type corresponding to the combined feature is the same as the feature type of the first feature; Based on the multiple combined features, the first user instance is determined.
5. The method according to claim 1, wherein There are multiple first user instances, and determining a first influence of the first feature on the intention based on the first user instances includes: For each of the first user instances, determining the intention of the first user instance with respect to the first event; Averaging the intentional degrees of the multiple first user instances to obtain an average intentional degree of the multiple first user instances; Determine the dispersion degree of the intentional degrees of the multiple first user instances relative to the average intentional degree, and use the dispersion degree as the first influence degree.
6. The method according to claim 5, characterized in that Determining the dispersion degree of the intentions of the plurality of first user instances relative to the average intention includes: Subtracting the intentionality of each first user instance from the average intentionality to obtain an intentionality difference; Summing the squares of the intentionality differences of each of the first user instances to obtain a sum of intentionality; A ratio of the sum of the intentional degrees to the number of the first user instances is determined, and a square root is performed on the ratio to obtain the degree of dispersion.
7. The method according to claim 1, characterized in that There is a one-to-one correspondence between the first influence level and the first user, and there is a one-to-one correspondence between the second influence level and the second user; The selecting, from the second user group, a third user whose influence degree is greater than an influence degree threshold based on the first influence degree and the second influence degree, includes: summing the first influence levels of each of the first users to obtain a sum of influence levels; determining a ratio of the sum of the influence levels to the number of the first users, and determining the influence level threshold based on the ratio; A second user whose corresponding second influence degree is greater than the influence degree threshold is selected from the second users as the third user.
8. The method according to claim 1, characterized in that The user's intention for the first event is predicted by the trained intention model; After selecting, based on the first influence level and the second influence level, a third user whose influence level is greater than an influence level threshold from the second user group, the method further includes: determining an execution result of the third user on the first event, where the execution result is used to indicate whether the third user executes the first event; determining, based on the execution result, a tag for the third user, where the tag is used to indicate a probability that the third user will execute the first event; The intentionality model is trained again based on the label and the intentionality of the third user.
9. An electronic device, characterized in that: include: Memory for storing computer-executable instructions or computer programs; The processor is configured to implement the user selection method according to any one of claims 1 to 8 when executing the computer executable instructions or computer program stored in the memory.
10. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the user selection method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Product recommendation method and device based on user intention recognition, computer equipment and storage medium
CN110738545A
Potential user identification method and device of target product and storage medium
CN118798936A