A crowd classification method and device, a storage medium and an electronic device
By identifying and clustering atomic population characteristics in population classification, the problem of low efficiency and inaccurate classification in existing technologies is solved, achieving highly discriminative and comprehensive clustered populations, and improving the accuracy and interpretability of marketing plans.
Patent Information
- Application Number
- CN202310484667.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-04-28
AI Technical Summary
Existing technologies rely on manual analysis for population classification, resulting in low efficiency and inaccurate classification results. They cannot meet the needs of population size and category differentiation, making it difficult to customize marketing plans.
By identifying the atomic demographic characteristics in user feature data, atomic demographic groups are segmented within the marketing target audience based on these characteristics, and clustering algorithms are used to cluster them, resulting in clustered demographic groups with high discriminative power and coverage.
It improves the accuracy and interpretability of population classification, enhances the efficiency of operational plan development, and is suitable for large-scale application.
Smart Images

Figure CN116644331B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the technical field of computer technology, and particularly relates to a crowd classification method and device, a storage medium and an electronic device. BACKGROUND
[0002] In most existing scenarios, in the process of generating marketing scripts for different crowds, experienced operators perform a large amount of manual analysis on historical data, and perform crowd grouping based on strong expert experience. This is more dependent on manpower, and the efficiency is relatively low. Moreover, the grouping result in the actual process may not have a distinguishing degree, and the classification may not be accurate, which is not suitable for large-scale expansion.
[0003] With the accumulation of a large amount of user online behavior data, many automatic methods based on data driving are proposed. Merchants can collect user feature data, and classify crowds by automatic clustering algorithms to mark corresponding classification labels. However, the crowds classified by the existing classification algorithms cannot meet the needs of crowd size, crowd category distinguishing degree, and interpretability. The existing classification algorithms do not have interpretability, and cannot distinguish the characteristics of each crowd. Some crowds have too many characteristics, and the size of some crowds cannot be guaranteed, which leads to inconvenience in subsequent customized production and delivery of marketing solutions for crowds. SUMMARY
[0004] The present specification provides a crowd classification method, device, storage medium and electronic device, and the technical solution is as follows:
[0005] In a first aspect, the present specification provides a crowd classification method, and the method comprises: determining a plurality of atomic crowd features according to user feature data; determining a user crowd corresponding to the atomic crowd features in a marketing target crowd as an atomic crowd based on the atomic crowd features; and clustering the atomic crowd to obtain a clustered crowd.
[0006] In a second aspect, the present specification provides a crowd classification device, and the device comprises: a feature determination module configured to determine a plurality of atomic crowd features according to user feature data; a crowd determination module configured to determine a user crowd corresponding to the atomic crowd features in a marketing target crowd as an atomic crowd based on the atomic crowd features; and a crowd clustering module configured to cluster the atomic crowd to obtain a clustered crowd.
[0007] In a third aspect, the present specification provides a computer storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded and executed by a processor to perform the method steps described above.
[0008] In a fourth aspect, the embodiments of the present specification provide a computer program product containing instructions, which, when executed on a computer or processor, cause the computer or processor to perform the method steps described above.
[0009] In a fifth aspect, the present specification provides an electronic device, which can include a processor and a memory; wherein the memory stores a computer program, which is adapted to be loaded by the processor and execute the method steps described above.
[0010] The technical solutions provided by some embodiments of the present specification have at least the following beneficial effects:
[0011] In one or more embodiments of the present specification, by first processing user feature data to determine atomic population features, each of which can affect a marketing target population, then dividing atomic populations based on atomic population features, the accuracy, discrimination, and coverage of atomic population division are ensured, and finally clustering the atomic populations to obtain the final clustered population. The final clustered population has both the discrimination of conversion rate and the discrimination in feature categories, and the accuracy is greatly improved compared to the results obtained by general clustering schemes. Compared with various existing population classification schemes, the embodiments of the present specification have good interpretability of the output population classification results under the premise of ensuring accuracy, which facilitates the operation personnel to formulate corresponding operation schemes and improves the efficiency of the overall process, and can be better landed and used on a large scale. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the present specification or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present specification, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0013] Figure 1 is a scene schematic diagram of a population classification system provided by the present specification
[0014] Figure 2 is a flowchart of a population classification method provided by the present specification.
[0015] Figure 3 is a specific implementation flowchart of step S100 in the population classification method according to Figure 2 is a specific implementation flowchart of step S100 in the population classification method according to
[0016] Figure 4 is a specific implementation flowchart of step S100 in the population classification method according to Figure 3 is a specific implementation flowchart of step S100 in the population classification method according to is a specific implementation flowchart of step S100 in the population classification method according to
[0017] Figure 5 is according to Figure 4 A specific implementation flow chart of step S112 in the crowd classification method shown in the corresponding embodiment.
[0018] Figure 6 is according to Figure 2 A specific implementation flow chart of step S300 in the crowd classification method shown in the corresponding embodiment.
[0019] Figure 7 is according to Figure 6 A specific implementation flow chart of step S320 in the crowd classification method shown in the corresponding embodiment.
[0020] Figure 8 is a total flow schematic diagram of a crowd classification method provided by the present specification.
[0021] Figure 9 is a structural schematic diagram of a crowd classification device provided by the present specification.
[0022] Figure 10 is a structural schematic diagram of an electronic device provided by the present specification.
[0023] Figure 11 is a structural schematic diagram of an operating system and user space provided by the present specification.
[0024] Figure 12 is Figure 11 is an architecture diagram of an Android operating system.
[0025] Figure 13 is Figure 11 is an architecture diagram of an IOS operating system.
[0026] Figure 14 is a structural schematic diagram of an electronic device provided by the present specification. DETAILED DESCRIPTION
[0027] The technical solutions in the present specification will be described clearly and completely in the present specification in combination with the drawings in the present specification. Obviously, the described embodiments are only some of the embodiments of the present specification, not all the embodiments. Based on the embodiments in the present specification, all the other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present specification.
[0028] In the description of the specification, it is understood that the terms "first", "second" and the like are used only for the purpose of description, and cannot be understood as indicating or implying relative importance. In the description of the specification, it is necessary to explain that, unless otherwise explicitly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units not listed, or optionally also includes other steps or units inherent to the process, method, product or device. The specific meaning of the above terms in the specification can be understood by the person skilled in the art. In addition, in the description of the specification, "a plurality of" means two or more, unless otherwise specified. The association relationship of the associated objects is described, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.
[0029] The specification will be described in detail below in conjunction with specific embodiments.
[0030] Please refer to Figure 1 , a scenario diagram of a population classification system provided by the specification. As Figure 1 indicated, the population classification system can at least include a client cluster and a service platform 100.
[0031] The client cluster can include at least one client, such as Figure 1 indicated, specifically including client 1 corresponding to user 1, client 2 corresponding to user 2,..., and client n corresponding to user n, n is an integer greater than 0.
[0032] Each client in the client cluster can be an electronic device with communication function, including but not limited to: wearable devices, handheld devices, personal computers, tablet computers, vehicle-mounted devices, smart phones, computing devices or other processing devices connected to wireless modems, etc. In different networks, the electronic device can be called by different names, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), 5G network or future evolution network electronic device, etc.
[0033] The service platform 100 can be a single server device, for example, a rack-mounted, blade, tower, or cabinet server device, or a hardware device with strong computing power such as a workstation, a mainframe computer, etc. The service platform 100 can also be a server cluster composed of multiple servers. The servers in the server cluster can be symmetrically composed, that is, each server is functionally and positionally equivalent in the transaction link, and each server can independently provide services. The independent service can be understood as not requiring the assistance of another server.
[0034] In one or more embodiments of the present specification, the service platform 100 can establish a communication connection with at least one client in the client cluster, and based on the communication connection, complete the interaction of data in the crowd classification process, such as online transaction data interaction. The transaction data interaction includes but is not limited to data interaction in aspects such as consumption, shopping, finance, and credit, and the specific transaction service type is determined based on actual application conditions. For example, the service platform 100 can implement targeted content recommendation to each client based on the clustered crowd obtained by the crowd classification method of the present specification.
[0035] In the process of the service platform 100 performing online transaction data interaction with a plurality of clients, in scenarios such as related transaction activity publishing, platform hot event, marketing activity, etc., the service platform 100 needs to divide the marketing target crowd, and then arrange personalized copy pushing, hot pushing, and product recommendation, etc. for the divided crowd. Based on this, the service platform 100 can divide the marketing target crowd by executing the crowd classification method of one or more embodiments of the present specification, obtain the final clustered crowd, and then perform copy pushing, hot pushing, and product recommendation, etc. subsequent processing for the final clustered crowd.
[0036] It should be noted that the service platform 100 and at least one client in the client cluster establish a communication connection through a network for interactive communication, where the network can be a wireless network or a wired network. The wireless network includes but is not limited to a cellular network, a wireless local area network, an infrared network or a Bluetooth network. The wired network includes but is not limited to an Ethernet, a universal serial bus (USB) or a controller area network. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML) and the like are used to represent data (such as target compressed packages) exchanged through the network. In addition, all or some links can be encrypted using conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec) and the like. In other embodiments, custom and / or dedicated data communication technologies can be used instead of or in addition to the above data communication technologies.
[0037] The population classification system provided by the specification and the population classification method in one or more embodiments belong to the same concept. The execution subject of the population classification method corresponding to the one or more embodiments of the specification can be the service platform 100 described above. The execution subject of the population classification method corresponding to the one or more embodiments of the specification can also be the electronic device corresponding to the client, which is determined based on the actual application environment. The implementation process of the population classification system embodiment can be seen from the method embodiment described below, which will not be described here.
[0038] Based on Figure 1 The scene diagram shown below describes the population classification method provided by one or more embodiments of the specification in detail.
[0039] Please refer to Figure 2 A flowchart of a population classification method provided by one or more embodiments of the specification is shown, which can be implemented by a computer program and can run on a population classification device based on the von Neumann architecture. The computer program can be integrated in an application or run as an independent tool application. The population classification device can be a service platform.
[0040] Specifically, the population classification method includes:
[0041] Step S100: Determine multiple atomic population characteristics based on user characteristic data.
[0042] Step S200: Based on the atomic audience characteristics, determine the user group corresponding to the atomic audience characteristics in the marketing target audience, and use it as the atomic audience.
[0043] Step S300: Cluster the atomic population to obtain a clustered population.
[0044] This specification first identifies multiple atomic audience characteristics based on user feature data. Each atomic audience characteristic is a relevant feature that impacts operational goals. Then, based on these atomic audience characteristics, the target audience is segmented to obtain atomic audiences. Each atomic audience corresponds to one atomic audience characteristic, representing the finest-grained atomic audience under that characteristic, possessing high discriminative power. Finally, the atomic audiences are clustered to obtain a final, well-defined clustered audience that meets coverage requirements and possesses discriminative power in both conversion rate and feature category.
[0045] In step S100, user characteristic data includes multiple aspects, such as profile tag data features, behavioral activity features, and preference features. Among these features, profile tag data features include occupation, age, and location features; behavioral activity features include login frequency, access frequency, and search frequency; and user preference features include historical search terms and preferred keywords. These features are processed and filtered to obtain the atomic user demographic features. The specific process for processing and filtering user characteristic data can be found in the following embodiment.
[0046] Specifically, in some embodiments, the specific implementation of step S100 can be found in [reference needed]. Figure 3 . Figure 3 It is based on Figure 2 According to the detailed description of step S100 in the population classification method shown in the corresponding embodiment, step S100 in the population classification method may include the following steps:
[0047] Step S110: Perform interpretable feature processing on the user feature data to obtain interpretable features.
[0048] Step S120: Input the interpretable feature into the transformation judgment model, and output the corresponding transformation result from the transformation judgment model.
[0049] Step S130: Determine the atomic population characteristics based on the result of whether or not the conversion has occurred.
[0050] In the present specification, for the user feature data, first, the interpretable feature processing is performed, the user feature data is changed into the interpretable feature, the interpretable feature has the interpretability and the computability, and is a feature that can be understood and calculated by a machine. Then, the conversion judgment model is used to determine the conversion result of each interpretable feature, and then the atomic population feature is determined by screening. The atomic population feature determined by screening through the conversion judgment model is a feature related to the operation target.
[0051] The interpretable feature has the interpretability and the computability, and is a feature that can be understood and calculated by a machine. In step S110, the interpretable feature is obtained by screening processing, and the specific manner can be referred to the following embodiments.
[0052] Specifically, in some embodiments, the specific implementation of step S110 can refer to Figure 4 . Figure 4 is according to Figure 3 The details of step S110 in the population classification method according to the corresponding embodiments are shown, and step S110 can include the following steps:
[0053] Step S111, screening candidate features from the user feature data, the candidate features having the interpretability.
[0054] Step S112, processing the candidate features to obtain the interpretable features.
[0055] In the present specification, the way to obtain the interpretable feature is to first screen the candidate features having the interpretability from the user feature data, and then process the candidate features to convert them into the interpretable features.
[0056] In step S111, the interpretability of the candidate feature means that the candidate feature can be understood by a machine, which is generally in the form of feature encoding plus corresponding feature value, for example, occupation, age, and address in portrait label data features; for example, login times, access times, and search times in behavior activity features; for example, various high-frequency search words, high-frequency click links, and other user preference words in preference class features. However, it should be noted that the commonly used embedding type matrix features in deep learning are user feature data, but they do not have interpretability because they are only continuous matrices, and therefore cannot be used as candidate features.
[0057] In step S112, the processing of the candidate features is to determine the feature encoding and the corresponding feature value of the candidate features, and to associate the feature encoding and the feature value to form the interpretable feature that can be understood and calculated by a machine. Specifically, the specific way to obtain the interpretable feature can be referred to the following embodiments.
[0058] In some embodiments, the implementation of step S112 can refer to the following detailed description Figure 5 . Figure 5 is according to Figure 4 The detailed description of step S112 in the people classification method shown in the corresponding embodiments can include the following steps:
[0059] Step S1121, dividing the user feature data into feature domains according to the categories of the alternative features.
[0060] Step S1122, defining feature categories in the feature domains and giving corresponding feature codes.
[0061] Step S1123, assigning feature values to the feature categories, associating the feature values with the feature codes, and forming interpretable features.
[0062] In this specification, the feature domains are first divided according to the categories of the alternative features, and then the features are defined in the feature domains, and the corresponding feature codes and feature values are given and associated to form interpretable features.
[0063] In step S1121, the feature domains are generally divided according to the usage or corresponding meaning, for example, can be divided into profile label data feature domain, behavior activity feature domain, preference class feature domain, etc. For example, for user u, the profile label data feature domain can be denoted as Profileu, and all features related to the profile of user u are divided into the feature domain Profileu. The behavior activity feature domain can be denoted as Activeu, and all features related to the behavior activity of user u are divided into the feature domain Activeu. The preference class feature domain can be denoted as Preferu, and all features related to the preference of user u are divided into the feature domain Preferu. u u u u
[0064] In step S1122, the corresponding interpretable features are sorted in different feature domains, the feature categories are defined, and the corresponding feature codes are given. The feature category is the name of each specific feature in the feature domain, and the feature code is the feature code corresponding to the feature name.
[0065] In step S1123, the feature value is the value corresponding to the feature code. One feature code and one feature value associated therewith constitute a complete interpretable feature.
[0066] For example, in the feature domain Profile u In this way, the occupation can be defined as a feature category, and the feature code of the occupation can be in the form of occupation, Occupation, Career, O, or 001, and the corresponding feature value can be responsible person, technical personnel, clerical staff, service personnel, production and auxiliary personnel, manufacturing personnel, armed personnel, and other personnel, which can also be coded as 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, and the like. Then, the corresponding occupation-responsible person, occupation-technical personnel, occupation-clerical staff, occupation-service personnel, occupation-production and auxiliary personnel, occupation-manufacturing personnel, occupation-armed personnel, and occupation-other personnel eight explainable features related to occupation can be obtained, which can be marked as Occupation-responsible person, Occupation-technical personnel, Occupation-clerical staff, Occupation-service personnel, Occupation-production and auxiliary personnel, Occupation-manufacturing personnel, Occupation-armed personnel, Occupation-other personnel, or O-10000, O-20000, O-30000, O-40000, O-50000, O-60000, O-70000, O-80000, respectively, as explainable features. It should be noted that the eight occupations listed in this embodiment are only used for illustration, and the most coarse-grained division method is used to enumerate the occupation for the convenience of illustration. In actual application, the occupation can be divided into more fine-grained categories according to actual needs, for example, 75 occupation categories, 434 occupation categories, 1481 occupation categories, and the like, which are not limited by the present disclosure.
[0067] For example, in the feature domain Prefer uIn some embodiments, the high-frequency search word can be defined as a feature category, and the feature code can be in the form of Hot Search, Hot Words, HW, or 006, and the corresponding feature value can be some search words with high user search frequency. Some search words with high user search frequency can be words with more than a predetermined number of search times, which are frequently used search words of the user and can reflect the user's preferences to some extent. Some search words with high user search frequency can also be the top few search words used by the user, and the number of specific search words can be limited according to requirements, for example, the top 5 high-frequency search words, the top 10 high-frequency search words, or the top 20 high-frequency search words, to avoid the problem of excessive data volume caused by too fine atomic group division of words with more than a predetermined number of search times, and avoid the problem of no distinction caused by too coarse atomic group division of words with less than a predetermined number of search times. Correspondingly, a plurality of explainable features related to the high-frequency search word can be obtained. For example, in an embodiment, the high-frequency search words of the user include wet wipes, masks, thermometers, hot pot seasonings, and milk, and as explainable features, they can be labeled as HW-wet wipes, HW-masks, HW-thermometers, HW-hot pot seasonings, and HW-milk, respectively.
[0068] The above embodiments only illustrate the processing method of the enumerated feature, which is a feature whose feature values can be enumerated completely, and the number of feature values can be enumerated completely. For example, the occupation feature in the above embodiment has only the above-mentioned 8 occupation features, and the feature values can be enumerated completely. In more fine-grained division, the number of occupation features is also limited, and even if it is divided into 1481 occupation categories, these occupation categories can also be completely enumerated.
[0069] The high-frequency search word in the above embodiment is also an enumerated feature, and the number of high-frequency search words of the user is always limited, which can be enumerated completely, but too many high-frequency search words can affect the division of atomic groups, so the number of high-frequency search words is generally limited, and the number of specific search words can be limited according to the use environment and use requirements, for example, the top 10 high-frequency search words, the top 20 high-frequency search words, or the top 50 high-frequency search words.
[0070] However, in the explainable feature, not only the enumerated feature whose feature values can be enumerated completely, but also the continuous feature whose feature values are continuous, and the feature value of the continuous feature is a range and cannot be enumerated completely, so it cannot be obtained by enumerating the feature value in the form of enumerated feature. At this time, in step S1123, the following embodiments need to be processed.
[0071] Specifically, in some embodiments, the specific implementation of step S112 can refer to the following embodiments. This embodiment is based onFigure 5 The corresponding embodiment shows the details of step S1123 in the people classification method, and step S1123 can include the following steps:
[0072] Assign a feature value to the feature category, and associate the feature value with the feature code to obtain an initial feature.
[0073] Bucket the initial feature to obtain an interpretable feature.
[0074] In this embodiment, for continuous features, the feature values are continuous, so a large range of feature values is assigned to the corresponding feature category first, and is associated with the feature code to form an initial feature, and the people corresponding to the initial feature is all users. Then the feature values in the initial feature are bucketed to obtain a plurality of small feature values, which are respectively associated with the aforementioned feature code to form a plurality of interpretable features.
[0075] For example, in the feature domain Profile u , the age can be defined as a feature category, and the feature code can be in the form of age, Age, A, YD, or 002, and the corresponding feature value can be determined according to the maximum age and the minimum age in the user feature data. If the maximum age in the user feature data is 87 years old and the minimum age is 23 years old, the feature value can be directly determined as 23-57, or the upper limit value and the lower limit value can be expanded to a certain extent to determine the range value of 20-60, which is an upper limit value and a lower limit value that is an integer of ten. The corresponding feature value can be directly defined as 0 or more according to the actual situation of human age. After determining the feature value, the feature value and the feature code can be corresponded to form an initial feature, and then the initial feature can be bucketed, and the bucketing method can use equal frequency bucketing, equal interval bucketing, model bucketing, etc.
[0076] For example, for feature values 20-90, they can be equally divided into four groups of equally spaced feature values 20-30, 30-40, 40-50, and 50-60, which can form four interpretable features, and can be labeled as Age-20-30, Age-30-40, Age-40-50, and Age-50-60, respectively, to avoid the situation that the number of some buckets is small and the number of some buckets is large. They can also be equally divided into five groups of equally spaced feature values 20-24, 24-32, 32-40, 40-50, and 50-60, which can form five interpretable features, and can be labeled as A-20-24, A-24-32, A-32-40, A-40-50, and A-50-60, respectively, to ensure that the number of each atomic population is evenly distributed, and to avoid the situation that the number of some atomic populations is large and the number of some atomic populations is small.
[0077] For example, if the feature value is greater than or equal to 0, the model can be used to divide the feature values into five groups: (0, 18], (18, 26], (26, 40], (40, 60], and (60, +∞). The five groups can form five explainable features, which can be labeled as YD~(0, 18], YD~(18, 26], YD~(26, 40], YD~(40, 60], and YD~(60, +∞), respectively. The model can be used to find the optimal feature segmentation point for continuous features, and the feature segmentation point can be used for discretization. The model can be used to divide the continuous features into a plurality of explainable features. The model can be used to divide the features directly, or can be used to establish a model for dividing the features according to a rule, or can be used to establish a model for dividing the features according to another rule.
[0078] Since the user feature data is complex, there may be missing data. Therefore, when assigning feature values to feature categories, the problem of missing data needs to be solved. Specifically, the specific steps of assigning feature values to feature categories can refer to the following embodiments.
[0079] Specifically, in some embodiments, the specific implementation of step S112 can refer to the following embodiments. The present embodiment is based on Figure 5 The details of step S1123 in the population classification method shown in the corresponding embodiments are described as follows.
[0080] If the user has a corresponding feature value under the feature category, the feature value is associated with the corresponding feature code to obtain an explainable feature.
[0081] If the user does not have a corresponding feature value under the feature category, a corresponding fill value is filled as a feature value and associated with the corresponding feature code to obtain an explainable feature.
[0082] In the present specification, for a user who has a corresponding feature value under a feature category, the feature value is normally associated with the corresponding feature code to obtain an explainable feature. For a user who does not have a corresponding feature value under a feature category, a corresponding fill value is found as a feature value and associated with the corresponding feature code to form an explainable feature.
[0083] For example, for the feature domain Profile uFor example, for the occupation feature in the feature domain Occupation, if the occupation of user u is known to be a person in charge, a technician, an office worker, a service worker, a production and auxiliary worker, a manufacturing worker, a force worker, and other workers, the corresponding value can be directly assigned to form an explainable feature. If the occupation data of user u is damaged or missing, the occupation of user u can be directly determined as unknown to form an explainable feature. That is, according to the occupation, there are nine explainable features of Occupation-person in charge, Occupation-technician, Occupation-office worker, Occupation-service worker, Occupation-production and auxiliary worker, Occupation-manufacturing worker, Occupation-force worker, Occupation-other worker, and Occupation-unknown. In other embodiments, for the two explainable features of Occupation-other worker and Occupation-unknown, the number of corresponding persons is generally small, and the two explainable features can be combined into one explainable feature, which can be directly assigned as Occupation-other worker or marked as Occupation-unknown. In other embodiments, the occupation feature can also be inferred according to other feature data of user u, and the corresponding occupation feature can be filled in.
[0084] For example, for the login times feature in the feature domain Active u , if the login times of user u are known, the corresponding value can be directly assigned to form an explainable feature. If the login times of user u are damaged or missing, the login times of user u can be directly filled in as 0 to form an explainable feature. Alternatively, the login times of user u can be determined according to the login times of users with more similar features to user u.
[0085] That is, for the missing feature value, the unknown can be filled in or the value 0 can be directly assigned according to the type of the actual feature and the characteristics of the feature value, or the corresponding feature value can be derived according to other feature data, which is not limited in the present specification.
[0086] In step S120, the conversion judgment model is a machine learning model that is pre-trained according to a large number of samples and has a predetermined accuracy requirement. The conversion judgment model can calculate and analyze the relevant information of the explainable feature and output the corresponding conversion result. By processing the explainable feature through the conversion judgment model, it is determined whether the explainable feature is converted, which can assist in judging and screening whether the corresponding explainable feature is suitable for being an atomic people feature. Whether the explainable feature is converted means whether the user corresponding to the explainable feature can be converted as a marketing target to perform operations such as visiting, watching, and purchasing.
[0087] Specifically, in some embodiments, the specific training method of the training of the conversion judgment model comprises:
[0088] obtaining a sample set of interpretable features, the sample set comprising a positive sample set and a negative sample set, each sample in the positive sample set comprising relevant information of an interpretable feature and a result that can be converted for the interpretable feature, and each sample in the negative sample set comprising relevant information of an interpretable feature and a result that cannot be converted for the interpretable feature.
[0089] inputting the samples in the sample set into the conversion judgment model to obtain the result of whether the conversion can be converted output by the conversion judgment model.
[0090] if there is a sample in the positive sample set that is input into the conversion judgment model to obtain a result that cannot be converted, and / or there is a sample in the negative sample set that is input into the conversion judgment model to obtain a result that can be converted, then adjusting the coefficients of the conversion judgment model.
[0091] when a predetermined proportion of samples in the positive sample set are input into the conversion judgment model to obtain a result that can be converted, and a predetermined proportion of samples in the negative sample set are input into the conversion judgment model to obtain a result that cannot be converted, the training is completed.
[0092] In an embodiment of the present specification, the conversion judgment model comprises an XGBoost (Extreme Gradient Boosting) algorithm model. The XGBoost algorithm is an efficient implementation of GBDT, which provides a gradient improvement framework, and its purpose is to provide a "scalable, portable and distributed gradient improvement library". XGBoost uses a boosting tree model, which has good adaptability to the conversion calculation scenario of the interpretable feature, can provide better fitting results, and forms an accurate risk index.
[0093] The XGBoost algorithm model comprises a plurality of optimized regression trees, each regression tree comprises a plurality of leaf nodes, and each leaf node corresponds to an index. After inputting the feature vector of the interpretable feature into the XGBoost algorithm model, each regression tree divides the interpretable feature into a leaf node according to the input feature by traversing the feature split point (for example, when a feature vector is less than A, it is divided into the left subtree, and when it is greater than A, it is divided into the right subtree). In this way, the index corresponding to the leaf node corresponding to the interpretable feature corresponding to each regression tree can be obtained, and the sum of the indices of all leaf nodes is the result of the XGBoost algorithm model predicting whether the interpretable feature can be converted.
[0094] In step S130, if the interpretable feature input conversion judgment model can be converted, the interpretable feature can be selected as the atomic population feature.
[0095] For tree models such as XGBoost algorithm models, the gain of the feature is calculated when the sub-tree is split, so when the tree model such as the XGBoost algorithm model determines that there are many interpretable features that can be converted, the atomic population feature can be determined according to the gain of each interpretable feature input into the tree model such as the XGBoost algorithm model.
[0096] Specifically, the first predetermined number of interpretable features with the largest gain can be selected as the atomic population feature, for example, the first 20 interpretable features with the largest gain can be selected as the atomic population feature.
[0097] These atomic population features determine whether the user will convert. This further ensures the accuracy of the population classification, rather than randomly selecting features or relying heavily on the experience of human experts to select features. In addition, selecting features with the largest gain also reduces the number of atomic populations that need to be calculated subsequently, improving overall computational efficiency.
[0098] In step S200, each atomic population feature can correspond to an atomic population. Based on the atomic population feature, all people in the marketing target population with the atomic population feature are gathered together, i.e., the atomic population corresponding to the atomic population feature is obtained. The atomic population is a population with the finest granularity and has a high degree of distinction.
[0099] In step S300, the atomic population with the finest granularity is clustered into several new clustered populations with common points by a clustering algorithm. Each clustered population is composed of multiple atomic populations and automatically has a feature meaning and a good degree of distinction, which meets the required coverage requirement and can be conveniently landed and applied, having the advantages of low implementation cost and good effect.
[0100] Specifically, in some embodiments, the specific implementation of step S300 can refer to Figure 6 . Figure 6 is according to Figure 2 the detailed description of step S300 in the population classification method according to the corresponding embodiment, wherein step S300 can include the following steps:
[0101] In step S310, the target group index is determined according to the total number of sample sets of the interpretable feature and the number of atomic populations.
[0102] In step S320, the atomic populations are clustered according to the target group index to obtain clustered populations.
[0103] In the embodiments of the present specification, the clustering manner is to calculate the target group index (TGI) corresponding to the atomic population first, and then cluster according to the TGI of each atomic population.
[0104] Specifically, in step S310, there are various ways to determine the target group index, which can be calculated step by step according to an algorithm, or directly calculated according to a formula.
[0105] Specifically, in some embodiments, the specific implementation of step S310 can refer to the following embodiments. The present embodiment is based on the following formula: Figure 6 The details of step S310 in the population classification method shown in the corresponding embodiments can include the following steps:
[0106] According to the total number of the sample set of the interpretable feature and the number of the atomic population, determine a first proportion and a second proportion, the first proportion being the proportion of people with the feature corresponding to the atomic population in the sample set of the interpretable feature, and the second proportion being the proportion of people with the feature corresponding to the atomic population in the positive sample set of the sample set of the interpretable feature.
[0107] According to the first proportion and the second proportion, determine the target group index.
[0108] In the embodiments of the present specification, the calculation method of TGI is to determine the proportion of people with the feature corresponding to the atomic population in the sample set of the interpretable feature, i.e. the first proportion, and the proportion of people with the feature corresponding to the atomic population in the positive sample set of the sample set of the interpretable feature, i.e. the second proportion. Then, according to the first proportion and the second proportion, determine the target group index.
[0109] The calculation method of the first proportion is the number of people with the feature corresponding to the atomic population in the sample set of the interpretable feature divided by the total number of people in the sample set of the interpretable feature.
[0110] The calculation method of the second proportion is the number of people with the feature corresponding to the atomic population in the positive sample set of the sample set of the interpretable feature divided by the total number of people in the positive sample set of the sample set of the interpretable feature.
[0111] After determining the first proportion and the second proportion, divide the second proportion by the first proportion, i.e. obtain the target group index, i.e. TGI.
[0112] In other embodiments, TGI can also be directly determined according to the following formula:
[0113]
[0114] wherein w is a set of atomic people, D + is a set of positive samples of the sample set of explainable features, D is a sample set of explainable features, ∩ is an intersection symbol, count() is a calculation symbol, and TGI w is a target group index corresponding to the atomic people w.
[0115] The TGI can help us analyze the performance of a feature in the target group relative to all users. If the TGI is partitioned, it can be mainly divided into three intervals. When TGI = 100%, the performance of the feature in the target group and all users is the same. When TGI > 100%, the feature performs more strongly in the target group, and the larger the number is, the stronger it is. When TGI < 100%, the feature performs more weakly in the target group, and the smaller the number is, the weaker it is. Therefore, the TGI can also be used to measure the similarity and discrimination between atomic people.
[0116] Understandably, in some embodiments, the TGI can be calculated when the atomic people are divided in step S200. That is, in the embodiments of the present specification, after step S200 is performed, not only the atomic people can be obtained, but also the feature category and feature value of the atomic people corresponding to the atomic people features, the TGI and the coverage of the atomic people. The calculation method of the coverage of the atomic people will be described in detail in subsequent embodiments, and will not be described here.
[0117] In step S320, clustering is performed according to the TGI of each atomic people. There are various ways of clustering, and the clustering requirements can also be diversified according to actual needs. The specific clustering method can be referred to in the following embodiments.
[0118] Specifically, in some embodiments, the specific implementation of step S320 can refer to Figure 7 . Figure 7 is based on Figure 6 the details of step S320 in the people classification method according to the corresponding embodiments, wherein step S320 can include the following steps:
[0119] Step S321, taking each atomic people as an initial clustering cluster, calculating the similarity between each clustering cluster according to the target group index of each clustering cluster.
[0120] Step S322, clustering according to the similarity between the clustering clusters until the clustering requirements are met, and outputting the clustered people.
[0121] In the embodiments of the present disclosure, the similarity between each initial clustering cluster is determined according to the TGI of each initial clustering cluster, two clustering clusters with high similarity are clustered together to obtain a new clustering cluster, and then the clustering step is repeated based on the clustered clustering cluster until the clustering requirement is met.
[0122] In step S321, the calculation formula of the similarity between each clustering cluster is as follows:
[0123]
[0124]
[0125]
[0126] wherein w1 and w2 are an atomic population respectively, Sim(w1, w2) is the similarity between the atomic population w1 and the atomic population w2, TGI w1 is the target group index of the atomic population w1, TGI w2 is the target group index of the atomic population w2, I is an indicator function, which is 1 when the condition is true, and 0 otherwise, δ is a penalty coefficient, is the number of feature categories of the two atomic populations w1 and w2. C1 and C2 are a clustering cluster respectively, Sim(C1, C2) is the similarity between the clustering cluster C1 and the clustering cluster C2, min(Sim(w i,j )) is the minimum value of the similarity between the atomic population in the clustering cluster C1 and the atomic population in the clustering cluster C2, e is the natural base, β is a constant coefficient, TypeLimit is a feature category threshold, and TypeCount(C1, C2) is the total number of feature categories of the clustering cluster C1 and the clustering cluster C2.
[0127] Through the above formula, the similarity between each clustering cluster can be obtained to perform step S322 for clustering.
[0128] Specifically, in some embodiments, the specific implementation of step S322 can refer to the following embodiments. The present embodiment is based on Figure 7 The details of step S322 in the population classification method shown in the corresponding embodiments are described as follows, wherein step S322 can include the following steps:
[0129] The two clustering clusters with the maximum similarity in each clustering cluster are selected and clustered into one type to form a new clustering cluster.
[0130] The number of clustering clusters is determined, and if the number of clustering clusters is not greater than a predetermined number threshold, the clustering requirement is met, the clustering is stopped, and the clustering cluster is output as a clustered population.
[0131] In the present specification, when clustering, the two clustering clusters with the largest similarity are selected to be clustered into a class, a new clustering cluster is formed, and then the new clustering cluster is added to the set of clustering clusters, the similarity with other existing clustering clusters is calculated for clustering, until the number of clustering clusters reaches a predetermined number threshold and below the predetermined number threshold, at which time the remaining clustering clusters are output as the clustered population. That is, in the present embodiment, the clustering requirement to be met is that the number of clustering clusters is not greater than the predetermined number threshold, at which time the number of clustering clusters obtained is suitable, the number of these sets of atomic populations before clustering is reduced, and at the same time the similarity between each atomic population is mined through clustering, the relevance of the characteristics of each atomic population is determined, and is directly used as a classified population, which can improve the pertinence of marketing to related user populations.
[0132] At the same time, in other embodiments, the clustering requirement to be met can also include the coverage of the clustering cluster and the number of feature categories contained. In the present embodiment, if the coverage of a certain clustering cluster exceeds the highest coverage threshold during clustering, the clustering cluster will no longer participate in subsequent clustering. Similarly, if the number of feature categories contained in a certain clustering cluster exceeds the feature category threshold (i.e., TypeLimit) during clustering, the clustering cluster will also no longer participate in subsequent clustering. It can be understood that the coverage of the atomic population can also be calculated according to the above method and formula.
[0133] It needs to be explained here that the number of feature categories refers to the types of feature categories contained in the clustering cluster, rather than the types of atomic population features. For example, in a clustering cluster, it contains three atomic population features of age feature 0 to 18 years old, age feature 18 to 26 years old, and age feature 26 to 30 years old, and the feature categories of the three atomic population features contained are all ages, so the number of feature categories contained is 1.
[0134] For another example, in a clustering cluster, it contains three atomic population features of age feature 0 to 18 years old, age feature 18 to 26 years old, and login times feature 12 to 21 times, and the feature categories of the three atomic population features contained are age features and login times features, so the number of feature categories contained is 2.
[0135] Among them, the highest coverage threshold can be set to 0.7, and the feature category threshold can be set to 2 or 3 according to actual needs.
[0136] For the coverage of the clustering cluster, it is the total proportion of the number of people in the clustering cluster in the entire marketing target population, and its determination method can be obtained by dividing the number of people in the clustering cluster by the number of people in the marketing target population.
[0137] In other embodiments, for the coverage of the clustering cluster, the calculation formula is as follows:
[0138]
[0139] Wherein, P is a cluster or an atomic crowd, U is a marketing target crowd, ∩ is an intersection symbol, count() is a calculation symbol, meaning the number of people in the set, coverage P is the coverage corresponding to the set P.
[0140] At the same time, if there is no cluster that can participate in clustering in the clustering process, but the number of clusters is still greater than the predetermined number threshold, the clustering is directly stopped.
[0141] In this specification, please refer to Figure 8 , the overall process of the crowd classification scheme can be summarized as first performing explainability processing on user feature data to obtain explainable features. The specific method of obtaining explainable features can refer to the previous Figure 4 and Figure 5 corresponding embodiments, which will not be repeated here. At the same time, a part of the user feature data is taken as a sample to train the conversion judgment model in the machine learning model aspect. The training process of the conversion judgment model has been described in the previous embodiments, which will not be repeated here. After the conversion judgment model is trained, the explainable features obtained before are input into the conversion judgment model to perform important feature mining. The important features, i.e. the atomic crowd features in the previous embodiments, are screened out. The specific method of important feature mining can refer to the corresponding embodiments of Figure 3 , which will not be repeated here. After obtaining the atomic crowd features, the atomic crowd mining of the TGI is performed on the marketing target crowd according to the atomic crowd features to obtain the atomic crowd and the parameters of the atomic crowd corresponding to the TGI, coverage, feature category and feature value, etc. Finally, the atomic crowd with the above parameters is clustered to obtain the classification result. The specific method of clustering can refer to the corresponding embodiments of Figure 6 and Figure 7 , which will not be repeated here.
[0142] The embodiment in the specification determines atomic people characteristics by processing user characteristic data first, each atomic people characteristic can influence a marketing target people, then divides atomic people based on the atomic people characteristics, ensures the accuracy, discrimination and coverage of the atomic people division, and finally clusters the atomic people to obtain the final clustered people. The final clustered people has discrimination of conversion rate and discrimination on feature categories at the same time, and the accuracy is greatly improved compared with the result obtained by a general clustering scheme. Compared with various people classification schemes, the embodiment of the specification has good interpretability of the output people classification result under the premise of ensuring accuracy, facilitates the operation personnel to formulate the corresponding operation scheme, improves the efficiency of the whole process, and can be better landed and used on a large scale.
[0143] The embodiment of the specification will be described in detail below. Figure 9 The people classification device provided in the specification will be described in detail. It should be noted that, Figure 9 The people classification device shown in the specification is used to execute the method of the embodiment Figures 1-8 The method of the embodiment shown in the specification, only the part related to the specification is shown for convenience of description, and the specific technical details not disclosed, please refer to the specification Figures 1-8 The embodiment shown in the specification.
[0144] Please refer to Figure 9 , which shows the structure diagram of the people classification device of the specification. The people classification device 1 can be realized by software, hardware or combination of the two to become all or part of the user terminal. According to some embodiments, the people classification device 900 includes a feature determination module 910, a people determination module 920 and a people clustering module 930, wherein the feature determination module 910 is configured to determine a plurality of atomic people characteristics according to user characteristic data. The people determination module 920 is configured to determine the user people corresponding to the atomic people characteristics as the atomic people in the marketing target people based on the atomic people characteristics. The people clustering module 930 is configured to cluster the atomic people to obtain clustered people.
[0145] Optionally, the feature determination module 910 specifically includes: a feature processing submodule, configured to perform interpretable feature processing on the user characteristic data to obtain interpretable features. A judgment model submodule is configured to input the interpretable features into a conversion judgment model, and output the corresponding whether conversion result from the conversion judgment model. A feature determination submodule is configured to determine the atomic people characteristics according to the whether conversion result.
[0146] Optionally, the crowd classification apparatus further comprises: a sample obtaining module, configured to obtain a sample set of the explainable feature, the sample set comprising a positive sample set and a negative sample set, each sample in the positive sample set comprising relevant information of the explainable feature and a transformable result calibrated for the explainable feature, and each sample in the negative sample set comprising relevant information of the explainable feature and a non-transformable result calibrated for the explainable feature; a sample input module, configured to input the samples in the sample set into the transformation judgment model to obtain the transformable result output by the transformation judgment model; a coefficient adjusting module, configured to adjust the coefficient of the transformation judgment model if there is a sample in the positive sample set input into the transformation judgment model to obtain the non-transformable result, and / or there is a sample in the negative sample set input into the transformation judgment model to obtain the transformable result; and a training end module, configured to end the training when a predetermined proportion of samples in the positive sample set are input into the transformation judgment model to obtain the transformable result, and a predetermined proportion of samples in the negative sample set are input into the transformation judgment model to obtain the non-transformable result.
[0147] Optionally, the feature processing submodule specifically comprises: a feature screening unit, configured to screen candidate features from the user feature data, the candidate features having explainability; and a feature processing unit, configured to process the candidate features to obtain the explainable features.
[0148] Optionally, the feature processing unit specifically comprises: a feature domain subunit, configured to divide the user feature data into feature domains according to the categories of the candidate features; a definition and coding subunit, configured to define feature categories in the feature domains and give corresponding feature codes; and an association subunit, configured to assign feature values to the feature categories and associate the feature values with the feature codes to form the explainable features.
[0149] Optionally, the association subunit is specifically configured to perform: assigning feature values to the feature categories and associating the feature values with the feature codes to obtain initial features; and performing bucketing on the initial features to obtain the explainable features.
[0150] Optionally, the association subunit is specifically further configured to perform the following steps: if the user has a corresponding feature value under the feature category, associating the feature value with the corresponding feature code to obtain the explainable features; and if the user does not have a corresponding feature value under the feature category, filling a corresponding padding value as a feature value to be associated with the corresponding feature code to obtain the explainable features.
[0151] Optionally, the population clustering module 930 specifically comprises: an index calculation submodule, configured to determine the target group index according to the total number of the sample set of the explainable feature and the number of the atomic populations. A population clustering submodule is configured to cluster the atomic populations according to the target group index to obtain clustered populations.
[0152] Optionally, the index calculation submodule specifically comprises: a proportion determination unit, configured to determine a first proportion and a second proportion according to the total number of the sample set of the explainable feature and the number of the atomic populations, the first proportion being a proportion of people having the feature corresponding to the atomic population in the sample set of the explainable feature, and the second proportion being a proportion of people having the feature corresponding to the atomic population in the positive sample set of the sample set of the explainable feature. An index determination unit is configured to determine the target group index according to the first proportion and the second proportion.
[0153] Optionally, the population clustering submodule specifically comprises: a similarity determination unit, configured to take each of the atomic populations as an initial clustering cluster, and calculate similarities between the clustering clusters according to target group indexes of the clustering clusters. A population clustering unit is configured to cluster according to the similarities between the clustering clusters until a clustering requirement is met, and output clustered populations.
[0154] Optionally, the population clustering unit specifically comprises: a clustering subunit, configured to select two clustering clusters with the largest similarity in each of the clustering clusters to form a new clustering cluster. An output subunit is configured to determine the number of the clustering clusters, and if the number of the clustering clusters is not greater than a predetermined number threshold, the clustering requirement is met, the clustering is stopped, and the clustering clusters are output as clustered populations.
[0155] It should be noted that the population classification device provided in the above embodiments is only used as an example to divide the above functional modules when performing the population classification method. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the population classification device and the population classification method provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments. Here, it will not be repeated.
[0156] The above serial numbers in the specification are only for description, and do not represent the advantages or disadvantages of the embodiments.
[0157] In one or more embodiments of the present specification, by processing user feature data first, atomic people features are determined, each of which can affect a marketing target people, then atomic people are divided based on the atomic people features, ensuring the accuracy, discrimination and coverage of the atomic people division, and finally the atomic people are clustered to obtain the final clustered people. The final clustered people have both the discrimination of conversion rate and the discrimination on feature categories, and the accuracy is greatly improved compared to the results obtained by general clustering schemes. Compared with various existing people classification schemes, the embodiments of the present specification have good interpretability of the output people classification results under the premise of ensuring accuracy, which facilitates the operation personnel to formulate corresponding operation schemes and improves the efficiency of the overall process, and can be better landed and used on a large scale.
[0158] The present specification also provides a computer storage medium, which can store a plurality of instructions, the instructions being suitable for being loaded and executed by a processor to perform the people classification method of the embodiments as shown in the above Figures 1-8 The specific implementation process can refer to the specific description of the embodiments as shown in the above Figures 1-8 The specific implementation process can refer to the specific description of the embodiments as shown in the above
[0159] The present specification also provides a computer program product, which stores at least one instruction, the at least one instruction being loaded and executed by the processor to perform the people classification method of the embodiments as shown in the above Figures 1-8 The specific implementation process can refer to the specific description of the embodiments as shown in the above Figures 1-8 The specific implementation process can refer to the specific description of the embodiments as shown in the above
[0160] Please refer to Figure 10 , which shows a structural block diagram of an electronic device provided by an example embodiment of the present specification. The electronic device in the present specification can include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140 and a bus 150. The processor 110, the memory 120, the input device 130 and the output device 140 can be connected through the bus 150.
[0161] The processor 110 can include one or more processing cores. The processor 110 connects various parts within the entire electronic device with various interfaces and lines, performs various functions of the electronic device 100 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Alternatively, the processor 110 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 110 can integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes an operating system, a user interface, and an application program, etc.; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 110, but be implemented by a separate communication chip.
[0162] The memory 120 can include a random access memory (RAM) and can also include a read-only memory (ROM). Alternatively, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 120 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing each of the following method embodiments, etc., the operating system can be an Android system, an IOS system developed by Apple Inc., a system developed based on the Android system or other systems. The data storage area can also store data created by the electronic device in use, such as a phone book, audio and video data, chat record data, etc.
[0163] Referring to Figure 11As shown, the memory 120 can be divided into an operating system space and a user space, the operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve good running effect, the operating system allocates corresponding system resources for different third-party applications. However, there are also differences in the demand for system resources in different application scenarios in the same third-party application. For example, in the local resource loading scenario, the third-party application has a higher requirement for the disk reading speed; in the animation rendering scenario, the third-party application has a higher requirement for the GPU performance. However, the operating system and the third-party application are independent of each other, and the operating system often cannot timely perceive the current application scenario of the third-party application, resulting in that the operating system cannot perform targeted system resource adaptation according to the specific application scenario of the third-party application.
[0164] In order to enable the operating system to distinguish the specific application scenario of the third-party application, it is necessary to open up the data communication between the third-party application and the operating system, so that the operating system can obtain the current scenario information of the third-party application at any time, and then perform targeted system resource adaptation based on the current scenario.
[0165] Taking the Android system as an example, the programs and data stored in the memory 120 are as follows Figure 12As shown, the memory 120 can store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380, wherein the Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to an operating system space, and the application layer 380 belongs to a user space. The Linux kernel layer 320 provides underlying drivers for various hardware of the electronic device, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, and the like. The system runtime library layer 340 provides main feature support for the Android system through some C / C++ libraries. For example, an SQLite library provides database support, an OpenGL / ES library provides 3D drawing support, a Webkit library provides browser kernel support, and the like. An Android runtime is also provided in the system runtime library layer 340, which mainly provides some core libraries to allow developers to use Java language to write Android applications. The application framework layer 360 provides various APIs that can be used when building an application, and developers can also build their own applications by using these APIs, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management. At least one application program is running in the application layer 380, which can be native applications provided by the operating system, such as a contact program, a message program, a clock program, a camera application, and the like, or third-party applications developed by third-party developers, such as game applications, instant messaging programs, photo beautification programs, and the like.
[0166] For example, taking an IOS system as the operating system, the programs and data stored in the memory 120 can include an IOS kernel 320, an IOS runtime library 340, an application framework 360, and an application layer 380. Figure 13As shown, the IOS system includes: a core operating system layer 420, a core service layer 440, a media layer 460, and a Cocoa Touch layer 480. The core operating system layer 420 includes an operating system kernel, drivers, and low-level hardware abstractions that provide more hardware-specific functionality to program frameworks in the core service layer 440. The core service layer 440 provides system services and / or program frameworks that applications need, such as a Foundation framework, an account framework, an advertisement framework, a data storage framework, a network connection framework, a geographic location framework, a motion framework, and the like. The media layer 460 provides interfaces for applications related to audio and video, such as interfaces related to graphics images, interfaces related to audio technology, interfaces related to video technology, an AirPlay interface for wireless audio and video transmission technology, and the like. The Cocoa Touch layer 480 provides various commonly used interface-related frameworks for application development, and is responsible for user touch interaction on the electronic device. For example, a local notification service, a remote push service, an advertisement framework, a game tool framework, a message user interface (UI) framework, a user interface UIKit framework, a map framework, and the like.
[0167] In Figure 13 In the framework shown, the frameworks related to most applications include, but are not limited to, the Foundation framework in the core service layer 440 and the UIKit framework in the Cocoa Touch layer 480. The Foundation framework provides many basic object classes and data types, and provides the most basic system services for all applications, and is UI-independent. The UIKit framework provides basic UI class libraries for creating touch-based user interfaces, and iOS applications can provide UIs based on the UIKit framework, so it provides the basic framework of the application for building user interfaces, drawing, processing and user interaction events, responding to gestures, and the like.
[0168] In the IOS system, the manner and principle of implementing data communication between a third-party application and an operating system can refer to the Android system, and the present specification will not be repeated here.
[0169] The input device 130 is configured to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is configured to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In an example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are a touch display screen configured to receive a touch operation of a user using a finger, a stylus, or any suitable object on or near the touch display screen, and display a user interface of each application. The touch display screen is usually arranged on a front panel of the electronic device. The touch display screen can be designed as a full screen, a curved screen, or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, a combination of a special-shaped screen and a curved screen, which is not limited in the present specification.
[0170] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above-described figures does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than those shown in the figures, or combine certain components, or different component arrangements. For example, the electronic device further includes a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WiFi) module, a power supply, a Bluetooth module, and the like, which are not described herein.
[0171] In the present specification, the execution subject of each step can be the electronic device described above. Alternatively, the execution subject of each step is an operating system of the electronic device. The operating system can be an Android system, an IOS system, or other operating systems, which are not limited in the present specification.
[0172] The electronic device of the present specification can further have a display device installed thereon. The display device can be various devices capable of realizing a display function, such as a cathode ray tube display (CRT), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), and the like. A user can use the display device on the electronic device 101 to view displayed text, images, videos, and the like. The electronic device can be a smartphone, a tablet computer, a game device, an augmented reality (AR) device, a car, a data storage device, an audio playback device, a video playback device, a notebook, a desktop computing device, a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, an electronic clothing, and the like.
[0173] In Figure 14 In the electronic device shown, the processor 110 can be configured to invoke a network optimization application stored in the memory 120, and specifically configured to: determine a plurality of atomic user group features according to user feature data; determine, based on the atomic user group features, a user group corresponding to the atomic user group features in a marketing target user group as an atomic user group; and cluster the atomic user groups to obtain clustered user groups.
[0174] In one embodiment, when the processor 110 determines a plurality of atomic user group features according to user feature data, the processor 110 is specifically configured to: perform interpretable feature processing on the user feature data to obtain interpretable features; input the interpretable features into a conversion judgment model to output a result of whether conversion corresponding to the interpretable features from the conversion judgment model; and determine the atomic user group features according to the result of whether conversion.
[0175] In one embodiment, before the processor 110 inputs the interpretable features into the conversion judgment model to output the result of whether conversion corresponding to the interpretable features from the conversion judgment model, the processor 110 is further configured to: obtain a sample set of interpretable features, the sample set including a positive sample set and a negative sample set, each sample in the positive sample set including relevant information of an interpretable feature and a conversion result marked for the interpretable feature, and each sample in the negative sample set including relevant information of an interpretable feature and a non-conversion result marked for the interpretable feature; input samples in the sample set into the conversion judgment model to obtain a result of whether conversion output from the conversion judgment model; adjust coefficients of the conversion judgment model if there is a sample in the positive sample set that, when input into the conversion judgment model, obtains a non-conversion result, and / or there is a sample in the negative sample set that, when input into the conversion judgment model, obtains a conversion result; and end training when a predetermined proportion of samples in the positive sample set, when input into the conversion judgment model, obtain a conversion result, and a predetermined proportion of samples in the negative sample set, when input into the conversion judgment model, obtain a non-conversion result.
[0176] In one embodiment, when the processor 110 performs interpretable feature processing on the user feature data to obtain interpretable features, the processor 110 is specifically configured to: screen candidate features from the user feature data, the candidate features having interpretability; and process the candidate features to obtain interpretable features.
[0177] In one embodiment, when the processor 110 processes the candidate features to obtain interpretable features, the processor 110 is specifically configured to:
[0178] According to the category of the alternative feature, the feature domain of the user feature data is divided; a feature category is defined in the feature domain, and a corresponding feature code is given; a feature value is assigned to the feature category, which is associated with the feature code to form an interpretable feature.
[0179] In one embodiment, when the processor 110 performs the operation of assigning a feature value to the feature category, which is associated with the feature code to form an interpretable feature, it specifically performs the following operation: assigning a feature value to the feature category, which is associated with the feature code to obtain an initial feature; and performing bucketing on the initial feature to obtain an interpretable feature.
[0180] In one embodiment, when the processor 110 performs the operation of assigning a feature value to the feature category, which is associated with the feature code to form an interpretable feature, it specifically performs the following operation: if the user has a corresponding feature value under the feature category, the feature value is associated with the corresponding feature code to obtain an interpretable feature; if the user does not have a corresponding feature value under the feature category, a corresponding padding value is filled as a feature value and associated with the corresponding feature code to obtain an interpretable feature.
[0181] In one embodiment, when the processor 110 performs the operation of clustering the atomic population to obtain a clustered population, it specifically performs the following operation: determining the target group index according to the total number of sample sets of the interpretable feature and the number of atomic populations; and clustering the atomic populations according to the target group index to obtain a clustered population.
[0182] In one embodiment, when the processor 110 performs the operation of determining the target group index according to the total number of sample sets of the interpretable feature and the number of atomic populations, it specifically performs the following operation: determining a first proportion and a second proportion according to the total number of sample sets of the interpretable feature and the number of atomic populations, the first proportion being the proportion of people with the corresponding feature of the atomic population in the sample set of the interpretable feature, and the second proportion being the proportion of people with the corresponding feature of the atomic population in the positive sample set of the sample set of the interpretable feature; and determining the target group index according to the first proportion and the second proportion.
[0183] In one embodiment, when the processor 110 performs the operation of clustering the atomic population according to the target group index to obtain a clustered population, it specifically performs the following operation: taking each atomic population as an initial clustering cluster, calculating the similarity between each clustering cluster according to the target group index of each clustering cluster; and clustering according to the similarity between the clustering clusters until the clustering requirement is met, and outputting the clustered population.
[0184] In one embodiment, when the processor 110 performs the clustering according to the similarity between the clustering clusters until the clustering requirement is met, and outputs the clustered population, the following operations are specifically performed: selecting two clustering clusters with the maximum similarity in each clustering cluster to form a new clustering cluster; determining the number of clustering clusters, and if the number of clustering clusters is not greater than a predetermined number threshold, the clustering requirement is met, and the clustering is stopped, and the clustering clusters are output as the clustered population.
[0185] In one or more embodiments of the present specification, by first processing the user feature data to determine atomic population features, each of which can affect the marketing target population, then dividing the atomic population based on the atomic population features, the accuracy, discrimination and coverage of the atomic population division are ensured, and finally the atomic population is clustered to obtain the final clustered population. The final clustered population has both the discrimination of conversion rate and the discrimination on the feature category, and the accuracy is greatly improved compared to the results obtained by general clustering schemes. Compared with various existing population classification schemes, the embodiments of the present specification ensure the accuracy, and the output population classification result also has good interpretability, which facilitates the operation personnel to formulate the corresponding operation scheme, improves the efficiency of the whole process, and can be better landed and used on a large scale.
[0186] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The program can be stored in a computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiments can be included. The storage medium can be a magnetic disc, an optical disc, a read-only memory or a random access memory, etc.
[0187] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the object features, interaction behavior features (such as user behavior activity features) and user information involved in the present specification are all obtained under sufficient authorization.
[0188] The above only describes the preferred embodiments of the present specification, and of course cannot limit the scope of the rights of the present specification, so equivalent changes made according to the claims of the present specification still fall within the scope covered by the present specification.
Claims
1. A method for classifying a population, the method comprising: Based on user characteristic data, multiple atomic group characteristics are determined; Based on the atomic audience characteristics, the user groups corresponding to the atomic audience characteristics are determined in the marketing target audience, and are regarded as atomic audiences; Cluster the atomic population to obtain clustered populations; Specifically, determining multiple atomic population characteristics based on user characteristic data includes: The user feature data is subjected to interpretable feature processing to obtain interpretable features; The interpretable features are input into the transformation judgment model, and the transformation judgment model outputs the corresponding transformation result. The characteristics of the atomic population are determined based on the result of whether or not the transformation occurs; Specifically, the process of performing interpretable feature processing on the user feature data to obtain interpretable features includes: Candidate features are selected from the user feature data, and the candidate features are interpretable; The candidate features are processed to obtain interpretable features; Specifically, the process of processing the candidate features to obtain interpretable features includes: Based on the types of the candidate features, the user feature data is divided into feature domains; Define feature categories within the feature domain and assign corresponding feature codes; The feature category is assigned a feature value and associated with the feature code to form an interpretable feature.
2. The method according to claim 1, before inputting the interpretable feature into the transformation judgment model and having the transformation judgment model output the corresponding result of whether or not to transform, the method further includes: Obtain a sample set of interpretable features, the sample set containing a positive sample set and a negative sample set, each sample in the positive sample set containing relevant information about the interpretable feature and a transformable result labeled for the interpretable feature, each sample in the negative sample set containing relevant information about the interpretable feature and a non-transformable result labeled for the interpretable feature; The samples in the sample set are input into the conversion judgment model to obtain the result of whether the conversion is possible, output by the conversion judgment model. If, after inputting samples from the positive sample set into the transformation judgment model, a result that cannot be transformed is obtained, and / or after inputting samples from the negative sample set into the transformation judgment model, a result that can be transformed is obtained, then the coefficients of the transformation judgment model are adjusted. When a predetermined proportion of samples from the positive sample set are input into the transformation judgment model, a result indicating that the samples can be transformed is obtained; and when a predetermined proportion of samples from the negative sample set are input into the transformation judgment model, a result indicating that the samples cannot be transformed is obtained, and the training ends.
3. The method according to claim 1, wherein assigning feature values to the feature category and associating them with the feature encoding to form interpretable features specifically includes: Assign feature values to the feature categories and associate them with the feature codes to obtain initial features; The initial features are bucketed to obtain interpretable features.
4. The method according to claim 1, wherein assigning feature values to the feature categories and associating them with the feature codes to form interpretable features specifically includes: If a user has a corresponding feature value under the feature category, then the feature value is associated with the corresponding feature code to obtain an interpretable feature; If a user does not have a corresponding feature value under the feature category, then the corresponding fill value is filled in and associated with the corresponding feature code to obtain an interpretable feature.
5. The method according to claim 2, wherein clustering the atomic population to obtain clustered populations specifically includes: The target population index is determined based on the total number of samples in the interpretable feature set and the number of atomic populations. The atomic population is clustered based on the target population index to obtain the clustered population.
6. The method according to claim 5, wherein determining the target population index based on the total number of samples of the interpretable features and the number of atomic populations specifically includes: Based on the total number of samples in the interpretable feature set and the number of atomic populations, a first proportion and a second proportion are determined. The first proportion is the proportion of people with the features corresponding to the atomic populations in the sample set of the interpretable feature set, and the second proportion is the proportion of people with the features corresponding to the atomic populations in the positive sample set of the sample set of the interpretable feature set. The target group index is determined based on the first proportion and the second proportion.
7. The method according to claim 5, wherein clustering the atomic population based on the target population index to obtain a clustered population specifically includes: Using each of the aforementioned atomic populations as the initial cluster, the similarity between each cluster is calculated based on the target population index of each cluster. Clustering is performed based on the similarity between the clusters until the clustering requirements are met, and the clustered population is output.
8. The method according to claim 7, wherein clustering is performed based on the similarity between the clusters until the clustering requirements are met, and the clustered population is output, specifically comprising: The two clusters with the highest similarity among the aforementioned clusters are selected and grouped into one class to form a new cluster. The number of clusters is determined. If the number of clusters is not greater than a predetermined threshold, the clustering requirement is met, the clustering is stopped, and the clusters are output as the clustered population.
9. A population classification device, the device comprising: The feature determination module is used to determine multiple atomic population features based on user feature data; The audience determination module is used to determine the user group corresponding to the atomic audience characteristics in the marketing target audience based on the atomic audience characteristics, and to use them as atomic audiences; The population clustering module is used to cluster the atomic population to obtain clustered populations. The feature determination module specifically includes: a feature processing submodule, used to perform interpretable feature processing on the user feature data to obtain interpretable features; a judgment model submodule, used to input the interpretable features into a transformation judgment model, and the transformation judgment model outputs the corresponding transformation result; and a feature determination submodule, used to determine the atomic population features based on the transformation result. The feature processing submodule specifically includes: a feature filtering unit, used to filter candidate features from the user feature data, wherein the candidate features are interpretable; and a feature processing unit, used to process the candidate features to obtain interpretable features. The feature processing unit specifically includes: a feature domain subunit, used to divide the user feature data into feature domains according to the types of the candidate features; a definition encoding subunit, used to define feature categories within the feature domains and assign corresponding feature codes; and an association subunit, used to assign feature values to the feature categories and associate them with the feature codes to form interpretable features.
10. A computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps of any one of claims 1 to 8.
11. A computer program product storing at least one instruction, said at least one instruction being loaded by a processor and executing the method steps of any one of claims 1 to 8.
12. An electronic device, comprising: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1 to 8.
Citation Information
Patent Citations
Activity information pushing method, device and equipment
CN112184332A
User grouping method, device and equipment and computer storage medium
CN115878870A