Crowd tag determination method and device, storage medium and electronic equipment
By combining the crowd clustering model with the crowd detection model, the problem of low accuracy of manual labeling is solved, efficient and accurate crowd label determination is achieved, and high-precision and multi-dimensional label description is provided.
Patent Information
- Application Number
- CN202510897719.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-03
AI Technical Summary
In existing technologies, manually determined crowd labels have low accuracy and cannot effectively consider multi-dimensional cross-features, resulting in large differences between operational results and expected results.
The crowd clustering model is used to cluster the initial user attribute data to obtain the representative feature data of the crowd clusters, and the crowd detection model is used to detect crowd labels to generate high-precision and multi-dimensional crowd label descriptions.
It achieves high-precision label description of individual individuals in the target population, improves labeling efficiency, enriches the information dimension of label description, and supports efficient and accurate population label determination.
Smart Images

Figure CN120744561A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method, device, storage medium, and electronic device for determining a crowd label. Background Art
[0002] In many fields, such as advertising, product / service promotion, and consumer finance, accurately labeling different groups of people is key to optimizing user operations strategies and improving user experience. Related technologies rely on business experts to segment groups based on their experience and understanding of the business. This process generates labels for different groups, which are then used to analyze their responses to business operations strategies.
[0003] However, because the cross-dimensionality considered by business experts is limited, while the groups classified by crowd labels are highly similar in some dimensions, the actual operational results during actual operations often differ significantly from the expected ones. Consequently, manually determined crowd labels suffer from low accuracy. Therefore, improving the accuracy of crowd labels is an urgent issue that needs to be addressed. Summary of the Invention
[0004] This specification provides a method, device, storage medium, and electronic device for determining a crowd tag. The technical solution is as follows: In a first aspect, this specification provides a method for determining a crowd label, the method comprising: Obtain initial user attribute data of the target population; Using a crowd clustering model to perform crowd clustering processing based on the initial user attribute data to obtain crowd cluster representative feature data corresponding to the target cluster; A large crowd detection model is used to perform crowd label detection processing based on the representative feature data of the crowd clusters to obtain crowd label description information corresponding to the target cluster.
[0005] In a second aspect, this specification provides a method for training a large crowd detection model, the method comprising: Create an initial crowd detection model; Determining sample user attribute data of a sample population, performing population clustering processing on the sample user attribute data to obtain representative characteristic data of the sample population cluster corresponding to the sample cluster cluster; Clustering representative characteristic data of the sample population and labeling the sample population labels; The sample population cluster representative feature data is used to perform at least one round of model training on the initial population detection large model. During the model training process, the initial population detection large model is called based on the sample population cluster representative feature data to perform population label prediction processing to obtain the predicted population label corresponding to the sample cluster cluster, and a comprehensive model loss value is determined based on the predicted population label and the sample population label. Based on the comprehensive model loss value, model parameters of the initial population detection large model are adjusted to obtain the population detection large model.
[0006] In a third aspect, this specification provides a device for determining a crowd label, the device comprising: Data acquisition module, used to obtain initial user attribute data of the target population; A crowd clustering module, configured to perform crowd clustering processing based on the initial user attribute data using a crowd clustering model to obtain crowd cluster representative feature data corresponding to a target cluster; The label detection module is used to use a large crowd detection model to perform crowd label detection processing based on the representative feature data of the crowd cluster to obtain crowd label description information corresponding to the target cluster.
[0007] In a fourth aspect, this specification provides a large-scale crowd detection model training device, the device comprising: Model creation module, used to create the initial crowd detection model; A sample acquisition module is used to determine sample user attribute data of a sample population, and perform population clustering processing on the sample user attribute data to obtain representative characteristic data of the sample population cluster corresponding to the sample cluster; A labeling module is used to label the sample population cluster representative feature data with sample population labels; A model training module is used to perform at least one round of model training on the initial population detection large model using the representative feature data of the sample population clusters. During the model training process, the initial population detection large model is called based on the representative feature data of the sample population clusters to perform population label prediction processing to obtain the predicted population label corresponding to the sample cluster cluster, a comprehensive model loss value is determined based on the predicted population label and the sample population label, and model parameter adjustment processing is performed on the initial population detection large model based on the comprehensive model loss value to obtain the population detection large model.
[0008] In a fifth aspect, this specification provides a computer storage medium having a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above method.
[0009] In a sixth aspect, this specification provides a computer program product, wherein the computer program product stores at least one instruction, and the at least one instruction is loaded by a processor to execute the above method.
[0010] In a seventh aspect, this specification provides an electronic device, which may include: a memory and a processor; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the memory and executing the above method.
[0011] The beneficial effects of the technical solutions provided in this specification include at least: The crowd label determination method provided in the embodiments of this specification first obtains the initial user attribute data of the target population, then uses a crowd clustering model to perform crowd clustering processing to obtain the crowd cluster representative feature data corresponding to the target cluster cluster, and then uses a crowd detection large model to perform crowd label detection processing to obtain the crowd label description information corresponding to the target cluster cluster. In this way, by inputting the crowd cluster representative feature data corresponding to different cluster clusters into the crowd detection large model for crowd label detection processing, it is possible to obtain high-precision crowd label description information for each individual in the target population, thereby improving the efficiency of annotating crowd labels. In addition, through the powerful understanding ability of the large model, the information dimension of the crowd label description information is also enriched. Therefore, the embodiments of this specification achieve the efficient provision of high-precision and multi-dimensional crowd label description information for the target population. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 This is a scenario diagram of a crowd label determination system provided by an embodiment of this specification; Figure 2 This is a flow chart of a method for determining a crowd label provided in an embodiment of this specification; Figure 3 This is a flowchart of another method for determining a crowd label provided in an embodiment of this specification; Figure 4 This is a flow chart of a method for training a large crowd detection model provided in an embodiment of this specification; Figure 5 This is a flow chart of a method for determining the loss value of a comprehensive model provided in the embodiment of this specification. Figure 6This is a schematic diagram of the structure of a crowd label determination device provided in an embodiment of this specification; Figure 7 This is a schematic diagram of the structure of a training device for a large crowd detection model provided in an embodiment of this specification; Figure 8 This is a structural diagram of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0014] In order to make the invention objectives, features, and advantages of the embodiments of this specification more obvious and easy to understand, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this specification.
[0015] In the description of this specification, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise clearly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood according to the specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0016] See Figure 1 , is a scene diagram of a crowd label determination system provided by an embodiment of this specification. Figure 1 As shown, the scenario diagram may include at least a terminal cluster and a server.
[0017] In some embodiments, the terminal cluster may include at least one terminal, such as Figure 1 As shown, it specifically includes terminal 1 corresponding to user 1, terminal 2 corresponding to user 2, ..., terminal n corresponding to user n, where n is an integer greater than 0.
[0018] Each terminal in the terminal cluster can be a smart device with communication capabilities, including but not limited to wearable devices, handheld devices, personal computers, tablet computers, smartphones, computing devices, or other processing devices connected to a wireless modem. Smart devices may be called different names in different networks, such as user equipment, access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolution networks.
[0019] In some embodiments, a server is a hardware device with strong computing capabilities. Specifically, the server can use a separate server device, such as a rack-mounted, blade, tower, or cabinet-mounted server device, or a workstation, mainframe computer, or other hardware device. A server cluster composed of multiple servers can also be used. The servers in the service cluster can be composed in a symmetrical manner, where each server has equivalent functions and status in the transaction link, and each server can provide services to the outside world independently. Providing services independently can be understood as not requiring the assistance of additional servers.
[0020] In some embodiments, the electronic device that executes the crowd label determination method may be a server, and the server may establish a communication connection with the terminals in the terminal cluster, and complete the data exchange in the crowd label determination process based on the communication connection. For example, in the crowd label determination method, the terminals in the terminal cluster generate initial user attribute data of at least one user in the target population. The server, after obtaining authorization from the terminal, obtains the initial user attribute data of the target population from at least two terminals. Afterwards, the server uses a crowd clustering model to perform crowd clustering processing based on the initial user attribute data to obtain crowd cluster representative feature data corresponding to the target cluster; the server uses a large crowd detection model to perform crowd label detection processing based on the crowd cluster representative feature data to obtain crowd label description information corresponding to the target cluster.
[0021] It should be noted that the server and the terminal establish a communication connection through a network for interactive communication, wherein the network can be a wireless network or a wired network. Wireless networks include but are not limited to cellular networks, wireless local area networks, infrared networks, or Bluetooth networks, and wired networks include but are not limited to Ethernet, universal serial bus (USB), or controller area network. In one or more embodiments of the specification, technologies and / or formats including Hypertext Markup Language (HTML) and Extensible Markup Language (XML) are used to represent data (such as target compressed packages) exchanged over the network. In addition, conventional encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0022] The crowd tag determination system embodiment provided in this specification and the crowd tag determination method described in one or more embodiments are based on the same concept. The execution subject corresponding to the crowd tag determination method involved in one or more embodiments of the specification can be an electronic device, and the electronic device can be the above-mentioned server. The specific implementation process of the crowd tag determination system embodiment can be found in the following method embodiment and will not be repeated here.
[0023] In one embodiment, Figure 2 As shown, a method for determining crowd labels is proposed. This method can be implemented by a computer program and can be run on a crowd label determination device based on the von Neumann architecture. The computer program can be integrated into an application or run as an independent tool application.
[0024] Specifically, the method for determining the crowd label includes: S102: Obtain initial user attribute data of the target population.
[0025] The target population refers to a group of users who, under certain conditions, have potential interest in or demand for a product, service, or brand. For example, in the advertising sector, the target population could be the user group currently using an advertising platform, which could include websites, software, mini-programs, and other platforms. Another example, in the consumer finance sector, the target population could be the user group interested in or currently using consumer financial services.
[0026] Initial user attribute data refers to raw user data that has not been processed and is used to describe the multi-dimensional characteristics of a user. In the embodiments of this specification, initial user attribute data may include feature data in multiple dimensions, such as user profiles, user behavior, and user preferences. A user profile is a collection of basic user information, which can include basic user information and some long-term, stable characteristics, such as age, gender, city, and registration authorization information. User behavior refers to the behavioral records generated by users when using products or services, such as visit frequency, transaction behavior, in-product behavior paths, product reviews, etc. User preferences refer to the tendencies or choices displayed by users during use, such as user preferences for product features and services.
[0027] In some embodiments, initial user attribute data of the target population can be obtained from a server or database. The server or database can store the initial user attribute data generated by the target population within a historical time period. The initial user attribute data is collected after authorization by all users in the target population.
[0028] S104: Performing crowd clustering processing based on the initial user attribute data using a crowd clustering model to obtain crowd cluster representative feature data corresponding to the target cluster.
[0029] The crowd clustering model is an unsupervised learning model that aims to divide crowd data in a crowd dataset into different groups. The crowd clustering model aims to group samples in a dataset based on the similarity of certain features. Each group is called a cluster. The data within each cluster is highly similar, while the data between clusters is less similar. In other words, the crowd clustering model aims to make elements within the same cluster as similar as possible, while elements between different clusters are as different as possible.
[0030] Target clusters are the clusters created by the crowd clustering model by dividing the initial user attribute data based on feature similarity. Within each target cluster, there is a cluster center, which can be a data point within the cluster or the mean of all data points within the cluster.
[0031] The representative characteristic data of a crowd cluster refers to the characteristic data of the cluster center user or the representative user in the target cluster.
[0032] In some embodiments, the initial user attribute data is divided into user feature data belonging to different target clusters using a population clustering model, and representative feature data of the population clusters is selected based on the user feature data within the target clusters. It is understood that the user feature data within each target cluster is highly similar, while the similarity of user feature data between different clusters is relatively low. The number of target clusters can be manually determined based on the business scenario and expert experience, for example, by manually setting the number of clusters desired to be divided by the population clustering model.
[0033] S106: Using a large crowd detection model, a crowd label detection process is performed based on the representative feature data of the crowd clusters to obtain crowd label description information corresponding to the target cluster.
[0034] Among them, the crowd detection large model refers to a large language model (LLM) applied to crowd detection scenarios.
[0035] The crowd label description information may include the crowd label and crowd interpretation information of the crowd label.
[0036] In the embodiments of this specification, the crowd detection model is intended to determine the crowd labels corresponding to different target clusters and the crowd interpretation information of the crowd labels through the crowd cluster representative feature data, thereby determining the crowd label description information based on the crowd labels and the label interpretation information of the crowd labels. It is understandable that since the target clusters can be determined by the crowd clustering model, the crowd clustering model can naturally determine the users belonging to different target clusters from the target population, then the crowd label description information corresponding to the target cluster is the crowd label description information of the users belonging to the target cluster. It is also understandable that the crowd label description information can be used for crowd segmentation processing of unknown populations, transaction processing of known populations, and so on.
[0037] The crowd label determination method provided in the embodiments of this specification first obtains the initial user attribute data of the target population, then uses a crowd clustering model to perform crowd clustering processing to obtain the crowd cluster representative feature data corresponding to the target cluster cluster, and then uses a crowd detection large model to perform crowd label detection processing to obtain the crowd label description information corresponding to the target cluster cluster. In this way, by inputting the crowd cluster representative feature data corresponding to different cluster clusters into the crowd detection large model for crowd label detection processing, it is possible to obtain high-precision crowd label description information for each individual in the target population, thereby improving the efficiency of annotating crowd labels. In addition, through the powerful understanding ability of the large model, the information dimension of the crowd label description information is also enriched. Therefore, the embodiments of this specification achieve the efficient provision of high-precision and multi-dimensional crowd label description information for the target population.
[0038] See Figure 3 , is a flow chart of another embodiment of a method for determining a crowd label provided in the embodiments of this specification. Specifically, the method may include the following steps: S202: Obtain initial user attribute data of the target population.
[0039] Specifically, the implementation of step S202 can be found in Figure 2 The description of the relevant steps in the illustrated embodiment will not be repeated in detail here.
[0040] S204 : Determine characteristic attributes of the initial user attribute data, and determine a crowd clustering model from at least two reference crowd clustering models based on the characteristic attributes.
[0041] It is understood that feature attributes include feature dimensions and feature types. Feature dimensions refer to the number of independent information units contained in the initial user attribute data of each user in the target population. These units are used to describe different aspects of a user's attributes or characteristics. Feature types refer to the types of independent information units contained in the initial user attribute data of each user in the target population. For example, feature types can include but are not limited to numerical features and categorical features.
[0042] In some embodiments, feature dimensions and feature types can be determined from initial user attribute data, and then a population clustering model that matches the feature dimensions and feature types can be selected from at least two reference population clustering models. Specifically, at least two clustering algorithms can be used as the at least two reference population clustering models. The clustering algorithms may include, but are not limited to, K-means clustering, K-Prototypes clustering, spectral clustering, Balanced Iterative Reducing and Clustering using Hierarchies (BIRCH), and other clustering algorithms.
[0043] For example, when the feature dimension is high, K-means can be used as a crowd clustering model. Alternatively, when the feature dimension is low, the feature types include both numerical and categorical features, and there are many categorical features, K-Prototypes can be used as a crowd clustering model.
[0044] S206: Performing crowd clustering processing on the initial user attribute data using a crowd clustering model to obtain crowd cluster representative feature data corresponding to the target cluster.
[0045] In some embodiments, executing step S206 may specifically include the following steps: A2: Determine the feature processing type of the initial user attribute data based on the feature dimension; A4: If the feature processing type is dimensionality reduction, the initial user attribute data is subjected to feature cleaning to obtain reference user feature data, the reference user feature data is subjected to feature dimensionality reduction to obtain first target user feature data, the first target user feature data is subjected to population clustering using a population clustering model to obtain reference population cluster representative feature data corresponding to the target cluster, reference population feature data belonging to the target cluster is determined from the first target user feature data, a feature distance between the reference population feature data and the reference population cluster representative feature data is determined, a target user identifier is determined based on the feature distance and a preset distance threshold, original population feature data corresponding to the target user identifier is determined from the reference user feature data, and the original population feature data is determined as the first population cluster representative feature data corresponding to the target cluster; A6: If the feature processing type is dimension-invariant, the initial user attribute data is subjected to feature cleaning processing to obtain second target user feature data. The second target user feature data is subjected to population clustering processing using a population clustering model to obtain second population cluster representative feature data corresponding to the target cluster cluster.
[0046] When executing step A2, the size of the feature dimension is compared with the preset dimension threshold. When the feature dimension is greater than or equal to the preset dimension threshold, the feature processing type of the initial user attribute data is determined to be the dimension reduction type; when the feature dimension is less than the preset dimension threshold, the feature processing type of the initial user attribute data is determined to be the dimension unchanged type.
[0047] When executing step A4, the initial user attribute data is subjected to feature cleaning processing to obtain reference user feature data. One feasible method is to perform missing value processing, outlier processing, data consistency verification, data standardization and other processing on the initial user attribute data to obtain reference user feature data.
[0048] When executing step A4, the reference user feature data is subjected to feature dimensionality reduction processing to obtain first target user feature data. One possible implementation method is: the encoder in the autoencoder (AE) performs low-dimensional space mapping processing on the reference user feature data to obtain low-dimensional representation data of the reference user feature data, and the decoder in the autoencoder performs feature reconstruction processing on the low-dimensional representation data of the reference user feature data to obtain the first target user feature data. It can be understood that the first target user feature data is data with a lower dimension than the original data that is reconstructed by the autoencoder based on the low-dimensional representation data and is as close as possible to the original data (i.e., the reference user feature data). The autoencoder used in the embodiment of this specification can be pre-trained. During the training process, the encoder and decoder are trained by minimizing the error between the original data and the reconstructed data (such as using the mean square error), so that the encoder learns how to retain the most important information by reducing the dimension (mapping the original data to the latent space), and the decoder learns how to reconstruct data with a lower dimension than the original data that is as close as possible to the original data based on these low-dimensional features.
[0049] When executing step A4, the first target user feature data is subjected to a crowd clustering process using a crowd clustering model to obtain reference crowd cluster representative feature data corresponding to the target cluster. One possible implementation method is: taking the crowd clustering algorithm as K-means as an example, the K value of the crowd clustering model is set according to expert experience and business requirements, and the first target user feature data is subjected to a crowd clustering process using the crowd clustering model with the set K value to obtain reference crowd cluster representative feature data corresponding to the target cluster. The crowd clustering model obtains the reference crowd cluster representative feature data corresponding to the target cluster, and the specific implementation process can be: Initialization step: randomly select K first target user feature data from all first target user feature data as the initial cluster centers of K target clusters; Assignment step: for each first target user feature data except the K first target user feature data in all first target user feature data, calculate its first cluster distance to the K initial cluster centers, arrange the K first cluster distances in ascending order, and assign it to the target cluster corresponding to the first cluster distance ranked first; wherein the first cluster distance can be selected from Euclidean distance or Manhattan distance; Update step: Recalculate the cluster center of each target cluster. The cluster center can be updated by calculating the mean of all the first target user feature data in each target cluster. For a target cluster, its cluster center is the mean of all the first target user feature data in each feature dimension.
[0050] Iteration step: Repeat the above “assignment step” and “update step” until the preset stopping condition is met, such as reaching the maximum number of iterations or the cluster center no longer changes; the cluster center after the iteration is completed is used as the representative characteristic data of the reference population clustering.
[0051] In step A4, the reference population characteristic data belonging to the target cluster is determined from the first target user characteristic data, and the characteristic distance between the reference population characteristic data and the representative characteristic data of the reference population cluster is determined. This can be understood as: determining the reference population characteristic data belonging to each target cluster from the first target user characteristic data, and for each target cluster, determining the characteristic distance between each reference population characteristic data therein and the representative characteristic data of the reference population cluster therein.
[0052] In step A4, the target user identifier is determined based on the feature distance and the preset distance threshold, the original population feature data corresponding to the target user identifier is determined from the reference user feature data, and the original population feature data is determined as the first population cluster representative feature data corresponding to the target cluster cluster. It can be understood as follows: for each target cluster cluster, the user identifier corresponding to the reference population feature data whose feature distance is greater than or equal to the preset distance threshold is determined, and a preset number of target user identifiers are selected from the user identifiers. The target user identifier is the identifier of a preset number of representative users that can represent a certain target cluster cluster selected from users who are closer to the cluster center. The reference user feature data includes the feature data of all users before dimensionality reduction. The feature data of these representative users before dimensionality reduction, that is, the original population feature data corresponding to the target user identifier, can be determined from the reference user feature data. Then, the original population feature data is determined as the first population cluster representative feature data corresponding to the target cluster cluster. The first population cluster representative feature data is the feature data of representative users in the target cluster cluster who are closer to the cluster center before dimensionality reduction.
[0053] In step A6, the initial user attribute data is subjected to feature cleaning processing to obtain second target user feature data. One possible implementation method is to perform missing value processing, outlier processing, data consistency verification, data standardization and other processing on the initial user attribute data to obtain second target user feature data.
[0054] In step A6, the second target user feature data is subjected to a crowd clustering process using a crowd clustering model to obtain second crowd cluster representative feature data corresponding to the target cluster. One possible implementation method is: taking the crowd clustering algorithm as K-Prototypes as an example, the K value of the crowd clustering model is set based on expert experience and business requirements, and the second target user feature data is subjected to a crowd clustering process using the crowd clustering model with the set K value to obtain second crowd cluster representative feature data corresponding to the target cluster. The crowd clustering model obtains the second crowd cluster representative feature data corresponding to the target cluster, and the specific implementation process can be: Initialization step: randomly select K second target user feature data from all second target user feature data as the initial cluster centers of K target clusters.
[0055] Assignment step: for each second target user feature data except the K second target user feature data in all the second target user feature data, comprehensively calculate the second clustering distance of each second target user feature data to the K initial cluster centers based on the numerical feature and the categorical feature, arrange the K second clustering distances in ascending order, and classify them into the target cluster corresponding to the second clustering distance ranked first; wherein the second clustering distance can be determined by: for each second target user feature data except the K second target user feature data, calculate the first distance of its numerical feature and the numerical feature of the K initial cluster centers, calculate the second distance of its categorical feature and the categorical feature of the K initial cluster centers, and weight the first distance and the second distance to obtain the second clustering distance of each second target user feature data to the K initial cluster centers.
[0056] Update step: recalculate the cluster center of each target cluster. The cluster center can be updated by calculating the mean of the numerical features of all the first target user feature data in each target cluster and determining the category feature that appears the most times in all the first target user feature data in each target cluster.
[0057] Iteration step: Repeat the above “assignment step” and “update step” until the preset stopping condition is met, such as reaching the maximum number of iterations or the cluster center no longer changes; the cluster center after the iteration is completed is used as the representative feature data of the second population clustering.
[0058] As such, for initial user attribute data with high feature dimensions, such as data with hundreds or thousands of feature dimensions, it is difficult for an unsupervised clustering model to directly learn, and there may even be a problem of excessive deviation in the results generated by different iteration rounds. To avoid the above problems, the embodiments of this specification use dimensionality reduction of the initial user attribute data, and then cluster the reduced feature data using an unsupervised crowd clustering model to obtain feature data of representative users corresponding to different clusters. For initial user attribute data with low feature dimensions, this embodiment uses an unsupervised crowd clustering model to cluster the feature data obtained by data cleaning to obtain feature data of cluster center users corresponding to different clusters.
[0059] S208: Determine a tag generation prompt word based on the representative characteristic data of the crowd cluster.
[0060] In some embodiments, executing step S208 may specifically include the following steps: B2: Determine the cluster center type corresponding to the representative feature data of the crowd cluster, and obtain the feature input prompt information corresponding to the cluster center type; B4: Generate labels and prompt words based on the representative feature data of the crowd cluster and the feature input prompt information.
[0061] In step B2, the cluster center type is used to characterize the user with representative characteristic data of a population cluster as the cluster center or representative user of the cluster. Optionally, the cluster center type may include a first type and a second type, wherein the first type is used to characterize the user with representative characteristic data of a population cluster as the cluster center of the cluster, and the second type is used to characterize the user with representative characteristic data of a population cluster as the representative user of the cluster.
[0062] When executing step B2, obtain the cluster identifier corresponding to the representative characteristic data of the crowd cluster, and determine the cluster center type corresponding to the representative characteristic data of the crowd cluster based on the cluster identifier. If the cluster center type indicates that the user of the representative characteristic data of the crowd cluster is the cluster center of the cluster cluster, then determine the feature input prompt information as the first input description information, and the first input description information includes indication information that the feature data input to the crowd detection large model is the feature data of the cluster center of the cluster cluster; if the cluster center type indicates that the user of the representative characteristic data of the crowd cluster is the representative user of the cluster cluster, then determine the feature input prompt information as the second input description information, and the second input description information includes indication information that the feature data input to the crowd detection large model is the feature data of the representative user of the cluster cluster.
[0063] During step B4, role description information, crowd detection task description information, background information, and output description information are obtained, and a label generation prompt is generated based on the crowd cluster representative feature data, feature input prompt information, role description information, crowd detection task description information, background description information, and output description information. It is understood that role description information refers to descriptive information about the identity, responsibilities, or specific persona of the role set for the large model. Role description information is used to indicate the perspective, stance, and style that the large model should follow when answering questions or performing tasks. Crowd detection task description information refers to task instructions that instruct the large model to perform crowd label detection processing to obtain crowd label description information corresponding to different target clusters. Background description information refers to some pre-condition information related to the crowd detection task, such as environmental descriptions, relevant knowledge, or historical conversations. Background description information is intended to provide the large model with the context and background of the problem, helping the model better understand the ins and outs of the problem and related factors. Output description information refers to the requirements and specifications for the output content to be generated by the large model, including specific instructions on the output format, structure, content depth, language style, and length. The output description information is intended to clarify the standards that the model should follow and the expected results when answering questions.
[0064] Specifically, a preset prompt word template can be obtained, and crowd cluster representative feature data, feature input prompt information, role description information, crowd detection task description information, background description information, and output description information can be filled in the prompt word template to obtain a label generation prompt word.
[0065] S210, input the label generation prompt word into the crowd detection large model, perform relevant feature combination processing through the crowd detection large model to obtain cluster representative joint feature data, and perform crowd label prediction processing based on the cluster representative joint feature data to obtain crowd label description information corresponding to the target cluster cluster.
[0066] In some embodiments, relevant feature combination processing is performed through a large crowd detection model to obtain cluster representative joint feature data. One feasible way is: relevant feature detection processing is performed on the crowd cluster representative feature data through a large crowd detection model to obtain cluster representative relevant feature data, and the cluster representative relevant feature data includes at least two relevant crowd cluster representative feature data. Afterwards, all crowd cluster representative feature data and cluster representative relevant feature data are respectively determined as cluster representative joint feature data, that is, the cluster representative joint feature data includes both individual crowd cluster representative feature data and relevant cluster representative relevant feature data.
[0067] For example, the crowd detection model considers that the two labels "like sports" and "like healthy eating" are correlated, and users who like sports also like healthy eating. Then the crowd detection model will detect from the crowd cluster representative feature data that the feature data reflecting the label "like sports" and the feature data reflecting the label "like healthy eating" are correlated feature data, and thus associate these two feature data to obtain cluster representative related feature data.
[0068] Therefore, by performing relevant feature detection processing on the relevant feature data of the population cluster representatives, relevant cluster representative relevant feature data can be obtained. The cluster representative relevant feature data can be used to characterize the intersection between features of different dimensions. After obtaining the cluster representative joint feature data including the cluster representative relevant feature data and the population cluster representative feature data, the population label description information corresponding to the target cluster cluster is obtained based on the cluster representative joint feature data. This can achieve a more accurate division of each population based on the intersection between features of different dimensions, thereby obtaining accurate population labels.
[0069] Performing crowd label prediction processing based on cluster representative joint feature data to obtain crowd label description information corresponding to the target cluster cluster may specifically include the following steps: C2: Crowd label classification is performed based on the joint feature data of cluster representatives to obtain the crowd labels corresponding to the target cluster; C4: Analyze the content of crowd labels to obtain crowd interpretation information; C6: Generate crowd label description information corresponding to the target cluster based on crowd labels and crowd interpretation information.
[0070] When executing step C2, the crowd detection model uses a built-in binary classifier to classify the cluster representative joint feature data into crowd labels, and obtains the crowd label of the target cluster cluster to which the cluster representative joint feature data belongs. Optionally, the binary classifier can be a one-to-one classifier for identifying two different types of crowd labels, or the binary classifier can be a one-to-many classifier for identifying a certain type of crowd label and a non-type of crowd label. It is understandable that the crowd label can include multi-dimensional labels such as age, gender, city, interests, and behavior.
[0071] In step C4, after determining the crowd label to which the target cluster cluster where the cluster representative joint feature data belongs, the crowd detection model uses natural language generation technology to perform content analysis on the crowd label, explain the meaning of the crowd label and the basis for its formation, and then generate crowd interpretation information based on the meaning of the crowd label and the basis for its formation.
[0072] For example, if the crowd labels include: high-value users, loyal users, and new users, then the crowd interpretation information may include the following information: Definition of population labels: High-value users refer to users who have spent more than 1,000 yuan in the past three months; loyal users refer to users who have purchased at least once a month for three consecutive months; new users refer to users who have been registered for no more than one month.
[0073] The basis for forming population labels: High-value users: According to analysis, this type of users has a high purchase frequency and relatively high brand loyalty, so we define them as high-value users; Loyal users: Surveys show that loyal users have a higher repurchase rate than ordinary users, so the importance of this label lies in identifying and maintaining such user groups; New users: The definition of new users is based on registration time and is used to formulate welcome activities and promotion strategies.
[0074] In step C6 , the crowd label and the crowd interpretation information corresponding to the crowd label are associated to generate crowd label description information including the crowd label and the crowd interpretation information corresponding to the crowd label.
[0075] S212: Perform indicator analysis on the transaction indicators corresponding to the target population using the population tag description information to obtain indicator analysis results, and perform transaction processing based on the indicator analysis results.
[0076] It is understandable that transaction indicators may include indicators related to other matters such as operational matters, advertising matters, etc., such as conversion rate, retention rate, purchase volume and other indicators.
[0077] In some embodiments, different population tags in the population tag description can be correlated with transaction indicators to obtain indicator analysis results. The indicator analysis results can be used to characterize the response effects of users with different population tags in different transactions. Afterwards, different transaction processing strategies can be adopted for different indicator analysis results.
[0078] For example, taking the transaction indicator as conversion rate, the population labels including "high-activity users" and "low-activity users", and the indicator analysis result as the conversion rate difference as an example, by analyzing the conversion rate difference between the two population labels of "high-activity users" and "low-activity users", the transaction processing performed according to the indicator analysis results is: judging which transaction strategies adopted by the operation transaction are more effective for different populations, and which transaction strategies need to be improved, so as to increase the adoption of effective transaction strategies and adjust the transaction strategies that need to be improved to achieve optimized processing of transaction strategies.
[0079] Optionally, the method for determining a crowd tag in the embodiment of this specification may further include: performing opportunity insight transaction processing on the target crowd using crowd tag description information.
[0080] Specifically, opportunity insight processing involves analyzing demographic data to identify potential market opportunities and predict future development trends. Identifying potential market opportunities can be understood as analyzing existing tagged groups to uncover unmet needs for specific groups. For example, if the group labeled "high-paid, travel-loving women" is found to be highly interested in customized travel services, then customized travel products and services could be developed specifically for this group.
[0081] Predicting future development trends can be understood as using target demographic data (such as consumption upgrades and increased health awareness) to predict future market trends. For example, based on the growth trends of the "young, environmentally conscious" demographic, you can plan ahead for green products or services.
[0082] In the method for determining a crowd label provided in this specification, first, the initial user attribute data of the target crowd is obtained, and then the characteristic attributes of the initial user attribute data are determined. Based on the characteristic attributes, a crowd clustering model is determined from at least two reference crowd clustering models, and the crowd clustering model is used to perform crowd clustering processing on the initial user attribute data to obtain crowd cluster representative characteristic data corresponding to the target cluster. In this way, by selecting a suitable crowd clustering model through characteristic attributes, the unsupervised crowd clustering model can achieve a better learning effect, thereby accurately dividing users into corresponding clusters and extracting accurate crowd cluster representative characteristic data from the clusters; further, based on Based on the representative characteristic data of the crowd clusters, label generation prompt words are determined, and the label generation prompt words are input into the crowd detection large model. The crowd detection large model performs relevant feature combination processing to obtain cluster representative joint characteristic data, and based on the cluster representative joint characteristic data, crowd label prediction processing is performed to obtain crowd label description information corresponding to the target cluster cluster. In this way, by generating accurate prompt words through the representative characteristic data of the crowd clusters, the crowd detection large model can obtain accurate output content according to requirements; then, the crowd label description information is used to perform indicator analysis processing on the transaction indicators corresponding to the target population to obtain indicator analysis results, and transaction processing is performed based on the indicator analysis results. In this way, the accurate crowd label description information corresponding to different populations output by the large model can be used to perform transaction indicator analysis and transaction processing, providing support for related transactions for the target population, thereby enabling more accurate transaction decision-making and strategy formulation.
[0083] See Figure 4 , provides a flow chart of a method for training a large crowd detection model according to an embodiment of the present application. Specifically, the method may include the following steps: S302, creating an initial crowd detection model.
[0084] In some embodiments, a large language model (LLM) can be used as the initial crowd detection model. An LLM is a deep learning model trained using large amounts of text data. It can generate natural language text or understand the meaning of text. An LLM can handle a variety of natural language tasks, such as text classification, question answering, and conversation. Alternatively, an LLM can be a generative artificial intelligence (AIGC) model, such as ChatGPT, Gemini, and Claude.
[0085] In some other embodiments, an initial crowd detection model can be created based on the large language model. For example, the architecture or parameters of the large language model can be adapted and adjusted to obtain the initial crowd detection model. Specifically, one possible implementation method for creating the initial crowd detection model based on the large language model is to obtain a basic large language model, create a crowd label detection scenario adaptation module for the crowd label detection scenario, and a large language generation module based on the basic large language model. The initial crowd detection model is composed of the large language generation module and the crowd label detection scenario adaptation module.
[0086] S304: Determine sample user attribute data of the sample population, perform population clustering processing on the sample user attribute data to obtain sample population cluster representative feature data corresponding to the sample cluster clusters.
[0087] It is understandable that a transaction scenario of a large crowd detection model can be determined, and a sample population can be obtained from the transaction scenario, where the sample population includes a large number of sample users.
[0088] Sample user attribute data refers to raw user data that has not been processed and is used to describe the multi-dimensional characteristics of sample users. In the embodiments of this specification, sample user attribute data may include feature data of multiple dimensions, such as sample user portraits, sample user behaviors, and sample user preferences. Among them, the sample user portrait is a collection of basic information of the sample user. The sample user portrait may include basic information of the sample user and some long-term, stable characteristics, such as age, gender, city, registration authorization information, etc. The sample user behavior refers to the behavioral records generated by the sample user when using a product or service, such as visit frequency, transaction behavior, behavioral flow within the product, product evaluation, etc. The sample user preferences refer to the tendencies or choices displayed by the sample user during use, such as the sample user's preferred product features, sample product services, etc.
[0089] In some embodiments, performing a population clustering process on the sample user attribute data to obtain representative characteristic data of the sample population cluster corresponding to the sample cluster may specifically include the following steps: D2: determining characteristic attributes of the sample user attribute data, and determining a population clustering model from at least two reference population clustering models based on the characteristic attributes; D4: using the population clustering model to perform a population clustering process on the sample user attribute data to obtain representative characteristic data of the sample population cluster corresponding to the sample cluster. Specifically, the implementation of steps D2 and D4 can be specifically referred to. Figure 3 The description of the relevant parts in the illustrated embodiment will not be repeated in detail here.
[0090] S306 , clustering representative characteristic data of the sample population and labeling the sample population.
[0091] In some embodiments, representative feature data of the sample population is clustered and labeled with sample population labels based on expert experience.
[0092] S308, use the representative feature data of the sample population clusters to perform at least one round of model training on the initial population detection large model. During the model training process, call the initial population detection large model based on the representative feature data of the sample population clusters to perform population label prediction processing to obtain the predicted population label corresponding to the sample cluster cluster, determine the comprehensive model loss value based on the predicted population label and the sample population label, and adjust the model parameters of the initial population detection large model based on the comprehensive model loss value to obtain the population detection large model.
[0093] In some embodiments, during the model training process, sample label generation prompt words can be generated based on the representative characteristic data of the sample population cluster, and the sample label generation prompt words can be input into the initial population detection large model. The population label prediction processing is performed through the initial population detection large model to obtain the predicted population label corresponding to the sample cluster cluster.
[0094] Optionally, the model parameters of the initial crowd detection large model are adjusted based on the comprehensive model loss value to obtain the crowd detection large model. One implementation method is: based on the comprehensive model loss value, the model parameters of the crowd label detection scene adaptation module in the initial crowd detection large model are adjusted, and the model parameters of the large language generation module are controlled to remain unchanged until the end conditions of the model training are met to obtain the large language generation module and the crowd label detection scene adaptation module, complete the model fusion of the large language generation module and the crowd label detection scene adaptation module, and obtain the trained crowd detection large model.
[0095] Specifically, the model fusion of the large language generation module and the crowd label detection scenario adaptation module can be: weight fusion of the model structure layer weight of the crowd label detection scenario adaptation module and the large language generation module, by determining the target model structure layer corresponding to the model structure layer weight in the large language generation module, parameter fusion of the model structure layer parameters of the target model structure layer and the model structure layer weight, the model structure layer weight of the crowd label detection scenario adaptation module can only partially correspond to and have model structure layer weights in all model structure layers in the basic large language model, by completing the parameter update of the model structure layer based on the model structure layer weight for this part of the target model structure layer, and so on to complete the reference update process of all model structure layer weights, thereby obtaining the crowd detection large model.
[0096] Optionally, the model training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset number threshold, etc. Specific model training termination conditions may be determined based on actual conditions and are not specifically limited here.
[0097] In some embodiments, see Figure 5 FIG. 1 is a flow chart of a method for determining a loss value of a comprehensive model, wherein the method specifically includes: S402: Determine at least two first population labels from the sample population labels, and determine second population labels corresponding to the first population labels from the predicted population labels; S404: Determine a label loss value based on the first group label and the second group label, and determine a comprehensive model loss value based on the label loss value.
[0098] In step S402, each label in the sample population label is determined as a first population label, and a second population label belonging to the same label dimension as the first population label is determined from the predicted population label. For example, the two first population labels in the sample population label are "like sports" and "like shopping", and the two second population labels in the predicted population label are "neither like nor dislike sports" and "like shopping". Then, "neither like nor dislike sports" is the second population label belonging to the same label dimension as "like sports", and "like shopping" is the second population label belonging to the same label dimension as "like shopping".
[0099] In step S404, a preset loss function may be used to calculate the first and second population labels to obtain a label loss value. The preset loss function may include a logistic regression loss function, a cross-entropy loss function, or the like. For each pair of labels consisting of the first and second population labels, a label loss value may be obtained. All label loss values may be weighted to obtain a comprehensive model loss value.
[0100] In the training method of the crowd detection large model provided in this specification, an initial crowd detection large model is created, sample user attribute data of the sample population is determined, the sample user attribute data is subjected to crowd clustering processing to obtain sample population cluster representative feature data corresponding to the sample cluster cluster, the sample population cluster representative feature data is labeled with a sample population label, and the sample population cluster representative feature data is used to perform at least one round of model training on the initial crowd detection large model. During the model training process, the initial crowd detection large model is called based on the sample population cluster representative feature data to perform crowd label prediction processing to obtain a predicted population label corresponding to the sample cluster cluster, a comprehensive model loss value is determined based on the predicted population label and the sample population label, and a model parameter adjustment processing is performed on the initial crowd detection large model based on the comprehensive model loss value to obtain a crowd detection large model. In this way, by pre-training the initial crowd detection large model using the feature data of users who are closer to the center of the cluster or the feature data of users who belong to the center of the cluster, the output effect of the trained crowd detection large model can be improved.
[0101] The following will be combined Figure 6 , the crowd label determination device provided in the embodiment of this specification is introduced in detail. It should be noted that, Figure 6 The crowd tag determination device shown is used to execute this instruction Figures 1 to 3 For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to this specification. Figures 1 to 3 The embodiment shown.
[0102] See Figure 6 , which shows a schematic diagram of the structure of a crowd label determination device according to an embodiment of the present specification. The crowd label determination device 1 can be implemented as all or part of the device through software, hardware, or a combination of both. According to some embodiments, the crowd label determination device 1 includes a data acquisition module 11, a crowd clustering module 12, and a label detection module 13, which are specifically used to: Data acquisition module 11, used to obtain initial user attribute data of the target population; A crowd clustering module 12 is configured to perform crowd clustering processing based on the initial user attribute data using a crowd clustering model to obtain crowd cluster representative feature data corresponding to a target cluster; The label detection module 13 is used to use a large crowd detection model to perform crowd label detection processing based on the representative feature data of the crowd clusters, and obtain crowd label description information corresponding to the target cluster.
[0103] Optionally, the tag detection module 13 includes: a prompt word generating unit, configured to determine a label and generate a prompt word based on the representative characteristic data of the crowd cluster; A model detection unit is used to input the label generation prompt word into the crowd detection large model, perform relevant feature combination processing through the crowd detection large model to obtain cluster representative joint feature data, and perform crowd label prediction processing based on the cluster representative joint feature data to obtain crowd label description information corresponding to the target cluster cluster.
[0104] Optionally, the prompt word generating unit is specifically used to: Determine the cluster center type corresponding to the representative feature data of the crowd cluster, and obtain feature input prompt information corresponding to the cluster center type; A label is generated based on the representative characteristic data of the crowd cluster and the characteristic input prompt information to generate prompt words.
[0105] Optional model checking unit, specifically used for: Performing crowd label classification processing based on the cluster representative joint feature data to obtain crowd labels corresponding to the target cluster; Performing label content analysis on the crowd labels to obtain crowd interpretation information; Crowd label description information corresponding to the target cluster is generated based on the crowd label and the crowd interpretation information.
[0106] Optionally, the crowd clustering module 12 includes: a model selection unit, configured to determine characteristic attributes of the initial user attribute data, and determine a crowd clustering model from at least two reference crowd clustering models based on the characteristic attributes; The model clustering unit is used to perform crowd clustering processing on the initial user attribute data using the crowd clustering model to obtain crowd cluster representative feature data corresponding to the target cluster.
[0107] Optional, model clustering unit, specifically used for: determining a feature processing type of the initial user attribute data based on the feature attribute; If the feature processing type is a dimensionality reduction type, feature cleaning processing is performed on the initial user attribute data to obtain reference user feature data, feature dimensionality reduction processing is performed on the reference user feature data to obtain first target user feature data, and population clustering processing is performed on the first target user feature data using the population clustering model to obtain reference population cluster representative feature data corresponding to the target cluster cluster, reference population feature data belonging to the target cluster cluster is determined from the first target user feature data, a feature distance between the reference population feature data and the reference population cluster representative feature data is determined, a target user identifier is determined based on the feature distance and a preset distance threshold, original population feature data corresponding to the target user identifier is determined from the reference user feature data, and the original population feature data is determined as the first population cluster representative feature data corresponding to the target cluster cluster; If the feature processing type is the dimension-invariant type, the initial user attribute data is subjected to feature cleaning processing to obtain second target user feature data, and the second target user feature data is subjected to population clustering processing through the population clustering model to obtain second population cluster representative feature data corresponding to the target cluster cluster.
[0108] Optionally, the crowd label determination device 1 further includes: The transaction processing module is used to perform indicator analysis on the transaction indicators corresponding to the target population using the population label description information to obtain indicator analysis results, and perform transaction processing according to the indicator analysis results.
[0109] The crowd label determination device provided in the embodiment of this specification first obtains the initial user attribute data of the target population, then uses the crowd clustering model to perform crowd clustering processing to obtain the crowd cluster representative feature data corresponding to the target cluster cluster, and then uses the crowd detection large model to perform crowd label detection processing to obtain the crowd label description information corresponding to the target cluster cluster. In this way, by inputting the crowd cluster representative feature data corresponding to different cluster clusters into the crowd detection large model for crowd label detection processing, it is possible to obtain high-precision crowd label description information of individual individuals in the target population, thereby improving the efficiency of annotating crowd labels. In addition, through the powerful understanding ability of the large model, the information dimension of the crowd label description information is also enriched. Therefore, the embodiment of this specification achieves the efficient provision of high-precision and multi-dimensional crowd label description information for the target population.
[0110] The following will be combined Figure 7 , the crowd detection large model training device provided in the embodiment of this specification is introduced in detail. It should be noted that, Figure 7 The large crowd detection model training device shown is used to execute this instruction Figure 4~Figure 5For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to this specification. Figure 4~Figure 5 The embodiment shown.
[0111] See Figure 7 , which shows a schematic diagram of the structure of a large-scale model training device for crowd detection according to an embodiment of this specification. The large-scale model training device 2 for crowd detection can be implemented as all or part of the device through software, hardware, or a combination of both. According to some embodiments, the large-scale model training device 1 for crowd detection includes a model creation module 21, a sample acquisition module 22, a labeling module 23, and a model training module 24, which are specifically used to: Model creation module 21, used to create an initial crowd detection model; The sample acquisition module 22 is used to determine the sample user attribute data of the sample population, and perform population clustering processing on the sample user attribute data to obtain the sample population cluster representative feature data corresponding to the sample cluster cluster; A labeling module 23 is used to label the sample population cluster representative feature data with sample population labels; The model training module 24 is used to perform at least one round of model training on the initial population detection large model using the representative feature data of the sample population clusters. During the model training process, the initial population detection large model is called based on the representative feature data of the sample population clusters to perform population label prediction processing to obtain the predicted population label corresponding to the sample cluster cluster, and the comprehensive model loss value is determined based on the predicted population label and the sample population label. Based on the comprehensive model loss value, the model parameters of the initial population detection large model are adjusted to obtain the population detection large model.
[0112] Please refer to Figure 8 , which shows a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of this specification. The electronic device described in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.
[0113] Processor 110 may include one or more processing cores. Using various interfaces and circuits, processor 110 connects various components within the terminal. It executes instructions, programs, code sets, or instruction sets stored in memory 120, as well as accesses data stored in memory 120, to perform various functions and process data for terminal 100. Optionally, processor 110 may be implemented in hardware using at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). Processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may also be implemented independently of the processor 110 via a separate communications chip.
[0114] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (e.g., a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The operating system may be an Android system, including systems deeply developed based on the Android system, an iOS system developed by Apple, including systems deeply developed based on the iOS system, or other systems.
[0115] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0116] The input device 130 is used to receive input commands or data and includes, but is not limited to, a keyboard, a mouse, a camera, a microphone, or a touch-sensitive device. The output device 140 is used to output commands or data and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 may be combined, and the input device 130 and the output device 140 may be a touch-sensitive display.
[0117] The touch display screen can be designed as a full screen, a curved screen or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in the embodiments of this specification.
[0118] In addition, those skilled in the art will understand that the structures of the electronic devices shown in the above figures do not limit the electronic devices. The electronic devices may include more or fewer components than shown, or may combine certain components or arrange the components differently. For example, the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, and other components, which will not be described in detail here.
[0119] In some embodiments, Figure 8 In the electronic device shown, the processor 110 may be configured to call a program for determining a crowd label stored in the memory 120 and specifically perform the following operations: Obtain initial user attribute data of the target population; Using a crowd clustering model to perform crowd clustering processing based on the initial user attribute data to obtain crowd cluster representative feature data corresponding to the target cluster; A large crowd detection model is used to perform crowd label detection processing based on the representative feature data of the crowd clusters to obtain crowd label description information corresponding to the target clusters.
[0120] Optionally, when executing the step of using the large crowd detection model to perform crowd label detection based on the crowd cluster representative feature data to obtain crowd label description information corresponding to the target cluster, the processor 110 specifically performs the following operations: Determining a label generation prompt word based on the representative characteristic data of the crowd cluster; The label generation prompt word is input into the crowd detection model, and the relevant feature combination processing is performed through the crowd detection model to obtain cluster representative joint feature data, and crowd label prediction processing is performed based on the cluster representative joint feature data to obtain the crowd label description information corresponding to the target cluster cluster.
[0121] Optionally, when executing the step of determining a tag generation prompt word based on the representative characteristic data of the crowd cluster, the processor 110 specifically performs the following operations: Determine the cluster center type corresponding to the representative feature data of the crowd cluster, and obtain feature input prompt information corresponding to the cluster center type; A label is generated based on the representative characteristic data of the crowd cluster and the characteristic input prompt information to generate prompt words.
[0122] Optionally, when executing the step of performing crowd label prediction processing based on the cluster representative joint feature data to obtain crowd label description information corresponding to the target cluster, the processor 110 specifically performs the following operations: Performing crowd label classification processing based on the cluster representative joint feature data to obtain crowd labels corresponding to the target cluster; Performing label content analysis on the crowd labels to obtain crowd interpretation information; Crowd label description information corresponding to the target cluster is generated based on the crowd label and the crowd interpretation information.
[0123] Optionally, when executing the step of using a crowd clustering model to perform crowd clustering processing based on the initial user attribute data to obtain crowd cluster representative feature data corresponding to the target cluster, the processor 110 specifically performs the following operations: determining characteristic attributes of the initial user attribute data, and determining a crowd clustering model from at least two reference crowd clustering models based on the characteristic attributes; The crowd clustering model is used to perform crowd clustering processing on the initial user attribute data to obtain crowd cluster representative feature data corresponding to the target cluster.
[0124] Optionally, when executing the step of performing crowd clustering processing on the initial user attribute data using the crowd clustering model to obtain crowd cluster representative feature data corresponding to the target cluster, the processor 110 specifically performs the following operations: determining a feature processing type of the initial user attribute data based on the feature attribute; If the feature processing type is a dimensionality reduction type, feature cleaning processing is performed on the initial user attribute data to obtain reference user feature data, feature dimensionality reduction processing is performed on the reference user feature data to obtain first target user feature data, and population clustering processing is performed on the first target user feature data using the population clustering model to obtain reference population cluster representative feature data corresponding to the target cluster cluster, reference population feature data belonging to the target cluster cluster is determined from the first target user feature data, a feature distance between the reference population feature data and the reference population cluster representative feature data is determined, a target user identifier is determined based on the feature distance and a preset distance threshold, original population feature data corresponding to the target user identifier is determined from the reference user feature data, and the original population feature data is determined as the first population cluster representative feature data corresponding to the target cluster cluster; If the feature processing type is the dimension-invariant type, the initial user attribute data is subjected to feature cleaning processing to obtain second target user feature data, and the second target user feature data is subjected to population clustering processing through the population clustering model to obtain second population cluster representative feature data corresponding to the target cluster cluster.
[0125] Optionally, the processor 110 further performs the following operations: The group label description information is used to perform indicator analysis processing on the transaction indicators corresponding to the target group to obtain indicator analysis results, and transaction processing is performed according to the indicator analysis results.
[0126] In another embodiment, Figure 8 In the electronic device shown, the processor 110 can be used to call the program of the training method of the crowd detection large model stored in the memory 120, and specifically perform the following operations: Create an initial crowd detection model; Determining sample user attribute data of a sample population, performing population clustering processing on the sample user attribute data to obtain representative characteristic data of the sample population cluster corresponding to the sample cluster cluster; Clustering representative characteristic data of the sample population and labeling the sample population labels; The sample population cluster representative feature data is used to perform at least one round of model training on the initial population detection large model. During the model training process, the initial population detection large model is called based on the sample population cluster representative feature data to perform population label prediction processing to obtain the predicted population label corresponding to the sample cluster cluster, and a comprehensive model loss value is determined based on the predicted population label and the sample population label. Based on the comprehensive model loss value, model parameters of the initial population detection large model are adjusted to obtain the population detection large model.
[0127] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the initial user attribute data, user portraits, user behaviors, user preferences, etc. involved in this specification are all obtained with the full authorization of the user.
[0128] An embodiment of this specification also provides a computer-readable storage medium storing at least one instruction, wherein the at least one instruction is used to be executed by a processor to implement the crowd label determination method and / or the crowd detection large model training method as described in the above embodiments.
[0129] An embodiment of this specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the crowd label determination method and / or the crowd detection large model training method as described in the above embodiments.
[0130] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of this specification can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0131] The above description is only an optional embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
[0132] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A method for determining a crowd label, the method comprising: Obtain initial user attribute data of the target population; Using a crowd clustering model to perform crowd clustering processing based on the initial user attribute data to obtain crowd cluster representative feature data corresponding to the target cluster; A large crowd detection model is used to perform crowd label detection processing based on the representative feature data of the crowd clusters to obtain crowd label description information corresponding to the target cluster.
2. The method according to claim 1, wherein the large crowd detection model is used to perform crowd label detection based on the representative characteristic data of the crowd clusters to obtain crowd label description information corresponding to the target clusters, including: Determining a tag generation prompt word based on the representative characteristic data of the crowd cluster; The label generation prompt word is input into the crowd detection model, and the relevant feature combination processing is performed through the crowd detection model to obtain cluster representative joint feature data, and crowd label prediction processing is performed based on the cluster representative joint feature data to obtain the crowd label description information corresponding to the target cluster cluster.
3. The method according to claim 2, wherein determining a label generation prompt word based on the representative characteristic data of the crowd cluster comprises: Determine the cluster center type corresponding to the representative feature data of the crowd cluster, and obtain feature input prompt information corresponding to the cluster center type; A label is generated based on the representative characteristic data of the crowd cluster and the characteristic input prompt information to generate prompt words.
4. The method according to claim 2, wherein the step of performing crowd label prediction based on the cluster representative joint feature data to obtain crowd label description information corresponding to the target cluster comprises: Performing crowd label classification processing based on the cluster representative joint feature data to obtain crowd labels corresponding to the target cluster; Performing label content analysis on the crowd labels to obtain crowd interpretation information; Crowd label description information corresponding to the target cluster is generated based on the crowd label and the crowd interpretation information.
5. The method according to claim 1, wherein the method of using a crowd clustering model to perform crowd clustering processing based on the initial user attribute data to obtain crowd cluster representative feature data corresponding to the target cluster comprises: determining characteristic attributes of the initial user attribute data, and determining a crowd clustering model from at least two reference crowd clustering models based on the characteristic attributes; The crowd clustering model is used to perform crowd clustering processing on the initial user attribute data to obtain crowd cluster representative feature data corresponding to the target cluster.
6. The method according to claim 5, wherein the method of performing crowd clustering processing on the initial user attribute data using the crowd clustering model to obtain crowd cluster representative feature data corresponding to the target cluster comprises: determining a feature processing type of the initial user attribute data based on the feature attribute; If the feature processing type is a dimensionality reduction type, feature cleaning processing is performed on the initial user attribute data to obtain reference user feature data, feature dimensionality reduction processing is performed on the reference user feature data to obtain first target user feature data, and population clustering processing is performed on the first target user feature data using the population clustering model to obtain reference population cluster representative feature data corresponding to the target cluster cluster, reference population feature data belonging to the target cluster cluster is determined from the first target user feature data, a feature distance between the reference population feature data and the reference population cluster representative feature data is determined, a target user identifier is determined based on the feature distance and a preset distance threshold, original population feature data corresponding to the target user identifier is determined from the reference user feature data, and the original population feature data is determined as the first population cluster representative feature data corresponding to the target cluster cluster; If the feature processing type is the dimension-invariant type, the initial user attribute data is subjected to feature cleaning processing to obtain second target user feature data, and the second target user feature data is subjected to population clustering processing through the population clustering model to obtain second population cluster representative feature data corresponding to the target cluster cluster.
7. The method according to claim 1, further comprising: The group label description information is used to perform indicator analysis processing on the transaction indicators corresponding to the target group to obtain indicator analysis results, and transaction processing is performed according to the indicator analysis results.
8. A method for training a large crowd detection model, the method comprising: Create an initial crowd detection model; Determining sample user attribute data of a sample population, performing population clustering processing on the sample user attribute data to obtain representative characteristic data of the sample population cluster corresponding to the sample cluster cluster; Clustering representative characteristic data of the sample population and labeling the sample population labels; The sample population cluster representative feature data is used to perform at least one round of model training on the initial population detection large model. During the model training process, the initial population detection large model is called based on the sample population cluster representative feature data to perform population label prediction processing to obtain the predicted population label corresponding to the sample cluster cluster, and a comprehensive model loss value is determined based on the predicted population label and the sample population label. Based on the comprehensive model loss value, model parameters of the initial population detection large model are adjusted to obtain the population detection large model.
9. The method according to claim 8, wherein determining the comprehensive model loss value based on the predicted population label and the sample population label comprises: Determine at least two first population labels from the sample population labels, determine a second population label corresponding to the first population label from the predicted population label, determine a label loss value based on the first population label and the second population label, and determine a comprehensive model loss value based on the label loss value.
10. A device for determining a crowd label, comprising: Data acquisition module, used to obtain initial user attribute data of the target population; A crowd clustering module, configured to perform crowd clustering processing based on the initial user attribute data using a crowd clustering model to obtain crowd cluster representative feature data corresponding to a target cluster; The label detection module is used to use a large crowd detection model to perform crowd label detection processing based on the representative feature data of the crowd cluster to obtain crowd label description information corresponding to the target cluster.
11. A training device for a large crowd detection model, comprising: Model creation module, used to create the initial crowd detection model; A sample acquisition module is used to determine sample user attribute data of a sample population, and perform population clustering processing on the sample user attribute data to obtain representative characteristic data of the sample population cluster corresponding to the sample cluster; A labeling module is used to label the sample population cluster representative feature data with sample population labels; A model training module is used to perform at least one round of model training on the initial population detection large model using the representative feature data of the sample population clusters. During the model training process, the initial population detection large model is called based on the representative feature data of the sample population clusters to perform population label prediction processing to obtain the predicted population label corresponding to the sample cluster cluster, a comprehensive model loss value is determined based on the predicted population label and the sample population label, and model parameter adjustment processing is performed on the initial population detection large model based on the comprehensive model loss value to obtain the population detection large model.
12. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 7 or 8 to 9.
13. A computer program product, wherein the computer program product stores at least one instruction, wherein the at least one instruction is loaded by a processor and executes the method according to any one of claims 1 to 7 or 8 to 9.
14. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method according to any one of claims 1 to 7 or 8 to 9.